多序列大规模多重检验的信号分类整合分析(项冬冬)

The integrative analysis of multiple data sets is becoming increasingly important in many fields of research. When the same features are studied in several independent experiments, it can often be useful to analyze jointly the multiple sequences of multiple tests that result. It is frequently necessary to classify each feature into one of several categories, depending on the null and non-null configuration of its corresponding test statistics. The paper studies this signal classification problem, motivated by a range of applications in large-scale genomics. Two new types of misclassification rate are introduced, and two oracle procedures are developed to control each type while also achieving the largest expected number of correct classifications. Corresponding data-driven procedures are also proposed, proved to be asymptotically valid and optimal under certain conditions and shown in numerical experiments to be nearly as powerful as the oracle procedures. In an application to psychiatric genetics, the procedures proposed are used to discover genetic variants that may affect both bipolar disorder and schizophrenia, as well as variants that may help to distinguish between these conditions.

 

Publication: 

Journal of the Royal Statistical Society – Statistical Methodology, Series B,  (2019) 81, Part 4, pp. 707–734

 

Authors: 

Dongdong Xiang, East China Normal University, Shanghai, People’s Republic of China Shanghai

Dave Zhao University of Illinois at Urbana–Champaign, USA

and

T. Tony Cai University of Pennsylvania, Philadelphia, USA

 

ddxiang@sfs.ecnu.edu.cn


来源:3044永利集团发布时间:2022-10-12浏览次数:188