Genome-wide rare variant analysis method based on time-event outcome phenotypes

By constructing the time-event outcome analysis model and using functional annotation information to calculate the weight of rare variant sites, the problem of high false positives in the existing technology is solved, and efficient correlation analysis of the time-event outcome phenotype is achieved, providing disease-related biological genetic information.

CN118887996BActive Publication Date: 2025-07-22SHANGHAI JIAOTONG UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202411028196.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-30
Publication Date
2025-07-22
Estimated Expiration
2044-07-30

AI Technical Summary

Technical Problem

The existing genome-wide rare variant analysis methods mainly target continuous or subtype outcomes, lack analysis methods for time-event outcome phenotypes, and the existing methods have high false positive rates in rare variant association analysis and are not highly effective in tests.

Method used

Analytical model of time-event outcome was constructed, and the correlation between common variant sites and phenotypes was detected by marginal score test statistics, the weight of rare variant sites was calculated using functional annotation information, and the significance of the unbalanced phenotype was estimated by saddle point approximation method, and the correlation analysis of rare variant sites was enhanced by constructing joint test statistics.

Benefits of technology

It improves the identification accuracy of rare variant sites, reduces the false positive rate, can effectively control the impact of sample correlation in large biological databases, and provides biogenetic information on disease diagnosis, progression and prognosis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118887996B_ABST
    Figure CN118887996B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for analyzing rare variants in the whole genome based on time-event outcome phenotypes, comprising the following steps: obtaining genomic data and adding functional annotation information to each variant site; based on the genomic data, performing genome-wide association analysis by constructing a test statistic, and judging whether the phenotype is a balanced type or an unbalanced type according to the event incidence rate of the time-event outcome phenotype, and respectively using corresponding methods to estimate the association significance between the variant site and the phenotype; for common variant sites, performing genome-wide association analysis on the phenotype of the time-event outcome, and detecting the strength of the correlation between each single nucleotide variant site and the phenotype by constructing a marginal score test statistic; for rare variant sites, detecting the strength of the correlation between the rare variant site set and the phenotype by constructing a set-based rare variant joint test statistic. Compared with the prior art, the present invention can detect rare variant sites related to time-event outcomes and has the advantages of high accuracy in identifying variant sites.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of whole genome sequencing data analysis, and particularly to a method for analyzing whole genome rare variants based on time-to-event outcome phenotypes. Background Art

[0002] In recent years, the development of whole genome sequencing (WGS) technology has greatly promoted researchers' understanding of human genetic diversity. Large-scale WGS research data has made it possible to analyze rare variants (minor allele frequency less than 1%), which play an important role in the genetic research of complex diseases and traits and are considered to be one of the main sources of the heritability of complex diseases.

[0003] Currently existing rare variant analysis methods mainly target continuous outcomes and categorical outcomes. CN110910955A discloses a method for establishing a longitudinal analysis model for rare variant sites of susceptibility genes, including the following steps: obtaining whole genome sequence variation data of patient samples to be analyzed; observing and statistically counting the number of gene mutations observed on genes in patient samples, performing truncated negative binomial regression on genes, and constructing a generalized linear regression function; calculating the coefficients of the truncated negative binomial regression and the expected value of the estimated number of rare variant alleles of genes by using the maximum likelihood estimation function; calculating the standardized offset residuals between the observed mutation values of genes and the regression estimated baseline mutation numbers; converting the standardized offset residuals into statistical significance levels; removing significant genes in genes according to a preset threshold, and then repeating the above steps to refit the truncated negative binomial regression coefficients until all significant genes in patient samples are removed, so as to obtain a longitudinal analysis model for the rare variant load of susceptibility genes. This method targets continuous outcomes. Usually, the diagnosis status of a disease is a time-to-event type outcome, and currently there is no whole genome rare variant analysis method for time-to-event outcome phenotypes. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for analyzing whole genome rare variants based on time-to-event outcome phenotypes.

[0005] The purpose of the present invention can be achieved by the following technical solutions:

[0006] The present invention constructs an analysis model for time-to-event outcomes, constructs a marginal score test statistic for each variant site on the basis of this model. In the detection of rare variants, a set of rare variant sites (for example, in units of genes) will be screened out according to known information conditions, and then a joint test statistic will be constructed for this set, and the correlation between this set and the phenotype will be estimated according to this test statistic.

[0007] In the construction of common rare variant combined statistics, the corresponding Beta distribution weight w is usually calculated according to the minor allele frequency (MAF) of each rare variant site. However, in actual situations, the rarer a variant is, the less likely it is to be a variable causally associated with the phenotype. Therefore, directly calculating the weight corresponding to a rare variant based on MAF easily leads to a high false positive rate in the test. To avoid the occurrence of the above problems, the present invention, based on the FAVOR database, annotates multiple functional annotation information for each variant site, and calculates the weight corresponding to the rare variant site according to the size ranking of the functional annotations, in order to find rare variants with more biological significance, thereby reducing the false positive rate of rare variant tests.

[0008] The present invention includes but is not limited to the following:

[0009] 1) Conduct a genome-wide association analysis for time-to-event outcome phenotypes;

[0010] 2) Conduct a joint analysis on the set of rare variants in the whole genome sequencing data; in the rare variant analysis, use functional annotation information to calculate the weights of different variants, enabling the present invention to find rare sets more likely to be causally associated with the time-to-event outcome phenotype, thereby reducing the false positive rate of the genome-wide association analysis.

[0011] 3) For unbalanced time-to-event outcome phenotypes (event incidence rate < 10%), the present invention uses the saddlepoint approximation method to estimate the significance of the association between the variant site and the phenotype. This approximation method can be applied to the analysis of common variant sites and rare variant sites.

[0012] Specifically, the present invention provides a genome-wide rare variant analysis method based on time-to-event outcome phenotypes, including the following steps:

[0013] Obtain genomic data and add functional annotation information to each variant site;

[0014] Based on the genomic data, conduct a genome-wide association analysis by constructing a test statistic, and determine whether the phenotype is balanced or unbalanced according to the event incidence rate of the time-to-event outcome phenotype. Assume that the marginal score test statistic follows a normal distribution for balanced phenotypes, and estimate the significance of the association between the variant site and the phenotype; for unbalanced phenotypes, use the saddlepoint approximation method to estimate the significance of the association between the variant site and the phenotype.

[0015] Among them, for common variant sites, genome-wide association analysis is performed on the phenotypes of time-to-event outcomes, and the correlation strength between each single nucleotide variant site and the phenotype is detected by constructing a marginal score test statistic; for rare variant sites, a combined test analysis of rare variant sites is performed, and the correlation strength between the set of rare variant sites and the phenotype is detected by constructing a combined test statistic for all rare variants in the set.

[0016] Performing genome-wide association analysis on the phenotypes of time-to-event outcomes can detect the correlation strength between each single nucleotide variant site and the phenotype by constructing a marginal score test statistic S. The method of constructing a marginal score test statistic is usually used for the association analysis of common variant sites (MAF > 0.01), and its test power for rare variant sites (MAF ≤ 0.01) is not high.

[0017] For the common variant sites, genome-wide association analysis is performed on the phenotypes of time-to-event outcomes, and the correlation strength between each single nucleotide variant site and the phenotype is detected by constructing a marginal score test statistic, including the following steps:

[0018] S101, Estimate the polygenic effect γ related to the phenotype for each sample i based on the genome-wide regression model i ;

[0019] S102, Construct a genome-wide association analysis model: λ(t) = λ0(t)exp(Xα + Gβ + γ), where λ0(t) is the baseline hazard function, X is the covariate, G is the genotype data, α is the effect coefficient corresponding to the covariate, β is the effect coefficient corresponding to the gene, and γ is the polygenic effect calculated by S101;

[0020] S103, Based on the null hypothesis H0: β = 0 in hypothesis testing, obtain the null model λ(t) = λ0(t)exp(Xα + γ), and estimate the martingale residual R of the model by fitting the null model;

[0021] S104, Construct a marginal score test statistic for each single nucleotide variant site according to the martingale residual

[0022] S105, Construct the variance of the marginal score test statistic according to the score equation and information matrix of the null model Among them, Q is the matrix calculated based on the information matrix. Standardize the marginal score test statistic according to the variance to obtain the standardized score test statistic S:

[0023] In a balanced phenotype, it can be assumed that S follows a normal distribution, then

[0024] In the unbalanced phenotype, the assumption that S follows a normal distribution cannot be satisfied, and the saddlepoint approximation method needs to be used to estimate the association significance between variant sites and phenotypes.

[0025] To solve the problem of low test power in the association analysis of rare variants in time-to-event outcome phenotypes, the present invention proposes a rare variant site joint test analysis method STAAR-survival, which enhances the test power by constructing a joint test statistic for all rare variants in the set.

[0026] In the construction of common rare variant joint statistics, usually according to each rare variant site G j 's minor allele frequency MAF j calculate its corresponding Beta distribution weight w j . The commonly used calculation method in rare variant research is w j =Beta(MAF j ; 1,1) and w j =Beta(MAF j ; 1,25). To detect more likely biologically significant rare variant sites, in constructing the joint test statistic, in addition to the first weight w calculated by the traditional method, the present invention additionally adds a second weight π calculated based on functional annotation principal component information. The calculation method of the second weight π is as follows: By annotating genomic data through the FAVOR database, the K annotation principal components A j corresponding to each variant site G jk can be obtained. Based on the empirical cumulative distribution function, the weight π jk corresponding to the functional annotation principal component A jk =rank(A jk ) / M, where M represents the number of sites in the whole genome sequencing data and rank is the ranking function. The weights used in the present invention combine the traditional weight w j with the weight π jk based on functional annotation principal components, denoted as π jk w j .

[0027] In constructing the joint test statistic, the present invention simultaneously uses three different construction methods to combine the marginal test results of all rare variants in the set, and the three test statistics are respectively denoted as: Q 1k =(∑π k wS) 2 , Q 2k =∑π k w 2 S 2 , Q3k = ∑π k w 2 MAF 1 - MAF tan(0.5 - p)π}, where π k represents the weight of the rare variant site calculated according to the k-th functional annotation principal component information, w represents the weight calculated from the minor allele frequency of the rare variant site following a Beta distribution, S represents the marginal score test statistic of the rare variant site, p represents the marginal association significance between the rare variant site and the phenotype, and MAF represents the minor allele frequency; where,

[0028] The test statistic Q 1k follows a chi-square distribution with 1 degree of freedom, and the corresponding joint significance p-value is denoted as p 1k ;

[0029] The test statistic Q 2k follows a mixture chi-square distribution, and its corresponding joint significance p is calculated based on the Davies method 2k ;

[0030] The test statistic Q 3k follows a Cauchy distribution. Based on the Cauchy association test, the marginal significance p-value of each rare variant site is transformed into a value following a Cauchy distribution, and then weighted and combined together to calculate its corresponding joint significance p 3k . Among them, considering that the marginal significance estimation of extremely rare variant sites is not accurate, before constructing the test statistic Q3, a joint test statistic is constructed for the subset composed of all extremely rare variant sites in the set Then the significance corresponding to the statistic is denoted as p 0k , and the mean value of the weights of extremely rare variants is obtained to get Q 3k :

[0031]

[0032] where p′ is the number of all non-extremely rare variants in this set, is the mean value of the weights of the extremely rare variant subset

[0033] When constructing the joint test statistic, based on K different functional annotations, 3K test statistics are constructed. Based on the Cauchy association test method, the K significance p-values corresponding to each test statistic are combined into a joint statistic Q . = ∑tan{(0.5 - p . )π} / K, and its corresponding significance p-value is calculated as: where Q .Including three combined test statistics Q 1k , Q 2k , Q 3k The calculated p 1k , p 2k , p 3k After merging respectively, the corresponding combined statistics Q1, Q2, Q3, p . Including the combined significance p-values p1, p2, p3 corresponding to the three combined test statistics Q1, Q2, Q3;

[0034] Based on the combined significance p-values p1, p2, p3 corresponding to the combined statistics, the STAAR-survival combined test statistic Q is obtained s and the associated significance p corresponding to this statistic s :

[0035]

[0036]

[0037] Generally, in relevant research work on time-to-event outcome analysis, the occurrence status of an event is defined as occurrence and censoring according to whether the outcome event of interest occurs. According to the event incidence rate of this outcome phenotype, the phenotype can be divided into a balanced type (event incidence rate > 10%) and an unbalanced type (event incidence rate ≤ 10%).

[0038] In balanced events, the above hypothesis test holds, and it can be considered that the marginal statistic follows a chi-square distribution. However, in unbalanced events, this hypothesis no longer holds. At this time, the present invention uses the method of saddlepoint approximation estimation to estimate the significance of the association between the variant site and the phenotype according to the marginal statistic.

[0039] In common variant association analysis, the significance corresponding to its marginal statistic can be directly estimated using the method of saddlepoint approximation.

[0040] In rare variant association analysis, the present invention simultaneously uses three different ways of constructing statistics, and the method details corresponding to each construction method will be introduced in turn below.

[0041] The statistic Q1 is actually the score test statistic corresponding to the linear combination of multiple rare variant sites, so the significance corresponding to it can also be directly estimated using the method of saddlepoint approximation.

[0042] When constructing the test statistic Q2, first calculate the association significance of the saddlepoint approximation for each rare variant site, denoted as Based on the inverse cumulative distribution function of the standard normal, calculate The following gives two hypotheses:

[0043] 1. Denote the original covariance matrix of the set of rare variant sites as Σ. Assume that the diagonal elements of non - diagonal elements

[0044] 2. Assume the rare variant sites

[0045] Based on the above assumptions, construct the joint test statistic which follows a mixture chi - square distribution, and thus calculate the association significance between the rare variant set and the phenotype.

[0046] For the extremely rare variant sites in the set, the statistic calculated based on the inverse cumulative distribution function of the standard normal is inaccurate. Therefore, in the case of phenotypic imbalance, similar to the construction method of statistic Q3, construct the joint test statistic for the subset composed of all extremely rare variant sites in the set, calculate the saddle - point approximation significance corresponding to this statistic, then take the mean of the weights of the extremely rare variants, and then perform the above steps to obtain the test statistic Q in the case of the existence of extremely rare variants 2k :

[0047]

[0048] where, is the mean of the weights of the subset of extremely rare variants in this set, is the variance of the score test after combining extremely rare variants, is the test statistic calculated according to the association significance after combining extremely rare variants, p′ is the number of all non - extremely rare variants in this set.

[0049] When constructing statistic Q3, the significance P - value estimated by the saddle - point approximation can be directly substituted, and the remaining calculation methods are the same as those for balanced traits.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] (1) By using functional annotation information with biological significance in the analysis of rare variant sites based on the set, the present invention makes the identification of variant sites more accurate.

[0052] (2) By using the saddle - point approximation method to estimate the marginal correlation significance in unbalanced phenotypes, the present invention can control false positives in the association analysis of low - prevalence (unbalanced) phenotypes.

[0053] (3) The present invention can be applied to large - scale biobank databases and can well control the estimation impact caused by the correlation between samples in large - scale data. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 is the flowchart of the method of the present invention;

[0055] Figure 2 is the flowchart of the combined test of rare variant sites of the present invention;

[0056] Figure 3 is the schematic diagram of the construction method of the combined statistic of rare variant sites of the present invention;

[0057] Figure 4 is the Manhattan plot of the time - event outcome association analysis of Alzheimer's disease in an embodiment;

[0058] Figure 5 is the Manhattan plot of the rare variant association analysis of the time - event outcome of Alzheimer's disease in an embodiment. Detailed implementation manners

[0059] The present invention will be described in detail below with reference to the accompanying drawings and specific embodiments. This embodiment is implemented on the premise of the technical solution of the present invention, and gives detailed implementation manners and specific operation processes, but the protection scope of the present invention is not limited to the following embodiments.

[0060] The present invention discloses a genome - wide rare variant analysis method for time - event outcome phenotypes. This method mainly includes: 1. Conduct a genome - wide data association analysis for time - event outcome phenotypes; 2. Conduct a combined analysis for the set of rare variants in the whole - genome sequencing data; 3. For unbalanced time - event outcome phenotypes (event incidence < 10%), use the saddle - point approximation method to more accurately estimate the significance level of the association between variant sites and phenotypes.

[0061] Time - event outcome phenotypes include both the disease diagnosis status and the diagnosis time. The analysis based on time - event outcomes can provide valuable information about disease diagnosis, progression, and patient prognosis. Combining rare variant analysis with time - event outcome data allows researchers to discover more rare genetic variants related to disease diagnosis time and disease prognosis, and can gain a deeper understanding of the disease occurrence and development mechanism, providing more bio - genetic information for disease prediction, prevention, treatment, and prognosis.

[0062] As Figure 1 shown, this embodiment provides a genome - wide rare variant analysis method based on time - event outcome phenotypes, including the following steps:

[0063] Step 1) Obtain genomic data and add functional annotation information to each variant site.

[0064] Required data: Quality-controlled whole-genome sequencing data, time-to-event clinical phenotype data (e.g., whether the disease is diagnosed and the corresponding diagnosis time), and clinical covariate data related to the phenotype (such as gender, age, blood pressure, etc.). In this example, Alzheimer's disease is selected as the outcome phenotype, and the confirmed samples account for about 0.85% of the total samples, belonging to an unbalanced phenotype.

[0065] Preparation: Use the Plink tool to perform quality control and trimming on the whole-genome sequencing data (obtaining approximately 500,000 variant data); use the FAVOR database to perform functional principal component annotation on the whole-genome sequencing data.

[0066] Step 2) Based on the genomic data, perform a genome-wide association analysis by constructing a test statistic, and determine whether the phenotype is balanced or unbalanced according to the event incidence rate of the time-to-event outcome phenotype. Assume that the marginal score test statistic follows a normal distribution for the balanced phenotype, and estimate the association significance between the variant locus and the phenotype; for the unbalanced phenotype, use the saddlepoint approximation method to estimate the association significance between the variant locus and the phenotype.

[0067] Among them, for common variant loci, perform a genome-wide association analysis on the time-to-event outcome phenotype, and detect the strength of the correlation between each single nucleotide variant locus and the phenotype by constructing a marginal score test statistic; for rare variant loci, perform a combined test analysis of rare variant loci, and detect the strength of the correlation between the set of rare variant loci and the phenotype by constructing a combined test statistic for all rare variants in the set.

[0068] Performing a genome-wide association analysis on the time-to-event outcome phenotype can detect the strength of the correlation between each single nucleotide variant locus and the phenotype by constructing a marginal score test statistic S. The method of constructing the marginal score test statistic is usually used for the association analysis of common variant loci (MAF > 0.01), and the test power for rare variant loci (MAF ≤ 0.01) is not high.

[0069] For the common variant loci, performing a genome-wide association analysis on the time-to-event outcome phenotype and detecting the strength of the correlation between each single nucleotide variant locus and the phenotype by constructing a marginal score test statistic includes the following steps:

[0070] S101, Based on the whole-genome regression model, use a part of the quality-controlled and trimmed gene data to estimate the polygenic effect γ related to the phenotype for each sample i i ; In this step, to improve the operation efficiency, simplify the time-to-event data to event data, that is, only retain the occurrence status of the event.

[0071] S102, Construct a genome-wide association analysis model: λi λ(y) = λ0(y) exp(X i α + G i β + γ i ), where i = 1, …, n represents samples, λ0(t) is the baseline hazard function; X is clinical covariate data, α is the effect coefficient corresponding to the covariates; G is genotype data, β is the effect coefficient corresponding to the gene; γ i is the polygenic effect for each sample.

[0072] S103, based on the null hypothesis H0: β = 0 in hypothesis testing, obtain the null model λ i (t) = λ0(t) exp X i α + γ i ), by fitting the null model, estimate the martingale residuals R i ;

[0073] S104, construct a marginal score test statistic for each single nucleotide variant site according to the martingale residuals where G j is a variant site data, j = 1, …, p, R i is the martingale residual estimated in the previous step.

[0074] S105, construct the variance σ j 2 = Var(S j ) = G j T QG j , where Q is a matrix calculated based on the information matrix, standardize the marginal score test statistic according to the variance, and obtain the standardized score test statistic S j :

[0075] In a balanced phenotype, it can be assumed that S follows a normal distribution, then

[0076] When the survival phenotype is unbalanced (event incidence rate ≤ 10%), at this time no longer follows a chi-square distribution, and at this time, saddlepoint estimation needs to be used to approximate the hypothesis test P-value.

[0077] Since the time consumed by saddlepoint approximation is relatively long, and the estimation effect of the saddlepoint approximation method is unstable when S j is close to the mean. Therefore, in this embodiment, a fixed threshold is set. When |S| < 2σ, the normal distribution is used to approximately estimate the association significance. When |S| ≥ 2σ, saddlepoint approximation is used. This scheme is used in the subsequent steps.

[0078] Gene-centered rare variant set analysis. In this step, according to the chromosomes and specific positions of known genes in the human genome, the present invention matches all rare variant sites contained in the gene in the existing data, and combines these sites into a set for testing. For example, the existing rare variant set contains p rare variant sites, among which there are q (q ≤ p) extremely rare variant sites. In this embodiment, an extremely rare site is defined as a variant site with a minor allele count (MAC) less than or equal to 80.

[0079] To solve the problem of low test power in the association analysis of rare variants with time-to-event outcome phenotypes, the present invention proposes a rare variant site joint test analysis method STAAR-survival, as Figure 2 shown, by constructing a joint test statistic for all rare variants in the set to enhance the test power.

[0080] In the construction of common rare variant joint statistics, usually according to each rare variant site G j of the minor allele frequency MAF j to calculate its corresponding Beta distribution weight w j . In this embodiment, two different Beta distributions are selected to generate the weight w j = Beta(MAF j ; a1, a2), (a1, a2) ∈ {(1, 1), (1, 25)}. That is, w j = Beta(MAF j ; 1, 1) and w j = Beta(MAF j ; 1, 25). In order to detect more rare variant sites that are more likely to have biological significance, in the construction of the joint test statistic, in addition to the first weight w calculated by the traditional method, the present invention additionally adds a second weight π calculated according to the functional annotation principal component information. The calculation method of the second weight π is: by annotating the genomic data through the FAVOR database, the K annotation principal components A j corresponding to each variant site G jk can be obtained. Based on the empirical cumulative distribution function, the weight π jk corresponding to the functional annotation principal component A jk = rank(A jk ) / M, where M represents the number of sites in the whole genome sequencing data, and rank is the ranking function. The weights used in the present invention combine the traditional weight w j with the weight π jk based on the functional annotation principal component, denoted as πjk w j 。

[0081] When constructing the combined test statistic, the present invention uses three different construction methods to combine the marginal test results of all rare variants in the set, and the three test statistics are respectively denoted as: Q 1k =(∑π k wS) 2 , Q 2k =∑π k w 2 S 2 , Q 3k =∑π k w 2 MAF 1-MAF tan(0.5-p)π}, where π k represents the weight of the rare variant site calculated according to the k-th functional annotation principal component information, w represents the weight calculated from the minor allele frequency of the rare variant site subject to the Beta distribution, S represents the marginal score test statistic of the rare variant site, p represents the marginal association significance between the rare variant site and the phenotype, and MAF represents the minor allele frequency.

[0082] (1) As shown in Figure (3a), construct the combined test statistic where k = 1,…,K represents the weight π calculated according to K functional annotations k , l = 1,2 represents MAF j The weight w generated based on two different Beta distributions l , S j represents the marginal score test statistic of each rare variant.

[0083] Since the phenotype in this embodiment belongs to an unbalanced phenotype, the saddlepoint approximation method needs to be used to estimate the association significance p corresponding to the Q 1lk statistic. 1lk .

[0084] Based on the Cauchy combination test method, the K×2 significance P-values are combined into a combined statistic: The corresponding significance P-value is:

[0085] (2) As shown in Figure (3b), construct the combined test statistic Q 2lk , and the definitions of each parameter are the same as those of Q1.

[0086] Since the phenotype in this embodiment belongs to an unbalanced phenotype, the saddlepoint approximation method needs to be used to estimate a single rare variant site to obtain the association significance.

[0087] However, in the next step, since extremely rare variants are inaccurate when calculating the test statistic based on the inverse cumulative distribution function, first, based on the joint test statistic Q1, q extremely rare variants are linearly added based on weights and combined into one. At this time, there are s + 1 rare variant sites, where s = p - q.

[0088] Calculate the marginal correlation significance p of the saddle point approximation for these s + 1 rare variants i , i = 0, …, s. i = 0 represents the site obtained by combining extremely rare variants.

[0089] Based on the inverse cumulative distribution function of the standard normal, calculate

[0090] Construct the joint test statistic where represents the variance of the marginal score test statistic. The statistic Q 2lk follows a mixed chi-square distribution, and the association significance p can be calculated 2lk .

[0091] Based on the Cauchy combination test method, combine K × 2 significance P-values into a joint statistic: Its corresponding significance P-value is:

[0092] Among them, for the extremely rare variant sites in the set, the statistic calculated based on the inverse cumulative distribution function of the standard normal is inaccurate. Therefore, in the case of phenotypic imbalance, construct the joint test statistic for the subset composed of all extremely rare variant sites in the set, calculate the saddle point approximation significance corresponding to this statistic, then take the mean of the weights of the extremely rare variants, and then perform the above steps to obtain the test statistic Q in the case of the existence of extremely rare variants 2k :

[0093]

[0094] where is the mean of the weights of the extremely rare variant subset in this set, is the score test variance after the combination of extremely rare variants, is the test statistic calculated based on the association significance after the combination of extremely rare variants, and p′ is the number of all non-extremely rare variants in this set.

[0095] (3) As shown in Figure (3c), construct the joint test statistic The definitions of the parameters are the same as the settings of the previous two statistics. When i = 0, it represents the corresponding result of the extremely rare variant subset.

[0096] Since the phenotypes in this embodiment belong to unbalanced phenotypes, the saddlepoint approximation method needs to be used to estimate a single rare variant site to obtain the association significance.

[0097] When constructing this statistic, first calculate the Q of q extremely rare variant subsets 1lk The association significance obtained by the statistic based on the saddlepoint approximation method is denoted as p 0lk 。

[0098] Calculate the marginal correlation significance p of the remaining s = p - q rare variants i , i = 1, …, s.

[0099] Calculate Q based on the Cauchy combination test 3lk The corresponding correlation significance p 3lk 。

[0100] Based on the Cauchy combination test method, combine the K × 2 significance P values into a joint statistic: The corresponding significance P value is:

[0101] Considering that the marginal significance estimation of extremely rare variant sites is not accurate, before constructing the test statistic Q3, construct its joint test statistic for the subset composed of all extremely rare variant sites in the set Then Denote the significance corresponding to the statistic as o 0k , take the mean of the weights of extremely rare variants to obtain Q 3k :

[0102]

[0103] where p′ is the number of all non - extremely rare variants in this set, is the mean of the weights of extremely rare variant subsets. Among them, the extremely rare variant sites can be variant sites with a minor allele count (MAC) less than 10.

[0104] (4) Based on the joint significance p values p1, p2, p3 corresponding to the joint statistic, obtain the STAAR - survival joint test statistic Q s and the association significance p corresponding to this statistic s :

[0105]

[0106]

[0107] Figure 4It shows the Manhattan plot of the genome-wide association analysis results of the diagnosis-time outcome of Alzheimer's disease for common variant sites in this embodiment. This phenotype is of the unbalanced type and usually has strong false positives in research. However, this method can well control the false positives in the association analysis by using saddlepoint approximation. As can be seen from the figure, there are a large number of significant single nucleotide variant sites only in the region of chromosome 19, and almost no significant sites on other chromosomes.

[0108] Figure 5 And Table 1 shows the results of rare variant site analysis in this embodiment. Consistent with Figure 4 this, this method well controls the false positives in the association analysis. What is shown in Table 1 are the significant gene sets in this embodiment, all of which are located in the region of chromosome 19. "The number of existing reported relevant studies" refers to the number of single nucleotide variant sites in this gene that are significantly associated with the Alzheimer's disease phenotype in all known research reports.

[0109] Table 1 Significant gene sets for rare variant association analysis of Alzheimer's disease time-event outcome

[0110]

[0111] The preferred specific embodiments of the present invention have been described in detail above. It should be understood that those of ordinary skill in the art can make many modifications and variations based on the concept of the present invention without creative labor. Therefore, all technical solutions that can be obtained by those skilled in the art in the technical field based on the concept of the present invention through logical analysis, reasoning, or limited experiments on the basis of the prior art should fall within the scope of protection determined by the claims.

Claims

1. A genome-wide rare variant analysis method based on time-event outcome phenotypes, characterized in that, Including the following steps: Obtain genomic data and add functional annotation information to each variant site; Based on the genomic data, conduct a genome-wide association study by constructing a test statistic. According to the event incidence rate of the time-to-event outcome phenotype, determine whether the phenotype is balanced or unbalanced. Assume that the marginal score test statistic for the balanced phenotype follows a normal distribution, and estimate the association significance between the variant site and the phenotype. For the unbalanced phenotype, use the saddlepoint approximation method to estimate the association significance between the variant site and the phenotype. The method for determining whether the phenotype is balanced or unbalanced is as follows: Define the occurrence status of the event as occurrence and censoring according to whether the outcome event of interest occurs. When the event incidence rate of the outcome phenotype > 10%, the corresponding phenotype is balanced. When the event incidence rate of the outcome phenotype ≤ 10%, the corresponding phenotype is unbalanced; Among them, for common variant sites, conduct a genome-wide association study on the time-to-event outcome phenotype, and detect the strength of the correlation between each single nucleotide variant site and the phenotype by constructing a marginal score test statistic. For rare variant sites, conduct a rare variant site joint test analysis, and detect the strength of the correlation between the rare variant site set and the phenotype by constructing a joint test statistic for all rare variants in the set.

2. The whole genome rare variant analysis method based on time-event outcome phenotypes according to claim 1, wherein For the common variant sites, conducting a genome-wide association study on the time-to-event outcome phenotype and detecting the strength of the correlation between each single nucleotide variant site and the phenotype by constructing a marginal score test statistic includes the following steps: S101, estimating the polygenic effects related to phenotypes for each sample based on the whole-genome regression model i ;​ S102, construct a genome-wide association analysis model: , where is the baseline risk function, is the covariate, is the genotype data, is the effect coefficient corresponding to the covariate, is the effect coefficient corresponding to the gene, is the polygenic effect calculated from S101; S103, based on the null hypothesis in hypothesis testing , obtain the null model , estimate the martingale residuals of the model by fitting the null model ; S104. Construct a marginal score test statistic for each single nucleotide variant site according to the martingale residual ; S105. Construct the variance of the marginal score test statistic according to the score equation and information matrix of the null model , where is a matrix calculated based on the information matrix. Standardize the marginal score test statistic according to the variance to obtain the standardized score test statistic : .

3. A method for analyzing rare variants across the genome based on time-event outcome phenotypes according to claim 1, wherein The combined test statistic is constructed based on weights, where the weights include a first weight calculated based on the minor allele frequency and a second weight calculated according to the principal component information of functional annotation .

4. The whole genome rare variant analysis method based on the time-event outcome phenotype according to claim 3, characterized in that, The second weight is calculated as follows: annotate the genomic data through the FAVOR database to obtain each corresponding annotated principal component ; based on the empirical cumulative distribution function, calculate the weight corresponding to the functionally annotated principal component , where M represents the number of sites in the whole genome sequencing data, and is the ranking function.

5. A method for analyzing rare variants in the whole genome based on a time-event outcome phenotype according to claim 1, characterized in that, When constructing the combined test statistic, three different construction methods are used to combine the marginal test results of all rare variants in the set. The three test statistics are denoted as follows: , , , where represents the weight of rare variant sites calculated based on the information of the th functional annotation principal component, represents the weight calculated from the minor allele frequency of rare variant sites following a Beta distribution, represents the marginal score test statistic of rare variant sites, represents the marginal association significance between rare variant sites and the phenotype, represents the minor allele frequency; among them, Test statistic follows a chi-square distribution with 1 degree of freedom, and its corresponding joint significance p value is denoted as ; Test statistic follows a mixture chi-square distribution, and its corresponding joint significance is calculated based on the Davies method ; Test statistic follows a Cauchy distribution. Based on the Cauchy combination test method, the marginal significance value of each rare variant site is transformed into a value following a Cauchy distribution, and then weighted and combined together to calculate its corresponding joint significance . Among them, before constructing the test statistic , a joint test statistic is constructed for the subset composed of all extremely rare variant sites in the set . Then, the significance corresponding to the statistic is denoted as . The mean value of the weights of extremely rare variants is obtained to get : where, is the number of all non-extremely rare variants in the set, is the mean of the weights of the extremely rare variant subset.

6. The whole genome rare variant analysis method based on a time-event outcome phenotype according to claim 5, wherein When constructing the combined test statistic, based on different functional annotations, test statistics are constructed. Based on the Cauchy combination test method, the K significance values corresponding to each test statistic are combined into a combined statistic , and its corresponding significance value is calculated as: , where includes three combined test statistics calculated the combined statistics corresponding to the respective combinations , includes the combined significance corresponding to three combined test statistics values ; Based on the joint significance corresponding to the joint statistic value obtain the STAAR - survival joint test statistic and the associated significance corresponding to this statistic : 。 7. A method for whole-genome rare variant analysis based on a time-event outcome phenotype according to claim 6, characterized in that For the test statistics in the association analysis of common variant sites and the association analysis of rare variant sites and , directly use the saddlepoint approximation method to estimate the significance corresponding to the marginal score test statistic 8. The whole genome rare variant analysis method based on time-event outcome phenotypes according to claim 6, characterized in that, Test statistic in the association analysis of rare variant sites , first calculate the association significance of the saddlepoint approximation for each rare variant site, denoted as , and calculate based on the inverse cumulative distribution function of the standard normal. The following gives two hypotheses: 1) Denote the original covariance matrix of the rare variant site set as , assume , The diagonal elements of , the off-diagonal elements , 2) Assume rare variant sites Based on the above assumptions, a combined test statistic is constructed , which follows a mixture chi-square distribution, and the association significance between the rare variant set and the phenotype is calculated therefrom.

9. The whole-genome rare variant analysis method based on a time-event outcome phenotype according to claim 8, wherein In the case of phenotypic imbalance, constructing a test statistic in the association analysis of rare variant sites Before, construct a joint test statistic for the subset composed of all extremely rare variant sites in the set, and then calculate the mean of the weights of the extremely rare variants to obtain the test statistic in the case of the existence of extremely rare variants : Among them, is the mean of the weights of the subsets of extremely rare variants in this set, is the variance of the score test after combining extremely rare variants, is the test statistic calculated based on the association significance after combining extremely rare variants, is the number of all non-extremely rare variants in this set.

Citation Information

Patent Citations

  • Method for establishing longitudinal analysis model of rare variation point of susceptible gene

    CN110910955A

  • Evaluation method for judging rare hereditary diseases

    CN112735599A

  • Analysis method suitable for gene-environment interaction of complex characters and storage medium

    CN114898809A