Statistical methods for horizontal pleiotropy correction in representative underrepresented populations using TWAS
By constructing a TWAS model for underrepresented populations and performing level pleiotropic correction, the problem of decreased gene expression prediction accuracy and statistical power in GWAS was solved. This enabled accurate statistical inference and effect estimation of gene-trait associations, controlled false findings, and improved the statistical power of TWAS studies.
Patent Information
- Application Number
- CN202610665885.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-14
- Publication Date
- 2026-08-25
AI Technical Summary
When existing genome-wide association studies (GWAS) models established in a single population are applied to other populations, the limited sample size leads to a significant decrease in the accuracy of gene expression prediction and the statistical power of gene association tests. Furthermore, the problem of level pleiotropy has not been fully resolved, resulting in the identification of spurious association genes.
The TWAS statistical method was adopted for underrepresented populations with level pleiotropic effect correction. By constructing a TWAS model for underrepresented populations, a smoothing truncation absolute bias penalty function was introduced to select non-zero effects. The least squares estimation method was combined to estimate and infer the effects and level pleiotropic effects of individual level data and summary statistics.
It improves the testing power of associated genes in underrepresented populations, effectively controls false findings, quantifies the contribution of expression prediction models to gene-trait associations in different populations, and theoretically proves the asymptotic consistency and oracle property of the estimator.
Smart Images

Figure CN122637869A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of genome-wide association studies, and in particular to a TWAS statistical method for underrepresented populations with level pleiotropy correction. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Genome-wide association studies (GWAS) identify genetic variations (generally single nucleotide polymorphisms, SNPs) associated with complex human traits or diseases. However, studies have shown that most of the associated SNPs identified by GWAS are difficult to verify experimentally directly, and the underlying molecular mechanisms remain unclear, limiting the biological interpretation of the results.
[0004] A key pathway by which genetic variations influence complex traits is through the regulation of neighboring gene expression. To this end, researchers often use transcriptome-wide association studies (TWAS) to integrate GWAS with expression quantitative trait locus (eQTL) studies to identify genes associated with complex traits or diseases, thereby further enhancing the interpretability of disease genetic mechanisms.
[0005] In TWAS, researchers first construct a gene expression prediction model for a specific gene using cis-SNPs (cis-SNPs, typically 500kb upstream of the transcription start site and 500kb downstream of the transcription termination site) from eQTL data. Subsequently, the effect weights of the cis-SNPs are applied to genotype data in GWAS studies to predict the gene's expression level in those studies. Finally, the gene is identified by an association test between the predicted gene expression level and the target phenotype. Crucially, this method achieves gene expression prediction in large-scale GWAS cohorts by integrating eQTL studies, avoiding reliance on actual transcriptome measurements and enabling expression-level association analysis even in large-scale GWAS studies with only genotype data.
[0006] Existing GWAS studies mainly focus on single populations. When TWAS models established in single populations are applied to other populations, the accuracy of gene expression prediction and the statistical power of gene association tests are significantly reduced due to the limited sample size.
[0007] To address this issue, large-scale GWAS studies can be conducted in underrepresented populations to increase the sample size of GWAS studies, thereby improving the statistical power of TWAS studies. However, this method requires significant human and financial resources. Furthermore, the critical challenge of horizontal pleiotropy remains largely unresolved. Horizontal pleiotropy occurs when cis-SNPs influence traits through pathways other than gene expression regulation, potentially leading to the identification of spurious associated genes in TWAS analyses. Summary of the Invention
[0008] To address the aforementioned issues, this invention proposes a TWAS statistical method for underrepresented populations with level pleiotropy correction. Within the TWAS framework, it accurately identifies cis-SNPs with level pleiotropy, enabling accurate statistical inference and effect estimation of gene expression prediction on phenotype.
[0009] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a TWAS statistical method for underrepresented populations with level pleiotropic correction, comprising: For each gene, consider the effect of its predicted expression value in GWAS studies on phenotype. And the level pleiotropic effect of cis-SNPs on phenotypes. Therefore, a TWAS model for underrepresented populations was constructed. Based on the TWAS model for underrepresented populations, this study introduces a smoothed absolute bias penalty function to select non-zero effects, and combines least squares estimation to evaluate the effects on individual level data and summary statistics. and level pleiotropic effects Estimates and inferences.
[0010] As an alternative implementation method, the TWAS model for underrepresented populations is as follows: ; ; in, It is the first In a group of people individual Dimensional gene expression vector; The total number of people; It is cis-SNPs 3D genotype matrix, This represents the number of cis-SNPs for this gene; It is the first cis-SNPs on gene expression in individual populations dimensional effect weight vector; yes 3D residual vector; yes 3D vector, representing Phenotypic measurements of each individual; For the corresponding cis-SNPs in GWAS studies 3D genotype matrix; It is the effect of cis-SNPs on gene expression in GWAS studies. dimensional effect vector, Indicates the first The effect weights of individual populations satisfy the following conditions: and ; This refers to the genetic regulatory expression component GReX constructed in the GWAS study, which is the predicted expression value of the gene in the GWAS study; It is the effect of GReX on phenotype. This indicates that, given the GReX effect in other populations, the first... The effect of GReX on phenotypes constructed in individual populations; yes A dimensional vector, representing the level pleiotropic effect of cis-SNPs on the phenotype; yes 3D residual vector.
[0011] As an alternative implementation method, effects are applied to individual-level data. and level pleiotropic effects The estimation process includes: ; ; in, Horizontal pleiotropic effect The estimated value, For effect The estimated value; yes 3D vector, representing Phenotypic measurements of each individual; , Let be the projection matrix. yes 3D identity matrix; For the corresponding cis-SNPs in GWAS studies 3D genotype matrix; It is the SCAD penalty function. It refers to adjusting parameters; The j-th cis-SNP of a gene represents the level pleiotropic effect on the phenotype. , express The genetic regulatory expression component GReX constructed in a GWAS study for an individual population is the gene expression prediction value in the GWAS study.
[0012] As an alternative implementation method, effects are applied to individual-level data. and level pleiotropic effects The inference process following estimation includes: Estimated quantity The asymptotic distribution is ; Among them, standard error for: ; Depend on It was estimated; Among them, set cis-SNP pairs It has a horizontal pleiotropic effect. , For the set The genotype matrix composed of cis-SNPs; For set The level of pleiotropic effects of cis-SNPs in China for The specific variance of each individual; For mathematical expectation operators; For the first Individuals are composed of sets The genotype vector composed of cis-SNPs; For the first The predicted expression values of genes in an individual in a GWAS study.
[0013] As an alternative implementation method, effects are applied to the aggregated statistical data. and level pleiotropic effects The estimation process includes: ; in,
[0014]
[0015] ; ; in, Marginal effect; Horizontal pleiotropic effect The estimated value, For effect The estimated value; yes 3D vector, representing Phenotypic measurements of each individual; , Let be the projection matrix. yes 3D identity matrix; For the corresponding cis-SNPs in GWAS studies 3D genotype matrix; It is the SCAD penalty function. It refers to adjusting parameters; The j-th cis-SNP of a gene represents the level pleiotropic effect on the phenotype. , express The genetic regulatory expression component GReX constructed in the GWAS study for an individual population, i.e. the gene expression prediction value in the GWAS study; The correlation matrix for the required cis-SNPs; for Weiquanwei ; , The effect weight vector of cis-SNPs on gene expression in the Mth population; and These are all intermediate parameters defined to simplify the formula. , .
[0016] As an alternative implementation method, effects are applied to the aggregated statistical data. and level pleiotropic effects The inference process following estimation includes: Estimated quantity The asymptotic distribution is ; Among them, standard error ; ; Estimate using the following formula: ; in, for The specific variance of each individual; set cis-SNP pairs It has a horizontal pleiotropic effect. , For the set The genotype matrix composed of cis-SNPs; for The covariance matrix; for and The covariance matrix; for and The covariance matrix; for and The covariance matrix.
[0017] Secondly, the present invention provides a TWAS statistical system for underrepresented populations with level pleiotropic correction, comprising: The modeling module is configured to, for each gene, consider the effect of the gene's predicted expression in GWAS studies on the phenotype. And the level pleiotropic effect of cis-SNPs on phenotypes. Therefore, a TWAS model for underrepresented populations was constructed. The estimation and inference module is configured to use a TWAS model based on an underrepresented population. It selects non-zero effects by introducing a smoothed absolute bias penalty function and combines least squares estimation to evaluate the effects on individual level data and summary statistics. and level pleiotropic effects Estimates and inferences.
[0018] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0019] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0020] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0021] Compared with the prior art, the beneficial effects of the present invention are as follows: For TWAS studies in underrepresented populations, and considering the sparsity of SNPs' phenotypic effects within genomic regions, this invention proposes a statistical method for TWAS in underrepresented populations corrected for horizontal pleiotropy. First, using large-sample population information, a horizontal pleiotropy-corrected TWAS model for underrepresented populations is constructed. Then, non-zero variables are selected for horizontal pleiotropy effects in the underrepresented population TWAS model to determine the pathways through which SNPs affect the phenotype without gene expression. Based on this, statistical inferences and estimates of the phenotypic effects of predicted gene expression are made based on theoretical properties. By utilizing cross-population genetic similarity between the underrepresented target population and the large-sample population, and correcting for horizontal pleiotropy, the contribution of expression prediction models to gene-trait associations in different populations can be quantified, thereby improving the testing power of associated genes in underrepresented populations and effectively controlling false findings. Furthermore, the asymptotic consistency of the proposed estimators and the Oracle property of selecting horizontal pleiotropy effect variables are theoretically proven.
[0022] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0024] Figure 1 The flowchart of the TWAS statistical method for level pleiotropic correction in underrepresented populations provided in Embodiment 1 of the present invention; Figures 2(A)-2(I) show the test results under different levels of pleiotropic effects provided in Embodiment 1 of the present invention. Detailed Implementation
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments.
[0026] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0027] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the scope of exemplary embodiments according to the invention. As used herein, unless the context clearly indicates otherwise, the singular form is intended to include the plural form as well. Furthermore, it should be understood that the terms “comprising” and “including”, and any variations thereof, are intended to cover non-exclusive inclusion, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0028] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0029] Example 1 This embodiment provides a TWAS statistical method for underrepresented populations with level pleiotropic correction, such as... Figure 1 As shown, it includes: For each gene, consider the effect of its predicted expression value in GWAS studies on phenotype. And the level pleiotropic effect of cis-SNPs on phenotypes. Therefore, a TWAS model for underrepresented populations was constructed. Based on the TWAS model for underrepresented populations, this study introduces a smoothed absolute bias penalty function to select non-zero effects, and combines least squares estimation to evaluate the effects on individual level data and summary statistics. and level pleiotropic effects Estimates and inferences.
[0030] The method of this embodiment will be described in detail below.
[0031] 1. Model building.
[0032] For those with Full rank matrix of rows and columns ,definition ,in yes The projection on the column space, yes 3D identity matrix.
[0033] For each specific gene, construct an underrepresented population TWAS model: (1); (2); in, It is the first ( ) in a group of people individual Dimensional gene expression vector, The total number of people; It is cis-SNPs 3D genotype matrix, This represents the number of cis-SNPs for this gene; It is the first cis-SNPs on gene expression in individual populations dimensional effect weight vector; yes A dimensional residual vector, wherein each element independently follows a normal distribution. , It is the first Individual population-specific variance.
[0034] in, yes 3D vector, representing Phenotypic measurements of each individual; For the corresponding cis-SNPs in GWAS studies Genotype matrix; It is the effect of cis-SNPs on gene expression in GWAS studies. dimensional effect vector, Indicates the first The effect weights of individual populations satisfy the following conditions: and ; This represents the genetic regulatory expression component (GReX) constructed in the GWAS study, i.e., the predicted expression value of the gene in the GWAS study; It is the effect of GReX on phenotype. This indicates that, given the GReX effect in other populations, the first... The effect of GReX on phenotypes constructed in individual populations; yes A dimensional vector representing the level pleiotropic effect of these cis-SNPs on the phenotype; yes A dimensional residual vector, wherein each element independently follows a normal distribution. , for The specific variance of each individual.
[0035] Without loss of generality, we assume , , and Each column has been standardized so that its sample mean is 0 and its sample variance is 1.
[0036] In the above model, the key parameter of interest in this embodiment is the effect of GReX on the phenotype. What I'm interested in is testing the null hypothesis. This is equivalent to the null hypothesis. With the null hypothesis true, GReX has no phenotypic effect in all populations. Therefore, the alternative hypothesis is that GReX has a non-zero phenotypic effect in at least one population. The Wald test is used for hypothesis testing, where the test statistic is based on the estimated effect size and its corresponding standard error. Constructed.
[0037] definition , Equation (2) is rewritten as: (3).
[0038] assumed This is given by a pre-trained prediction model obtained through eQTL research, ignoring estimation errors. For simplicity and without ambiguity, let... Alternative ,in ; for A dimensional vector, representing the 3rd dimension vector. The genotype of each individual.
[0039] make and They are respectively and The true value, set cis-SNP pairs It has a direct effect, that is, a horizontal pleiotropic effect; in other words... .
[0040] Oracle estimates as follows: (4); in, The estimated effect; For set Estimates of the level pleiotropic effects of cis-SNPs in China; For the set The genotype matrix composed of cis-SNPs; For set The level pleiotropic effect of cis-SNPs in China.
[0041] therefore, The least squares estimate is: (5); in, , This is the projection matrix.
[0042] The least squares estimate is: (6); in, , This is the projection matrix.
[0043] Under the standard assumptions, we have: (7).
[0044] in, (8); in, For mathematical expectation operators; For the first Individuals are composed of sets The genotype vector is composed of cis-SNPs.
[0045] because It is unknown, and can be determined by... It was estimated.
[0046] 2. Statistical inference based on individual data.
[0047] In equation (4), it is assumed that It is known. However, in general, it is unknown which cis-SNPs belong to the set. Therefore, a variable selection method based on SCAD (smoothly clipped absolute deviation) is proposed for estimation. and The specific formula is as follows: (9); in, It is the SCAD penalty function. It refers to adjusting parameters; yes 3D vector, representing Phenotypic measurements of an individual; j represents the j-th cis-SNP of a gene; This represents the number of cis-SNPs for this gene; The j-th cis-SNP of a gene represents the level pleiotropic effect on the phenotype.
[0048] Equation (9) is solved using the following algorithm; (10).
[0049] because: (11).
[0050] make ,have: (12).
[0051] Solve Equivalent to solving equation (13): (13).
[0052] make (14).
[0053] therefore, The estimated value can be equivalently estimated by equation (15): (15).
[0054] Correspondingly, The estimated value is: (16).
[0055] 3. Theoretical nature (statistically guaranteed).
[0056] The asymptotic properties of the estimate are examined below.
[0057] Define the following symbols: make for The true value, without loss of generality, It can be decomposed into ,in All elements are non-zero. All elements are zero.
[0058] akin, The corresponding decomposition is as follows ,in and .
[0059] Therefore, the result obtained from equation (15) Decomposed into .
[0060] make , ,in, The first derivative of the penalty function; The second derivative of the penalty function; For the first The true value of the pleiotropic effect at the cis-SNP level.
[0061] make ,in and They are respectively and The covariates.
[0062] Assume the following regularization conditions: Assumption 1: There exists a positive definite matrix Make Furthermore, assuming .Will Top left corner The submatrix is represented as , the bottom right corner The submatrix is represented as .
[0063] Assumption 2: , .
[0064] Assumption 3: Adjusting parameters Satisfy when hour, and The penalty function satisfies .
[0065] Assumption 1 ensures that the principal term in the Taylor expansion is greater than the remainder term.
[0066] Assumption 2 ensures the unbiasedness of large parameters, and Existence of consistent penalized least squares estimators.
[0067] Assumption 3 makes the penalty function singular at the origin, thus making the penalty estimator sparse.
[0068] Theorem 1: Assuming that assumptions 1-3 hold, there exists a local minimum. Make .
[0069] Theorem 2: Assuming that assumptions 1-3 are true, satisfy: (a) Sparsity: The probability is 1.
[0070] (b) Asymptotic normality: ; in , , and .
[0071] Inference 1: Under the conditions of Theorem 2, the estimator The asymptotic distribution is ,in, As defined in equation (8).
[0072] 4. Summarize statistics and make statistical inferences.
[0073] Due to data management and privacy concerns, individual-level genotypic and phenotypic data are typically not publicly available. However, publicly available summary statistics, such as the size of marginal effects (i.e.,...), are available in research. In addition, the correlation matrix of the required cis-SNPs (i.e. Reference panels (such as the 1000 Genomes Project) are needed to approximate the estimation in order to infer the linkage disequilibrium matrix.
[0074] For the sake of simplicity, let for Weiquanwei vector: ; ; .
[0075] Among them, the correlation matrix of cis-SNPs ; This represents the effect weight vector of cis-SNPs on gene expression in the first population. Let be the effect weight vector of cis-SNPs on gene expression in the Mth population. and These are intermediate parameters, defined to simplify the formula.
[0076] Considering With some simplification, we have: (17).
[0077] in,
[0078]
[0079] ; Among them, marginal effect .
[0080] further, (18).
[0081] Similar to statistical inference methods for individual data, the summary statistic method also yields: (19); in, , These are intermediate parameters, defined to simplify the formula. The estimated effect.
[0082] Notes on subsets , yes The Line 1 Column elements, yes The covariance matrix.
[0083] Define matrix: (20).
[0084] in, for The covariance matrix; for and The covariance matrix; for and The covariance matrix; for and The covariance matrix.
[0085] For summary statistics, please refer to the reference panel. .
[0086] because It is unknown, but can be estimated using equation (21): (twenty one).
[0087] To comprehensively evaluate the performance of the method in this embodiment and compare it with existing methods, a series of simulation studies were conducted based on real genotype data. In the simulations, real genotype data were used to generate gene expression and outcome traits. To better approximate the real-world scenario, the same 4646 genes used in MA-FOCUS were simulated. For each gene, genotype data was obtained by matching the TWAS weight information provided by MA-FOCUS with genotype data from the UK Biobank located within the range of 500 kb upstream of the transcription start site to 500 kb downstream of the transcription termination site. Subsequently, SNPs with an absolute correlation greater than 0.9 between each pair of genes were removed. The genotype data was restricted to high-quality HapMap SNPs, and genotype loss rates >1%, minor allele frequencies <1%, and Hardy-Weinberg equilibrium tests were removed. Value < 1×10 -6 SNPs.
[0088] First, gene expression data is generated. For each gene, it is randomly extracted from the real genotype data. and Each sample serves as the genotype data for two population groups. For each population group, the genotypes are analyzed using a normal distribution. Simulate the genetic effects of SNPs, among which It represents the proportion of variance explained by SNPs in gene expression. This represents the number of SNPs for each gene. Furthermore, it is determined through a normal distribution. Residual errors are generated. Finally, genetic effects and residual errors are combined to generate population-specific gene expression.
[0089] Then, phenotypic data for the GWAS study were generated. For each gene, 40,000 samples were randomly extracted from the real genotype data as genotype data for the population with a smaller GWAS sample size. The ratio of the expression of specific genes to the phenotypic effect between the two races was set as follows: And set the proportion of variance explained by overall gene expression in the outcome trait as For the level of pleiotropic effects, the following settings are made: The SNPs exhibit sparse level pleiotropic effects, and the proportion of variance explained for the outcome trait is set to be... Furthermore, from the variance setting The residual error of the simulated trait in a normal distribution is calculated. The gene expression effect, level pleiotropic effect, and residual error are combined to obtain the resulting trait.
[0090] Simulation results.
[0091] Gene expression in two populations was simulated using real genotype data, and then the outcome traits of an underrepresented population were generated within an independent GWAS framework. These simulated datasets were used to evaluate the method in this embodiment and other TWAS methods.
[0092] Figures 2(A)-2(I) show the results of gene-trait association testing under different levels of pleiotropy. The QQ diagrams are shown; Figure 2(A) shows the case where there is no level pleiotropy; Figure 2(B) shows the case where level pleiotropy satisfies Egger's hypothesis; Figure 2(C) shows... , The situation is as shown in Figure 2(D). , The situation is as shown in Figure 2(E). , The situation is as shown in Figure 2(F). , The situation is as shown in Figure 2(G). , The remaining parameters are set to the baseline case; Figure 2(H) shows the situation in... Under the condition of no level pleiotropy, the QQ diagram; Figure 2(I) shows the QQ diagram under the condition of no level pleiotropy. Under the condition that the level of pleiotropic effect is set as follows: , QQ image in the scene. P-value less than The setting is .
[0093] The comparison methods include: the individual data version of METRO and the individual data version of this method, and the summary statistics version of METRO and the summary statistics version of this method.
[0094] First, in null simulations, the control capabilities of different methods for Type I errors were evaluated under different levels of pleiotropy. The results showed that when pleiotropy was absent, both our method and METRO produced well-calibrated p-values (A). However, when pleiotropy was present, METRO failed to control Type I errors, while our method maintained its robust performance (BF). Notably, when the pleiotropy effects were very small (G), the p-value of our method began to inflate slightly, but its performance still outperformed other methods. Importantly, regardless of the proportion of variance explained by SNPs in gene expression (…),… Regardless of how the level of pleiotropy changes, this method can robustly control Type I errors both with and without level pleiotropy. In contrast, when level pleiotropy is present, METRO loses control over Type I errors (HI).
[0095] Example 2 This embodiment provides a TWAS statistical system for underrepresented populations with level pleiotropic correction, including: The modeling module is configured to, for each gene, consider the effect of the gene's predicted expression in GWAS studies on the phenotype. And the level pleiotropic effect of cis-SNPs on phenotypes. Therefore, a TWAS model for underrepresented populations was constructed. The estimation and inference module is configured to use a TWAS model based on an underrepresented population. It selects non-zero effects by introducing a smoothed absolute bias penalty function and combines least squares estimation to evaluate the effects on individual level data and summary statistics. and level pleiotropic effects Estimates and inferences.
[0096] It should be noted that the above modules correspond to the steps described in Embodiment 1, and the examples and application scenarios implemented by the above modules and the corresponding steps are the same, but are not limited to the content disclosed in Embodiment 1. It should also be noted that the above modules, as part of the system, can be executed in a computer system such as a set of computer-executable instructions.
[0097] In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0098] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0099] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0100] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0101] The method in Example 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the field, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0102] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0103] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0104] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0105] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0106] Those skilled in the art will recognize that the units and algorithm steps described in connection with the various examples of this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this invention.
[0107] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A TWAS statistical method for underrepresented populations with pleiotropic level correction, characterized in that, include: For each gene, consider the effect of its predicted expression value in GWAS studies on phenotype. And the level pleiotropic effect of cis-SNPs on phenotypes. Therefore, a TWAS model for underrepresented populations was constructed. Based on the TWAS model for underrepresented populations, this study introduces a smoothed absolute bias penalty function to select non-zero effects, and combines least squares estimation to evaluate the effects on individual level data and summary statistics. and level pleiotropic effects Estimates and inferences.
2. The TWAS statistical method for underrepresented populations with multi-effects correction as described in claim 1, characterized in that, The TWAS model for underrepresented populations is as follows: ; ; in, It is the first In a group of people individual Dimensional gene expression vector; The total number of people; It is cis-SNPs 3D genotype matrix, This represents the number of cis-SNPs for this gene; It is the first cis-SNPs on gene expression in individual populations dimensional effect weight vector; yes 3D residual vector; yes 3D vector, representing Phenotypic measurements of each individual; For the corresponding cis-SNPs in GWAS studies 3D genotype matrix; It is the effect of cis-SNPs on gene expression in GWAS studies. dimensional effect vector, Indicates the first The effect weights of individual populations satisfy the following conditions: and ; This refers to the genetic regulatory expression component GReX constructed in the GWAS study, which is the predicted expression value of the gene in the GWAS study; It is the effect of GReX on phenotype. This indicates that, given the GReX effect in other populations, the first... The effect of GReX on phenotypes constructed in individual populations; yes A dimensional vector, representing the level pleiotropic effect of cis-SNPs on the phenotype; yes 3D residual vector.
3. The TWAS statistical method for underrepresented populations with level pleiotropic correction as described in claim 1, characterized in that, Effects on individual level data and level pleiotropic effects The estimation process includes: ; ; in, Horizontal pleiotropic effect The estimated value, For effect The estimated value; yes 3D vector, representing Phenotypic measurements of each individual; , Let be the projection matrix. yes 3D identity matrix; For the corresponding cis-SNPs in GWAS studies 3D genotype matrix; It is the SCAD penalty function. It refers to adjusting parameters; The j-th cis-SNP of a gene represents the level pleiotropic effect on the phenotype. , express The genetic regulatory expression component GReX constructed in a GWAS study for an individual population is the gene expression prediction value in the GWAS study.
4. The TWAS statistical method for underrepresented populations with multi-effects correction as described in claim 3, characterized in that, Effects on individual level data and level pleiotropic effects The inference process following estimation includes: Estimated quantity The asymptotic distribution is ; Among them, standard error for: ; Depend on It was estimated; Among them, set cis-SNP pairs It has a horizontal pleiotropic effect. , For the set The genotype matrix composed of cis-SNPs; For set The level of pleiotropic effects of cis-SNPs in China for The specific variance of each individual; For mathematical expectation operators; For the first Individuals are composed of sets The genotype vector composed of cis-SNPs; For the first The predicted expression values of genes in an individual in a GWAS study.
5. The TWAS statistical method for underrepresented populations with multi-effects correction as described in claim 1, characterized in that, Effects on summary statistics and level pleiotropic effects The estimation process includes: ; in, ; ; in, Marginal effect; Horizontal pleiotropic effect The estimated value, For effect The estimated value; yes 3D vector, representing Phenotypic measurements of each individual; , Let be the projection matrix. yes 3D identity matrix; For the corresponding cis-SNPs in GWAS studies 3D genotype matrix; It is the SCAD penalty function. It refers to adjusting parameters; The j-th cis-SNP of a gene represents the level pleiotropic effect on the phenotype. , express The genetic regulatory expression component GReX constructed in the GWAS study for an individual population, i.e. the gene expression prediction value in the GWAS study; The correlation matrix for the required cis-SNPs; for Weiquanwei ; , The effect weight vector of cis-SNPs on gene expression in the Mth population; and These are all intermediate parameters defined to simplify the formula. , .
6. The TWAS statistical method for underrepresented populations with level pleiotropic correction as described in claim 5, characterized in that, Effects on summary statistics and level pleiotropic effects The inference process following estimation includes: Estimated quantity The asymptotic distribution is ; Among them, standard error ; ; Estimate using the following formula: ; in, for The specific variance of each individual; set cis-SNP pairs It has a horizontal pleiotropic effect. , For the set The genotype matrix composed of cis-SNPs; for The covariance matrix; for and The covariance matrix; for and The covariance matrix; for and The covariance matrix.
7. A TWAS statistical system for underrepresented populations with pleiotropic level correction, characterized in that, include: The modeling module is configured to, for each gene, consider the effect of the gene's predicted expression in GWAS studies on the phenotype. And the level pleiotropic effect of cis-SNPs on phenotypes. Therefore, a TWAS model for underrepresented populations was constructed. The estimation and inference module is configured to use a TWAS model based on an underrepresented population. It selects non-zero effects by introducing a smoothed absolute bias penalty function and combines least squares estimation to evaluate the effects on individual level data and summary statistics. and level pleiotropic effects Estimates and inferences.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-6.