Non-centralized whole genome association aggregation analysis method and system
By using a decentralized genome-wide association analysis method, individual-level data is transformed into comprehensive report data, and least squares formula calculations are performed. This solves the communication bottlenecks and data security problems in traditional methods, and achieves efficient data privacy protection and improved computational efficiency.
Patent Information
- Application Number
- CN202510881894.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-11-21
AI Technical Summary
Traditional genome-wide association studies (GWAS) methods have increasingly revealed communication bottlenecks, computational costs, transmission costs, and data security issues between data sources and organizational centers, making them unable to meet the needs of multi-party computation.
A decentralized genome-wide association analysis (GWAS) approach is adopted, which transforms individual-level data into comprehensive report data through each data source and performs least squares calculations in a distributed architecture. The organization center performs aggregate analysis to synthesize effect values, thus avoiding the sharing of raw data.
It achieves efficient data privacy protection, reduces the risk of data leakage, improves computing efficiency, solves the problem of model selection and parameter adjustment between data sources and organizational centers, and supports the expansion needs of large-scale big data.
Smart Images

Figure CN120998309A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of bioinformatics, and particularly relates to a whole genome association meta-analysis method and system. BACKGROUND
[0002] Genome-wide association studies (GWAS) is the main statistical method for positioning disease-related genes at present. Although a single GWAS data set is the basis for gene positioning, in order to increase the statistical power, a strategy of aggregating multiple data is often used, which is called Genome-Wide Association Meta-Analysis (GWAMA). In the traditional GWAMA workflow, according to the model required by the organization center, the data source generates report data, each data source completes the specified calculation in each data source according to the instructions of the organization center, and then sends the calculation results to the organization center, and finally the organization center finally synthesizes the final result through meta-analysis. The problem is that the data source must respond to each instruction of the organization center, the organization center needs to send (k+1)2 k models to each data center, and each time the analyst needs to respond, which can traverse all possible calculations. With the increase of data source connections, the bottleneck gradually emerges, and neither the analyst of the data source nor the scheduling ability of the organization center can gradually meet the demand of multi-party calculation, and there are defects in computing power consumption, transmission consumption and data security. SUMMARY
[0003] In view of the above problems, the purpose of the present application is to provide a non-centralized whole genome association meta-analysis method;
[0004] The purpose of the present application is also to provide a non-centralized whole genome association meta-analysis system.
[0005] A non-centralized whole genome association meta-analysis method comprises:
[0006] Step S1, each data source converts individual-level data of the data source into all-purpose report data;
[0007] Step S2, importing the all-purpose report data into a least square formula to obtain a whole genome association analysis model;
[0008] Step S3, the organization center performs meta-analysis to synthesize the required effect value.
[0009] The decentralized genome-wide association analysis method of the present invention, wherein the individual-level data includes genotype data and phenotypic data, and the genotype data is represented as G = [g1, g2, ..., g m G is an n×m matrix, where n is the number of samples and m is the number of genetic loci; the phenotypic data is represented as Y = [y, x1, ..., xm]. k Y is an n×k matrix, y is the target trait vector, and k represents the number of phenotypic data; in step S1, omnipotent report data is constructed based on the genotype data and the phenotypic data:
[0010]
[0011] Where g l It contains the genotype data for the l-th genetic locus of interest.
[0012] The decentralized genome-wide association analysis method of the present invention, wherein the omnipotent report data Ω l =Ω A +Ω B +Ω C ;in,
[0013]
[0014] Ω A It is Ω l The element in the second row and second column is the variance of the l-th column of the genotype data X;
[0015]
[0016] Ω B It is Ω l The remaining elements in the second column and second row are the covariances between the l-th column of the genotype data G and each column of the phenotype data Y.
[0017]
[0018] Ω C It is Ω l Excluding the elements in the second column and the second row, the remaining elements constitute the variance-covariance matrix of the phenotypic data Y.
[0019] The decentralized genome-wide association analysis method of the present invention includes the following least squares formula in step S2:
[0020] Λ l It is Ω l The diagonal part; b l It is the effect value of the genetic locus;
[0021] wherein the molecule is y and g l the covariance of genetic loci, is the variance of genetic loci l, for then similarly the molecule denotes the covariance of y and x j , and denotes the variance of x j , x j is the jth phenotypic data.
[0022] The non-centralized genome-wide association meta-analysis method of the present application, the meta-analysis in step S3 is obtained by the following formula:
[0023]
[0024] n j denotes the size of a certain data set j participating in the meta-analysis, b o.j denotes the first element of the β l vector, that is, the estimated value associated with g l in each data set.
[0025] The non-centralized genome-wide association meta-analysis method of the present application further comprises step S4, which comprises SNP heritability estimation, which is obtained by the following formula:
[0026]
[0027] wherein b 0.l is the estimated value of β l in the least squares formula and x l , wherein and tr(K 2 ) is the trace of K 2 , h 2 is the SNP heritability estimate.
[0028] The step S4 further comprises a polygenic risk score, which is obtained by the following formula:
[0029]
[0030] is the polygenic risk score prediction value of the jth individual, b 0.l is the estimated value of β l in the least squares formula and g l , x jiis the element of the jth row and the ith column of the genotype data G.
[0031] A non-centralized whole genome association aggregated analysis system comprises:
[0032] A full-ability report data conversion module converts individual-level data of each data source into full-ability report data.
[0033] A whole genome association analysis model obtaining module, connected with the full-ability report data conversion module, imports the full-ability report data into a least square formula to obtain a whole genome association analysis model.
[0034] An organization center, connected with the whole genome association analysis model obtaining module, performs aggregated analysis to synthesize required effect values.
[0035] The non-centralized whole genome association aggregated analysis system of the application, the individual-level data comprises genotype data and phenotype data, the genotype data is represented as G=[g1, g2, …, g m , X is an n*m matrix, n is the number of samples, and m is the number of genetic loci; the phenotype data is represented as Y=[y, x1, …, x k , Y is an n*k matrix, y is a target trait vector, and k represents the number of phenotype data; in the step S1, the full-ability report data is constructed according to the genotype data and the phenotype data:
[0036]
[0037] Wherein g l is the genotype data of the lth genetic locus of interest.
[0038] The non-centralized whole genome association aggregated analysis system of the application, the least square formula in the step S2 is as follows:
[0039] Λ l is the diagonal part of Ω l ; b l is the effect value of the genetic locus;
[0040] Wherein the molecule of y and g l genetic locus is the covariance of y and g is the variance of the genetic locus l, and for the molecule indicates the covariance of y and x j , indicates the variance of x j , jThe jth phenotype data.
[0041] The non-centralized whole genome association analysis method of the present application obtains the required effect value by the following formula:
[0042]
[0043] n j The size of the jth data set participating in the aggregate analysis, b o.j The size of the jth data set participating in the aggregate analysis, b The first element of the vector, that is, the estimated value associated with g l in each data set.
[0044] Beneficial effects: the present application converts the data of each data source into all-in-one report data, then reconstructs the least square algorithm, imports the all-in-one report data into the least square formula to obtain a whole genome association analysis model, the non-centralized design can meet the data privacy requirements, reduce the risk of data leakage, and the organization center performs aggregate analysis to synthesize the required effect value, solves the model selection and parameter adjustment problem between the data source and the organization center, and maximizes the efficiency of the aggregate analysis. BRIEF DESCRIPTION OF DRAWINGS
[0045] Figure 1 The overall workflow of the present application;
[0046] Figure 2 The genotype data of the present application;
[0047] Figure 3 The phenotype data of the present application;
[0048] Figure 4 The method flowchart of the present application;
[0049] Figure 5 The system module schematic diagram of the present application;
[0050] Figure 6 The all-in-one report data generation schematic diagram of the present application. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0052] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0053] The present application will be further described below in conjunction with the drawings and specific embodiments, but not as a limitation of the present application.
[0054] Referring to Figures 1 to 6 A non-centralized whole genome association aggregation analysis method, comprising:
[0055] Step S1, each data source converts individual-level data of the data source into full-capability report data;
[0056] Step S2, the full-capability report data is imported into a least square formula to obtain a whole genome association analysis model;
[0057] Step S3, the organization center performs aggregation analysis to synthesize required effect values.
[0058] Each data source of the present application converts its own data into full-capability report data summarystatistics, nss), the elements of the full-capability report data are variance and covariance, then the least square algorithm is reconstructed, the full-capability report data is imported into a least square formula to obtain a whole genome association analysis model, the non-centralized design can meet the data privacy requirement and reduce the data leakage risk, the organization center performs aggregation analysis to synthesize required effect values, if the organization center finds that the model exported by the data source is unreasonable after evaluation, the model can be quickly reconstructed again through the full-capability report data, the model selection and parameter adjustment problem between the data source and the organization center is solved, and the efficiency of the aggregation analysis is maximized.
[0059] Figure 1 is a whole workflow diagram of the present application, wherein data source 1 to data source 4 are data providers (Biobank), each data source includes genomic data (GWAS) and phenotype data or electronic health data (eHealth), each data source converts individual-level data into de-privacy full-capability report data and transmits the data to an organization center, the organization center performs aggregation analysis through a high-throughput data aggregation method, and finally provides data application and service for a terminal.
[0060] The non-centralized whole genome association aggregation analysis method of the present application, the individual-level data includes genotype data and phenotype data, the genotype data is represented as G=[g1,g2,…,g m ], G is an n×m matrix, n is the number of samples, and m is the number of genetic loci; the phenotype data is represented as Y=[y,x1,…,x kY is an n x k matrix, y is a target trait vector, and k represents the number of phenotype data; in step S1, the full genotype report data is constructed according to the genotype data and the phenotype data:
[0061]
[0062] wherein g l is the genotype data of the lth genetic locus of interest.
[0063] Figure 2 is a schematic diagram of genotype data, the first column is a sample (Sample) number, and the second column gives the genotype of each genetic locus.
[0064] Figure 3 is a schematic diagram of phenotype data, the first column is a sample (Sample) individual number, and the second column is a specific numerical value corresponding to each individual.
[0065] The present application provides a data basis for multivariate joint analysis or complex models by integrating genotype data and phenotype data, supports parallel computing, is more suitable for the expansion needs of large data scale, and in addition, the original genotype data and phenotype data do not need to be centrally stored or transmitted, promotes cross-institutional collaborative research, and at the same time reduces the risk of data leakage.
[0066] The non-centralized whole genome association aggregation analysis method of the present application, the full genotype report data Ω l = Ω A + Ω B + Ω C ; wherein,
[0067]
[0068] Ω A is the second row and second column element of Ω l , and the substituted is the variance of the lth column of genotype data X;
[0069]
[0070] Ω B is the remaining elements of the second column and the second row of Ω l , which is the covariance of the lth column of genotype data G and each column of phenotype data Y;
[0071]
[0072] Ω C is the elements of Ω l except the second column and the second row, and the composition is the variance-covariance matrix of the phenotype data Y.
[0073] According to ΩA , Ω B and Ω C structure, each time the least square estimation model calculation, replace g l sliding to g l ', as long as the update Ω A , Ω B element, Ω C remains unchanged, can be directly from Ω l update to Ω l′ .
[0074] The application avoids repeated calculation by adjusting complex genotype data and phenotype data into a single data structure, and significantly improves the calculation efficiency.
[0075] The non-centralized whole genome association aggregation analysis method of the application, the least square formula in the step S2 is as follows:
[0076] Λ l =tr(Ω l ), Λ l is the diagonal part of Ω l ; b l is the effect value of the genetic locus;
[0077] Wherein the molecule is the covariance of y and g l genetic locus, is the variance of the genetic locus l, for then the same molecule indicates the covariance of y and x j , x j indicates the variance of x j , x j is the jth phenotype data.
[0078] The application supports privacy protection calculation in a distributed architecture, can complete effect value estimation without original data, avoids direct sharing of original data, and meets data security requirements.
[0079] Figure 6 The leftmost red part in the figure indicates the variance vector of SNP, which is calculated from the genotype data G matrix such as g j , the purple part indicates SNP and general phenotype data Y, such as Y1, Y2, Y3, Yi, and derived phenotype data such as v1, v2, v3; and the green part is the full genotype report data.
[0080] The non-centralized whole genome association aggregation analysis method of the application, the aggregation analysis in the step S3 obtains the required effect value through the following formula:
[0081]
[0082] n j b represents the size of a dataset j participating in the aggregation analysis. o.j express beta l The first element of the vector, which is the element in each dataset that is related to g. l The associated estimated value.
[0083] The decentralized genome-wide association analysis method of the present invention further includes step S4, wherein step S4 includes SNP heritability estimation, and the SNP heritability estimation is obtained by the following formula:
[0084]
[0085] Where b 0.l β in the least squares formula l With x l The relevant estimates, of which And tr(K) 2 ) is K 2 traces, h 2 This is an estimate of the heritability of the SNP.
[0086] Step S4 further includes a polygenic risk score (PRS), which is obtained using the following formula:
[0087]
[0088] b is the predicted value of the polygenic risk score for the j-th individual. 0.l β in the least squares formula l With g l The relevant estimated value, x ji It is the element in the j-th row and i-th column of the genotype data G.
[0089] Reference Figure 5 A decentralized genome-wide association analysis system, comprising:
[0090] The all-in-one report data conversion module 11 converts individual-level data from each data source into all-in-one report data;
[0091] The genome-wide association analysis model acquisition module 12 is connected to the all-purpose report data conversion module 11, and imports the all-purpose report data into the least squares formula to obtain the genome-wide association analysis model;
[0092] The organization center 13 is connected to the genome-wide association analysis model acquisition module 12 to perform aggregation analysis to synthesize the required effect values.
[0093] The decentralized genome-wide association analysis system of the present invention includes individual-level data comprising genotype data and phenotypic data, wherein the genotype data is represented as G = [g1, g2, ..., g m X is an n×m matrix, where n is the number of samples and m is the number of genetic loci; the phenotypic data is represented as Y = [y, x1, ..., xm]. k Y is an n×k matrix, y is the target trait vector, and k represents the number of phenotypic data; in step S1, omnipotent report data is constructed based on the genotype data and the phenotypic data:
[0094]
[0095] Where g l It contains the genotype data for the l-th genetic locus of interest.
[0096] In the decentralized genome-wide association analysis system of the present invention, the least squares formula in step S2 is as follows:
[0097] Λ l It is Ω l The diagonal part; b l It is the effect value of the genetic locus;
[0098] in molecules It is y and g l Covariance of genetic loci It is the variance of the genetic locus l, for Then the numerator Indicate y and x j covariance, x represents j The variance, x j This represents the j-th phenotypic data.
[0099] The present invention provides a decentralized genome-wide association analysis method, wherein the desired effect size is obtained through the following formula:
[0100]
[0101] n j b represents the size of a dataset j participating in the aggregation analysis. o.j express beta l The first element of the vector, which is the element in each dataset that is related to g.l the estimate of the association.
[0102] Each data source of the present application converts its own data into all- powerful report data, and then reconstructs a least square algorithm to obtain a full-genotype association analysis model. The full-genotype association analysis model does not depend on first-hand data, and a non-centralized design can meet data privacy requirements, reduce data leakage risks, and organize a center to aggregate and analyze effect values required, solve the model selection and parameter adjustment problems between the data source and the organization center, realize high-throughput aggregate analysis, and the aggregate results can be introduced into the clinical medical field, such as multi-gene risk assessment, SNP heritability estimation, or Mendelian randomization.
[0103] The specific structure of the specific embodiments of the specific embodiments is given by the description and the drawings, and other conversions can be made based on the spirit of the present application. Although the above-mentioned application presents the preferred embodiments, these contents are not as limitations.
[0104] For those skilled in the art, various changes and modifications will undoubtedly be apparent after reading the above description. Therefore, the appended claims should be considered as covering all changes and modifications within the true intent and scope of the present application. Any and all equivalent ranges and contents within the scope of the claims should be considered as still within the intent and scope of the present application.
Claims
1. A decentralized genome-wide association analysis method, characterized in that, include: Step S1: Each data source converts the individual-level data of the data source into universal report data; Step S2: Import the all-in-one report data into the least squares formula to obtain the genome-wide association analysis model; Step S3: The tissue center performs a polymerase chain reaction (PCR) analysis to synthesize the required effect value.
2. The decentralized genome-wide association analysis method according to claim 1, characterized in that, The individual-level data includes genotype data and phenotypic data, wherein the genotype data is represented as G = [g1, g2, ..., g m G is an n×m matrix, where n is the number of samples and m is the number of genetic loci; the phenotypic data is represented as Y = [y, x1, ..., xm]. k Y is an n×k matrix, y is the target trait vector, and k represents the number of phenotypic data; in step S1, omnipotent report data is constructed based on the genotype data and the phenotypic data: Where g l It contains the genotype data for the l-th genetic locus of interest.
3. The decentralized genome-wide association analysis method according to claim 2, characterized in that, The all-in-one report data Ω l =Ω A +Ω B +Ω C ; in, Ω A It is Ω l The element in the second row and second column is the variance of the l-th column of the genotype data X; Ω B It is Ω l The remaining elements in the second column and second row are the covariances between the l-th column of the genotype data G and each column of the phenotype data Y. Ω C It is Ω l Excluding the elements in the second column and the second row, the remaining elements constitute the variance-covariance matrix of the phenotypic data Y.
4. The decentralized genome-wide association analysis method according to claim 2, characterized in that, The least squares formula in step S2 is as follows: Λ l It is Ω l The diagonal part; b l It is the effect value of the genetic locus; in molecules It is y and g l Covariance of genetic loci It is the variance of the genetic locus l, for Then the numerator Indicate y and x j covariance, x represents j The variance, x j This represents the j-th phenotypic data.
5. The decentralized genome-wide association analysis method according to claim 1, characterized in that, In step S3, the aggregate analysis yields the desired effect value using the following formula: n j b represents the size of a dataset j participating in the aggregation analysis. o.j express beta l The first element of the vector, which is the element in each dataset that is related to g. l The associated estimated value.
6. The decentralized genome-wide association analysis method according to claim 4, characterized in that, The method also includes step S4, which includes SNP heritability estimation, obtained by the following formula: Where b 0.l β in the least squares formula l With x l The relevant estimates, of which And tr(K) 2 ) is K 2 traces, h 2 This is an estimate of the heritability of the SNP. Step S4 further includes a polygenic risk score, which is obtained using the following formula: b is the predicted value of the polygenic risk score for the j-th individual. 0.l β in the least squares formula l With g l The relevant estimated value, x ji It is the element in the j-th row and i-th column of the genotype data G.
7. A decentralized genome-wide association analysis system, characterized in that, include: The all-in-one report data conversion module converts individual-level data from each data source into all-in-one report data. The genome-wide association analysis model acquisition module is connected to the all-purpose report data conversion module, and imports the all-purpose report data into the least squares formula to obtain the genome-wide association analysis model; The organization center connects to the module obtained by the genome-wide association analysis model to perform aggregation analysis and synthesize the required effect values.
8. The decentralized genome-wide association analysis system according to claim 7, characterized in that, The individual-level data includes genotype data and phenotypic data, wherein the genotype data is represented as G = [g1, g2, ..., g m X is an n×m matrix, where n is the number of samples and m is the number of genetic loci; the phenotypic data is represented as Y = [y, x1, ..., xm]. k Y is an n×k matrix, y is the target trait vector, and k represents the number of phenotypic data; in step S1, omnipotent report data is constructed based on the genotype data and the phenotypic data: Where g l It contains the genotype data for the l-th genetic locus of interest.
9. The decentralized genome-wide association analysis system according to claim 8, characterized in that, The least squares formula in step S2 is as follows: Λ l It is Ω l The diagonal part; b l It is the effect value of the genetic locus; in molecules It is y and g l Covariance of genetic loci It is the variance of the genetic locus l, for Then the numerator Indicate y and x j covariance, x represents j The variance, x j This represents the j-th phenotypic data.
10. The decentralized genome-wide association analysis system according to claim 8, characterized in that, The aggregate analysis yields the desired effect size using the following formula: n j b represents the size of a dataset j participating in the aggregation analysis. o.j express beta l The first element of the vector, which is the element in each dataset that is related to g. l The associated estimated value.