A Federated Analysis Method for Non-Independent Gene Data
The federated analysis method for GWAS data addresses privacy and sample overlap issues by encrypting and anonymizing data locally, enabling secure cross-institutional research and improving genetic variant detection.
Patent Information
- Application Number
- CN202510241928.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-03
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-03-03
AI Technical Summary
The prior art is difficult to solve the problem of data privacy protection and sample overlap when merging multiple genome-wide association studies (GWAS) data, resulting in the impact of the accuracy and privacy of the analysis results.
The federal analysis method for non-independent gene data is adopted, identity encryption and anonymization is performed through the client, data privacy is protected using identification code generation algorithm and hash function, and data preprocessing and model training is performed locally, and data aggregation is performed in combination with the federated averaging algorithm of the main server to ensure that the data is processed locally and only the encryption model is transmitted to update.
Effectively protect data privacy, comply with data protection regulations, achieve cross-institutional and cross-country research cooperation, and improve statistical efficacy and identify the probability of genetic mutation related to diseases.
Smart Images

Figure CN119741976B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a method for federated analysis of gene data, and particularly to a method for federated analysis of non-independent gene data, belonging to the technical field of federated learning. Background Art
[0002] First of all, Federated Learning is a machine learning setting in which multiple participants (which can be mobile devices or entire organizations) collaborate to train a model without sharing their data. This method is particularly suitable for scenarios with strict requirements for data privacy. Applying the federated learning method to genome-wide association studies (GWAS) can aggregate GWAS data from multiple sources.
[0003] Secondly, genome-wide association studies (GWAS) aim to identify genetic variants associated with specific phenotypes (such as diseases). A single GWAS study may not be able to detect small effects or statistically significant associations due to limited sample size. By combining the results of multiple GWAS studies, the statistical power can be increased, and the ability to discover true genetic variants can be improved.
[0004] However, genomic data is highly sensitive because it contains detailed descriptions of personal genetic information. Since these data are usually subject to strict privacy protection regulations, it may not be feasible or permitted to directly share GWAS data for traditional centralized meta-analysis. At the same time, if there is sample overlap between multiple GWAS data to be combined, it is necessary to adjust the GWAS statistics for sample overlap, otherwise incorrect GWAS statistics will be obtained, affecting the analysis results and causing false positive results.
[0005] Therefore, there is an urgent need to develop a method for federated analysis of non-independent gene data. Summary of the Invention
[0006] Aiming at the above existing technical problems, the present invention provides a method for federated analysis of non-independent gene data to achieve the technical purpose of solving the data privacy problem and sample overlap problem when combining multiple GWAS analyses.
[0007] To achieve the above technical purpose, the present invention provides a method for federated analysis of non-independent gene data, using a federated analysis system for non-independent gene data mainly composed of multiple clients and a master server, including the following steps:
[0008] S1. The client uses an identification code generation algorithm to perform identity encryption and anonymization processing on the user gene sequence data set to obtain an anonymized data set;
[0009] S2. Based on the anonymized dataset, preprocess the phenotypic data and genotypic data respectively to obtain a standardized phenotypic dataset and a genotypic dataset, and perform principal component analysis on the standardized genotypic data to obtain the first 5 principal components;
[0010] S3. Based on the standardized phenotypic dataset, perform model training locally, select other required features except genotypic data in model training, and use a regression model. Taking the specific phenotype determined according to the research purpose as the dependent variable, taking other required features and the first 5 principal components as covariates, and taking the genotype of gene loci as the independent variable, calculate the genetic effect value of each gene locus on the specific phenotype respectively to obtain the GWAS data of the specific phenotype;
[0011] S4. Based on the GWAS data of the specific phenotype, calculate the list of identification codes of the corresponding samples, as well as the control sample ratio and case ratio, to obtain the dataset of the client and transmit it to the main server;
[0012] S5. Based on the datasets from multiple clients and itself, the main server uses the federated averaging algorithm to calculate the variance matrix and weight vector of different genetic effect values for each gene locus respectively, as well as the aggregated genetic effect value.
[0013] Furthermore, in the present invention, in S1, the client uses an identification code generation algorithm to perform identity encryption and anonymization processing on the user gene sequence dataset, including the following steps:
[0014] S1-1. Clean and preprocess the dataset to remove information that can directly identify personal identity;
[0015] S1-2. Group the dataset according to non-sensitive attributes to ensure that each equivalence class meets the requirements of L-diversity;
[0016] S1-3. Generate a unique salt value for each equivalence class;
[0017] S1-4. Combine the personal ID of each record with the corresponding salt value and apply a hash function to generate a unique identification code;
[0018] S1-5. Use the generated identification code to replace the original personal ID.
[0019] Furthermore, in the present invention, in S2, the data preprocessing of the phenotypic data includes the following steps:
[0020] Perform quality control screening on the gene data, store the phenotypic data in a.csv file, use the identification code as the first column, each phenotype as other columns, and represent missing values with NA.
[0021] Further, in the present invention, the data preprocessing of the genotype data in S2 includes the following steps:
[0022] S2-2-1. Perform quality control screening on the genotype data to obtain a quality-controlled genotype data set;
[0023] S2-2-2. Perform standardization processing on the quality-controlled genotype data set, subtract the original value of the gene locus genotype from its mean and then divide by its standard deviation.
[0024] Furthermore, in the present invention, the quality control screening of the gene data in S2-2-1 includes the following steps:
[0025] S2-2-11. Remove gene loci with a non-detection rate higher than 2%;
[0026] S2-2-12. Remove gene loci with a P-value of the Hardy-Weinberg test less than ;
[0027] S2-2-13. Remove gene loci with a minor allele frequency lower than 1%;
[0028] S2-2-14. Remove gene loci with a gene imputation quality lower than 0.3;
[0029] S2-2-15. Remove samples with an overall gene locus deletion rate greater than 5%.
[0030] Further, in the present invention, the GWAS data of a specific phenotype in S3 is represented in the form of a data table on locus information, and its fields include the locus ID, coordinate information, genetic effect value and its standard deviation, regression intercept, and P-value of each gene locus.
[0031] Further, in the present invention, the calculation of the variance matrix of different genetic effect values for each gene locus in S5 includes the following steps:
[0032] Denote the different genetic effect values of any gene locus in multiple client and master server data sets as , and , represents a total of data sets, represents the th data set, , , respectively represent the genetic effect values of this gene locus in the 1st, th, and th data sets;
[0033] Denote the standard deviation of the corresponding different genetic effect values as , , , respectively represent the standard deviations of the genetic effect values of the gene locus in the 1st, th, and th datasets;
[0034] The variance matrix corresponding to different genetic effect values is denoted as ;
[0035] Then, the calculation formula for the value in the th row and th column of the variance matrix is:
[0036]
[0037]
[0038] In the formula, the genetic effect value of the gene locus in the th dataset, and ;
[0039] , , respectively represent the number of cases, the number of controls, and the total sample size of the th dataset;
[0040] , , respectively represent the number of cases, the number of controls, and the total sample size of the th dataset;
[0041] , respectively represent the number of overlapping controls and cases in the th and th datasets.
[0042] Furthermore, in the present invention, the calculation method of is as follows: Based on the list of identification codes and the case proportion transmitted by multiple clients, by comparing whether the identification codes overlap, the total number of
[0043] is obtained, and then multiplied by the corresponding control sample proportion and case proportion respectively to obtain and and .
[0044] Furthermore, in the present invention, calculating the weight vector of different genetic effect values for each gene locus in S5 includes the following steps:
[0045] Denote the weight vector of the different genetic effect values of any gene locus in multiple client and master server datasets as , , , respectively represent the weights of the genetic effect values of this gene locus in the 1st, th, th datasets;
[0046] Then the calculation formula of the weight vector is:
[0047]
[0048] In the formula, represents the unit vector; represents the variance matrix of the corresponding different genetic effect values.
[0049] Furthermore, in the present invention, calculating the aggregated genetic effect value for each gene locus in S5 includes the following steps:
[0050] Denote the aggregated genetic effect value of any gene locus in multiple client and master server datasets as ;
[0051] Then the calculation formula of the aggregated genetic effect value is:
[0052]
[0053] In the formula, represents the genetic effect value of this gene locus in the th dataset, represents the weight of the genetic effect value of this gene locus in the th dataset.
[0054] In summary, the present invention solves the data privacy problem and the sample overlap problem when combining multiple GWAS analyses, and has the following technical advantages:
[0055] 1. Privacy protection: The method of the present invention ensures data localization, that is, data is processed at its original location (such as a hospital or a research institution), and only the updates of the model (such as gradient information) are shared under encryption or privacy protection measures, thus effectively protecting the privacy of users.
[0056] 2. Compliance: Since the data does not leave its original storage location, the method of the present invention helps to comply with data protection regulations, which makes cross-institutional and cross-national GWAS research possible and promotes international cooperation and data sharing.
[0057] 3. Improving statistical power: By integrating data from multiple studies, the method of the present invention can increase statistical power. This helps to improve the probability of identifying genetic variations related to diseases, thereby accelerating the discovery and research process of disease-related genes.
[0058] 4. Heterogeneous data integration: Different studies may have used different genotyping platforms, sample processing methods, or analysis procedures. The method of the present invention can handle this data heterogeneity because it allows local computing and model updates, thus enabling effective integration and analysis of data from different sources. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] Figure 1 is a flowchart of the steps of the method of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0060] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0061] As Figure 1 shown, taking the aggregation algorithm of binary phenotype data of the WeGene cohort genome-wide association study (GWAS) as an example, the present invention provides a federated analysis method for non-independent gene data, using a federated analysis system for non-independent gene data mainly composed of multiple clients and a master server, including the following specific steps.
[0062] S1. The client generates an algorithm according to the identification code (unique id), performs identity encryption and anonymization processing on the user gene sequence dataset thereon, ensures data privacy, and publishes the anonymized dataset to provide a secure data basis for subsequent analysis.
[0063] It should be noted that the user gene sequence data is the gene detection result obtained based on chip detection or next-generation sequencing technology. Technicians perform identity encryption and anonymization processing on the sample information according to the gene sequence dataset output by the sequencer or chip detection instrument. The technicians performing subsequent analysis only come into contact with the encrypted gene sequence data and do not come into contact with the specific identity information of the sample.
[0064] It should also be noted that in the process of implementing the identification code generation algorithm, the present invention adopts a composite method of combining the L-diversity algorithm with the hash function and the salt value algorithm to ensure data privacy protection and anonymity when the data is published, which is introduced as follows.
[0065] First, L-diversity is a privacy protection technique designed to prevent the unique identification of individuals through query results, thereby protecting data privacy. In this step, the L-diversity algorithm is the core of privacy protection. It effectively prevents the inference of sensitive information by ensuring that each equivalence class in the dataset contains at least L different sensitive values. An equivalence class is a grouping consisting of one or more records that are the same in non-sensitive attributes, such as age, gender, postal code, etc. In this way, even if an attacker can determine that an individual belongs to a specific equivalence class, it is not easy to infer the specific sensitive information of that individual, such as health status, income level, etc. In addition, L-diversity can also protect minority groups and prevent the leakage of their sensitive information due to data publication, which is crucial for maintaining social fairness and respecting individual privacy.
[0066] Secondly, the use of a hash function combined with a salt value algorithm provides an additional layer of security for the generation of identification codes. A hash function is a one-way encryption technique that can convert input data (such as a personal ID) into a fixed-length string. This process is irreversible, meaning that the original data cannot be deduced from the hash value. The salt value is a randomly generated string that, when combined with the original data and then hashed, further increases the complexity and randomness of the hash value, such that even the same input data will produce completely different hash values due to different salt values.
[0067] When specifically implemented, the composite method combining the L-diversity algorithm with the hash function and salt value algorithm is introduced in the following specific steps.
[0068] S1-1, Data preprocessing: Clean and preprocess the user gene sequence dataset, removing information that can directly identify personal identity. Specifically, remove all information directly associated with the sample identity, such as name, address, age, health status, income level, etc.
[0069] S1-2, Apply L-diversity: Group the dataset according to non-sensitive attributes to ensure that each equivalence class meets the requirements of L-diversity. Specifically, ensure that there are at least L different sensitive values in each equivalence class.
[0070] S1-3, Generate salt value: Generate a unique salt value for each equivalence class. Specifically, this salt value is random and confidential and will not be made public.
[0071] S1-4. Hash processing: Combine the personal ID of each record with the corresponding salt value and apply a hash function to generate a unique identification code. Specifically, assign a unique identification code to each sample. This identification code is randomly generated and does not contain any information that can be traced back to the sample's identity. This identification code will serve as the unique identifier for each record in the dataset and the sole reference for all subsequent data processing and analysis. Moreover, all clients use the same identification code generation algorithm, and the identification codes for the same sample are the same.
[0072] S1-5. Data release: Finally, replace the original personal ID with the generated identification code. Specifically, use the AES encryption algorithm for data transmission to ensure the security of data during transmission.
[0073] Specifically in implementation, examples of the encrypted unique identification codes are shown in Table 1 below.
[0074] Table 1. Examples of the encrypted unique identification codes
[0075]
[0076] As can be seen from the above, in order to ensure user privacy and security, this step uses advanced identity encryption technology to anonymize the user's genetic data. Specifically, first perform desensitization processing on the gene sequence, using a composite method that combines the L-diversity algorithm, the hash function, and the salt value algorithm to assign a unique identification code to each user. This identification code is completely separated from the user's true identity information and is only used for subsequent data storage, analysis, and research. Then release the anonymized dataset and use an encryption algorithm to encrypt the data during transmission to ensure the security of data during transmission.
[0077] In particular, through the above composite method, this step not only achieves the L-diversity of the dataset and protects sensitive information from being inferred, but also ensures the irreversibility and uniqueness of the identification code through the hash function and the salt value algorithm, thus protecting the privacy of the data to the greatest extent. This method is particularly important in the application of sensitive data fields such as medical data and demographic survey data. It helps organizations to comply with data protection regulations while also providing valuable data resources for social research and policy making.
[0078] S2. Based on the anonymized dataset, perform data preprocessing on the phenotypic data and the genotypic data respectively to obtain standardized phenotypic data and genotypic datasets, ensuring the consistency and comparability of the data. Calculate the principal components (PC values) for the standardized genotypic dataset and select the top 5 principal components as important features for subsequent analysis.
[0079] It should be noted that in this step, preprocessing is performed on genotype data and phenotype data respectively to achieve data quality control and standardization, which is specifically introduced as follows.
[0080] S2-1. Based on the anonymized dataset, preprocess the phenotype data, store the phenotype data in a.csv file, with the identification code as the first column and each phenotype as other columns, and represent missing values with NA, to obtain a standardized phenotype dataset.
[0081] In this embodiment, the standardized phenotype dataset includes 2,501 phenotypes such as gender, age, smoking, cancer, etc. Among them, the two phenotypes of gender and age (or date of birth) must be provided, and other phenotypes are collected and sorted according to the needs of the research project. Table 2 below lists 31 phenotypes among the 2,501 phenotypes.
[0082] Table 2. 31 Phenotypes among 2,501 Phenotypes
[0083]
[0084]
[0085]
[0086] S2-2. Based on the anonymized dataset, preprocess the genotype data to obtain a standardized genotype dataset to ensure data consistency and comparability.
[0087] S2-2-1. Perform quality control screening on the genotype data to obtain a quality-controlled genotype dataset, including the following steps:
[0088] S2-2-11. Remove loci with too high non-detection rate. Set the threshold to 2%, which means removing loci with a non-detection rate higher than 2% because a higher non-detection rate may indicate lower quality of genotype data.
[0089] S2-2-12. Remove loci with too small P-values in the Hardy-Weinberg test. Set the threshold to , because such loci may indicate errors in genotyping or sequencing.
[0090] S2-2-13. Remove loci with too low minor allele frequency (MAF). Set the threshold to 1% because such loci may cause large errors in statistical tests due to too low variation.
[0091] S2-2-14. Remove loci with too low imputation quality. Set the threshold to 0.3 because such loci may cause errors in statistical test results due to too low imputation quality.
[0092] S2-2-15. Remove samples with an overall site deletion rate that is too high, and set the threshold to 5%. Because a high undetected rate may indicate low-quality genotype data.
[0093] In specific implementation, the genotype data is used as the original input data, including sites such as single nucleotide polymorphisms (SNPs), insertions and deletions (indels). And, the genotype dataset is stored in a.vcf file or a.bed file, including site location information, variant detection results (genotype information), and detection quality information. An example of the quality-controlled genotype dataset obtained in this step is shown in Table 3 below.
[0094] Table 3. Example of the quality-controlled genotype dataset
[0095]
[0096] S2-2-2. Perform standardization on the quality-controlled genotype dataset, that is, standardize the site (column) data. Subtract the mean of the original values of the gene site genotypes and then divide by their standard deviation to obtain a standardized genotype dataset, preparing for subsequent analysis.
[0097] In specific implementation, an example of the standardized genotype dataset obtained in this step is shown in Table 4 below.
[0098] Table 4. Example of the standardized genotype dataset
[0099]
[0100] S2-3. Calculate the principal components (PC values) for the standardized genotype data to obtain the first 5 principal components (PCs).
[0101] It should be noted that principal component analysis (PCA) is a commonly used technique for adjusting subgroup effects in data. Its basic principle lies in identifying mathematical vectors that can maximize the gene frequency differences, thereby obtaining the principal components (i.e., eigenvectors) of the data. Through these principal components, the population structure of the samples can be evaluated. The purpose of applying principal component analysis is to correct the systematic bias of allele frequencies caused by non-random mating between different subgroups within a population. This bias is an important confounding factor in genome-wide association studies and may lead to false positive results.
[0102] In specific implementation, use the PLINK software to perform principal component analysis on the standardized genotype data and select the first 5 principal components as important features for subsequent analysis.
[0103] S3. Based on the standardized phenotypic dataset, conduct model training locally. Select other required features except genotype data during model training, and adopt a regression model. Using the specific phenotype determined according to the research purpose as the dependent variable, other required features and the first 5 principal components as covariates, and the genotype of the gene locus as the independent variable, calculate the genetic effect value of each gene locus on the specific phenotype respectively to obtain the GWAS data of the specific phenotype. The specific steps are introduced as follows.
[0104] S3-1. Feature selection: Select other required features except genotype data during model training.
[0105] Specifically, when implementing, select gender and age as other required features, and use them together with the first 5 principal components as covariates. When conducting GWAS, it is a convention to put gender, age and 5 principal components into the model, and usually no other variables need to be added. If there are special research purposes, other required variables can also be added.
[0106] S3-2. Model construction: Adopt a regression model. Using the specific phenotype determined according to the research purpose as the dependent variable (such as phenotypes like smoking, cancer, etc.), other required features (gender, age), and the first 5 principal components as covariates, and the genotype of the gene locus as the independent variable. After conducting regression analysis on each gene locus respectively, obtain the genetic effect value (regression slope), its standard deviation, regression intercept, and the corresponding P-value (used to evaluate the significance of the association), and organize them together with other relevant information (such as the locus ID, coordinate information, etc. of each gene locus) into the GWAS data of the specific phenotype.
[0107] It should be noted that if the dependent variable is a continuous variable, such as height, weight, blood glucose value, etc., the regression model adopts an ordinary multiple linear regression model; if the dependent variable is a binary variable, such as whether having diabetes, whether being myopic, whether smoking, etc., the regression model adopts a logistic regression model.
[0108] Specifically, when implementing, in this embodiment, whether being myopic is used as the dependent variable, age, gender, and the first 5 principal components are used as covariates, and the genotype of the gene locus is used as the independent variable, and a logistic regression model is adopted for local model training.
[0109] In particular, the GWAS data of the specific phenotype is represented in the form of a data table about locus information, and its fields include the locus ID, coordinate information corresponding to each gene locus, as well as the corresponding genetic effect value, regression intercept, standard deviation of the genetic effect value, and P-value. And subsequently, the local GWAS data is transmitted to the main server.
[0110] S4. Based on the GWAS data of specific phenotypes, calculate the list of identification codes of the corresponding samples, as well as the control sample ratio and case ratio, organize the dataset of the client, and securely transmit it to the main server.
[0111] It should be noted that the control sample ratio and case sample ratio refer to the ratios of different samples of the dependent variable. For example, if the dependent variable is whether one is nearsighted, the control samples are non-nearsighted samples, and the case samples are nearsighted samples. The control sample ratio and case sample ratio are obtained by dividing the number of the two types of samples by the total number of samples. Moreover, since only the ratios of the control and pathological samples in the whole sample are statistically analyzed, rather than directly using the phenotypic data corresponding to the samples, the security of the sample information is further protected.
[0112] In specific implementation, the list of identification codes of the samples is stored as a.csv file, represented in a single column, and each row represents the identification code of a sample. Regarding whether one is nearsighted, in this embodiment, the total number of samples is 6133, among which there are 975 control samples and 5158 case samples. The control sample ratio and the ratio of cases are statistically analyzed, and the statistical data is stored in the form of a.csv table. Finally, the client transmits the local dataset to the main server, including the GWAS data of specific phenotypes, the corresponding list of sample identification codes, and the control sample ratio and case ratio information.
[0113] S5. Based on the datasets from multiple clients and the main server, the main server uses the federated averaging algorithm to calculate the variance matrix and weight vector of different genetic effect values, as well as the aggregated genetic effect value for each gene locus respectively.
[0114] In specific implementation, the federated analysis system for non-independent gene data of the present invention includes multiple clients and one main server. The dataset output by each client includes the GWAS data of specific phenotypes, the corresponding list of sample identification codes, and the control sample ratio and case ratio information. The main server itself stores a dataset with the same data format as that of the clients, mainly the data transmitted by each client in the past, also including the GWAS data of specific phenotypes, the corresponding list of sample identification codes, and the control sample ratio and case ratio information. It's just that the main server is a larger database, encompassing more phenotypes and more samples.
[0115] Moreover, when the main server receives the dataset transmitted by the client, it performs an aggregation operation on the dataset of the client and its own dataset. And, in order to consider data correlation, the aggregation algorithm is optimized, and the federated averaging algorithm is used to perform weighted summation on the parameters transmitted by the client. However, when aggregating the parameters transmitted by different clients, or merging the parameters of the client with the aggregated parameters, it is necessary to consider the impact of duplicate samples on the model, that is, it is necessary to adjust for the non-independence problem and perform the aggregation of the model parameters of non-independent data processing.
[0116] Note that in federated learning, Federated Averaging is a commonly used method for collaborative computing among multiple clients without the need for centralized data storage. This method is particularly suitable for privacy-sensitive scenarios such as genome-wide association studies. The following describes how to calculate the variance matrix, weight vector, and aggregated genetic effects of genetic effect values for each genetic locus based on the Federated Averaging algorithm.
[0117] S5-1. Calculate the variance matrix of different genetic effect values of each genetic locus from datasets of multiple clients and the main server.
[0118] Specifically, denote the different genetic effect values of any genetic locus in datasets of multiple clients and the main server as , and , represents that there are a total of datasets, represents the th dataset, , , respectively represent the genetic effect values of this genetic locus in the 1st, th, and th datasets.
[0119] Denote the standard deviation of the corresponding different genetic effect values as , , , respectively represent the standard deviations of the genetic effect values of this genetic locus in the 1st, th, and th datasets. Moreover, the genetic effect values and their standard deviations are obtained from the corresponding fields in the GWAS data of specific phenotypes in the client and the main server.
[0120] Denote the variance matrix of the corresponding different genetic effect values as .
[0121] Then, the formula for the value in the th row and th column of the variance matrix is:
[0122]
[0123] where
[0124]
[0125] In the formula, the genetic effect value of this genetic locus in the The genetic effect values in a dataset, and ; , , respectively represent the number of cases, the number of controls, and the total number of samples in the th dataset; , , respectively represent the number of cases, the number of controls, and the total number of samples in the th dataset; , respectively represent the number of overlapping controls and cases in the th and th datasets.
[0126] In specific implementation, , is calculated approximately based on the list of identification codes transmitted by multiple clients and the case proportion row. Since the identification codes and phenotype type data are not included in the data transmission, the total number of can be obtained by comparing whether the identification codes overlap, and then multiplied by the corresponding control sample proportion and case proportion respectively to finally obtain and .
[0127] S5-2. Calculate the weight vector of different genetic effect values from multiple clients and the main server dataset for each gene locus.
[0128] Specifically, the weight vector corresponding to different genetic effect values from multiple clients and the main server dataset for any locus is denoted as , , , respectively represent the weights of the genetic effect values of this gene locus in the 1st, th, th datasets.
[0129] Then the calculation formula of the weight vector is:
[0130]
[0131] Among them, is the unit vector; represents the variance matrix of the corresponding different genetic effect values.
[0132] In specific implementation, when calculating the weight for each gene locus, an inverse operation needs to be performed on , and parallel operations are used to improve the overall operation efficiency.
[0133] S5-3. Calculate the aggregated genetic effect value for each gene locus respectively.
[0134] Taking any gene locus as an example, the genetic effect value after aggregation of this locus is obtained by aggregating and calculating the data of this gene locus in datasets from multiple clients and the main server.
[0135] Specifically, the aggregated genetic effect value of any gene locus in datasets from multiple clients and the main server is denoted as ;
[0136] Then the calculation formula for the aggregated genetic effect value is:
[0137]
[0138] In the formula, represents the genetic effect value of this gene locus in the th dataset, represents the weight of the genetic effect value of this gene locus in the th dataset.
[0139] In specific implementation, aggregation calculations are performed for each gene locus, and finally the aggregated genetic effect values of each gene locus are obtained. Moreover, the aggregated genetic effect values are saved in the main server, and each client can also pull the data on the main server for research and analysis.
[0140] In summary, the entire calculation process of the present invention only uses summary statistics to protect the individual data of actual users, and it is impossible to obtain the individual data of each individual from the summary statistics. That is, the data is processed at the original location (such as a hospital or a research institution), and only the updates of the model (such as gradient information) are shared under encryption or privacy protection measures, solving the problems of data privacy and sample overlap when merging multiple GWAS analyses. And by integrating the data of multiple studies, the federated learning method of the present invention can increase the statistical power and improve the probability of identifying genetic variations related to diseases.
[0141] The technical solutions provided in the embodiments of the present invention have been introduced in detail above. Specific examples are used in this article to elaborate on the principle and implementation manner of the present invention. The description of the above embodiments is only used to help understand the method and its core idea of the present invention; at the same time, for those of ordinary skill in the art, according to the idea of the present invention, the technical solutions recorded in the foregoing embodiments can be modified, or some of the technical features can be equivalently replaced; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the idea and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A federated analysis method for non-independent gene data, characterized in that, Using a federated analysis system for non-independent gene data mainly composed of multiple clients and a main server, including the following steps: S1. The client uses an identification code generation algorithm to perform identity encryption and anonymization processing on the user gene sequence dataset to obtain an anonymized dataset; S2. Based on the anonymized dataset, perform data preprocessing on the phenotype data and genotype data respectively to obtain a standardized phenotype dataset and genotype dataset, and perform principal component analysis on the standardized genotype data to obtain the first 5 principal components; S3. Based on the standardized phenotype dataset, perform model training locally, select other required features except genotype data in the model training, and use a regression model. Using a specific phenotype determined according to the research purpose as the dependent variable, using other required features and the first 5 principal components as covariates, and using gene locus genotypes as independent variables, calculate the genetic effect value of each gene locus on the specific phenotype respectively to obtain GWAS data of the specific phenotype; S4. Based on the GWAS data of the specific phenotype, calculate the list of identification codes of the corresponding samples, as well as the control sample ratio and case ratio, to obtain the dataset of the client and transmit it to the main server; S5. The main server, based on the datasets from multiple clients and itself, uses the federated averaging algorithm to calculate the variance matrix and weight vector of different genetic effect values for each gene locus, as well as the aggregated genetic effect value; In step S5, calculating the variance matrix of different genetic effect values for each gene locus includes the following steps: Denote the different genetic effect values of any gene locus in multiple client and master server datasets as , and , indicates that there are a total of datasets, represents the th dataset, , , respectively represent the genetic effect values of this gene locus in the 1st, the th, and the th datasets; Denote the standard deviations of the corresponding different genetic effect values as , , , respectively represent the standard deviations of the genetic effect values of this gene locus in the 1st, th, th datasets; Denote the variance matrix of the corresponding different genetic effect values as ; Then, the calculation formula for the value at the -th row and the -th column of the variance matrix is: In the formula, the genetic effect value of this gene locus in the th dataset, and ; , , respectively represent the number of cases, the number of controls, and the total number of samples of the th dataset; , , respectively represent the number of cases, the number of controls, and the total number of samples of the th dataset; , respectively represent the number of overlapping controls and cases in the -th and -th datasets; The said , The calculation method is as follows: Based on the list of identification codes transmitted by multiple clients and the case proportions, by comparing whether the identification codes overlap, obtain the total number of, and then multiply by the corresponding control sample proportion and case proportion respectively to obtain and .
2. The federated analysis method for non-independent gene data according to claim 1, wherein In step S1, the client uses an identification code generation algorithm to perform identity encryption and anonymization processing on the user gene sequence dataset, including the following steps: S1-1. Clean and preprocess the dataset to remove information that can directly identify personal identity; S1-2. Group the dataset according to non-sensitive attributes to ensure that each equivalence class meets the requirements of L-diversity; S1-3. Generate a unique salt value for each equivalence class; S1-4. Combine the personal ID of each record with the corresponding salt value and apply a hash function to generate a unique identification code; S1-5. Use the generated identification code to replace the original personal ID.
3. A federated analysis method for non-independent gene data according to claim 1 or 2, characterized in that In step S2, performing data preprocessing on the phenotype data includes the following steps: Store the phenotype data in a.csv file, with the identification code as the first column, each phenotype as other columns, and missing values represented by NA.
4. The federated analysis method for non-independent gene data according to claim 1 or 2, characterized in that In step S2, performing data preprocessing on the genotype data includes the following steps: S2-2-1. Perform quality control screening on the genotype data to obtain a quality-controlled genotype dataset; S2-2-2. Perform standardization processing on the quality-controlled genotype dataset, subtracting the original value of the gene locus genotype from its mean and then dividing by its standard deviation.
5. The federated analysis method for non-independent gene data according to claim 4, characterized in that In step S2-2-1, performing quality control screening on the gene data includes the following steps: S2-2-11. Remove gene loci with an undetected rate higher than 2%; S2-2-12. Remove the gene loci with P-values in Hardy-Weinberg test less than ; S2-2-13. Remove gene loci with a minor allele frequency lower than 1%; S2-2-14. Remove gene loci with a gene imputation quality lower than 0.3; S2-2-15. Remove samples with an overall gene locus deletion rate greater than 5%.
6. A federated analysis method for non-independent gene data according to claim 1, characterized in that, The GWAS data of the specific phenotype in S3 is represented in the form of a data table on locus information, and its fields include the locus ID, coordinate information, genetic effect value and its standard deviation, regression intercept, and P-value of each gene locus.
7. A federated analysis method for non-independent gene data according to claim 1, characterized in that In S5, calculating the weight vector of different genetic effect values for each gene locus includes the following steps: Denote the weight vectors of the different genetic effect values of any gene locus in multiple client and master server datasets as , , , respectively represent the weights of the genetic effect values of this gene locus in the 1st, th, th datasets; Then the calculation formula for the weight vector is: In the formula, represents a unit vector; represents the variance matrix of the corresponding different genetic effect values.
8. A federated analysis method for non-independent gene data according to claim 7, characterized in that In S5, calculating the aggregated genetic effect value for each gene locus includes the following steps: Denote the aggregated genetic effect value of any gene locus in multiple client and master server datasets as ; Then the calculation formula for the aggregated genetic effect value is: In the formula, represents the genetic effect value of this gene locus in the th data set, represents the weight of the genetic effect value of this gene locus in the th data set.
Citation Information
Patent Citations
Privacy protection-based alliance learning system and method for realizing whole genome association analysis
CN113517027A
AI model private domain and public domain cooperative processing system based on data security and privacy protection
CN119106450A