Multi-agent data joint method, system, server and terminal device
By processing genomic data with encryption keys, cross-institutional data sharing and fusion were achieved, solving the problem of ineffective data aggregation in the field of genomics and improving the accuracy and efficiency of Mendelian randomization analysis.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- MGI TECH CO LTD
- Filing Date
- 2022-08-16
- Publication Date
- 2026-06-02
AI Technical Summary
Due to data privacy and legal policy restrictions, genomic data held by various parties cannot be effectively aggregated, affecting the performance of machine learning models. In particular, it is difficult to achieve cross-institutional data sharing and joint analysis in the field of genomics.
The original data is encrypted by generating first and second encryption keys. Instrumental variables and exposure factors are encrypted using random orthogonal matrices and random seeds. Data fusion and regression coefficient calculation are performed on the server side, enabling cross-institutional data sharing and Mendelian randomization analysis while ensuring data privacy.
While ensuring data privacy, it has enabled cross-institutional data sharing and integration, improved the accuracy and efficiency of Mendelian randomization analysis, and solved the problem of data ineffective aggregation.
Smart Images

Figure CN117634634B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data processing technology, and in particular to a method, system, server, and terminal device for multi-entity data collaboration. Background Technology
[0002] Data used to train machine learning models is often owned by different organizations and institutions, especially highly sensitive genomic data. Due to data security and user privacy concerns, this data cannot be simply aggregated for training or use. Furthermore, due to legal and regulatory concerns, as well as data privacy and security worries, data owners are unwilling to directly exchange raw data, resulting in ineffective data aggregation, impacting machine learning performance, and hindering the improvement of AI models. In the field of genomics, the amount of genomic data held by various parties is even more limited. To promote better data mining, large-scale sample aggregation and joint analysis are needed. Therefore, how to jointly train machine learning models by multiple data owners while ensuring data confidentiality has become a major challenge in data sharing and unlocking data value. Summary of the Invention
[0003] This disclosure aims to at least partially address one of the technical problems in the related art. To this end, one objective of this disclosure is to propose a multi-agent data collaboration method to achieve cross-institutional data sharing and fusion, and to complete Mendelian randomization analysis, while ensuring data privacy.
[0004] The second objective of this disclosure is to propose an alternative method for multi-agent data collaboration.
[0005] The third objective of this disclosure is to propose a server.
[0006] The fourth objective of this disclosure is to provide a terminal device.
[0007] The fifth objective of this disclosure is to propose a multi-agent data synergy system.
[0008] To achieve the above objectives, a first aspect of this disclosure proposes a multi-agent data joint method. The original data is distributed among K data owners, where K is an integer greater than or equal to 2. The original data includes instrumental variables, exposure factors, and outcome variables. The method includes: generating a first encryption key and a second encryption key; sending the first encryption key to the K data owners, enabling them to encrypt their own instrumental variables and exposure factors using the first encryption key; receiving the encrypted instrumental variables and exposure factors sent by the K data owners, and obtaining a first regression coefficient based on the encrypted instrumental variables and exposure factors sent by the K data owners; and then... The regression coefficient and the second encryption key are sent to the K data owners, so that the K data owners can use the first regression coefficient and their own instrumental variables to obtain the predicted values of their own exposure factors, and use the second encryption key to encrypt the predicted values of their own exposure factors and the outcome variables. The system receives the encrypted predicted values and outcome variables sent by the K data owners, and obtains the second regression coefficient based on the encrypted predicted values and outcome variables sent by the K data owners. The system sends the second regression coefficient to the K data owners, so that the K data owners can decrypt the second regression coefficient to obtain the regression coefficient between the predicted values of the exposure factors and the outcome variables in the original data.
[0009] The multi-subject data fusion method of this disclosure involves a server generating a first encryption key and a second encryption key, and sending the first encryption key to K data owners. Each data owner uses the first encryption key to encrypt their instrumental variables and exposure factors. The server receives the encrypted instrumental variables and exposure factors and uses them to obtain a first regression coefficient. Each data owner uses the first regression coefficient and their own instrumental variables to obtain predicted values for their own exposure factors, and then uses the second encryption key to encrypt the predicted values for their own exposure factors and the outcome variable. The server receives the encrypted predicted values and the outcome variable and uses them to obtain a second regression coefficient. Each data owner decrypts the second regression coefficient to obtain the regression coefficients between the predicted values of the exposure factors and the outcome variable in the original data. This method enables cross-institutional data sharing and fusion, and completes Mendelian randomization analysis, while ensuring data privacy.
[0010] In addition, the multi-subject data joint method proposed in the above embodiments of this disclosure may also have the following additional technical features:
[0011] According to one embodiment of this disclosure, the first encryption key includes a first random orthogonal matrix and a first random seed, and the second encryption key includes a second random orthogonal matrix and a second random seed. Before sending the first encryption key and the second encryption key to each of the data owners, the method further includes: dividing the first random orthogonal matrix into K first sub-matrices and dividing the second random orthogonal matrix into K second sub-matrices according to the original data quantity of each data owner; wherein, sending the first encryption key and the second encryption key to the K data owners includes: sending the K first sub-matrices and the K second sub-matrices to the corresponding data owners, and sending the first random seed and the second random seed to the K data owners respectively.
[0012] According to one embodiment of this disclosure, the first regression coefficient is obtained by the following formula:
[0013]
[0014] in, Let G′ be the first regression coefficient, and G′ = G′1 + G′2 + … + G′ k +…+G′ K ,X′=X′1+X′2+…+X′ k +…+X′ K , G′ k Let X′ be the encrypted instrumental variable of the k-th data owner. k The encrypted exposure factor for the kth data owner.
[0015] According to one embodiment of this disclosure, the second regression coefficient is obtained by the following formula:
[0016]
[0017] in, The second regression coefficient, Y′ is the encrypted prediction value of the kth data owner. k This is the encrypted ending variable for the kth data owner.
[0018] To achieve the above objectives, a second aspect of this disclosure proposes another multi-agent data joint method. The original data is distributed among K data owners, where K is an integer greater than or equal to 2. The original data includes instrumental variables, exposure factors, and outcome variables. The method includes: receiving a first encryption key and a second encryption key sent by a server, and encrypting its own instrumental variables and exposure factors using the first encryption key; sending the encrypted instrumental variables and exposure factors to the server, so that the server obtains a first regression coefficient based on the encrypted instrumental variables and exposure factors sent by the K data owners; receiving the first regression coefficient sent by the server and decrypting the first regression coefficient; using the decrypted first regression coefficient and its own instrumental variables to obtain a predicted value for its own exposure factors, and encrypting the predicted value for its own exposure factors and outcome variables using the second encryption key; sending the encrypted predicted value and outcome variables to the server, so that the server obtains a second regression coefficient based on the encrypted predicted value and outcome variables sent by the K data owners; receiving the second regression coefficient sent by the server and decrypting the second regression coefficient to obtain the regression coefficient between the predicted value of the exposure factors and the outcome variable in the original data.
[0019] In addition, the multi-subject data joint method proposed in the above embodiments of this disclosure may also have the following additional technical features:
[0020] According to one embodiment of this disclosure, the first encryption key includes a first random orthogonal matrix and a first random seed, and the second encryption key includes a second random orthogonal matrix and a second random seed. The server, based on the amount of original data for each data owner, divides the first random orthogonal matrix into K first sub-matrices and the second random orthogonal matrix into K second sub-matrices, and sends the K first sub-matrices and the K second sub-matrices to the corresponding data owners. The server also sends the first random seed and the second random seed to each of the K data owners. The step of encrypting its own instrumental variables and exposure factors using the first encryption key includes: generating a third orthogonal matrix based on the first random seed, and encrypting its own instrumental variables and exposure factors using the following formula:
[0021]
[0022] X′ k = k X k ,
[0023] Among them, G′ k G is the encrypted instrumental variable for the k-th data owner. kLet X′ be an instrumental variable in the original data of the k-th data owner. k X is the encrypted exposure factor for the k-th data owner. k For the exposure factors in the original data of the k-th data owner, P k Let Q be the first submatrix corresponding to the kth data owner, and let Q be the third orthogonal matrix.
[0024] According to one embodiment of this disclosure, encrypting the predicted values and outcome variables of the self-exposure factors using the second encryption key includes: generating a fourth orthogonal matrix based on the second random seed, and encrypting the predicted values and outcome variables of the self-exposure factors using the following formula:
[0025]
[0026] Y′ k =U k Y k ,
[0027] in, The encrypted predicted value for the k-th data owner. Y′ is the predicted value of the k-th data owner before encryption. k Y is the encrypted outcome variable for the kth data owner. k U is the outcome variable in the original data of the kth data owner. k V is the second submatrix corresponding to the kth data owner, and V is the fourth orthogonal matrix.
[0028] According to one embodiment of this disclosure, the first regression coefficient is decrypted using the third orthogonal matrix, and the second regression coefficient is decrypted using the fourth orthogonal matrix.
[0029] According to one embodiment of this disclosure, before sending the encrypted instrumental variables and exposure factors to the server, the method further includes: cooperating with any of the K data owners other than itself to determine a random seed; generating a random matrix based on the random seed; and using the random matrix to add random perturbations to the encrypted instrumental variables and exposure factors.
[0030] To achieve the above objectives, a third aspect of this disclosure provides a server comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the multi-subject data association method as described above.
[0031] To achieve the above objectives, a fourth aspect of this disclosure provides a terminal device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the multi-subject data association method as described above.
[0032] To achieve the above objectives, a fifth aspect of this disclosure proposes a multi-entity data federation system, the system comprising: a server as described above and K terminal devices as described in claim 11, wherein K is an integer greater than or equal to 2.
[0033] Additional aspects and advantages of this disclosure will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description
[0034] Figure 1 This is a flowchart of a multi-subject data association method according to an embodiment of this disclosure;
[0035] Figure 2 This is a flowchart illustrating an embodiment of the present disclosure before the encrypted instrumental variables and exposure factors are sent to the server.
[0036] Figure 3 This is a flowchart of a multi-subject data federation method according to another embodiment of this disclosure;
[0037] Figure 4 This is a schematic diagram of the structure of a server according to an embodiment of this disclosure;
[0038] Figure 5 This is a schematic diagram of the structure of a multi-subject data collaboration system according to an embodiment of this disclosure. Detailed Implementation
[0039] Embodiments of this disclosure are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this disclosure, and should not be construed as limiting this disclosure.
[0040] The following will refer to the instruction manual appendix. Figure 1-5 The present disclosure provides a detailed description of the multi-subject data association method, system, server, and terminal device according to specific implementation methods.
[0041] Mendelian randomization (MR) is a data analysis method that has been primarily applied in epidemiological etiology inference in recent years. Following Mendelian inheritance principles, where parental alleles are randomly assigned to offspring, if genotype determines phenotype and genotype is associated with disease through phenotype, then genotype can be used as an instrumental variable to infer the association between phenotype and disease. With the continuous development of statistical methods, genome-wide association studies, epigenetics, and omics technologies, MR is increasingly widely used in exploring the causal association between complex exposure factors and disease outcomes. Mendelian randomization analysis, using genetic variation as an instrumental variable to infer the causal relationship between exposure factors and outcomes, can effectively overcome the bias caused by heterozygous reverse causality.
[0042] In previous applications, one-sample Mendelian randomization (MR) analyzed a single study sample, requiring the availability of the original genomic data for the study. A large number of samples with the original genomic data were needed to obtain more confident analytical results, thus limiting its use. Therefore, to achieve sufficient power in MR analysis, it is necessary to increase the sample size. For example, to explore the association between exposure and disease using genetic tools that represent only 1% of the effect, at least tens of thousands of case and control samples are needed to achieve 80% power and obtain a 50% increase in biological effect.
[0043] Traditional MapReduce methods typically require a unified aggregation of all data before centralized computation. This process necessitates the transfer and storage of the original data. Once the data leaves the data domain, it becomes uncontrollable, potentially leading to data privacy breaches and creating data security vulnerabilities.
[0044] Figure 1 This is a flowchart of a multi-subject data association method according to an embodiment of this disclosure.
[0045] In embodiments of this disclosure, the raw data is distributed among K data owners, where K is an integer greater than or equal to 2. The raw data includes instrumental variables, exposure factors, and outcome variables. The execution entity in this embodiment may be a server, which may include a generation server and a decomposition server. This server may be owned by a third party, or a portion of the server may be owned by a third party, while another portion may be owned by one or more data owners, wherein the third party may not be any of the data owners. Figure 1 As shown, multi-agent data joint methods include:
[0046] S1, generate the first encryption key and the second encryption key.
[0047] Specifically, there are K original data owners, where K is an integer greater than or equal to 2. These K data owners participate in joint MR calculations to achieve privacy protection. The original data (such as genetic data) includes instrumental variables, exposure factors, and outcome variables, denoted as G, X, and Y, respectively. The instrumental variables, exposure factors, and outcome variables in the original data of the k-th data owner among the K data owners are denoted as G. k X k Y k k∈[1,K]. Each column in the original data can represent a feature space, and each row can represent a sample. The original data of the K data owners can be horizontal data, that is, the original data of the K data owners have the same feature space, that is, the number of columns in the original data is the same, but the sample space (i.e. the number of samples) can be different, that is, the number of rows in the original data of each data owner is different.
[0048] Among them, the instrumental variables in the original data of the K data owners Exposure factors Ending variables n k Let d be the number of samples from the k-th data owner, and d be the instrumental variable G. k The number of feature spaces is K, where the exposure factors and outcome variables in the original data of the K data owners are all single-feature-space data.
[0049] More specifically, when K data owners perform joint MR calculations, in order to protect data privacy, the original data is encrypted before being sent to the server. This disclosure uses two encryption keys generated by the server to encrypt the original data, wherein the two encryption keys are a first encryption key and a second encryption key.
[0050] In one embodiment of this disclosure, the first encryption key includes a first random orthogonal matrix and a first random seed, and the second encryption key includes a second random orthogonal matrix and a second random seed. Before sending the first encryption key and the second encryption key to each data owner, the multi-subject data joint method may further include: dividing the first random orthogonal matrix into K first sub-matrices and dividing the second random orthogonal matrix into K second sub-matrices according to the amount of original data of each data owner; wherein, sending the first encryption key and the second encryption key to the K data owners includes: sending the K first sub-matrices and the K second sub-matrices to the corresponding data owners, and sending the first random seed and the second random seed to the K data owners respectively.
[0051] Specifically, a first encryption key and a second encryption key can be generated by a generation server. The first encryption key includes a first random orthogonal matrix and a first random seed, and the second encryption key includes a second random orthogonal matrix and a second random seed. The first random orthogonal matrix and the second random orthogonal matrix are denoted as P and U, respectively. Both the first random orthogonal matrix P and the second random orthogonal matrix U can be n-dimensional square matrices, that is, the number of rows and columns of the first random orthogonal matrix P and the second random orthogonal matrix U are both n, where P∈R. n×n ,U∈R n×n n can be 5000.
[0052] In some implementations, the methods for generating the first random orthogonal matrix and the second random orthogonal matrix are the same. Taking the generation of the first random orthogonal matrix as an example, the generation method may include the following steps:
[0053] Step A: Create a 0 matrix and randomly generate a random number matrix of size a*a, where a is a preset constant.
[0054] Step B: Perform QR decomposition on the random number matrix to obtain an orthogonal submatrix, and write the orthogonal submatrix from the (i*a, i*a) position of the 0 matrix, where 0≤i≤a.
[0055] Step C: Increment the counter variable i and repeat the above steps until no new a*a orthogonal submatrix can be written into the 0 matrix, thus obtaining the first random orthogonal matrix.
[0056] It should be noted that the increment counter variable i is a positive integer, the initial value of the increment counter variable i can be set to 0, and the increment counter variable i will increase sequentially as the orthogonal submatrix is written into the 0 matrix.
[0057] As an example, let the value of a be 5000 and the initial value of i be 0. The server randomly generates a 5000*5000 random number matrix. This matrix is then subjected to QR decomposition to obtain an orthogonal submatrix. This orthogonal submatrix is written from position (0*5000, 0*5000) of the zero matrix, which is the (0, 0) position of the zero matrix. After writing the orthogonal submatrix from position (0*5000, 0*5000), the counter variable i is incremented to 1. The orthogonal submatrix is then written from position (1*5000, 1*5000) of the zero matrix. This process is repeated, continuously filling the zero matrix with 5000*5000 orthogonal submatrixes along its diagonal, until the zero matrix can no longer hold another 5000*5000 orthogonal submatrix and can no longer be filled with orthogonal submatrixes of other sizes along its diagonal. This yields the first random orthogonal matrix.
[0058] If the zero matrix cannot be filled with another 5000*5000 orthogonal submatrix, but can be filled with orthogonal submatrices of other sizes, the generator server generates a smaller random number matrix to obtain a new orthogonal submatrix smaller than the 5000*5000 orthogonal submatrix, and writes the new orthogonal submatrix from the (i*a, i*a) position of the zero matrix, where i is the incremented value. When the zero matrix cannot be filled with any other orthogonal submatrix of other sizes along its diagonal after this new orthogonal submatrix is filled, the first random orthogonal matrix is obtained.
[0059] More specifically, before sending the first random orthogonal matrix P and the second random orthogonal matrix U to each data owner, in order to match the data dimensions of each data owner, the first random orthogonal matrix P and the second random orthogonal matrix U are vertically partitioned into K first sub-matrices P. k and K second submatrices U k Ensure that the number of columns in the resulting submatrix is the same as the number of samples from the corresponding data owner. Where, n k Let be the number of samples of the k-th data owner. For example, if data owner 1 has 100 original data samples and data owner 2 has 200 original data samples, then the generation server can divide the first random orthogonal matrix into a 100*1 first submatrix and a 200*1 first submatrix.
[0060] S2, the first encryption key is sent to K data owners so that the K data owners can use the first encryption key to encrypt their own instrumental variables and exposure factors.
[0061] Specifically, the server generates the first submatrix P after splitting the first encryption key. k The first random seed from the first encryption key is sent to the corresponding data owner, and then sent to the K data owners. The K data owners use the first sub-matrix P after partitioning. k Encrypt the instrumental variables and exposure factors of the first random seed itself.
[0062] In one embodiment of this disclosure, encrypting one's own instrumental variables and exposure factors using a first encryption key may include:
[0063] A third orthogonal matrix is generated based on the first random seed, and its instrumental variables and exposure factors are encrypted using the following formula:
[0064]
[0065] X′ k =P k X k ,
[0066] Among them, G′ k G is the encrypted instrumental variable for the k-th data owner. k Let X′ be an instrumental variable in the original data of the k-th data owner. k X is the encrypted exposure factor for the k-th data owner. k For the exposure factors in the original data of the k-th data owner, P k Let Q be the first submatrix corresponding to the k-th data owner, and let Q be the third orthogonal matrix.
[0067] Specifically, the K data owners can use the same first random seed to generate a third orthogonal matrix Q. The third orthogonal matrix Q can be a square matrix, and its number of rows and columns can both be the number of feature spaces of the data owners plus 1, Q∈R. (d +1)×(d+1) The instrumental variable G of the kth data owner for itself. k Add 1 to the eigenvector in the first column and multiply it by the first submatrix P. k Right-multiplying by the third orthogonal matrix Q yields the encrypted instrumental variable G′ of the k-th data owner. k Exposure factor X in the original data of the k-th data owner k Left multiplying the first submatrix P k Encryption is performed. The obtained encrypted instrumental variable G′ from the k-th data owner is... k ∈R n×(d+1) The sample size is n, and the encrypted exposure factor X′ of the k-th data owner is... k ∈R n×1 The number of samples is also n, where n is greater than n k The number of samples for both the encrypted instrumental variables and exposure factors has increased, thereby improving the accuracy of the final analysis results by expanding the sample size of the original data from the data owner.
[0068] More specifically, after encrypting their instrumental variables and exposure factors using a first encryption key, the K data owners send the encrypted instrumental variables and exposure factors to their respective servers using a secure aggregation algorithm. The secure aggregation algorithm sends data of the same type and dimension to each data owner. Each pair of data owners secretly negotiates a random seed and uses the random seed to generate a random matrix with the same dimension as the sent data. This random matrix then adds random perturbation to the sent data.
[0069] In one embodiment of this disclosure, such as Figure 2 As shown, before sending the encrypted instrumental variables and exposure factors to the server, the multi-agent data joint approach may also include:
[0070] S201, cooperate with any of the K data owners other than itself to determine the random seed.
[0071] S202 generates a random matrix based on a random seed, and uses the random matrix to add random perturbations to the encrypted instrumental variables and exposure factors.
[0072] Specifically, the k data owners will encrypt the instrumental variable G′ k and exposure factor X′ k The above-described secure aggregation algorithm is used to send the data to the decomposition server. The following uses exposure factor X′ as an example. k Taking a secure aggregation algorithm as an example, k data owners possess encrypted exposure factors {X′1,…,X′}. k Each encrypted exposure factor has a dimension of R. n×1 K data owners cooperate with any other data owner (excluding themselves) to determine a random seed s. i,j (1≤i≤j≤K), generate a random matrix M based on a random seed. i,j Random matrix M i,j The dimension is also R n×1 The sum of the random matrices generated by the two data owners with the random seed is 0. The i-th data owner adds a random perturbation to its own encrypted exposure factor, and the exposure factor after adding the random perturbation is X′. i =X′ i +∑ i≠j sign(i,j)M i,j If i > j, then sign(i,j) = 1; otherwise, sign(i,j) = -1, and the exposure factor X′ after random perturbation will be increased. i Send to the decomposition server.
[0073] More specifically, based on the random seed s i,j Generate a random matrix M i,j Hash algorithms, such as Sha256, or stream key generation algorithms can be used as pseudo-random number generators for secure cryptography. Using a hash function H (such as Sha256), n*1 random numbers r are calculated. t =H(s) i,j The binary random number is converted into data of a specific data structure (such as a floating-point number) according to rules, and is still denoted as r. t After rearranging, an n*1 dimensional random matrix M is obtained. i,j .
[0074] It should be noted that if i > j, then a random seed s is used. j,iTo reduce communication and computational load, a fixed number of data owners can be selected to negotiate a random seed, eliminating the need for negotiation between any two data owners and avoiding negotiation with all other data owners to generate a random matrix. To further reduce communication and computational load, a symmetric random matrix can be used to generate the random matrix. A secure aggregation algorithm is then used to process the encrypted instrumental variable G′. k The processing method and the exposure factor X′ after encryption k The processing method is the same, so I will not repeat it here.
[0075] S3 receives encrypted instrumental variables and exposure factors from K data owners and obtains the first regression coefficient based on the encrypted instrumental variables and exposure factors from the K data owners.
[0076] Specifically, the k data owners will encrypt the instrumental variable G′ k and exposure factor X′ k The above-described secure aggregation algorithm is used to send the data to the decomposition server. The decomposition server receives K encrypted instrumental variables G′ from the data owners. k and exposure factor X′ k The complete instrumental variables and exposure factors are summed together, and the first regression coefficient is obtained using the calculation formula.
[0077] In one embodiment of this disclosure, the first regression coefficient is obtained by the following formula:
[0078]
[0079] Where G′=G′1+G′2+…+G′ k +…+G′ K ,X′=X′1+X′2+…+X′ k +…+X′ K , G′ is the first regression coefficient. k Let X′ be the encrypted instrumental variable of the k-th data owner. k The encrypted exposure factor for the kth data owner.
[0080] Specifically, the decomposition server receives the encrypted instrumental variable G sent by each data owner. k and exposure factor X′ k The instrumental variable G′ sent by each data owner k With the same dimension, n*(d+1), the decomposition server encrypts the instrumental variables G′ of each data owner. k Summing these together yields the complete instrumental variable G′ = G′1 + G′2 + ... + G′ k +…+G′ KSimilarly, the exposure factors X′ sent by each data owner k The dimensions are the same, all n*1. The decomposition server encrypts the exposure factors X′ of each data owner. k The complete exposure factors X′ are obtained by summing them up: X′1 + X′2 + … + X′ k +…+X′ K Furthermore, utilizing Obtain the first regression coefficient
[0081] S4. Send the first regression coefficient and the second encryption key to K data owners so that the K data owners can use the first regression coefficient and their own instrumental variables to obtain the predicted values of their own exposure factors, and use the second encryption key to encrypt the predicted values of their own exposure factors and the outcome variables.
[0082] Specifically, the decomposition server will decompose the first regression coefficient. The data is sent to K data owners, who then use the first regression coefficient and their own instrumental variables to obtain predicted values for their own exposure factors. Specifically, the first regression coefficient sent by the decomposition server... For each data owner, this is the encrypted first regression coefficient. It cannot be used directly; the first regression coefficient needs to be adjusted. Decryption processing is performed because of the first regression coefficient. The first regression coefficient after decryption can be derived. The first regression coefficient is about to be obtained Left multiplication by the third orthogonal matrix yields the decrypted first regression coefficient. Then, the decrypted first regression coefficients are used. and its own instrumental variables Obtain the predicted values of self-exposure factors. Predicted values of self-exposure factors obtained For n k *1-dimensional matrix.
[0083] More specifically, the K data owners obtain predicted values for their own exposure factors. Then, the predicted values of the self-exposure factors and the outcome variables are encrypted using the second encryption key. The second encryption key includes a second random orthogonal matrix and a second random seed. Before sending the second random orthogonal matrix U to the K data owners, the second random orthogonal matrix U is vertically partitioned into K second sub-matrices U to match the data dimensions of each data owner. k Ensure that the number of columns in the resulting submatrix is the same as the number of samples from the corresponding data owner. Where, nk Let be the number of samples owned by the k-th data owner.
[0084] In one embodiment of this disclosure, the predicted values and outcome variables of self-exposure factors are encrypted using a second encryption key, including: generating a fourth orthogonal matrix based on a second random seed, and encrypting the predicted values and outcome variables of self-exposure factors using the following formula:
[0085]
[0086] Y′ k =U k Y k ,
[0087] in, This is the encrypted predicted value for the k-th data owner. Y′ is the predicted value before encryption by the k-th data owner. k Y is the encrypted outcome variable for the kth data owner. k U is the outcome variable in the original data of the kth data owner. k Let V be the second submatrix corresponding to the k-th data owner, and let V be the fourth orthogonal matrix.
[0088] Specifically, the K data owners can use the same second random seed to generate a fourth orthogonal matrix V. The fourth orthogonal matrix V can be a square matrix with 2 rows and 2 columns, where V∈R. 2×2 The predicted value of the k-th data owner for their own exposure factors. Add 1 to the eigenvector in the last column and multiply it by the second submatrix U on the left. k Right-multiplying by the fourth orthogonal matrix V yields the encrypted prediction value for the k-th data owner. For the outcome variable Y in the original data of the kth data owner k Left multiplying the second submatrix U k Encryption is performed. The resulting encrypted predicted value of the self-exposure factor of the k-th data owner is obtained. The number of samples is n, and the encrypted final variable Y′ of the kth data owner is... k ∈R n×1 The number of samples is also n, where n is greater than n k The encrypted predicted values of self-exposure factors and the sample size of outcome variables both increased. Therefore, by increasing the sample size of the original data from the data owner, the accuracy of the final analysis results can be improved.
[0089] Furthermore, the K data owners send the encrypted predicted values and outcome variables to the server using a secure aggregation algorithm, which has been described in detail above and will not be repeated here.
[0090] It should be noted that the random seed used when transmitting the encrypted predicted value and the outcome variable through the secure aggregation algorithm can be a reused random seed from the previous stage, or a new random seed can be generated.
[0091] S5 receives encrypted predicted values and outcome variables from K data owners, and obtains the second regression coefficient based on the encrypted predicted values and outcome variables from the K data owners.
[0092] Specifically, the decomposition server receives encrypted predicted values of the self-exposed factors from each data owner. and the outcome variable Y′ k The predicted values and outcome variables are summed to obtain the complete predicted values and outcome variables, and the second regression coefficient is obtained using the calculation formula.
[0093] In embodiments of this disclosure, the second regression coefficient is obtained by the following formula:
[0094]
[0095] in, The second regression coefficient, Y′ is the encrypted prediction value of the kth data owner. k This is the encrypted ending variable for the kth data owner.
[0096] Specifically, the decomposition server receives encrypted predicted values of its own exposure factors from each data owner. and the outcome variable Y′ K The predicted values sent by each data owner The dimensions are the same, both being n*2 dimensional matrices. The decomposition server encrypts the predicted values from each data owner. Sum to obtain the complete predicted value Similarly, the final variable Y′ sent by each data owner k The dimensions are all the same, n*1, and the decomposition server encrypts the final variable Y′ of each data owner. k The summation yields the complete outcome variable Y′ = Y1′ + Y2′ + ... + Y′ k +…+Y′ K Furthermore, utilizing Obtain the second regression coefficient
[0097] S6. Send the second regression coefficient to K data owners so that they can decrypt the second regression coefficient and obtain the regression coefficient between the predicted values of the exposure factors and the outcome variables in the original data.
[0098] Specifically, the decomposition server will decompose the second regression coefficient. Sending data to K data owners, and the K data owners' responses to the second regression coefficients sent. Decryption processing is performed due to the second regression coefficient. The regression coefficients between the predicted values of the exposure factors and the outcome variables in the decrypted original data can be derived. The regression coefficients between the predicted values of the exposure factors and the outcome variables in the obtained raw data. This is a 2x1 dimensional matrix. It represents the regression coefficients between the predicted values of the exposure factors and the outcome variables in the original data. This is the result of a multi-faceted Mendelian randomization analysis.
[0099] The multi-agent data fusion method of this disclosure involves a server generating a first encryption key and a second encryption key, and sending the first encryption key to K data owners. Each of the K data owners uses the first encryption key to encrypt their instrumental variables and exposure factors. The server receives the encrypted instrumental variables and exposure factors and uses them to obtain a first regression coefficient. Each of the K data owners uses the first regression coefficient and their own instrumental variables to obtain predicted values for their own exposure factors, and then uses the second encryption key to encrypt the predicted values for their own exposure factors and the outcome variable. The server receives the encrypted predicted values and the outcome variable and uses them to obtain a second regression coefficient. The K data owners decrypt the second regression coefficient to obtain the regression coefficients between the predicted values of the exposure factors and the outcome variable in the original data. This multi-agent data fusion method can achieve cross-institutional data sharing and fusion, and complete Mendelian randomization analysis, while ensuring data privacy.
[0100] This disclosure also proposes another method for multi-agent data collaboration.
[0101] In one embodiment of this disclosure, such as Figure 3 As shown, the original data is distributed among K data owners, where K is an integer greater than or equal to 2. The original data includes instrumental variables, exposure factors, and outcome variables. The implementing entity of this multi-agent data consortium method can be the data owner, and the multi-agent data consortium method may include:
[0102] S10, receive the first encryption key and the second encryption key sent by the server, and use the first encryption key to encrypt its own instrumental variables and exposure factors.
[0103] S20, send the encrypted instrumental variables and exposure factors to the server so that the server can obtain the first regression coefficient based on the encrypted instrumental variables and exposure factors sent by the K data owners.
[0104] S30: Receive the first regression coefficient sent by the server and decrypt the first regression coefficient.
[0105] S40, using the decrypted first regression coefficient and its own instrumental variables to obtain the predicted values of its own exposure factors, and using the second encryption key to encrypt the predicted values of its own exposure factors and the outcome variables.
[0106] S50 sends the encrypted predicted values and outcome variables to the server so that the server can obtain the second regression coefficients based on the encrypted predicted values and outcome variables sent by the K data owners.
[0107] S60: Receive the second regression coefficient sent by the server, decrypt the second regression coefficient, and obtain the regression coefficient between the predicted value of the exposure factor and the outcome variable in the original data.
[0108] In some embodiments of this disclosure, the first encryption key includes a first random orthogonal matrix and a first random seed, and the second encryption key includes a second random orthogonal matrix and a second random seed. The server, based on the amount of original data from each data owner, divides the first random orthogonal matrix into K first sub-matrices and the second random orthogonal matrix into K second sub-matrices, and sends the K first sub-matrices and K second sub-matrices to the corresponding data owners. The first random seed and the second random seed are each sent to one of the K data owners. The encryption of the server's instrumental variables and exposure factors using the first encryption key includes: generating a third orthogonal matrix based on the first random seed, and encrypting the server's instrumental variables and exposure factors using the following formula:
[0109]
[0110] X′ k =P k X k ,
[0111] Among them, G′ k G is the encrypted instrumental variable for the k-th data owner. k Let X′ be an instrumental variable in the original data of the k-th data owner. k X is the encrypted exposure factor for the k-th data owner. k For the exposure factors in the original data of the k-th data owner, P k Let Q be the first submatrix corresponding to the k-th data owner, and let Q be the third orthogonal matrix.
[0112] In some embodiments of this disclosure, the predicted values and outcome variables of self-exposure factors are encrypted using a second encryption key, including: generating a fourth orthogonal matrix based on a second random seed, and encrypting the predicted values and outcome variables of self-exposure factors using the following formula:
[0113]
[0114] Y′ k =U k Y k ,
[0115] in, This is the encrypted predicted value for the k-th data owner. Y′ is the predicted value before encryption by the k-th data owner. k Y is the encrypted outcome variable for the kth data owner. k U is the outcome variable in the original data of the kth data owner. k Let V be the second submatrix corresponding to the k-th data owner, and let V be the fourth orthogonal matrix.
[0116] In some embodiments of this disclosure, the first regression coefficient is decrypted using a third orthogonal matrix, and the second regression coefficient is decrypted using a fourth orthogonal matrix.
[0117] In some embodiments of this disclosure, before sending the encrypted instrumental variables and exposure factors to the server, the method further includes: cooperating with any one of the K data owners other than itself to determine a random seed; generating a random matrix based on the random seed; and using the random matrix to add random perturbations to the encrypted instrumental variables and exposure factors.
[0118] It should be noted that for other specific implementations of the multi-subject data federation method for data owners in this embodiment, please refer to the specific implementation of the multi-subject data federation method for servers in the first aspect of this disclosure.
[0119] The multi-subject data collaboration method disclosed in this embodiment can achieve cross-institutional data sharing and fusion, and complete Mendelian randomization analysis, while ensuring data privacy.
[0120] This disclosure also proposes a server.
[0121] Figure 4 This is a structural block diagram of a server according to an embodiment of the present disclosure.
[0122] like Figure 4As shown, server 400 includes a processor 401 and a memory 403. The processor 401 and memory 403 are connected, for example, via a bus 402. Optionally, server 400 may also include a transceiver 404. It should be noted that in practical applications, the transceiver 404 is not limited to one type, and the structure of server 400 does not constitute a limitation on the embodiments of this disclosure.
[0123] Processor 401 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It may implement or execute the various exemplary logic blocks, modules, and circuits described in connection with this disclosure. Processor 401 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.
[0124] Bus 402 may include a pathway for transmitting information between the aforementioned components. Bus 402 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 402 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 4 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0125] The memory 403 may be a ROM (Read Only Memory) or other type of static storage device capable of storing static information and instructions, RAM (Random Access Memory) or other type of dynamic storage device capable of storing information and instructions, or an EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media or other magnetic storage devices, or any other medium capable of carrying or storing desired program code in the form of instructions or data structures and accessible by a computer, but not limited thereto.
[0126] The memory 403 is used to store application code that executes the present disclosure, and its execution is controlled by the processor 401. The processor 401 is used to execute the application code stored in the memory 403 to implement the content shown in the foregoing method embodiment for a server.
[0127] Figure 4 The server 400 shown is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.
[0128] This disclosure also proposes a terminal device.
[0129] In this embodiment, the terminal device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements the multi-subject data association method as described in the second aspect of this disclosure.
[0130] This disclosure also proposes a multi-entity data synergy system.
[0131] In this embodiment, such as Figure 5 As shown, the multi-entity data joint system 1000 includes: the server 400 as described above and K terminal devices 200 as described above, where K is an integer greater than or equal to 2.
[0132] It should be noted that the logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.
[0133] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.
[0134] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this disclosure. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0135] In the description of this disclosure, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," "outer," "clockwise," "counterclockwise," "axial," "radial," and "circumferential" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. They are used only for the convenience of describing this disclosure and simplifying the description, and are not intended to indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on this disclosure.
[0136] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this disclosure, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0137] In this disclosure, unless otherwise expressly specified and limited, the terms "installation," "connection," "linking," "fixing," etc., should be interpreted broadly. For example, they can refer to a fixed connection, a detachable connection, or an integral part; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; they can refer to the internal communication of two components or the interaction between two components, unless otherwise expressly limited. Those skilled in the art can understand the specific meaning of the above terms in this disclosure according to the specific circumstances.
[0138] In this disclosure, unless otherwise expressly specified and limited, "above" or "below" the second feature can mean that the first and second features are in direct contact, or that the first and second features are in indirect contact through an intermediate medium. Furthermore, "above," "on top of," and "over" the second feature can mean that the first feature is directly above or diagonally above the second feature, or simply that the first feature is at a higher horizontal level than the second feature. "Below," "below," and "under" the second feature can mean that the first feature is directly below or diagonally below the second feature, or simply that the first feature is at a lower horizontal level than the second feature.
[0139] Although embodiments of the present disclosure have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present disclosure. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present disclosure.
Claims
1. A multi-agent data joint method, characterized in that, The raw data is distributed among K data owners, where K is an integer greater than or equal to 2. The raw data includes instrumental variables, exposure factors, and outcome variables. The method includes: Generate the first encryption key and the second encryption key; The first encryption key is sent to the K data owners so that the K data owners can use the first encryption key to encrypt their own instrumental variables and exposure factors; Receive encrypted instrumental variables and exposure factors sent by the K data owners, and obtain the first regression coefficient based on the encrypted instrumental variables and exposure factors sent by the K data owners; The first regression coefficient and the second encryption key are sent to the K data owners so that the K data owners can use the first regression coefficient and their own instrumental variables to obtain the predicted values of their own exposure factors, and use the second encryption key to encrypt the predicted values of their own exposure factors and the outcome variables. Receive encrypted predicted values and outcome variables sent by the K data owners, and obtain the second regression coefficient based on the encrypted predicted values and outcome variables sent by the K data owners; The second regression coefficient is sent to the K data owners so that the K data owners can decrypt the second regression coefficient to obtain the regression coefficient between the predicted value of the exposure factor and the outcome variable in the original data.
2. The multi-entity data joint method according to claim 1, characterized in that, The first encryption key includes a first random orthogonal matrix and a first random seed; the second encryption key includes a second random orthogonal matrix and a second random seed; before sending the first encryption key and the second encryption key to each of the data owners, the method further includes: Based on the original amount of data for each data owner, the first random orthogonal matrix is divided into K first sub-matrices and the second random orthogonal matrix is divided into K second sub-matrices; The step of sending the first encryption key and the second encryption key to the K data owners includes: The K first sub-matrices and the K second sub-matrices are sent to their respective data owners, and the first random seed and the second random seed are sent to the K data owners respectively.
3. The multi-subject data joint method according to claim 1, characterized in that, The first regression coefficient is obtained using the following formula: in, Let G′ be the first regression coefficient, and G′ = G′1 + G′2 + … + G′ k +…+G′ K ,X′=X′1+X′2+…+X′ k +…+X′ K , G′ k Let X′ be the encrypted instrumental variable of the k-th data owner. k The encrypted exposure factor for the kth data owner.
4. The multi-entity data joint method according to claim 1, characterized in that, The second regression coefficient is obtained using the following formula: in, The second regression coefficient, Y′=Y′1+Y′2+…+Y′ k +…+Y′ K , Y′ is the encrypted prediction value of the kth data owner. k This is the encrypted ending variable for the kth data owner.
5. A multi-agent data joint method, characterized in that, The raw data is distributed among K data owners, where K is an integer greater than or equal to 2. The raw data includes instrumental variables, exposure factors, and outcome variables. The method includes: Receive the first encryption key and the second encryption key sent by the server, and use the first encryption key to encrypt its own instrumental variables and exposure factors; The encrypted instrumental variables and exposure factors are sent to the server so that the server can obtain the first regression coefficient based on the encrypted instrumental variables and exposure factors sent by the K data owners; Receive the first regression coefficient sent by the server and decrypt the first regression coefficient; The predicted values of the self-exposure factors are obtained by using the decrypted first regression coefficient and the instrumental variable of the self, and the predicted values of the self-exposure factors and the outcome variable are encrypted by using the second encryption key; The encrypted predicted values and outcome variables are sent to the server so that the server can obtain the second regression coefficient based on the encrypted predicted values and outcome variables sent by the K data owners; The system receives the second regression coefficient sent by the server and decrypts the second regression coefficient to obtain the regression coefficient between the predicted value of the exposure factor and the outcome variable in the original data.
6. The multi-entity data syndication method according to claim 5, characterized in that, The first encryption key includes a first random orthogonal matrix and a first random seed; the second encryption key includes a second random orthogonal matrix and a second random seed. The server, based on the amount of original data from each data owner, divides the first random orthogonal matrix into K first sub-matrices and the second random orthogonal matrix into K second sub-matrices, and sends the K first sub-matrices and K second sub-matrices to the corresponding data owners. The server also sends the first random seed and the second random seed to each of the K data owners. The step of encrypting its own instrumental variables and exposure factors using the first encryption key includes: A third orthogonal matrix is generated based on the first random seed, and its instrumental variables and exposure factors are encrypted using the following formula: X′ k =P k X k , Among them, G′ k G is the encrypted instrumental variable for the k-th data owner. k Let X′ be an instrumental variable in the original data of the k-th data owner. k X is the encrypted exposure factor for the k-th data owner. k For the exposure factors in the original data of the k-th data owner, P k Let Q be the first submatrix corresponding to the kth data owner, and let Q be the third orthogonal matrix.
7. The multi-entity data joint method according to claim 6, characterized in that, The step of encrypting the predicted values of the self-exposure factors and the outcome variables using the second encryption key includes: A fourth orthogonal matrix is generated based on the second random seed, and the predicted values of the self-exposure factors and the outcome variables are encrypted using the following formula: AND' k =U k AND k , in, The encrypted predicted value for the k-th data owner. Y′ is the predicted value of the k-th data owner before encryption. k Y is the encrypted outcome variable for the kth data owner. k U is the outcome variable in the original data of the kth data owner. k V is the second submatrix corresponding to the kth data owner, and V is the fourth orthogonal matrix.
8. The multi-entity data joint method according to claim 7, characterized in that, The first regression coefficient is decrypted using the third orthogonal matrix, and the second regression coefficient is decrypted using the fourth orthogonal matrix.
9. The multi-entity data joint method according to claim 5, characterized in that, Before sending the encrypted instrumental variables and exposure factors to the server, the method further includes: The random seed is determined by cooperating with any of the K data owners other than itself. A random matrix is generated based on the random seed, and the random matrix is used to add random perturbations to the encrypted instrumental variables and exposure factors.
10. A server, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor, when executing the program, implements the multi-subject data association method as described in any one of claims 1-4.
11. A terminal device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, it implements the multi-subject data association method as described in any one of claims 5-9.
12. A multi-entity data collaboration system, characterized in that, The system includes: a server as described in claim 10 and K terminal devices as described in claim 11, wherein K is an integer greater than or equal to 2.