Joint modeling method, device, electronic device and storage medium

By performing public sample ID identification processing on the first sample set in federated learning and calculating gradient values ​​for parameter updates, the problem of low joint model accuracy when participants lack sample characteristics and sample count is solved, and a higher joint model accuracy is achieved.

CN114298321BActive Publication Date: 2025-05-27WELAB INFORMATION TECH SHENZHEN LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111607572.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-24
Publication Date
2025-05-27
Estimated Expiration
2041-12-24

AI Technical Summary

Technical Problem

In federated learning, if the participants lack both sample characteristics and sample size, the joint model constructed using existing horizontal or vertical federated learning schemes is not very accurate.

Method used

By receiving the public key in the homomorphic encryption key pair sent by each second participant, a public sample ID identification process is performed on the local first sample set and the second sample set of each second participant, and the first sample set is split into the first sub-sample set corresponding to each second participant. Then, based on the initial parameters of the preset model of each first sub-sample set and the second sample set of the corresponding second participants, the gradient value is calculated and the parameter update is performed. Finally, when the preset model converges, the target parameters are determined and sent to other participants to complete joint modeling.

Benefits of technology

In the absence of sample features and sample count, the accuracy of the joint model is improved, through the modeling of lateral participants with the same sample characteristics and different sample objects and vertical participants with the same sample objects and different sample characteristics.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114298321B_ABST
    Figure CN114298321B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data processing, and discloses a joint modeling method, including: performing public sample ID recognition processing on the first sample set and each second sample set respectively, and splitting the first sample set into first sub-sample sets corresponding to each second participant based on the recognition results; calculating gradient values corresponding to each first sub-sample set based on the initial parameters of the preset model corresponding to each first sub-sample set and the second sample set of the corresponding second participant; determining first parameters corresponding to each first sub-sample set based on the gradient values; receiving second parameters sent by each second participant and third parameters sent by other first participants; when it is determined that the preset model converges, determining target parameters based on the first parameters, second parameters and third parameters, and sending the target parameters to other participants to complete joint modeling. The present invention also provides a joint modeling device, an electronic device and a storage medium. The present invention improves the accuracy of the joint model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and in particular, to a joint modeling method, apparatus, electronic device, and storage medium. Background Art

[0002] To eliminate data islands and ensure data security, federated learning has been widely applied to joint modeling. During the federated learning process, each participating party does not share data, trains a model using local data respectively, and updates the model parameters by exchanging encrypted model parameters to complete the modeling.

[0003] Federated learning includes horizontal federated learning and vertical federated learning. Generally, if participating parties have the same sample features but insufficient sample quantities, a horizontal federated learning scheme is adopted; if participating parties have sufficient sample quantities but lack sample features, a vertical federated learning scheme is adopted. However, for the situation where both sample features and sample quantities are lacking, no matter whether a horizontal federated learning scheme or a vertical federated learning scheme is used, the accuracy of the constructed joint model is not high. Therefore, there is an urgent need for a joint modeling method to improve the accuracy of the joint model in the case of lacking sample features and sample quantities. Summary of the Invention

[0004] In view of the above, it is necessary to provide a joint modeling method aimed at improving the accuracy of the joint model.

[0005] The joint modeling method provided by the present invention is applied to any first participating party in a joint modeling system. The joint modeling system includes a plurality of first participating parties and a plurality of second participating parties that are communicatively connected. There are the same sample objects and different sample features between each first participating party and each second participating party, and there are the same sample features and different sample objects between each second participating party. The method includes:

[0006] Receiving the public key in the homomorphic encryption key pair sent by each second participating party in the joint modeling system, performing public sample ID recognition processing on the first sample set without label information stored locally and the second sample set with label information of each second participating party based on the public key, and splitting the first sample set into first sub-sample sets corresponding to each second participating party based on the public sample ID recognition result;

[0007] Obtaining the initial parameters of a preset model corresponding to each first sub-sample set, and calculating the gradient value corresponding to each first sub-sample set based on the public key, the initial parameters, and the second sample set of the corresponding second participating party;

[0008] Performing parameter update processing on the preset model corresponding to each first sub-sample set based on the gradient value to obtain the first parameter corresponding to each first sub-sample set;

[0009] Receive the second parameters and loss values processed by the secure aggregation algorithm corresponding to its second sample set sent by each second participant, and receive the third parameters processed by the secure aggregation algorithm corresponding to each sub-sample set sent by other first participants;

[0010] Based on the loss value, determine whether the preset model converges. When the determination is yes, determine the target parameters based on the first parameter, the second parameter, and the third parameter, and send the target parameters to other participants in the joint modeling system respectively to complete the joint modeling.

[0011] Optionally, the performing public sample ID identification processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant based on the public key includes:

[0012] Select a second participant, calculate the first hash values of the sample IDs in the first sample set, encrypt the first hash values with the public key in the homomorphic encryption key pair corresponding to the selected second participant to obtain first ciphertexts, and establish a mapping relationship between the first ciphertexts and the sample IDs;

[0013] Receive the second ciphertexts sent by the selected second participant, where the second ciphertexts are the second hash values of the sample IDs in its second sample set encrypted by the public key in the same homomorphic encryption key pair by the selected second participant;

[0014] Calculate the intersection of the first ciphertexts and the second ciphertexts to obtain the public sample ID ciphertexts, and determine the plaintext data of the public sample ID ciphertexts based on the mapping relationship.

[0015] Optionally, the calculating the gradient value corresponding to each first sub-sample set based on the public key, the initial parameters, and the second sample set of the corresponding second participant includes:

[0016] Start multiple processes according to the number of the first sub-sample sets. Each process calculates the first feature matrix corresponding to each first sub-sample set according to the corresponding first sub-sample set and its initial parameters;

[0017] Send the first feature matrix to the corresponding second participant, and receive the error value encrypted by the public key sent by the corresponding second participant, where the error value is calculated by the corresponding second participant according to the second feature matrix of its second sample set and the first feature matrix;

[0018] Substitute the encrypted error value into the gradient value calculation formula to obtain the encrypted gradient value corresponding to each first sub-sample set, and send the encrypted gradient value to the corresponding second participant to obtain the plaintext data of the encrypted gradient value.

[0019] Optionally, the process by which the corresponding second party calculates the error value based on the second feature matrix of its second sample set and the first feature matrix includes:

[0020] The corresponding second party calculates the eigenvalue of its second sample set based on the second feature matrix of its second sample set and the first feature matrix;

[0021] Input the eigenvalue into a preset model to obtain the predicted value of its second sample set;

[0022] Determine the true value of its second sample set based on the label information, calculate the error value based on the true value and the predicted value, and encrypt the error value with the public key in the corresponding homomorphic encryption key pair and send it to the corresponding first party.

[0023] Optionally, the sending the encrypted gradient value to the corresponding second party to obtain the plaintext data of the encrypted gradient value includes:

[0024] Generate a third random number for each second party, encrypt the corresponding third random number with the corresponding public key, calculate the sum of the encrypted gradient value and the encrypted third random number to obtain an encrypted sum, and send the encrypted sum to the corresponding second party;

[0025] Receive the value decrypted by the corresponding second party from the encrypted sum, and subtract the corresponding third random number from the obtained value to obtain the decrypted gradient value corresponding to the corresponding first sub-sample set.

[0026] Optionally, each process calculates the first feature matrix corresponding to each first sub-sample set according to the corresponding first sub-sample set and its initial parameters, including:

[0027] Select a process, obtain the first sub-sample set and the initial parameters corresponding to the process, determine the initial feature matrix of the obtained first sub-sample set, and calculate the first feature matrix corresponding to the obtained first sub-sample set based on the initial feature matrix and the initial parameters.

[0028] Optionally, the calculation formula of the loss value is:

[0029]

[0030] where L i is the loss value corresponding to the i-th second party, y ij is the true value of the j-th sample in the second sample set of the i-th second party, h θ (x ijis the predicted value of the j-th sample in the second sample set of the i-th second participant, and n is the total number of samples in the second sample set of the i-th second participant.

[0031] To solve the above problems, the present invention also provides a joint modeling device, which includes:

[0032] A receiving module, configured to receive the public keys in the homomorphic encryption key pairs sent by each second participant in the joint modeling system, perform public sample ID recognition processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant based on the public keys, and split the first sample set into first sub-sample sets corresponding to each second participant based on the public sample ID recognition results;

[0033] A calculation module, configured to obtain the initial parameters of the preset model corresponding to each first sub-sample set, and calculate the gradient values corresponding to each first sub-sample set based on the public keys, the initial parameters, and the second sample sets of the corresponding second participants;

[0034] An update module, configured to perform parameter update processing on the preset model corresponding to each first sub-sample set based on the gradient values to obtain the first parameters corresponding to each first sub-sample set;

[0035] A receiving module, configured to receive the second parameters and loss values corresponding to the second sample sets of each second participant processed by the secure aggregation algorithm, and receive the third parameters corresponding to each sub-sample set of other first participants processed by the secure aggregation algorithm sent by other first participants;

[0036] A determination module, configured to determine whether the preset model converges based on the loss value. When it is determined that the model converges, determine the target parameters based on the first parameters, the second parameters, and the third parameters, and send the target parameters to other participants in the joint modeling system respectively to complete the joint modeling.

[0037] To solve the above problems, the present invention also provides an electronic device, which includes:

[0038] At least one processor; and,

[0039] A memory communicatively connected to the at least one processor; wherein,

[0040] The memory stores a joint modeling program executable by the at least one processor. The joint modeling program is executed by the at least one processor so that the at least one processor can execute the above joint modeling method.

[0041] To solve the above problems, the present invention also provides a computer-readable storage medium, on which a joint modeling program is stored. The joint modeling program can be executed by one or more processors to implement the above joint modeling method.

[0042] Compared with the prior art, the present invention first performs public sample ID recognition processing on the local first sample set and the second sample sets of each second participant respectively, and splits the first sample set into first sub-sample sets corresponding to each second participant based on the recognition results; then, based on the initial parameters of the preset model corresponding to each first sub-sample set and the second sample sets of the corresponding second participants, calculates the gradient values corresponding to each first sub-sample set, and determines the first parameters corresponding to each first sub-sample set based on the gradient values; then, receives the second parameters sent by each second participant, and receives the third parameters corresponding to each sub-sample set sent by other first participants; finally, when it is determined that the preset model converges, determines the target parameters based on the first parameters, the second parameters and the third parameters, and sends the target parameters to other participants to complete the joint modeling. The present invention jointly models horizontal participants with the same sample characteristics but different sample objects and vertical participants with the same sample objects but different sample characteristics, improving the accuracy of the joint model. BRIEF DESCRIPTION OF THE DRAWINGS

[0043] Figure 1 It is a schematic flowchart of the joint modeling method provided by an embodiment of the present invention;

[0044] Figure 2 It is a schematic module diagram of the joint modeling device provided by an embodiment of the present invention;

[0045] Figure 3 It is a schematic structural diagram of an electronic device for implementing the joint modeling method provided by an embodiment of the present invention;

[0046] The realization, functional characteristics and advantages of the object of the present invention will be further described with reference to the embodiments and the accompanying drawings. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0047] In order to make the object, technical solution and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention, and are not used to limit the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts fall within the protection scope of the present invention.

[0048] It should be noted that in the present invention, the descriptions involving "first", "second", etc. are only for descriptive purposes, and cannot be construed as indicating or implying their relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one such feature. Additionally, the technical solutions between various embodiments can be combined with each other, but it must be based on the ability of those of ordinary skill in the art to implement. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0049] The present invention provides a joint modeling method, which is applied to any one of the first participants in a joint modeling system. The joint modeling system includes a plurality of first participants and a plurality of second participants that are communicatively connected. Referring to Figure 1 As shown, it is a schematic flowchart of the joint modeling method provided by an embodiment of the present invention. This method can be executed by an electronic device (the electronic device corresponding to the first participant executing this solution), and the electronic device can be implemented by software and / or hardware.

[0050] In this embodiment, each first participant and each second participant contain the same sample objects and different sample features, and each second participant contains the same sample features and different sample objects. The joint modeling method includes:

[0051] S1. Receive the public keys in the homomorphic encryption key pairs sent by each second participant in the joint modeling system. Based on the public keys, perform public sample ID recognition processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant respectively. Based on the public sample ID recognition results, split the first sample set into first sub-sample sets corresponding to each second participant.

[0052] The characteristics of the homomorphic encryption algorithm are as follows: After multiple homomorphically encrypted data are operated using a pre-designed calculation formula to obtain an operation result, the operation result is decrypted, and the decrypted result is the same as the result obtained by operating the unencrypted original data using the same calculation formula. By using the homomorphic encryption algorithm, the normal operation of data can be ensured without leaking the original data.

[0053] In this embodiment, each second participant generates multiple pairs of homomorphic encryption key pairs according to the number of first participants, and sends the public keys in each homomorphic encryption key pair to the corresponding first participant. Among them, the second participant holds the public key and the private key, and the first participant holds the public key. The public key is used for encryption, and the private key is used for decryption. Thus, the first participant cannot decrypt the ciphertext data sent by the second participant, fully ensuring the security of the data.

[0054] In this embodiment, the same sample objects and different sample features are included between the first participating party and the second participating party, and the same sample features and different sample objects are included between the second participating parties. The samples in the second sample sets of each second participating party carry label information.

[0055] For example, assume that the first participating party includes a social security bureau and a shopping platform, and the second participating parties include Bank 1, Bank 2, and Bank 3. Among them, the first sample set of the social security bureau includes sample data with sample IDs from 1 to 10,000, and the sample features corresponding to this sample data include the social security contribution base, the social security contribution company, the social security contribution duration, etc.; the first sample set of the shopping platform includes sample data with sample IDs from 1 to 9,500, and the sample features corresponding to this sample data include the number of shopping times, shopping types, and shopping amounts of the user in the recent six months; the second sample set of Bank 1 includes sample data with sample IDs from 1 to 1,000, and this sample data includes the number of deposit times, borrowing times, and the quantity and types of financial products purchased by the user in the recent six months; the second sample sets of Bank 2 and Bank 3 respectively include sample data with sample IDs from 1,000 to 6,000 and 6,001 to 9,000, and the sample features of the sample data are the same as those of Bank 1.

[0056] If the first participating party implementing this solution is the social security bureau, the social security bureau takes the sample data with sample IDs from 1 to 1,000 (the common sample IDs with Bank 1) in its first sample set as the 1st first sub-sample set, takes the sample data with sample IDs from 1,001 to 6,000 (the common sample IDs with Bank 2) in its first sample set as the 2nd first sub-sample set, and takes the sample data with sample IDs from 6,001 to 9,000 (the common sample IDs with Bank 3) in its first sample set as the 3rd first sub-sample set. Since the remaining sample data in this first sample set will not be used in this modeling process, this part of the remaining data can be transferred from this first sample set to other places for storage.

[0057] For the social security bureau, the corresponding relationship between its first sub-sample sets and the second participating parties is as follows:

[0058] The 1st first sub-sample set corresponds to Bank 1;

[0059] The 2nd first sub-sample set corresponds to Bank 2;

[0060] The 3rd first sub-sample set corresponds to Bank 3.

[0061] Using the above method, other first participating parties also split their local sample sets into sub-sample sets corresponding to each second participating party.

[0062] Performing public sample ID recognition processing on the first sample set without tag information stored locally and the second sample sets with tag information of each second participant based on the public key respectively, includes:

[0063] A11. Select a second participant, calculate the first hash value of each sample ID in the first sample set, encrypt the first hash value using the public key in the homomorphic encryption key pair corresponding to the selected second participant to obtain a first ciphertext, and establish a mapping relationship between the first ciphertext and the sample ID;

[0064] The sample ID can be a user ID, and the user ID can be other information such as the user's mobile phone number, ID card number, etc. that identifies the user's identity.

[0065] By performing hash operation and encryption processing on the sample ID, the security of the sample ID is fully guaranteed.

[0066] A12. Receive the second ciphertext sent by the selected second participant, where the second ciphertext is obtained by the selected second participant encrypting the second hash value of each sample ID in its second sample set using the public key in the same homomorphic encryption key pair;

[0067] The selected second participant also performs hash operation and encryption processing on its second sample set.

[0068] A13. Calculate the intersection of the first ciphertext and the second ciphertext to obtain the public sample ID ciphertext, and determine the plaintext data of the public sample ID ciphertext based on the mapping relationship.

[0069] The above mapping relationship includes the mapping relationship between each sample ID in the first sample set and the corresponding first ciphertext, and the plaintext data of the public sample ID ciphertext can be queried from this mapping relationship.

[0070] The first participant executing this solution also sends the public sample ID ciphertext to the selected second participant for the selected second participant to obtain the public sample ID.

[0071] S2. Obtain the initial parameters of the preset model corresponding to each first sub-sample set, and calculate the gradient value corresponding to each first sub-sample set based on the public key, the initial parameters, and the second sample set of the corresponding second participant.

[0072] In this embodiment, multiple preset models (with the same structure for each preset model) are configured according to the number of first sub-sample sets, and the preset models are initialized to obtain the initial parameters corresponding to each first sub-sample set.

[0073] Assume that the execution entity of this solution is the social security bureau, which configures three preset models in total and generates three initial parameters θ A1 、θ A2, θ A3 , each first sub - sample set corresponds to an initial parameter. In this embodiment, the number of dimensions of the initial parameter is determined according to the feature dimensions of the samples in the sample set. If each sample has 10 features, the initial parameter is a 10 * 1 vector.

[0074] For the shopping platform, it also configures three preset models to generate three initial parameters θ B1 , θ B2 , θ B3 ; for Bank 1, Bank 2, and Bank 3, each only configures one preset model to generate one initial parameter, which are θ C1 , θ C2 , θ C3 .

[0075] In this embodiment, each first participant starts multiple processes according to the number of its sub - sample sets. Each process performs gradient value calculation based on the local sub - sample set and the second sample set of the second participant with the same sample ID. The following describes the calculation process of the gradient value corresponding to each first sub - sample set through the first participant executing this solution.

[0076] Calculating the gradient value corresponding to each first sub - sample set based on the public key, the initial parameter, and the second sample set of the corresponding second participant includes steps B11 - B13:

[0077] B11. Start multiple processes according to the number of first sub - sample sets. Each process calculates the first feature matrix corresponding to each first sub - sample set according to the corresponding first sub - sample set and its initial parameter;

[0078] For example, for the social security bureau, three processes are started, and each process corresponds to processing one first sub - sample set.

[0079] Each process calculating the first feature matrix corresponding to each first sub - sample set according to the corresponding first sub - sample set and its initial parameter includes:

[0080] Select a process, obtain the first sub - sample set and the initial parameter corresponding to this process, determine the initial feature matrix of the obtained first sub - sample set, and calculate the first feature matrix corresponding to the obtained first sub - sample set based on the initial feature matrix and the initial parameter.

[0081] Suppose the selected first sub - sample set is the first sub - sample set A1 of the social security bureau. If there are 1000 samples in A1 and each sample has 10 features, the initial matrix of the obtained first sub - sample set is a 1000 * 10 matrix. Calculate the product of the initial feature matrix and the initial vector to obtain the first feature matrix corresponding to the selected first sub - sample set.

[0082] The calculation formula for the first feature matrix is as follows:

[0083] (θx) A1-1 =θ A1 x A1

[0084] Among them, (θx) A1-1 is the first feature matrix corresponding to the selected first sub-sample set, x A1 is the initial feature matrix corresponding to the selected first sub-sample set, and θ A1 is the initial parameter corresponding to the selected first sub-sample set.

[0085] If θ A1 is a 10*1 vector (a column vector with 10 rows), and x A1 is a 1000*10 matrix, then (θx) A1-1 is a 1000*1 vector.

[0086] B12. Send the first feature matrix to the corresponding second participant, and receive the error value encrypted by the public key sent by the corresponding second participant. The error value is calculated by the corresponding second participant according to the second feature matrix of its second sample set and the first feature matrix;

[0087] Since the first feature matrix is a product and does not disclose the initial feature matrix and initial parameters of the corresponding first sub-sample set, there is no need for encrypted transmission.

[0088] If the second participant corresponding to A1 is Bank 1, the calculation formula for the second feature matrix of the second sample set of Bank 1 is:

[0089] (θx) A1-2 =θ A2 x A2

[0090] Among them, (θx) A1-2 is the second feature matrix of the second sample set of the corresponding second participant, x A2 is the initial feature matrix of the second sample set of the corresponding second participant, and θ A2 is the initial parameter of the preset model of the corresponding second participant.

[0091] The process by which the corresponding second participant calculates the error value according to the second feature matrix of its second sample set and the first feature matrix includes steps C11 - C13:

[0092] C11. The corresponding second participant calculates the eigenvalue of its second sample set based on the second feature matrix of its second sample set and the first feature matrix;

[0093] The calculation formula for the eigenvalue is:

[0094] (θx) A1 =(θx) A1-1 +(θx) A1-2

[0095] where (θx) A1 is the eigenvalue of the second sample set of the corresponding second participating party, and (θx) A1-1 is the first feature matrix of the corresponding first sub-sample set, and (θx) A1-2 is the second feature matrix of the second sample set of the corresponding second participating party.

[0096] C12. Input the eigenvalue into a preset model to obtain the predicted value of its second sample set;

[0097] The calculation formula for the predicted value is:

[0098]

[0099] where h A1 (x) is the predicted value of the second sample set of the corresponding second participating party, and (θx) A1 is the eigenvalue of the second sample set of the corresponding second participating party.

[0100] C13. Determine the true value of its second sample set based on the label information, calculate the error value based on the true value and the predicted value, and send the encrypted error value using the public key in the corresponding homomorphic encryption key pair to the corresponding first participating party.

[0101] In this embodiment, the difference between the true value and the predicted value is used as the error value, and the error value is encrypted using the public key in the homomorphic encryption public and private key pair and sent to the first participating party (i.e., the social security bureau) that executes this solution.

[0102] Since (θx) A1-1 , (θx) A1-2 are both 1000*1 column vectors, then (θx) A1 and h A1 (x) are also 1000*1 column vectors. For h A1 (x), the value of each row in its column vector corresponds to the predicted value of a sample in the second sample set.

[0103] B13. Substitute the encrypted error value into the gradient value calculation formula to obtain the encrypted gradient value corresponding to each first sub-sample set, and send the encrypted gradient value to the corresponding second participating party to obtain the plaintext data of the encrypted gradient value.

[0104] The gradient value calculation formula is:

[0105]

[0106] Among them, is the gradient value corresponding to the i-th first sub-sample set, and x ij is the original value of the j-th sample in the i-th first sub-sample set, and y ij -h θ (x ij ) is the error value between the true value and the predicted value of the j-th sample in the i-th first sub-sample set, and n is the total number of samples in the i-th first sub-sample set.

[0107] In this embodiment, the encrypted error value is substituted into the gradient value calculation formula, and the homomorphic encryption algorithm is used to calculate the encrypted gradient value corresponding to each first sub-sample set.

[0108] The sending the encrypted gradient value to the corresponding second party to obtain the plaintext data of the encrypted gradient value includes:

[0109] D11. Generate a third random number for each second party, encrypt the corresponding third random number using the corresponding public key, calculate the sum of the encrypted gradient value and the encrypted third random number to obtain an encrypted sum, and send the encrypted sum to the corresponding second party;

[0110] According to the corresponding relationship, encrypt the corresponding third random number using the public key in the homomorphic encryption key pair of the corresponding second party.

[0111] D12. Receive the value obtained by the corresponding second party decrypting the encrypted sum, and subtract the corresponding third random number from the obtained value to obtain the decrypted gradient value corresponding to the corresponding first sub-sample set.

[0112] In this embodiment, the purpose of generating the random number is to ensure the security of the data. After the corresponding second party obtains the encrypted sum and decrypts it, what is obtained is the sum of the gradient value and the corresponding third random number. Since it does not know the specific value of the third random number, it cannot know the specific value of the gradient value either.

[0113] S3. Perform parameter update processing on the preset model corresponding to each first sub-sample set based on the gradient value to obtain the first parameter corresponding to each first sub-sample set.

[0114] After obtaining the plaintext gradient value, each process can update the parameters of the preset model corresponding to each first sub-sample to obtain the first parameter corresponding to each first sub-sample set.

[0115] S4. Receive the second parameters and loss values corresponding to their second sample sets processed by the secure aggregation algorithm sent by each second participant, and receive the third parameters corresponding to their respective sub-sample sets processed by the secure aggregation algorithm sent by other first participants.

[0116] Each second participant inputs the predicted value of each sample in its second sample set and the true value in the label information into the gradient value calculation formula, and the corresponding gradient value can be obtained. According to this gradient value, the parameters of its local preset model can be updated to obtain the second parameters. At the same time, by inputting the predicted value and the true value into the loss function, the loss value can be obtained.

[0117] The calculation formula for the loss value is:

[0118]

[0119] where L i is the loss value corresponding to the i-th second participant, y ij is the true value of the j-th sample in the second sample set of the i-th second participant, h θ (x ij ) is the predicted value of the j-th sample in the second sample set of the i-th second participant, and n is the total number of samples in the second sample set of the i-th second participant.

[0120] Other first participants can obtain the third parameters corresponding to each sub-sample set according to the content described in steps S1 - S4, process the third parameters using the secure aggregation algorithm, and send them to the first participant implementing this solution.

[0121] The secure aggregation algorithm is a technology among different participants that generates addition and subtraction random numbers through the Difie - Hellmanm key exchange technology, thereby achieving the technology of hiding plaintext information from a third party. The third party will cancel the random numbers during aggregation, which does not affect the operation result. That is: other participants add or subtract random numbers generated based on the Difie - Hellmanm key exchange technology in the data they send, and the first participant implementing this solution will cancel all the random numbers during aggregation (i.e., summation).

[0122] S5. Based on the loss value, determine whether the preset model converges. When the determination is yes, determine the target parameters based on the first parameter, the second parameter, and the third parameter, and send the target parameters to other participants in the joint modeling system respectively to complete the joint modeling.

[0123] In this embodiment, the average value of the loss values of each second participant is calculated to obtain the average loss value (since the sum is calculated first and then averaged during the calculation of the average value, and the random numbers are cancelled out during the summation, the obtained average loss value is accurate). If the difference between the obtained average loss value and the average loss value of the previous iteration is less than the preset threshold (for example, 0.01%), it is determined that the preset model has converged.

[0124] If it is determined that the preset model has converged, the average values of the first parameter, the second parameter, and the third parameter are calculated (also by summing first and then averaging, and the random numbers are cancelled out during the summation), to obtain the target parameter, and the target parameter is distributed to other participants in the joint modeling system, thus completing the joint modeling.

[0125] In this embodiment, step S5 can also be executed by a third party, which can be a public server or a cloud server that is communicatively connected to each participant in the joint modeling system.

[0126] As can be seen from the above embodiments, for the joint modeling method proposed by the present invention, first, the public sample ID recognition process is respectively performed on the local first sample set and the second sample sets of each second participant, and based on the recognition results, the first sample set is split into the first sub-sample sets corresponding to each second participant; then, based on the initial parameters of the preset model corresponding to each first sub-sample set and the second sample sets of the corresponding second participants, the gradient values corresponding to each first sub-sample set are calculated, and based on the gradient values, the first parameters corresponding to each first sub-sample set are determined; then, the second parameters sent by each second participant are received, and the third parameters corresponding to each sub-sample set sent by other first participants are received; finally, when it is determined that the preset model has converged, the target parameter is determined based on the first parameter, the second parameter, and the third parameter, and the target parameter is sent to other participants to complete the joint modeling. The present invention jointly models horizontal participants with the same sample characteristics but different sample objects and vertical participants with the same sample objects but different sample characteristics, improving the accuracy of the joint model.

[0127] As Figure 2 shown, it is a schematic diagram of the modules of a joint modeling device provided by an embodiment of the present invention.

[0128] The joint modeling device 100 according to the present invention can be installed in an electronic device. According to the functions to be realized, the joint modeling device 100 can include a receiving module 110, a calculating module 120, an updating module 130, a receiving module 140, and a determining module 150. The modules in the present invention can also be referred to as units, which refer to a series of computer program segments that can be executed by a processor of an electronic device and can complete fixed functions, and are stored in the memory of the electronic device.

[0129] In this embodiment, the functions of each module / unit are as follows:

[0130] A receiving module 110, configured to receive public keys in the homomorphic encryption key pairs sent by each second participant in the joint modeling system, perform public sample ID recognition processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant based on the public keys, and split the first sample set into first sub-sample sets corresponding to each second participant based on the public sample ID recognition results.

[0131] The performing public sample ID recognition processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant based on the public keys includes:

[0132] A21. Select a second participant, calculate first hash values of sample IDs in the first sample set, encrypt the first hash values using the public key in the homomorphic encryption key pair corresponding to the selected second participant to obtain first ciphertexts, and establish a mapping relationship between the first ciphertexts and the sample IDs;

[0133] A22. Receive second ciphertexts sent by the selected second participant, where the second ciphertexts are obtained by the selected second participant encrypting second hash values of sample IDs in its second sample set using the public key in the same homomorphic encryption key pair;

[0134] A23. Calculate the intersection of the first ciphertexts and the second ciphertexts to obtain public sample ID ciphertexts, and determine the plaintext data of the public sample ID ciphertexts based on the mapping relationship.

[0135] A calculation module 120, configured to obtain initial parameters of a preset model corresponding to each first sub-sample set, and calculate gradient values corresponding to each first sub-sample set based on the public keys, the initial parameters, and the second sample sets of the corresponding second participants.

[0136] The calculating gradient values corresponding to each first sub-sample set based on the public keys, the initial parameters, and the second sample sets of the corresponding second participants includes steps B21 - B23:

[0137] B21. Start multiple processes according to the number of first sub-sample sets, and each process calculates a first feature matrix corresponding to each first sub-sample set according to the corresponding first sub-sample set and its initial parameters;

[0138] The calculating, by each process, a first feature matrix corresponding to each first sub-sample set according to the corresponding first sub-sample set and its initial parameters includes:

[0139] Select a process, obtain the corresponding first subset of samples and initial parameters for the process, determine the initial feature matrix of the obtained first subset of samples, and calculate the first feature matrix corresponding to the obtained first subset of samples based on the initial feature matrix and the initial parameters.

[0140] B22. Send the first feature matrix to the corresponding second party, and receive the error value encrypted with the public key sent by the corresponding second party. The error value is calculated by the corresponding second party according to the second feature matrix of its second sample set and the first feature matrix.

[0141] The process by which the corresponding second party calculates the error value according to the second feature matrix of its second sample set and the first feature matrix includes steps C21 - C23:

[0142] C21. The corresponding second party calculates the eigenvalue of its second sample set based on the second feature matrix of its second sample set and the first feature matrix.

[0143] C22. Input the eigenvalue into a preset model to obtain the predicted value of its second sample set.

[0144] C23. Determine the true value of its second sample set based on the label information, calculate the error value based on the true value and the predicted value, and encrypt the error value with the public key in the corresponding homomorphic encryption key pair and send it to the corresponding first party.

[0145] B23. Substitute the encrypted error value into the gradient value calculation formula to obtain the encrypted gradient value corresponding to each first subset of samples, and send the encrypted gradient value to the corresponding second party to obtain the plaintext data of the encrypted gradient value.

[0146] The sending the encrypted gradient value to the corresponding second party to obtain the plaintext data of the encrypted gradient value includes:

[0147] D21. Generate a third random number for each second party, encrypt the corresponding third random number with the corresponding public key, calculate the sum of the encrypted gradient value and the encrypted third random number to obtain an encrypted sum, and send the encrypted sum to the corresponding second party.

[0148] D22. Receive the value obtained by the corresponding second party decrypting the encrypted sum, and subtract the corresponding third random number from the obtained value to obtain the decrypted gradient value corresponding to the corresponding first subset of samples.

[0149] The update module 130 is used to perform parameter update processing on the preset model corresponding to each first subset of samples based on the gradient value to obtain the first parameter corresponding to each first subset of samples.

[0150] A receiving module 140, configured to receive the second parameters and loss values processed by the secure aggregation algorithm corresponding to its second sample set sent by each second participant, and receive the third parameters processed by the secure aggregation algorithm corresponding to each sub-sample set sent by other first participants.

[0151] The calculation formula of the loss value is:

[0152]

[0153] where L i is the loss value corresponding to the i-th second participant, y ij is the true value of the j-th sample in the second sample set of the i-th second participant, h θ (x ij ) is the predicted value of the j-th sample in the second sample set of the i-th second participant, and n is the total number of samples in the second sample set of the i-th second participant.

[0154] A determination module 150, configured to determine whether the preset model converges based on the loss value. When the determination result is yes, determine target parameters based on the first parameter, the second parameter, and the third parameter, and send the target parameters to other participants in the joint modeling system respectively to complete joint modeling.

[0155] As Figure 3 shown, it is a schematic structural diagram of an electronic device for implementing the joint modeling method provided by an embodiment of the present invention.

[0156] The electronic device 1 is a device capable of automatically performing numerical calculations and / or information processing according to pre-set or stored instructions. The electronic device 1 can be a computer, a single network server, a server group composed of multiple network servers, or a cloud composed of a large number of hosts or network servers based on cloud computing, where cloud computing is a type of distributed computing and consists of a group of loosely coupled computer sets forming a super virtual computer.

[0157] In this embodiment, the electronic device 1 includes, but is not limited to, a memory 11, a processor 12, and a network interface 13 that can communicate with each other through a system bus. The memory 11 stores a joint modeling program 10, and the joint modeling program 10 can be executed by the processor 12. Figure 3 Only the electronic device 1 with components 11-13 and the joint modeling program 10 is shown. Those skilled in the art can understand that Figure 3 the shown structure does not limit the electronic device 1, and it may include fewer or more components than shown, or combine some components, or have different component arrangements.

[0158] Among them, the memory 11 includes a memory and at least one type of readable storage medium. The memory provides a cache for the operation of the electronic device 1; the readable storage medium can be a non-volatile storage medium such as flash memory, hard disk, multimedia card, card-type memory (for example, SD or DX memory, etc.), random access memory (RAM), static random access memory (SRAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), programmable read-only memory (PROM), magnetic memory, magnetic disk, optical disk, etc. In some embodiments, the readable storage medium can be an internal storage unit of the electronic device 1, such as the hard disk of the electronic device 1; in other embodiments, the non-volatile storage medium can also be an external storage device of the electronic device 1, such as a plug-in hard disk equipped on the electronic device 1, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc. In this embodiment, the readable storage medium of the memory 11 is generally used to store the operating system and various application software installed in the electronic device 1, such as storing the code of the joint modeling program 10 in an embodiment of the present invention. In addition, the memory 11 can also be used to temporarily store various data that have been output or will be output.

[0159] In some embodiments, the processor 12 can be a Central Processing Unit (CPU), a controller, a microcontroller, a microprocessor, or other data processing chips. The processor 12 is generally used to control the overall operation of the electronic device 1, such as performing control and processing related to data interaction or communication with other devices. In this embodiment, the processor 12 is used to run the program code stored in the memory 11 or process data, such as running the joint modeling program 10, etc.

[0160] The network interface 13 can include a wireless network interface or a wired network interface, and this network interface 13 is used to establish a communication connection between the electronic device 1 and a client (not shown in the figure).

[0161] Optionally, the electronic device 1 may further include a user interface, which may include a display, an input unit such as a keyboard. Optionally, the user interface may further include a standard wired interface and a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch liquid crystal display, and an OLED (Organic Light-Emitting Diode) toucher, etc. Among them, the display may also be appropriately referred to as a display screen or a display unit, and is used to display the information processed in the electronic device 1 and to display a visual user interface.

[0162] It should be understood that the above embodiments are only for illustration purposes and are not limited by this structure in the scope of the patent application.

[0163] The joint modeling program 10 stored in the memory 11 of the electronic device 1 is a combination of multiple instructions, and when running in the processor 12, it can implement the steps in the above joint modeling method.

[0164] Specifically, for the specific implementation method of the above joint modeling program 10 by the processor 12, reference can be made to Figure 1 the description of the relevant steps in the corresponding embodiments, which will not be elaborated here.

[0165] Furthermore, if the modules / units integrated in the electronic device 1 are implemented in the form of software function units and sold or used as independent products, they can be stored in a computer-readable storage medium. The computer-readable medium can be non-volatile or non-volatile. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard disk, a magnetic disk, an optical disc, a computer memory, a read-only memory (ROM, Read-Only Memory).

[0166] The joint modeling program 10 is stored on the computer-readable storage medium, and the joint modeling program 10 can be executed by one or more processors to implement the steps in the above joint modeling method.

[0167] In several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are only illustrative. For example, the division of the modules is only a logical function division, and there may be other division methods in actual implementation.

[0168] The module described as a separation component may or may not be physically separated. The component shown as a module may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0169] In addition, in each embodiment of the present invention, each functional module may be integrated in a processing unit, or each unit may exist physically alone, or two or more units may be integrated in one unit. The above integrated unit can be implemented in the form of hardware, or in the form of a combination of hardware and software functional modules.

[0170] Therefore, from any point of view, the embodiments should be regarded as exemplary and non-limiting. The scope of the present invention is defined by the appended claims rather than the above description. Therefore, all changes falling within the meaning and scope of the equivalent elements of the claims are intended to be embraced within the present invention. Any reference signs in the claims should not be construed as limiting the claims involved.

[0171] The blockchain referred to in the present invention is a new application mode of computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanism, and encryption algorithms. Blockchain, essentially a decentralized database, is a series of data blocks generated by using cryptographic methods. Each data block contains information about a batch of network transactions, which is used to verify the validity (anti-counterfeiting) of the information and generate the next block. The blockchain can include a blockchain underlying platform, a platform product service layer, an application service layer, etc.

[0172] In addition, obviously, the word "including" does not exclude other units or steps, and the singular does not exclude the plural. The multiple units or devices stated in the system claims can also be implemented by one unit or device through software or hardware. Words such as "second" are used to denote names and do not denote any particular order.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention.

Claims

1. A joint modeling method is applied to any one of the first participants in a joint modeling system. The joint modeling system includes a plurality of first participants and a plurality of second participants that are communicatively connected. Characterized in that, There are the same sample objects and different sample features between each first participant and each second participant, and there are the same sample features and different sample objects between each second participant. The method includes: Receiving the public keys in the homomorphic encryption key pairs sent by each second participant in the joint modeling system, and respectively performing public sample ID identification processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant based on the public keys. Based on the public sample ID identification results, splitting the first sample set into first sub-sample sets corresponding to each second participant; Obtaining the initial parameters of the preset model corresponding to each first sub-sample set, and calculating the gradient value corresponding to each first sub-sample set based on the public key, the initial parameters, and the second sample set of the corresponding second participant; Performing parameter update processing on the preset model corresponding to each first sub-sample set based on the gradient value to obtain the first parameters corresponding to each first sub-sample set; Receiving the second parameters and loss values after being processed by the secure aggregation algorithm corresponding to the second sample sets sent by each second participant, and receiving the third parameters after being processed by the secure aggregation algorithm corresponding to each sub-sample set sent by other first participants; Judging whether the preset model converges based on the loss value. When the judgment is yes, determining the target parameters based on the first parameters, the second parameters, and the third parameters, and sending the target parameters to other participants in the joint modeling system respectively to complete the joint modeling.

2. The joint modeling method according to claim 1, Characterized in that, The performing public sample ID identification processing on the first sample set without label information stored locally and the second sample sets with label information of each second participant based on the public key respectively includes: Selecting a second participant, calculating the first hash values of each sample ID in the first sample set, encrypting the first hash values using the public key in the homomorphic encryption key pair corresponding to the selected second participant to obtain first ciphertexts, and establishing a mapping relationship between the first ciphertexts and the sample IDs; Receiving the second ciphertexts sent by the selected second participant. The second ciphertexts are the second hash values of each sample ID in its second sample set encrypted by the public key in the same homomorphic encryption key pair by the selected second participant; Calculating the intersection of the first ciphertexts and the second ciphertexts to obtain the public sample ID ciphertexts, and determining the plaintext data of the public sample ID ciphertexts based on the mapping relationship.

3. The joint modeling method according to claim 1, Characterized in that, The calculating the gradient value corresponding to each first sub-sample set based on the public key, the initial parameters, and the second sample set of the corresponding second participant includes: Starting multiple processes according to the number of the first sub-sample sets. Each process calculates the first feature matrix corresponding to each first sub-sample set according to the corresponding first sub-sample set and its initial parameters; Send the first feature matrix to the corresponding second participant, and receive the error value encrypted with the public key sent by the corresponding second participant. The error value is calculated by the corresponding second participant according to the second feature matrix of its second sample set and the first feature matrix; Substitute the encrypted error value into the gradient value calculation formula to obtain the encrypted gradient value corresponding to each first sub-sample set, and send the encrypted gradient value to the corresponding second participant to obtain the plaintext data of the encrypted gradient value.

4. The joint modeling method according to claim 3, characterized in that The process by which the corresponding second participant calculates the error value according to the second feature matrix of its second sample set and the first feature matrix includes: The corresponding second participant calculates the eigenvalue of its second sample set based on the second feature matrix of its second sample set and the first feature matrix; Input the eigenvalue into a preset model to obtain the predicted value of its second sample set; Determine the true value of its second sample set based on the label information, calculate the error value based on the true value and the predicted value, and encrypt the error value with the public key in the corresponding homomorphic encryption key pair and send it to the corresponding first participant.

5. The joint modeling method according to claim 3, characterized in that The step of sending the encrypted gradient value to the corresponding second participant to obtain the plaintext data of the encrypted gradient value includes: Generate a third random number for each second participant, encrypt the corresponding third random number with the corresponding public key, calculate the sum of the encrypted gradient value and the encrypted third random number to obtain an encrypted sum, and send the encrypted sum to the corresponding second participant; Receive the value obtained by the corresponding second participant decrypting the encrypted sum, and subtract the corresponding third random number from the obtained value to obtain the decrypted gradient value corresponding to the corresponding first sub-sample set.

6. The joint modeling method according to claim 3, characterized in that Each process calculates the first feature matrix corresponding to each first sub-sample set according to the corresponding first sub-sample set and its initial parameters, including: Select a process, obtain the first sub-sample set and the initial parameters corresponding to the process, determine the initial feature matrix of the obtained first sub-sample set, and calculate the first feature matrix corresponding to the obtained first sub-sample set based on the initial feature matrix and the initial parameters.

7. The joint modeling method according to claim 1, characterized in that The calculation formula of the loss value is: Among them, L i is the loss value corresponding to the i-th second participant, y ij is the true value of the j-th sample in the second sample set of the i-th second participant, h θ (x ij ) is the predicted value of the j-th sample in the second sample set of the i-th second participant, and n is the total number of samples in the second sample set of the i-th second participant.

8. A joint modeling device, characterized in that The device includes: A receiving module, configured to receive the public key in the homomorphic encryption key pair sent by each second participant in the joint modeling system, perform public sample ID recognition processing on the first sample set without label information stored locally and the second sample set with label information of each second participant based on the public key, and split the first sample set into the first sub-sample sets corresponding to each second participant based on the public sample ID recognition result; A calculation module, configured to obtain the initial parameters of a preset model corresponding to each first sub-sample set, and calculate the gradient value corresponding to each first sub-sample set based on the public key, the initial parameters, and the second sample set of the corresponding second participant; An update module, configured to perform parameter update processing on the preset model corresponding to each first sub-sample set based on the gradient value to obtain the first parameter corresponding to each first sub-sample set; A receiving module, configured to receive the second parameter and the loss value corresponding to its second sample set processed by the secure aggregation algorithm sent by each second participant, and receive the third parameter processed by the secure aggregation algorithm corresponding to each sub-sample set of other first participants sent by other first participants; A determination module, configured to determine whether the preset model converges based on the loss value. When the determination result is yes, determine the target parameter based on the first parameter, the second parameter, and the third parameter, and send the target parameter to other participants in the joint modeling system respectively to complete the joint modeling.

9. An electronic device, characterized in that, the electronic device includes: at least one processor; and, a memory communicatively connected to the at least one processor; wherein, the memory stores a joint modeling program executable by the at least one processor, and the joint modeling program is executed by the at least one processor so that the at least one processor can execute the joint modeling method according to any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, the computer-readable storage medium stores a joint modeling program, and the joint modeling program can be executed by one or more processors to implement the joint modeling method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Decision model training method, prediction method and device based on longitudinal federation learning

    CN111598186A

  • Longitudinal federation learning modeling method and system, medium and equipment

    CN112241537A