A method and device for data processing in federated learning

Through the method of grouping and marginal contribution degree calculation, the problem of low efficiency of contribution measurement in federated learning is solved, and more efficient and accurate data quality evaluation is achieved.

CN114021735BActive Publication Date: 2025-07-18CHINA UNIONPAY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111231870.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-22
Publication Date
2025-07-18
Estimated Expiration
2041-10-22

AI Technical Summary

Technical Problem

The existing federated learning incentive mechanism quantifies the contribution by calculating the approximate Shapril value of each participant, resulting in large time overhead and affecting the efficiency of data quality evaluation.

Method used

By receiving encrypted data from multiple participants, grouping according to data similarity and selecting representatives of the participating group, the sub-data contribution of each round of training is calculated based on the marginal contribution degree, and the data contribution of each participant is finally determined.

Benefits of technology

It reduces time overhead, improves data contribution detection efficiency, and improves the accuracy and accuracy of contribution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114021735B_ABST
    Figure CN114021735B_ABST
Patent Text Reader

Abstract

The embodiments of the present application provide a method and device for data processing in federated learning, which relate to the field of artificial intelligence technology. The method includes: receiving encrypted data sent by multiple participants, grouping the multiple participants according to the similarity of the encrypted data to obtain multiple first participant groups, and selecting one participant from each first participant group as the representative of the participant group. Then, based on the marginal contribution of the target participant group representative who participates in the federated training in each round among the multiple participant group representatives, the sub-data contribution degree of the target participant group representative is obtained. Based on the sub-data contribution degrees obtained after multiple rounds of federated training, the data contribution degrees corresponding to the multiple participant group representatives are obtained. Finally, according to the data contribution degrees corresponding to the multiple participant group representatives, the data contribution degrees corresponding to each of the multiple participants are determined. Since one participant is selected from each first participant group as the representative of the participant group to participate in the federated training, the detection efficiency of the data contribution degree is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of artificial intelligence technology, and in particular, to a method and device for data processing in federated learning. Background Art

[0002] In the scenario of the Internet of Everything, the data association among different institutions and departments will form a huge data alliance. Against this background, federated learning, as a solution to solve the data circulation bottleneck, has attracted wide attention. Federated learning realizes joint modeling by all participating parties on the basis of ensuring data privacy security and legal compliance, so as to improve the effect of the federated learning model.

[0003] When all participating parties participate in federated learning, it will inevitably consume the device resources of all participating parties. Without sufficient rewards, all participating parties may not be willing to participate in the training process of the model in federated learning. The federated learning incentive mechanism can quantify the contribution degree of each participating party to the effect of the federated learning model according to the data provided by each participating party, and reward each participating party according to the contribution degree of each participating party. Therefore, the federated learning incentive mechanism, as an important link in federated learning, can not only ensure the stable operation of the federated learning system, but also encourage more data owners to participate in federated learning, and finally improve the effect of the federated learning model.

[0004] The existing federated learning incentive mechanism mainly quantifies the contribution degree of each participating party based on the approximate Shapley value calculated for each participating party. Since calculating the approximate Shapley value of each participating party brings a large time overhead, it affects the efficiency of data quality evaluation. Summary of the Invention

[0005] The embodiments of the present application provide a method, device, device and storage medium for data processing in federated learning, which is used to improve the efficiency of data quality evaluation.

[0006] On the one hand, the embodiments of the present application provide a method for data processing in federated learning, and the method includes:

[0007] Receiving encrypted data sent by multiple participating parties;

[0008] Grouping the multiple participating parties according to the similarity of the obtained encrypted data to obtain multiple first participating groups, and selecting one participating party from each first participating group as the representative of the participating group;

[0009] Obtaining the sub-data contribution degree of the target participating group representative participating in the federated training in each round based on the marginal contribution of the target participating group representative participating in the federated training in each round among the multiple participating group representatives;

[0010] Obtain the data contribution degrees corresponding to each of the multiple participating group representatives based on the sub-data contribution degrees obtained after multiple rounds of federated training;

[0011] Determine the data contribution degrees corresponding to each of the multiple participating parties based on the data contribution degrees corresponding to each of the multiple participating group representatives.

[0012] On the one hand, an embodiment of the present application provides a device for data processing in federated learning, and the device includes:

[0013] A receiving module, configured to receive encrypted data sent by multiple participating parties;

[0014] A grouping module, configured to group the multiple participating parties according to the similarity of the obtained encrypted data, obtain multiple first participating groups, and select one participating party from each first participating group as a participating group representative;

[0015] A sub-data contribution degree obtaining module, configured to obtain the sub-data contribution degrees of the target participating group representatives in each round of participating in federated training based on the marginal contributions of the target participating group representatives in each round of participating in federated training among multiple participating group representatives;

[0016] A data contribution degree obtaining module, configured to obtain the data contribution degrees corresponding to each of the multiple participating group representatives based on the sub-data contribution degrees obtained after multiple rounds of federated training;

[0017] The data contribution degree obtaining module is further configured to determine the data contribution degrees corresponding to each of the multiple participating parties based on the data contribution degrees corresponding to each of the multiple participating group representatives.

[0018] Optionally, the grouping module is specifically configured to:

[0019] Determine the similarity between the encrypted data sent by any two participating parties among the multiple participating parties;

[0020] Divide any two participating parties that satisfy the similarity being greater than the similarity threshold into one first participating group, and obtain multiple first participating groups.

[0021] Optionally, the sub-data contribution degree obtaining module is specifically configured to:

[0022] For the target participating group representatives in each round of participating in federated training, respectively perform the following steps:

[0023] Obtain at least one joining order of multiple target participating group representatives joining the federated training;

[0024] During the federated training process, determine the respective marginal contributions of the multiple target participating group representatives under each joining order;

[0025] The marginal contributions of each obtained representative of the target participating group are weighted and averaged to obtain the sub-data contribution degree of each representative of the target participating group.

[0026] Optionally, the sub-data contribution degree obtaining module is further configured to:

[0027] For the first representative of the target participating group located at the first position in the joining order, use the first model improvement value corresponding to the first representative of the target participating group as the marginal contribution of the first representative of the target participating group;

[0028] For the second representative of the target participating group located at other positions in the joining order, determine at least one representative of the target participating group before the second representative of the target participating group;

[0029] Determine the second model improvement value obtained when the second representative of the target participating group and the at least one representative of the target participating group participate in the federated training, and the third model improvement value obtained when the at least one representative of the target participating group participates in the federated training;

[0030] Use the difference between the second model improvement value and the third model improvement value as the marginal contribution of the second representative of the target participating group.

[0031] Optionally, the sub-data contribution degree obtaining module is further configured to:

[0032] The multiple representatives of the target participating groups in the first round of participating in the federated training are obtained from the multiple first participating groups;

[0033] The M representatives of the target participating groups in the T-th round of participating in the federated training are determined according to the sub-data contribution degrees of the N representatives of the target participating groups in the (T - 1)-th round of participating in the federated training, where T > 1 and N >= M.

[0034] Optionally, the sub-data contribution degree obtaining module is further configured to:

[0035] Re-group the N representatives of the target participating groups according to the similarity of the sub-data contribution degrees of the N representatives of the target participating groups to obtain multiple second participating groups, and obtain M representatives of the target participating groups from the multiple second participating groups.

[0036] Optionally, the sub-data contribution degree obtaining module is further configured to:

[0037] Sort the N representatives of the target participating groups in descending order according to the sub-data contribution degrees;

[0038] According to the sorting result, obtain the representatives of the target participating groups whose sub-data contribution degrees are greater than or equal to the contribution degree threshold from the N representatives of the target participating groups;

[0039] Group the obtained representative of each target participating group according to the similarity of sub - data contribution degrees to obtain multiple second participating groups.

[0040] Optionally, the data contribution degree obtaining module is further configured to:

[0041] For the multiple representatives of the participating groups, respectively perform the following steps:

[0042] After multiple rounds of training, perform weighted average on multiple sub - data contribution degrees corresponding to a representative of a participating group to obtain the data contribution degree corresponding to the representative of the participating group.

[0043] Optionally, the encrypted data is obtained by the participating party encrypting the original data through a locality - sensitive hashing algorithm, and the data contribution degree is the Shapley value.

[0044] On the one hand, an embodiment of the present application provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the steps of the method for data processing in the above - mentioned federated learning are implemented.

[0045] On the one hand, an embodiment of the present application provides a computer - readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is enabled to execute the steps of the method for data processing in the above - mentioned federated learning.

[0046] In the embodiment of the present application, since one participating party is selected from each first participating group as the representative of the participating group to participate in the federated training instead of all participating parties participating in the federated training, the time overhead is greatly reduced and the data contribution degree detection efficiency is improved. At the same time, considering that the sub - data contribution degrees corresponding to the representatives of each participating group are also different during different rounds of federated training, therefore, the data contribution degree corresponding to each representative of the participating group is determined based on the sub - data contribution degrees obtained from multiple rounds of federated training, and the accuracy of the data contribution degree of each representative of the participating group is improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0048] Figure 1 It is a schematic diagram of a system architecture provided by an embodiment of the present application;

[0049] Figure 2Schematic flowchart of a method for data processing in federated learning provided by an embodiment of this application;

[0050] Figure 3 Schematic flowchart of a method for data processing in federated learning provided by an embodiment of this application;

[0051] Figure 4 Schematic structural diagram of a device for data processing in federated learning provided by an embodiment of this application;

[0052] Figure 5 Schematic structural diagram of a computer device provided by an embodiment of this application. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and beneficial effects of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] For the convenience of understanding, the nouns involved in the embodiments of the present invention will be explained below.

[0055] Federated learning: A machine learning framework that can effectively help multiple institutions to perform data usage and machine learning model building while meeting the requirements of user privacy protection, data security, and government regulations. Federated learning includes three categories: horizontal federated learning, vertical federated learning, and federated transfer learning. The method for data processing in federated learning provided by the embodiments of this application can be applied to horizontal federated scenarios, and can also be adapted to vertical federated learning and federated transfer learning scenarios.

[0056] Locality Sensitive Hashing (LSH): A fast nearest neighbor search algorithm for massive high-dimensional data.

[0057] Shapely Value (SV): Mainly used to solve the problem of interest distribution among all parties in a cooperative game, reflecting the contribution degree of each cooperative party to the overall cooperative goal and avoiding egalitarianism in distribution.

[0058] Reference Figure 1 , which is a system architecture diagram applicable to the embodiments of this application. The system architecture at least includes Party 101~1 to Party 101~X, and a data contribution degree detection system 102, where X is an integer greater than 1.

[0059] Parties 101~1 to Parties 101~X respectively provide encrypted data to the data contribution degree detection system 102. Each party can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0060] The data contribution degree detection system 102 provides data processing services. The data contribution degree detection system 102 can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms.

[0061] Parties 101~1 to Parties 101~X are respectively connected to the data contribution degree detection system 102, and can be directly or indirectly connected through wired or wireless communication methods, which are not limited in this application.

[0062] Each party respectively sends encrypted data to the data contribution degree detection system 102. The data contribution degree detection system 102 receives the encrypted data sent by multiple parties, groups the multiple parties according to the similarity of the obtained encrypted data to obtain multiple first participating groups, and selects one party from each first participating group as the participating group representative. The data contribution degree detection system 102 obtains the sub-data contribution degree of the target participating group representative in each round of participating in federated training based on the marginal contribution of the target participating group representative in each round of participating in federated training among multiple participating group representatives. Then, based on the sub-data contribution degrees obtained after multiple rounds of federated training, the data contribution degrees corresponding to each of the multiple participating group representatives are obtained. Finally, the data contribution degree detection system 102 determines the data contribution degrees corresponding to each of the multiple parties according to the data contribution degrees corresponding to each of the multiple participating group representatives.

[0063] Based on Figure 1 the system architecture diagram described above, the embodiments of this application provide a process of a data processing method, as Figure 2 shown. The process of this method is executed by a computer device, and this computer device can be Figure 1 the data contribution degree detection system 102 shown above, and includes the following steps:

[0064] Step S201: Receive the encrypted data sent by multiple participants.

[0065] Specifically, the encrypted data is obtained by the participants encrypting the original data through the locality - sensitive hashing algorithm.

[0066] Step S202: Group the multiple participants according to the similarity of the obtained encrypted data to obtain multiple first - participant groups, and select one participant from each first - participant group as the representative of the participant group.

[0067] Specifically, determine the similarity between the encrypted data sent by any two participants among the multiple participants, and then divide any two participants whose similarity is greater than the similarity threshold into a first - participant group, and finally obtain multiple first - participant groups.

[0068] For example, the participants include Participant 1, Participant 2, and Participant 3. The original data corresponding to each participant is Data 1, Data 2, and Data 3 respectively. Each participant encrypts its respective original data through the locality - sensitive hashing algorithm to obtain encrypted data Encrypted Data 1, Encrypted Data 2, and Encrypted Data 3 respectively.

[0069] Pairwise combine the above three participants to form three pairs of participants. Pair 1 includes Encrypted Data 1 and Encrypted Data 2, Pair 2 includes Encrypted Data 1 and Encrypted Data 3, and Pair 3 includes Encrypted Data 2 and Encrypted Data 3. Calculate the similarity between the two encrypted data in each pair of participants respectively, as shown in Table 1.

[0070] Table 1.

[0071] Involved in the Similarity Involved pair 1 (encrypted data 1 and encrypted data 2) 0.6 Involved pair 2 (encrypted data 1 and encrypted data 3) 0.7 Involved pair 3 (encrypted data 2 and encrypted data 3) 0.8

[0072] Set the similarity threshold to 0.5. The similarity between Encrypted Data 1 and Encrypted Data 2 is 0.6, which is greater than the similarity threshold 0.5. Therefore, Encrypted Data 1 and Encrypted Data 2 are divided into a first - participant group. The similarity between Encrypted Data 1 and Encrypted Data 3 is 0.7, which is greater than the similarity threshold 0.5. Therefore, Encrypted Data 1 and Encrypted Data 3 are divided into a first - participant group. The similarity between Encrypted Data 2 and Encrypted Data 3 is 0.8, which is greater than the similarity threshold 0.5. Therefore, Encrypted Data 2 and Encrypted Data 3 are divided into a first - participant group.

[0073] Therefore, Encrypted Data 1, Encrypted Data 2, and Encrypted Data 3 are divided into the same first - participant group.

[0074] Step S203: Based on the marginal contribution of the target participant group representative participating in each round of federated training among the multiple participant group representatives, obtain the sub - data contribution degree of the target participant group representative participating in each round of federated training.

[0075] Specifically, in the K-th (where K > 0) round of federated training, the marginal contribution of the target participating group representative i is obtained using the following formula (1):

[0076] δ i (s) = v(s ∪ {i}) - v(s)…………………(1)

[0077] Where v(s) represents the contribution of the set s in federated training, s ∪ {i} represents the addition of the target participating group representative i to the set s, and v(s ∪ {i}) represents the contribution of the set s ∪ {i} in federated training. δ i (s) represents the marginal contribution of the target participating group representative i in federated training after the target participating group representative i joins the set s.

[0078] Since there are multiple types of sets s, in the K-th round of federated training, the target participating group representative i corresponds to multiple marginal contributions. The multiple marginal contributions corresponding to the target participating group representative i are weighted and averaged to determine the sub-data contribution degree of the target participating group representative i in the K-th round of federated training.

[0079] In the above method, the sub-data contribution degree of the target participating group representative obtained by weighted averaging the multiple marginal contributions corresponding to the target participating group representative is the Shapley value, where different learning rates can be set for multiple rounds of federated training.

[0080] Step S204, based on the sub-data contribution degrees obtained after multiple rounds of federated training, obtain the data contribution degrees corresponding to each of the multiple participating group representatives.

[0081] Optionally, for the multiple participating group representatives, the following steps are respectively executed:

[0082] After multiple rounds of training, the multiple sub-data contribution degrees corresponding to a participating group representative are weighted and averaged to obtain the data contribution degree corresponding to a participating group representative.

[0083] Specifically, in one possible implementation, as the number of training times increases, the weight of the sub-data contribution degree in the later round is greater than the weight of the sub-data contribution degree in the previous round.

[0084] In one possible implementation, as the number of training times increases, the weight of the sub-data contribution degree in the later round is smaller than the weight of the sub-data contribution degree in the previous round.

[0085] In one possible implementation, the weights of the sub-data contribution degrees in each round are equal.

[0086] Step S205, based on the data contribution degrees corresponding to each of the multiple participating group representatives, determine the data contribution degrees corresponding to each of the multiple participating parties.

[0087] Specifically, the following steps are respectively performed for multiple participating group representatives:

[0088] Use the data contribution degree of one participating group representative as the data contribution degree of other participants in the corresponding first participating group.

[0089] In the embodiments of the present application, since one participant is selected from each first participating group as the participating group representative to participate in the federated training, rather than all participants participating in the federated training, the time overhead is greatly reduced and the data contribution degree detection efficiency is improved. At the same time, considering that the sub-data contribution degrees corresponding to each participating group representative are also different in different rounds of federated training, therefore, the data contribution degree corresponding to each participating group representative is determined based on the sub-data contribution degrees obtained from multiple rounds of federated training, improving the accuracy of the data contribution degree of each participating group representative.

[0090] Optionally, in the above step S203, the following steps are respectively performed for each target participating group representative participating in the federated training in each round:

[0091] Obtain at least one joining order of multiple target participating group representatives joining the federated training; then, during the federated training process, determine the respective marginal contributions of multiple target participating group representatives under each joining order; finally, perform a weighted average on the obtained marginal contributions of each target participating group representative to obtain the sub-data contribution degree of each target participating group representative.

[0092] Specifically, determining the respective marginal contributions of multiple target participating group representatives under each joining order includes the following two implementation manners.

[0093] A possible implementation manner is that for the first target participating group representative located at the first position in the joining order, use the first model improvement value corresponding to the first target participating group representative as the marginal contribution of the first target participating group representative.

[0094] Another possible implementation manner is that for the second target participating group representative located at other positions in the joining order, determine at least one target participating group representative before the second target participating group representative. Then determine the second model improvement value obtained when the second target participating group representative and at least one target participating group representative participate in the federated training, and the third model improvement value obtained when at least one target participating group representative participates in the federated training. Use the difference between the second model improvement value and the third model improvement value as the marginal contribution of the second target participating group representative.

[0095] For example, there are 3 target participating group representatives participating in the K-th (where K > 0) round of federated training, namely target participating group representative 1, target participating group representative 2, and target participating group representative 3. When target participating group representative 1 independently trains the federated learning model, the model improvement value obtained is v1. When target participating group representative 2 independently trains the federated learning model, the model improvement value obtained is v2. When target participating group representative 3 independently trains the federated learning model, the model improvement value obtained is v3.

[0096] When target participating group representative 1 and target participating group representative 2 jointly train the federated learning model, the model improvement value obtained is v 12 When target participating group representative 1 and target participating group representative 3 jointly train the federated learning model, the model improvement value obtained is v 13 When target participating group representative 2 and target participating group representative 3 jointly train the federated learning model, the model improvement value obtained is v 23 When target participating group representative 1, target participating group representative 2, and target participating group representative 3 jointly train the federated learning model, the model improvement value obtained is v 123

[0097] There are 6 joining orders for the 3 target participating group representatives to join the federated training. Joining order 1 is target participating group representative 1, target participating group representative 2, and target participating group representative 3. Joining order 2 is target participating group representative 1, target participating group representative 3, and target participating group representative 2. Joining order 3 is target participating group representative 2, target participating group representative 1, and target participating group representative 3. Joining order 4 is target participating group representative 2, target participating group representative 3, and target participating group representative 1. Joining order 5 is target participating group representative 3, target participating group representative 1, and target participating group representative 2. Joining order 6 is target participating group representative 3, target participating group representative 2, and target participating group representative 1, as shown in Table 2.

[0098] Table 2.

[0099]

[0100]

[0101] For joining order 1, target participating group representative 1 is in the first position of joining order 1. The model improvement value v1 obtained by target participating group representative 1 when independently training the federated learning model is the first model improvement value. Therefore, the model improvement value v1 is used as the marginal contribution of target participating group representative 1.

[0102] Target participating group representative 2 is in other positions of joining order 1. The target participating group representative before target participating group representative 2 is target participating group representative 1. The model improvement value v obtained by target participating group representative 1 and target participating group representative 2 when jointly training the federated learning model​12 is the improvement value of the second model. The improvement value v1 obtained by the representative 1 of the target participating group training the federated learning model alone is the improvement value of the third model. Therefore, the difference between the improvement value v of the second model 12 and the improvement value v1 of the third model is used as the marginal contribution of the representative 2 of the target participating group.

[0103] The representative 3 of the target participating group is in other positions of the joining order 1. The representatives of the target participating group before the representative 3 of the target participating group are the representative 1 of the target participating group and the representative 2 of the target participating group. The improvement value v obtained by the representative 1, the representative 2, and the representative 3 of the target participating group jointly training the federated learning model 123 is the improvement value of the second model. The improvement value v obtained by the representative 1 and the representative 2 of the target participating group jointly training the federated learning model 12 is the improvement value of the third model. The difference between the improvement value v of the second model 123 and the improvement value v of the third model 12 is the marginal contribution of the representative 3 of the target participating group.

[0104] In other joining orders, the method for calculating the respective marginal contributions of the representative 1, the representative 2, and the representative 3 of the target participating group is the same as the above method and will not be elaborated here.

[0105] The marginal contributions obtained by the representative 1, the representative 2, and the representative 3 of the target participating group under different joining orders are shown in Table 3.

[0106] Table 3.

[0107]

[0108]

[0109] Assuming that the probability of each joining order is 1 / 6, then the contribution degree of the sub-data of the representative 1 of the target participating group is (v1 + v1 + (v 12 - v2) + (v 123 - v 23 ) + (v 13 - v3) + (v 123 - v 23 )) / 6, and the contribution degree of the sub-data of the representative 2 of the target participating group is ((v 12 - v1) + (v 123 - v 13 ) + v2 + v2 + (v 123 - v 13 ) + (v 23 - v3)) / 6, and the contribution degree of the sub-data of the representative 3 of the target participating group is ((v 123 - v12 )+(v 13 -v1)+(v 123 -v 12 )+(v 23 -v2)+v3+v3) / 6。

[0110] In the embodiments of the present application, for each round of federated training, the target participating group representatives participating in the federated training are permuted to obtain multiple joining orders. Then, the marginal contribution of each target participating group representative participating in the federated training is calculated under various joining orders, and then the sub-data contribution degree of the target participating group representative is determined based on the marginal contribution of the target participating group representative under various joining orders, so as to more accurately measure the sub-data contribution degree of each target participating group representative in each round of federated training.

[0111] Optionally, there are two ways to determine multiple target participating group representatives:

[0112] The first way, if it is the first round of federated training currently, then multiple target participating group representatives participating in the first round of federated training are obtained from multiple first participating groups.

[0113] Specifically, the multiple participating parties are grouped according to the similarity of the obtained encrypted data to obtain multiple first groupings, and one participating party is selected from each first participating group as the participating group representative.

[0114] The selection rule can be set according to the actual situation, and then one participating party is selected from each first participating group as the participating group representative according to the selection rule. For example, one participating party can be randomly selected from each first participating group as the participating group representative.

[0115] The second way, if it is the T-th round of federated training currently, then M target participating group representatives participating in the T-th round of federated training are determined according to the sub-data contribution degrees of N target participating group representatives participating in the (T - 1)-th round of federated training, where T > 1 and N >= M.

[0116] Specifically, according to the similarity of the sub-data contribution degrees of N target participating group representatives, the N target participating group representatives are regrouped to obtain multiple second participating groups, and M target participating group representatives are obtained from the multiple second participating groups.

[0117] Specifically, for regrouping the N target participating group representatives to obtain multiple second participating groups, the following steps are included:

[0118] A possible implementation manner is to directly regroup the N target participating group representatives to obtain multiple second participating groups.

[0119] In another possible implementation, the N target participating group representatives are sorted in descending order of sub-data contribution degree. Then, according to the sorting result, the target participating group representatives with sub-data contribution degrees greater than or equal to the contribution degree threshold are obtained from the N target participating group representatives. Finally, the obtained target participating group representatives are grouped according to the similarity of the sub-data contribution degrees to obtain multiple second participating groups.

[0120] Grouping the retained target participating group representatives to obtain multiple second participating groups includes the following two possible implementation manners:

[0121] Implementation manner one: Compare the sub-data contribution degrees of each target participating group representative with multiple preset intervals, and group the target participating group representatives whose sub-data contribution degrees fall into the same preset interval into one group, that is, multiple second participating groups are obtained.

[0122] Implementation manner two: Group the sub-data contribution degrees of each target participating group representative through a clustering algorithm, that is, multiple second participating groups are obtained.

[0123] The selection rule can be set according to the actual situation, and then one or zero participants are selected from each second participating group as the participating group representative according to the selection rule. For example, one or zero participants can be randomly selected from each second participating group as the participating group representative.

[0124] For example, if there are 3 target participating group representatives participating in the federated training in the (T - 1)th round, namely target participating group representative 1, target participating group representative 2, and target participating group representative 3. The sub-data contribution degrees corresponding to target participating group representative 1, target participating group representative 2, and target participating group representative 3 are 0.2, 0.6, and 0.1 respectively. Three preset intervals are set as [0, 0.3), [0.3, 0.9], and (0.9, 1]. The sub-data contribution degree 0.2 of target participating group representative 1 falls into the interval [0, 0.3), the sub-data contribution degree 0.6 of target participating group representative 2 falls into the interval [0.3, 0.9], and the sub-data contribution degree 0.1 of target participating group representative 3 falls into the interval [0, 0.3). Therefore, target participating group representative 1 and target participating group representative 3 are grouped into one group as the second participating group 1, and target participating group representative 2 is grouped into one group as the second participating group 2.

[0125] Select target participating group representative 1 from the second participating group 1, select target participating group representative 2 from the second participating group 2, and use target participating group representative 1 and target participating group representative 2 as the target participating group representatives participating in the federated training in the Tth round.

[0126] In the embodiments of the present application, the multiple target participating group representatives participating in the first round of federated training are obtained from multiple first participating groups, and the M target participating group representatives participating in the T-th round of federated training are determined according to the sub-data contribution degrees of the N target participating group representatives participating in the (T - 1)-th round of federated training, where T > 1 and N >= M. Since the target participating group representatives participating in each round of federated training are continuously adjusted, the sub-data contribution degrees of the target participating group representatives can be determined more accurately. As the number of rounds of federated training increases, the number of target participating group representatives gradually decreases, and at the same time, the parties with high contribution degrees are screened out in each round of training, improving the efficiency of calculating the sub-data contribution degrees corresponding to the target participating group representatives.

[0127] To better explain the embodiments of the present application, the following describes a method for data processing in federated learning provided by the embodiments of the present application in combination with a specific implementation scenario, as Figure 3 shown, including the following steps:

[0128] Step S301: Receive encrypted data sent by multiple parties.

[0129] Step S302: Group the multiple parties according to the similarity of the encrypted data to obtain multiple first participating groups.

[0130] Step S303: Select one party from each first participating group as the participating group representative.

[0131] Step S304: Set the variable i = 1.

[0132] Step S305: Determine whether i is less than or equal to the preset number of training times T. If so, execute Step S306; otherwise, execute Step S318.

[0133] Step S306: Use the selected multiple participating group representatives as the target participating group representatives.

[0134] Step S307: Determine the marginal contribution of each target participating group representative in the i-th round of federated training.

[0135] Step S308: Determine the sub-data contribution degree of each target participating group representative in the i-th round of federated training according to the obtained marginal contributions.

[0136] Step S309: Set the variable j = 1.

[0137] Step S310: Determine whether j is less than or equal to the total number of target participating group representatives. If so, execute Step S311; otherwise, execute Step S315.

[0138] Step S311: Obtain the sub-data contribution degree of the j-th target participating group representative in the i-th round of federated training.

[0139] Step S312: Determine whether the sub - data contribution degree of the j - th target participating group representative in the i - th round of federated training is less than the contribution degree threshold. If so, execute Step S314; otherwise, execute Step S313.

[0140] Step S313: Add the j - th target participating group representative to the target set.

[0141] Step S314: Set j = j + 1, and jump to Step S310.

[0142] Step S315: Group the target participating group representatives in the target set to obtain multiple second participating groups.

[0143] Step S316: Select one participant from each second participating group as the participating group representative.

[0144] Step S317: Set i = i + 1, and jump to Step S305.

[0145] Step S318: Weight - average the multiple sub - data contribution degrees corresponding to each participating group representative to obtain the data contribution degree corresponding to each participating group representative.

[0146] Step S319: Determine the data contribution degree corresponding to each participant according to the data contribution degree corresponding to each participating group representative.

[0147] In the embodiment of the present application, since one participant is selected from each first participating group as the participating group representative to participate in federated training instead of all participants participating in federated training, the time overhead is greatly reduced and the data contribution degree detection efficiency is improved.

[0148] For each target participating representative group participating in federated training in each round, calculate the marginal contribution of each target participating group representative under all joining orders, and more accurately measure the sub - data contribution degree of the target participating group representative. At the same time, as the number of rounds of federated training increases, the number of target participating group representatives gradually decreases, and the participants with high contribution degrees are screened out in each round of training, improving the efficiency of calculating the sub - data contribution degree corresponding to the target participating group representative.

[0149] Finally, during the multi - round federated training process, the sub - data contribution degrees corresponding to each participating group representative are not constant, but change correspondingly with the change of the federated training rounds, that is, the sub - data contribution degree has temporal variation in the federated training of different rounds, so as to realize the dynamic evaluation of the sub - data contribution degrees corresponding to each participating group representative in each round of federated training. At the same time, determining the data contribution degree corresponding to each participating group representative based on the sub - data contribution degrees obtained from multi - round federated training improves the accuracy of the data contribution degree of each participating group representative.

[0150] Based on the same technical concept, an embodiment of the present application provides a device for data processing in federated learning, as Figure 4 shown. The device 400 includes:

[0151] A receiving module 401, configured to receive encrypted data sent by multiple participants;

[0152] A grouping module 402, configured to group the multiple participants according to the similarity of the obtained encrypted data, obtain multiple first participant groups, and select one participant from each first participant group as a representative of the participant group;

[0153] A sub-data contribution degree obtaining module 403, configured to obtain the sub-data contribution degree of the target participant group representative participating in each round of federated training based on the marginal contribution of the target participant group representative participating in each round of federated training among multiple participant group representatives;

[0154] A data contribution degree obtaining module 404, configured to obtain the respective data contribution degrees of the multiple participant group representatives based on the sub-data contribution degrees obtained after multiple rounds of federated training;

[0155] The data contribution degree obtaining module 404 is further configured to determine the respective data contribution degrees of the multiple participants according to the respective data contribution degrees of the multiple participant group representatives.

[0156] Optionally, the grouping module 402 is specifically configured to:

[0157] Determine the similarity between the encrypted data sent by any two participants among the multiple participants;

[0158] Divide any two participants whose similarity is greater than the similarity threshold into a first participant group, and obtain multiple first participant groups.

[0159] Optionally, the sub-data contribution degree obtaining module 403 is specifically configured to:

[0160] For each target participant group representative participating in each round of federated training, respectively perform the following steps:

[0161] Obtain at least one joining order of multiple target participant group representatives joining the federated training;

[0162] During the federated training process, determine the respective marginal contributions of the multiple target participant group representatives under each joining order;

[0163] Weight and average the obtained marginal contributions of each target participant group representative to obtain the sub-data contribution degree of each target participant group representative.

[0164] Optionally, the sub-data contribution degree obtaining module 403 is further configured to:

[0165] For the first target participating group representative located at the first position in the joining order, use the first model improvement value corresponding to the first target participating group representative as the marginal contribution of the first target participating group representative;

[0166] For the second target participating group representative located at other positions in the joining order, determine at least one target participating group representative before the second target participating group representative;

[0167] Determine the second model improvement value obtained when the second target participating group representative and the at least one target participating group representative participate in federated training, and the third model improvement value obtained when the at least one target participating group representative participates in federated training;

[0168] Use the difference between the second model improvement value and the third model improvement value as the marginal contribution of the second target participating group representative.

[0169] Optionally, the sub-data contribution degree acquisition module 403 is further configured to:

[0170] The multiple target participating group representatives in the first round of participating in federated training are obtained from the multiple first participating groups;

[0171] The M target participating group representatives in the T-th round of participating in federated training are determined according to the sub-data contribution degrees of the N target participating group representatives in the (T - 1)-th round of participating in federated training, where T > 1 and N >= M.

[0172] Optionally, the sub-data contribution degree acquisition module 403 is further configured to:

[0173] Re-group the N target participating group representatives according to the similarity of their sub-data contribution degrees to obtain multiple second participating groups, and obtain M target participating group representatives from the multiple second participating groups.

[0174] Optionally, the sub-data contribution degree acquisition module 403 is further configured to:

[0175] Sort the N target participating group representatives in descending order of sub-data contribution degree;

[0176] According to the sorting result, obtain the target participating group representatives with sub-data contribution degrees greater than or equal to the contribution degree threshold from the N target participating group representatives;

[0177] Group the obtained target participating group representatives according to the similarity of their sub-data contribution degrees to obtain multiple second participating groups.

[0178] Optionally, the data contribution degree acquisition module 404 is further configured to:

[0179] For each of the multiple participating group representatives, the following steps are respectively performed:

[0180] After multiple rounds of training, the data contribution degrees corresponding to a participating group representative are weighted and averaged for multiple sub-data contribution degrees to obtain the data contribution degree corresponding to the participating group representative.

[0181] Optionally, the encrypted data is obtained by the participating party encrypting the original data through a locality-sensitive hashing algorithm, and the data contribution degree is a Shapley value.

[0182] Based on the same technical concept, an embodiment of the present application provides a computer device, which can be a terminal or a server. For example, Figure 5 as shown, it includes at least one processor 501 and a memory 502 connected to at least one processor. In the embodiment of the present application, the specific connection medium between the processor 501 and the memory 502 is not limited. Figure 5 Taking the example that the processor 501 and the memory 502 are connected through a bus. The bus can be divided into an address bus, a data bus, a control bus, etc.

[0183] In the embodiment of the present application, the memory 502 stores instructions executable by at least one processor 501. By executing the instructions stored in the memory 502, at least one processor 501 can execute the steps included in the method for data processing in the above-mentioned federated learning.

[0184] Among them, the processor 501 is the control center of the computer device. It can connect various parts of the computer device through various interfaces and lines. By running or executing the instructions stored in the memory 502 and calling the data stored in the memory 502, data processing in federated learning can be performed. Optionally, the processor 501 may include one or more processing units. The processor 501 may integrate an application processor and a modulation and demodulation processor. Among them, the application processor mainly processes the operating system, user interface, application programs, etc., and the modulation and demodulation processor mainly processes wireless communication. It can be understood that the above-mentioned modulation and demodulation processor may not be integrated into the processor 501. In some embodiments, the processor 501 and the memory 502 can be implemented on the same chip, and in some embodiments, they can also be separately implemented on independent chips.

[0185] The processor 501 may be a general-purpose processor, such as a central processing unit (CPU), a digital signal processor, an application specific integrated circuit (ASIC), a field programmable gate array or other programmable logic device, discrete gate or transistor logic device, discrete hardware component, and can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or executed by a combination of hardware and software modules in the processor.

[0186] The memory 502, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs and modules. The memory 502 may include at least one type of storage medium, for example, it may include flash memory, a hard disk, a multimedia card, a card-type memory, a random access memory (RAM), a static random access memory (SRAM), a programmable read-only memory (PROM), a read-only memory (ROM), an electrically erasable programmable read-only memory (EEPROM), a magnetic memory, a magnetic disk, an optical disk, and so on. The memory 502 is any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer, but is not limited thereto. The memory 502 in the embodiments of the present application may also be a circuit or any other device capable of implementing a storage function, for storing program instructions and / or data.

[0187] Based on the same inventive concept, the embodiments of the present application provide a computer-readable storage medium, which stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the method for data processing in the above-mentioned federated learning.

[0188] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0189] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, as well as the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices produce means for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0190] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory produce a manufactured article including instruction means that implement the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0191] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one flow or multiple flows and / or blocks Figure 1 one block or multiple blocks.

[0192] Obviously, those skilled in the art can make various modifications and variations to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalent technologies, the present application is also intended to include these modifications and variations.

Claims

1. A method for data processing in federated learning, characterized in that Including: Receiving encrypted data sent by multiple participants; Determining the similarity between the encrypted data sent by any two of the multiple participants, dividing any two participants that meet the condition that the similarity is greater than the similarity threshold into a first participant group, obtaining multiple first participant groups, and selecting one participant from each first participant group as the representative of the participant group; Obtaining the sub-data contribution degree of the target participant group representative in each round of participating in federated training based on the marginal contribution of the target participant group representative in each round of participating in federated training among multiple participant group representatives; Wherein, the M target participant group representatives in the T-th round of participating in federated training are determined according to the sub-data contribution degrees of the N target participant group representatives in the (T - 1)-th round of participating in federated training, where T > 1 and N >= M; Obtaining the respective data contribution degrees corresponding to the multiple participant group representatives based on the sub-data contribution degrees obtained after multiple rounds of federated training; Determining the respective data contribution degrees corresponding to the multiple participants according to the respective data contribution degrees corresponding to the multiple participant group representatives.

2. The method according to claim 1, wherein The obtaining the sub-data contribution degree of the target participant group representative in each round of participating in federated training based on the marginal contribution of the target participant group representative in each round of participating in federated training among multiple participant group representatives includes: For the target participant group representative in each round of participating in federated training, respectively perform the following steps: Obtaining at least one joining order of multiple target participant group representatives joining federated training; During the process of federated training, determining the respective marginal contributions corresponding to the multiple target participant group representatives under each joining order; Taking the weighted average of the obtained marginal contributions of each target participant group representative to obtain the sub-data contribution degree of each target participant group representative.

3. The method according to claim 2, wherein The determining the respective marginal contributions corresponding to the multiple target participant group representatives under each joining order includes: For the first target participant group representative located at the first position in the joining order, taking the first model improvement value corresponding to the first target participant group representative as the marginal contribution of the first target participant group representative; For the second target participant group representative located at other positions in the joining order, determining at least one target participant group representative before the second target participant group representative; Determining the second model improvement value obtained when the second target participant group representative and the at least one target participant group representative participate in federated training, and the third model improvement value obtained when the at least one target participant group representative participates in federated training; Taking the difference between the second model improvement value and the third model improvement value as the marginal contribution of the second target participant group representative.

4. The method according to claim 1, wherein The determining that the M target participant group representatives in the T-th round of participating in federated training are determined according to the sub-data contribution degrees of the N target participant group representatives in the (T - 1)-th round of participating in federated training includes: Re-grouping the N target participant group representatives according to the similarity of the sub-data contribution degrees of the N target participant group representatives to obtain multiple second participant groups, and obtaining M target participant group representatives from the multiple second participant groups.

5. The method according to claim 4, wherein Re-grouping the N target participating group representatives according to the similarity of the sub-data contribution degrees thereof to obtain a plurality of second participating groups, including: Sorting the N target participating group representatives in descending order of the sub-data contribution degrees; According to the sorting result, obtaining the target participating group representatives with sub-data contribution degrees greater than or equal to the contribution degree threshold from the N target participating group representatives; Grouping the obtained target participating group representatives according to the similarity of the sub-data contribution degrees to obtain a plurality of second participating groups.

6. The method according to claim 1, wherein Obtaining the respective corresponding data contribution degrees of the plurality of participating group representatives based on the sub-data contribution degrees obtained after multiple rounds of federated training, including: For the plurality of participating group representatives, respectively performing the following steps: After multiple rounds of training, performing weighted averaging on multiple sub-data contribution degrees corresponding to one participating group representative to obtain the data contribution degree corresponding to the one participating group representative.

7. The method according to claim 6, wherein The encrypted data is obtained by the participating party encrypting the original data through a locality-sensitive hashing algorithm, and the data contribution degree is the Shapley value.

8. An apparatus for data processing in federated learning, characterized in that, Including: A receiving module, configured to receive encrypted data sent by a plurality of participating parties; A grouping module, configured to determine the similarity between the encrypted data sent by any two of the plurality of participating parties, divide any two participating parties that satisfy that the similarity is greater than the similarity threshold into a first participating group to obtain a plurality of first participating groups, and select one participating party from each first participating group as a participating group representative; A sub-data contribution degree obtaining module, configured to obtain the sub-data contribution degrees of the target participating group representatives participating in each round of federated training based on the marginal contributions of the target participating group representatives participating in each round of federated training among the plurality of participating group representatives; Wherein, the M target participating group representatives in the T-th round of participating in federated training are determined according to the sub-data contribution degrees of the N target participating group representatives in the (T-1)-th round of participating in federated training, where T>1 and N>=M; A data contribution degree obtaining module, configured to obtain the respective corresponding data contribution degrees of the plurality of participating group representatives based on the sub-data contribution degrees obtained after multiple rounds of federated training; The data contribution degree obtaining module is further configured to determine the respective corresponding data contribution degrees of the plurality of participating parties according to the respective corresponding data contribution degrees of the plurality of participating group representatives.

9. A computer device, comprising a memory, a processor, and a computer program stored on the memory and executable on the processor, characterized in that, When the processor executes the program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that, It stores a computer program executable by a computer device. When the program runs on the computer device, the computer device is caused to execute the steps of the method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Federation learning method and device based on evolutionary computation, central server and medium

    CN111709535A

  • Federal model training method and device, certificate detection method and device, equipment and medium

    CN113239879A