Data Processing Method, System and Related Devices for Federated Learning of Group Features

By grouping and computing the sample sets, group samples are generated, and the federated logistic regression algorithm is used to solve the problem of unauthorized use of individual feature data in vertical federated learning, which improves the model training efficiency and calculation speed.

CN117313897BActive Publication Date: 2025-07-08INSIGHT TECHNOLOGY (XIONGAN) CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202311406055.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-10-26
Publication Date
2025-07-08
Estimated Expiration
2043-10-26

AI Technical Summary

Technical Problem

In vertical federated learning, individual characteristic data cannot be used directly without authorization, and the existing fully anonymous federated learning still involves the use of unauthorized historical individual data, which has compliance issues.

Method used

After the sample alignment of the sample set is performed by the initiator and the participant, the samples are grouped using preset grouping rules and performed operations to generate group samples. The specified federated logistic regression algorithm is used until the model converges, avoiding the direct use of individual feature data.

Benefits of technology

On the premise of ensuring the accuracy of the model, the training efficiency and calculation speed of federated learning are improved, the problem of unauthorized use of individual data is solved, and the model calculation speed in practical application scenarios is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117313897B_ABST
    Figure CN117313897B_ABST
Patent Text Reader

Abstract

The present application discloses a data processing method, system and related device for federated learning of group features. The method includes: sample alignment of a first sample set and a second sample set by an initiator and participants; grouping the aligned first sample set by the initiator using a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1; grouping the aligned second sample set by the participants using a preset grouping rule to obtain b groups of samples, where b is an integer greater than 1; performing operations on the a groups of samples by the initiator to obtain a group samples; performing operations on the b groups of samples by the participants to obtain b group samples; running a specified federated logistic regression algorithm by the initiator and the participants based on the a group samples and the b group samples until the federated learning model converges. The embodiments of the present application can achieve unauthorized direct use of individual feature data in federated learning.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of privacy computing technology and computer technology, and particularly relates to a data processing method, system and related device for federated learning of group features. Background Art

[0002] With the further development of big data, attaching importance to data privacy and security has become a worldwide trend. The increasingly strict management of user data privacy and security will be a world trend, which brings unprecedented challenges to data analysis work. In practical applications, data in different fields often has great complementarity, and there is a great demand for data fusion between different organizations. However, due to factors such as privacy protection, self-interest and policy supervision, it is very difficult for organizations to directly aggregate data. This data silo problem poses great challenges to artificial intelligence researchers.

[0003] In recent years, the academic and industrial communities have begun to use the federated learning solution to solve such problems. Federated learning utilizes knowledge such as cryptography and machine learning, and strives to improve the actual effect of the AI model on the premise of ensuring data privacy, security, legality and compliance. Among them, vertical federated learning (with consistent data samples and complementary feature dimensions) is particularly widely used among enterprises, but it often involves access rights and privacy issues of detailed data.

[0004] In the process of vertical federated learning calculation, the feature data of the aligned individuals needs to be used. The existing fully anonymized federated learning solves the problem of completing model training without the need for alignment results, but still involves the use of unauthorized historical individual data, and the problem of legal use of individual data requiring authorization still exists. Summary of the Invention

[0005] The embodiments of this application provide a data processing method, system and related device for federated learning of group features, which can realize the unauthorized direct use of individual feature data in federated learning.

[0006] In a first aspect, the embodiments of this application provide a data processing method for federated learning of group features, which is applied to a two-party computing system. The two-party computing system includes: an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. The method includes:

[0007] Align the first sample set and the second sample set through the initiator and the participant;

[0008] Group the aligned first sample set by the initiator using a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1;

[0009] The participating party groups the aligned second sample set according to the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1;

[0010] The initiating party performs operations on the a groups of samples to obtain a group of cluster samples;

[0011] The participating party performs operations on the b groups of samples to obtain b groups of cluster samples;

[0012] The initiating party and the participating party run a specified federated logistic regression algorithm based on the a groups of cluster samples and the b groups of cluster samples until the federated learning model converges.

[0013] In a second aspect, an embodiment of the present application provides a two-party computing system, where the two-party computing system includes: an initiating party and a participating party. The initiating party includes a first sample set, and the participating party includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. Among them,

[0014] The initiating party and the participating party are configured to align the first sample set and the second sample set;

[0015] The initiating party is configured to group the aligned first sample set according to a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1;

[0016] The participating party is configured to group the aligned second sample set according to the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1;

[0017] The initiating party is configured to perform operations on the a groups of samples to obtain a groups of cluster samples;

[0018] The participating party is configured to perform operations on the b groups of samples to obtain b groups of cluster samples;

[0019] The initiating party and the participating party are further configured to run a specified federated logistic regression algorithm based on the a groups of cluster samples and the b groups of cluster samples until the federated learning model converges.

[0020] In a third aspect, an embodiment of the present application provides an electronic device, including a processor, a memory, a communication interface, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above processor. The above programs include instructions for performing the steps in the first aspect of the embodiment of the present application.

[0021] Fourthly, an embodiment of the present application provides a computer-readable storage medium. The computer-readable storage medium stores a computer program for electronic data exchange. The computer program enables a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application.

[0022] Fifthly, an embodiment of the present application provides a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to enable a computer to execute some or all of the steps described in the first aspect of the embodiment of the present application. The computer program product can be a software installation package.

[0023] Implementing the embodiments of the present application has the following beneficial effects:

[0024] It can be seen that the data processing method, system and related devices for federated learning of group features described in the embodiments of the present application are applied to a two-party computing system. The two-party computing system includes an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. By the initiator and the participant, the first sample set and the second sample set are sample-aligned. By the initiator, the aligned first sample set is grouped according to a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1. By the participant, the aligned second sample set is grouped according to a preset grouping rule to obtain b groups of samples, where b is an integer greater than 1. By the initiator, operations are performed on the a groups of samples to obtain a group samples. By the participant, operations are performed on the b groups of samples to obtain b group samples. By the initiator and the participant, a specified federated logistic regression algorithm is run based on the a group samples and the b group samples until the federated learning model converges. This can not only solve the problem that individual feature data in federated learning cannot be directly used without authorization, but also improve the calculation speed of the model in actual application scenarios. Description of the Drawings

[0025] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0026] Figure 1 It is a schematic diagram of the architecture of a two-party computing system for implementing the data processing method of federated learning of group features provided by an embodiment of the present application;

[0027] Figure 2It is a schematic flowchart of a data processing method for federated learning of group features provided by an embodiment of the present application;

[0028] Figure 3 It is another schematic flowchart of a data processing method for federated learning of group features provided by an embodiment of the present application;

[0029] Figure 4 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0030] To enable those skilled in the art to better understand the solutions of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0031] The terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish different objects, rather than to describe a specific order. In addition, the terms "include" and "have" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but optionally further includes steps or units not listed, or optionally further includes other steps or units inherent to these processes, methods, products or devices.

[0032] Referring to "embodiment" herein means that a specific feature, structure or characteristic described in connection with the embodiment may be included in at least one embodiment of the present application. The phrase appears in various places in the specification does not necessarily refer to the same embodiment, nor is it an independent or alternative embodiment mutually exclusive with other embodiments. Those skilled in the art will explicitly and implicitly understand that the embodiments described herein may be combined with other embodiments.

[0033] The initiator and participants described in the embodiments of the present application may both include electronic devices. The electronic devices may include smart phones (such as Android phones, iOS phones, Windows Phone phones, etc.), tablet computers, palmtop computers, driving recorders, servers, laptop computers, mobile Internet devices (MIDs), or wearable devices (such as smart watches, Bluetooth headsets), etc. The above are only examples, not an exhaustive list, including but not limited to the above electronic devices. The electronic device may also be a cloud server, or the electronic device may also be a computer cluster.

[0034] The embodiments of the present application will be introduced in detail below.

[0035] Please refer to Figure 1 , Figure 1 , which is a schematic diagram of the architecture of a two-party computing system for a data processing method of federated learning for implementing group features provided by an embodiment of the present application. As shown in the figure, the two-party computing system includes: an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. The following functions can be implemented based on this two-party computing system:

[0036] The initiator and the participant are used to align the first sample set and the second sample set.

[0037] The initiator is used to group the aligned first sample set according to a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1.

[0038] The participant is used to group the aligned second sample set according to the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1.

[0039] The initiator is used to perform operations on the a groups of samples to obtain a group samples.

[0040] The participant is used to perform operations on the b groups of samples to obtain b group samples.

[0041] The initiator and the participant are also used to run a specified federated logistic regression algorithm according to the a group samples and the b group samples until the federated learning model converges.

[0042] Optionally, in running the specified federated logistic regression algorithm according to the a group samples and the b group samples, it includes:

[0043] The initiator is used to shuffle the a group samples and synchronize the shuffled order to the participant.

[0044] The participant is used to shuffle the b group samples according to the shuffled order.

[0045] The initiator and the participant are used to run the specified federated logistic regression algorithm according to the shuffled a group samples and the shuffled b group samples until the federated learning model converges.

[0046] Optionally, in terms of performing operations on the a groups of samples to obtain a group samples, it includes:

[0047] The initiator is used to perform an averaging operation on the a groups of samples to obtain the a group samples;

[0048] Performing an operation on the b groups of samples to obtain b group samples includes:

[0049] The participant is used to perform an averaging operation on the b groups of samples to obtain the b group samples.

[0050] Optionally, the following functions can be implemented based on this two-party computing system:

[0051] Obtain p first test group features through the initiator; predict the p first test group features through the converged federated learning algorithm to obtain p first sample prediction probabilities, where p is a positive integer;

[0052] Obtain q second test group features through the participant; predict the q second test group features through the converged federated learning algorithm to obtain q second sample prediction probabilities, where q is a positive integer;

[0053] The initiator determines the target model evaluation value of the federated learning model according to the p first sample prediction probabilities and the q second sample prediction probabilities.

[0054] Optionally, the aligning of the first sample set and the second sample set includes:

[0055] The initiator and the participant are used to align the first sample set and the second sample set by using a private intersection algorithm.

[0056] Please refer to Figure 2 , Figure 2 FIG. is a schematic flowchart of a data processing method for federated learning of group features provided by an embodiment of the present application, which is applied to a two-party computing system. The two-party computing system includes an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. As shown in the figure, the data processing method for federated learning of group features of the present application includes:

[0057] 201. Align the first sample set and the second sample set through the initiator and the participant.

[0058] In an embodiment of the present application, a two-party computing system may include: an initiator and a participant. The initiator includes a first sample set, which may include multiple samples. The participant includes a second sample set, which may also include multiple samples. The number of samples in the first sample set and the number of samples in the second sample set may be the same or different. Both the first sample set and the second sample set may be referred to as training data.

[0059] In an embodiment of the present application, a sample can be a sample for which federated learning needs to be performed, and the sample may include at least one of the following: images, age data, consumption records, identity data, credit record data, etc., which are not limited herein. Different labels can correspond to different categories, and the label can be a number, a phrase, a graph, etc., which are not limited herein. For example, the label can be 0 or 1.

[0060] For example, taking Bank A and Bank B as an example, if Bank A is the initiator and Bank B is the participant, then the samples of Bank A and Bank B can be sample-aligned.

[0061] Among them, the first sample set and the second sample set have the same type of labels, that is, if the first sample set has 2 types of labels, then the second sample set can also have 2 types of labels.

[0062] In a specific implementation, the initiator and the participant can sample-align the first sample set and the second sample set through a private intersection algorithm.

[0063] Optionally, in step 201 above, the sample alignment of the first sample set and the second sample set by the initiator and the participant can be implemented in the following manner:

[0064] The first sample set and the second sample set are sample-aligned by the initiator and the participant using a private intersection algorithm.

[0065] In an embodiment of the present application, the initiator and the participant can sample-align the first sample set and the second sample set using a private intersection algorithm.

[0066] For example, assume that the initiator is A and the participant is B. Then, without providing detailed features, A and B first complete the alignment of the training data IDs through a private intersection technology (PSI).

[0067] 202. The initiator uses a preset grouping rule to group the aligned first sample set to obtain a groups of samples, where a is an integer greater than 1.

[0068] In the embodiments of the present application, the preset grouping rule can be set in advance or be the system default. For example, the preset grouping rule can be to group based on the sequential ID of the samples and the sample labels. First, classify the first sample set according to the sample labels, and then group the samples in each category based on the sequential ID of the samples.

[0069] In a specific implementation, the initiator can group the aligned first sample set by using the preset grouping rule to obtain a groups of samples, where a is an integer greater than 1.

[0070] Illustrating with an example, the initiator A can set the grouping rule to group the aligned IDs on its own side. For example, a feasible solution is to group the samples with the same label (positive / negative), with every 5 samples as a group. For example, five positive samples with IDs 1, 5, 8, 13, and 19 are grouped into one group, and five negative samples with IDs 2, 7, 22, 26, and 78 are grouped into one group... and so on.

[0071] 203. The participating party uses the preset grouping rule to group the aligned second sample set to obtain b groups of samples, where b is an integer greater than 1.

[0072] In the implementation of the present application, the grouping methods of the initiator and the participating party are the same.

[0073] In the embodiments of the present application, the initiator can send the grouping method of the initiator to the participating party. The participating party can use the preset grouping rule to group the aligned second sample set to obtain b groups of samples, where b is an integer greater than 1. Thus, it can only know which IDs are in a group, but cannot know the specific sample information, thereby solving the problem that individual feature data in federated learning cannot be directly used without authorization.

[0074] In a specific implementation, a can be equal to b. For example, in the context of vertical federation, the grouping methods of both parties are the same, the samples are the same, but the held sample features are different.

[0075] Illustrating with an example, the initiator A synchronizes the grouping method to B. After receiving the grouping method, the participating party B synchronously groups the corresponding aligned IDs into one group. That is, five samples with IDs 1, 5, 8, 13, and 19 are grouped into one group, and five samples with IDs 2, 7, 22, 26, and 78 are grouped into one group,... The participating party B can only know which IDs are in a group, but cannot know the positive and negative sample information.

[0076] 204. The initiator performs operations on the a groups of samples to obtain a group of group samples.

[0077] In the embodiments of the present application, the initiator can perform operations on the a groups of samples to obtain a group of group samples, that is, each group corresponds to one group sample.

[0078] 205. The participating party performs operations on the b groups of samples to obtain b group samples.

[0079] In the embodiment of the present application, the participating party can perform operations on the b groups of samples to obtain b group samples, that is, each group corresponds to one group sample.

[0080] Optionally, in step 204 above, the initiator performs operations on the a groups of samples to obtain a group samples, which can be implemented in the following manner:

[0081] The initiator performs an averaging operation on the a groups of samples to obtain the a group samples;

[0082] Then, in step 205 above, the participating party performs operations on the b groups of samples to obtain b group samples, which can be implemented in the following manner:

[0083] The participating party performs an averaging operation on the b groups of samples to obtain the b group samples.

[0084] In the embodiment of the present application, the initiator can perform an averaging operation on the a groups of samples to obtain a group samples. Similarly, the participating party can also perform an averaging operation on the b groups of samples to obtain b group samples.

[0085] For example, A and B calculate the average of the grouped samples to obtain group samples. Denote the binary classification label as 0 (negative sample) or 1 (positive sample). Samples with label 0 are denoted as Samples with label 1 are denoted as Every m data with the same label are divided into a group, and the mean value is calculated for them to obtain group samples Specifically as follows:

[0086]

[0087]

[0088] Among them, represents the group feature, m is the number of data with the same label in each group, represents the samples of A or B with label 0, represents the samples of A or B with label 1, and k represents the starting serial number of each group.

[0089] Among them, for each type of label in A and B, there can be corresponding group samples. For each type of label, there are as many corresponding group samples as there are groups.

[0090] 206. The initiator and the participant run a specified federated logistic regression algorithm based on the a group samples and the b group samples until the federated learning model converges.

[0091] In the embodiments of the present application, running the specified federated logistic regression algorithm can be preset or the system default. For example, the specified federated logistic regression algorithm can include a federated logistic regression algorithm with a batchsize of 1, that is, the parameter batchsize = 1 in the federated logistic regression algorithm.

[0092] In the embodiments of the present application, the federated learning model can be optimized by constructing group features. On the premise of ensuring the model accuracy, the training efficiency of federated learning is improved, and the problem that the use of the federated learning model is restricted when individual data is unauthorized under the existing laws and regulations is solved. In addition, the calculation speed of the model in the actual application scenario can be improved, and the calculation efficiency is significantly improved after compressing the samples.

[0093] In practical applications, the optimization method in the embodiments of the present application has a significant effect and has a good performance in the prediction results of a large sample size.

[0094] Optionally, in step 206 above, running the specified federated logistic regression algorithm by the initiator and the participant based on the a group samples and the b group samples may include the following steps:

[0095] 61. The initiator shuffles the a group samples and synchronizes the shuffled order to the participant;

[0096] 62. The participant shuffles the b group samples according to the shuffled order;

[0097] 63. The initiator and the participant run the specified federated logistic regression algorithm based on the shuffled a group samples and the shuffled b group samples until the federated learning model converges.

[0098] In the embodiments of the present application, the initiator can shuffle the a group samples and synchronize the shuffled order to the participant. The participant shuffles the b group samples according to the shuffled order, that is, both the initiator and the participant shuffle according to the same shuffled order, and then the initiator and the participant run the specified federated logistic regression algorithm based on the shuffled a group samples and the shuffled b group samples until the federated learning model converges.

[0099] For example, the initiator A shuffles the group samples of its own party and synchronizes the shuffled order to party B. The participating party B receives the shuffled order and shuffles the group samples of its own party according to the received order. Both parties A and B perform the federated logistic regression algorithm with a batch size of 1 until the model converges and the model parameters are obtained.

[0100] Optionally, the following steps may further be included:

[0101] A1. Obtain p first test group features through the initiator; predict the p first test group features through the converged federated learning algorithm to obtain p first sample prediction probabilities, where p is a positive integer;

[0102] A2. Obtain q second test group features through the participating party; predict the q second test group features through the converged federated learning algorithm to obtain q second sample prediction probabilities, where q is a positive integer;

[0103] A3. Determine the target model evaluation value of the federated learning model by the initiator according to the p first sample prediction probabilities and the q second sample prediction probabilities.

[0104] In the embodiments of the present application, the p first test group features can be understood as the group features of the test set samples of the initiator, which can be obtained by the above-mentioned method of generating group samples, and p is a positive integer. Similarly, correspondingly, the q second test group features can be understood as the group features of the test set samples of the participating party, and q is a positive integer.

[0105] In the embodiments of the present application, the initiator can obtain p first test group features, predict the p first test group features through the converged federated learning algorithm to obtain p first sample prediction probabilities. The participating party can obtain q second test group features, and then predict the q second test group features through the converged federated learning algorithm to obtain q second sample prediction probabilities. Then, the initiator determines the target model evaluation value of the federated learning model according to the p first sample prediction probabilities and the q second sample prediction probabilities. In specific implementation, for example, in the fields such as bank risk control credit, indicators such as the area under the Roc curve (Area under Curve, AUC), KS (Kolmogorov-Smirnov), etc. are often used to evaluate the quality of the binary classification model. Here, after averaging the samples of both parties, the different features of both parties can be integrated together through the vertical federated learning algorithm for modeling.

[0106] In specific implementation, the evaluation index of TOP-K recall applicable to groups can also be adopted, and this index is more instructive in actual business scenarios such as banks.

[0107] In the embodiments of the present application, considering that in the process of model training in vertical federated learning, there is generally a situation where feature data cannot be directly used without authorization, the main idea of the embodiments of the present application comes from the connection between the stochastic gradient descent method and the batch gradient descent method in the optimization method. The stochastic gradient descent method calculates the gradient once for one sample and then updates the parameters; the batch gradient descent method calculates the gradient once for multiple samples (referred to as a batch) and then updates the parameters. Under the condition of selecting an appropriate learning rate, both can reach the optimal solution of the model. Therefore, in federated learning, the model can be trained only using the fused features of one batch (i.e., the group features in the present invention), without providing individual features for the stochastic gradient descent method. However, the premise is that the features of both parties need to come from the same group, which can be achieved through the private set intersection technology (PSI). After the two parties align the user IDs using the PSI private set intersection technology without providing feature data, the initiator (including the label party) locally groups the samples according to the labels. After averaging the samples within each group, a new sample is obtained. The new sample is named the group sample, and the label of the new sample is the label corresponding to the group. Then, the grouping method is synchronized to the participant (feature provider). The participant groups and averages the subsequent samples in the same way, thus obtaining a set of group samples. Finally, the group samples are used for the federated learning method to obtain the model.

[0108] For example, taking the vertical logistic regression model based on group features of two parties (A and B) as an example, where A is the initiator (including the modeling label party) and B is the participant. The specific implementation of the invention can be roughly divided into three parts: the first part is to group the samples to generate group samples, the second part is to train the federated model, and the third part is to evaluate the effect of the group model.

[0109] First, generate group samples, specifically as follows:

[0110] 1) Without providing detailed features, A and B first complete the alignment of the training data IDs through the private set intersection technology (PSI).

[0111] 2) The initiator A sets the grouping rules and groups the aligned IDs on its own side. For example, a feasible solution is to group the samples with the same label (positive / negative), with every 5 samples as a group. For example, the five positive samples with IDs 1, 5, 8, 13, and 19 are grouped into one group, and the five negative samples with IDs 2, 7, 22, 26, and 78 are grouped into one group... and so on.

[0112] 3) The initiator A synchronizes the grouping method to B. After receiving the grouping method, the participant B synchronously divides the corresponding aligned IDs into groups. That is, a total of five samples with IDs 1, 5, 8, 13, and 19 are divided into one group, a total of five samples with IDs 2, 7, 22, 26, and 78 are divided into one group,... B can only know which IDs are in a group, but cannot know the positive and negative sample information.

[0113] 4) A and B calculate the average of the grouped samples to obtain the group samples. Denote the label of binary classification as 0 (negative sample) or 1 (positive sample). The samples with label 0 are denoted as The samples with label 1 are denoted as Every m data with the same label are divided into one group, and the mean value is calculated for them to obtain the group samples Specifically as follows:

[0114]

[0115]

[0116] Among them, represents the group feature, m is the number of data with the same label in each group, represents the samples of A or B with label 0, represents the samples of A or B with label 1, and k represents the starting serial number of each group.

[0117] Secondly, the federated model training is specifically as follows:

[0118] 1) The initiator A shuffles its own group samples and synchronizes the shuffled order to Party B.

[0119] 2) The participant B receives the shuffled order and shuffles its own group samples according to the received order.

[0120] 3) Both A and B perform the federated logistic regression algorithm with a batch size of 1 until the model converges and obtain the model parameters.

[0121] Finally, the evaluation of the group model effect can include the following steps 1) and 2):

[0122] 1) When the test set samples do not have feature data authorization, then according to the steps in the first part, generate the group features of the test set, and both A and B use the obtained federated learning model parameters to perform model prediction on the test set group samples.

[0123] 2) Calculate the binary classification evaluation index based on the predicted values.

[0124] That is, the final evaluation of the group model effect can be determined based on the above binary classification evaluation index.

[0125] Of course, finally, the evaluation of the group model effect can include the following steps 3) and 4):

[0126] 3) On the premise that the test set samples have feature data authorization, both Party A and Party B use the obtained federated learning model parameters to perform model prediction on the test set samples to obtain the prediction probability of each sample.

[0127] 4) Calculate the TOP-K recall value for evaluation, that is, calculate the recall rate of 0%-100% positive samples, and view the recall value of the top K% positive examples as an indicator.

[0128] It should be noted that the test set generally simulates the real prediction scenario, that is to say, it targets new samples. And new samples can declare data authorization during access. Therefore, steps 3) and 4) are more instructive in actual business scenarios such as banks. Calculate the TOP-K recall value using the prediction probability of each sample for evaluation, specifically: calculate the recall rate of the top k% positive examples, so as to obtain the final evaluation result of the group model effect.

[0129] To illustrate with another example, as Figure 3 shown, the initiator Party A and the participant Party B perform intersection through the PSI technology; the initiator Party A groups the samples and synchronizes the grouping method B to the participant Party B; both parties average the grouped samples to obtain group features; the initiator Party A shuffles its own group samples and synchronizes the shuffled order to Party B; the participant Party B receives the shuffled order and shuffles its own group samples according to the received order; both Party A and Party B perform federated logistic regression algorithm modeling with a batch size of 1; evaluate the group model effect.

[0130] In the embodiments of the present application, without sacrificing the model accuracy and accuracy rate, by designing the idea of using group samples instead of individual samples for modeling, the privacy of individual data is effectively protected. And a new model evaluation index is given, which has reference significance in the bank scenario.

[0131] It can be seen that the data processing method of federated learning for group features described in the embodiments of the present application is applied to a two-party computing system. The two-party computing system includes an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. The first sample set and the second sample set are sample-aligned by the initiator and the participant. The aligned first sample set is grouped by the initiator using a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1. The aligned second sample set is grouped by the participant using the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1. The a groups of samples are operated on by the initiator to obtain a group samples, and the b groups of samples are operated on by the participant to obtain b group samples. The initiator and the participant run a specified federated logistic regression algorithm based on the a group samples and the b group samples until the federated learning model converges. This can not only solve the problem that individual feature data in federated learning cannot be directly used without authorization, but also improve the calculation speed of the model in actual application scenarios.

[0132] Consistent with the above embodiments, please refer to Figure 4 , Figure 4 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present application. As shown in the figure, the electronic device includes a processor, a memory, a communication interface, and one or more programs. The above one or more programs are stored in the above memory and are configured to be executed by the above processor. It is applied to a two-party computing system. The two-party computing system includes an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of label types owned by both the first sample set and the second sample set is the same. In the embodiments of the present application, the above program includes instructions for performing the following steps:

[0133] The first sample set and the second sample set are sample-aligned by the initiator and the participant;

[0134] The aligned first sample set is grouped by the initiator using a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1;

[0135] The aligned second sample set is grouped by the participant using the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1;

[0136] The a groups of samples are operated on by the initiator to obtain a group samples;

[0137] The b groups of samples are operated on by the participant to obtain b group samples;

[0138] The initiator and the participant run a specified federated logistic regression algorithm based on the a group samples and the b group samples until the federated learning model converges.

[0139] Optionally, in the aspect of the initiator and the participant running a specified federated logistic regression algorithm based on the a group samples and the b group samples, the above program includes instructions for performing the following steps:

[0140] The initiator shuffles the a group samples and synchronizes the shuffled order to the participant;

[0141] The participant shuffles the b group samples according to the shuffled order;

[0142] The initiator and the participant run the specified federated logistic regression algorithm based on the shuffled a group samples and the shuffled b group samples until the federated learning model converges.

[0143] Optionally, in the aspect of the initiator performing operations on the a group of samples to obtain a group samples, the above program includes instructions for performing the following steps:

[0144] The initiator performs an averaging operation on the a group of samples to obtain the a group samples;

[0145] Then, in the aspect of the participant performing operations on the b group of samples to obtain b group samples, the above program includes instructions for performing the following steps:

[0146] The participant performs an averaging operation on the b group of samples to obtain the b group samples.

[0147] Optionally, the above program further includes instructions for performing the following steps:

[0148] The initiator obtains p first test group features; the converged federated learning algorithm predicts the p first test group features to obtain p first sample prediction probabilities, where p is a positive integer;

[0149] The participant obtains q second test group features; the converged federated learning algorithm predicts the q second test group features to obtain q second sample prediction probabilities, where q is a positive integer;

[0150] The initiator determines the target model evaluation value of the federated learning model based on the p first sample prediction probabilities and the q second sample prediction probabilities.

[0151] Optionally, in the aspect of aligning the first sample set and the second sample set by the initiator and the participant, the above program includes instructions for performing the following steps:

[0152] The initiator and the participant use the private set intersection algorithm to align the first sample set and the second sample set.

[0153] It can be seen that the electronic device described in the embodiments of the present application is applied to a two-party computing system. The two-party computing system includes an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of tag types owned by both the first sample set and the second sample set is the same; the first sample set and the second sample set are aligned by the initiator and the participant, the aligned first sample set is grouped by the initiator using a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1, the aligned second sample set is grouped by the participant using a preset grouping rule to obtain b groups of samples, where b is an integer greater than 1, the initiator performs operations on the a groups of samples to obtain a group samples, the participant performs operations on the b groups of samples to obtain b group samples, and the initiator and the participant run a specified federated logistic regression algorithm based on the a group samples and the b group samples until the federated learning model converges. This can not only solve the problem that individual feature data in federated learning cannot be directly used without authorization, but also improve the calculation speed of the model in actual application scenarios.

[0154] The embodiments of the present application further provide a computer storage medium. The computer storage medium stores a computer program for electronic data exchange, and the computer program causes the computer to execute part or all of the steps of any of the methods described in the above method embodiments. The above computer includes an electronic device.

[0155] The embodiments of the present application further provide a computer program product. The computer program product includes a non-transitory computer-readable storage medium storing a computer program. The computer program is operable to cause the computer to execute part or all of the steps of any of the methods described in the above method embodiments. The computer program product can be a software installation package, and the above computer includes an electronic device.

[0156] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application.

[0157] In the above embodiments, the descriptions of the respective embodiments have their own focuses. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0158] In several embodiments provided in the present application, it should be understood that the disclosed device can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the above division of units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces. The indirect coupling or communication connection of devices or units can be in an electrical or other form.

[0159] The units described above as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0160] In addition, in each embodiment of the present application, the functional units can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0161] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable memory. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the respective embodiments of the present application. And the aforementioned memory includes: various media such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs that can store program codes.

[0162] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructing relevant hardware through a program, and the program can be stored in a computer-readable memory, which may include: a flash drive, a read-only memory (abbreviation: ROM), a random access memory (abbreviation: RAM), a magnetic disk, an optical disc, etc.

[0163] The above embodiments of the present application have been introduced in detail. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation on the present application.

Claims

1. A data processing method for federated learning of group features, characterized in that Applied to a two-party computing system, the two-party computing system includes: an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of tag types owned by both the first sample set and the second sample set is the same. The method includes: The initiator and the participant perform sample alignment on the first sample set and the second sample set; The initiator groups the aligned first sample set according to a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1; The participant groups the aligned second sample set according to the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1; The initiator performs operations on the a groups of samples to obtain a group samples; The participant performs operations on the b groups of samples to obtain b group samples; The initiator and the participant run a specified federated logistic regression algorithm according to the a group samples and the b group samples until the federated learning model converges.

2. The method according to claim 1, wherein The step of the initiator and the participant running the specified federated logistic regression algorithm according to the a group samples and the b group samples includes: The initiator shuffles the a group samples and synchronizes the shuffled order to the participant; The participant shuffles the b group samples according to the shuffled order; The initiator and the participant run the specified federated logistic regression algorithm according to the shuffled a group samples and the shuffled b group samples until the federated learning model converges.

3. The method according to claim 1 or 2, characterized in that, The step of the initiator performing operations on the a groups of samples to obtain a group samples includes: The initiator performs an averaging operation on the a groups of samples to obtain the a group samples; Then the step of the participant performing operations on the b groups of samples to obtain b group samples includes: The participant performs an averaging operation on the b groups of samples to obtain the b group samples.

4. The method according to claim 1 or 2, characterized in that, The method further includes: The initiator obtains p first test group features; the converged federated learning algorithm predicts the p first test group features to obtain p first sample prediction probabilities, where p is a positive integer; The participant obtains q second test group features; the converged federated learning algorithm predicts the q second test group features to obtain q second sample prediction probabilities, where q is a positive integer; The initiator determines the target model evaluation value of the federated learning model according to the p first sample prediction probabilities and the q second sample prediction probabilities.

5. The method according to claim 1 or 2, characterized in that The step of the initiator and the participant performing sample alignment on the first sample set and the second sample set includes: The initiator and the participant perform sample alignment on the first sample set and the second sample set by using a private set intersection algorithm.

6. A two-party computing system, characterized in that, The two-party computing system includes: an initiator and a participant. The initiator includes a first sample set, and the participant includes a second sample set. The number of types of labels owned by both the first sample set and the second sample set is the same; wherein, the initiator and the participant are used to align the first sample set and the second sample set; the initiator is used to group the aligned first sample set according to a preset grouping rule to obtain a groups of samples, where a is an integer greater than 1; the participant is used to group the aligned second sample set according to the preset grouping rule to obtain b groups of samples, where b is an integer greater than 1; the initiator is used to perform operations on the a groups of samples to obtain a group samples; the participant is used to perform operations on the b groups of samples to obtain b group samples; the initiator and the participant are further used to run a specified federated logistic regression algorithm based on the a group samples and the b group samples until the federated learning model converges.

7. The system according to claim 6, wherein When running the specified federated logistic regression algorithm based on the a group samples and the b group samples, it includes: the initiator is used to shuffle the a group samples and synchronize the shuffled order to the participant; the participant is used to shuffle the b group samples according to the shuffled order; the initiator and the participant are used to run the specified federated logistic regression algorithm based on the shuffled a group samples and the shuffled b group samples until the federated learning model converges.

8. The system according to claim 6 or 7, characterized in that When performing operations on the a groups of samples to obtain a group samples, it includes: the initiator is used to perform an averaging operation on the a groups of samples to obtain the a group samples; When performing operations on the b groups of samples to obtain b group samples, it includes: the participant is used to perform an averaging operation on the b groups of samples to obtain the b group samples.

9. An electronic device, characterized in that, It includes a processor and a memory. The memory is used to store one or more programs and is configured to be executed by the processor. The programs include instructions for performing the steps in the method according to any one of claims 1-5.

10. A computer-readable storage medium, characterized in that, A computer program for electronic data exchange is stored, wherein the computer program causes a computer to execute the method according to any one of claims 1-5.

Citation Information

Patent Citations

  • Federal learning-based data processing method and system and related equipment

    CN117010529A