Federated Integrated Learning Method, Apparatus, Device, and Storage Medium
By adopting a differential privacy method based on an exponential mechanism in federated learning, participants can send model parameters in plain text, solving the problems of large communication overhead and imbalance in data distribution in traditional methods, and achieving more efficient and stable federated learning.
Patent Information
- Application Number
- CN202111261571.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2041-10-27
AI Technical Summary
While ensuring data privacy and security, traditional federated learning methods have high communication overhead, high requirements for network bandwidth and stability, and it is difficult to generate effective model parameters when data distribution is unbalanced.
Using a differential privacy method based on the index mechanism, participants can send local model parameters in plain text to reduce communication overhead, and perform feature selection and model selection locally to generate effective model parameters.
It realizes the federated integration model with better performance while ensuring data privacy and security, reduces communication overhead and improves the stability of federated learning methods.
Smart Images

Figure CN114330756B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the fields of artificial intelligence and robotics, and more particularly, to a federated integrated learning method, apparatus, device, and storage medium. Background Art
[0002] For reasons such as protecting user privacy and commercial competition, data cooperation faces many difficulties, and the potential value of dispersed data sources has not been fully utilized. In recent years, Federated Learning (FL) technology has developed rapidly, providing a new solution for cross-departmental, cross-organizational, and cross-industry data cooperation. Its goal is to achieve joint modeling and improve the performance of the training model on the basis of ensuring data privacy, security, legality, and compliance.
[0003] Since an attacker may infer the training data information or even the original training data used when training the model from the trained model parameters, in traditional federated learning methods, the participating parties cannot directly send the model parameters trained locally in plaintext to the federated server or other participating parties. Instead, they send the encrypted model parameters for secure model fusion through a method based on cryptography (or secret sharing), or perturb the generated model through a random gradient descent method based on the Gaussian mechanism to protect the security and privacy of the data. However, for the method based on cryptography (or secret sharing), its communication overhead is high, and it has high requirements for network communication bandwidth and stability. For the random gradient descent method based on the Gaussian mechanism, it is difficult to generate effective model parameters to ensure the performance of the fusion model generated by the federated server when the data distributions of the participating parties are unbalanced.
[0004] Therefore, an efficient and secure federated learning method is needed, which can achieve data cooperation applicable to various scenarios with low communication overhead. Summary of the Invention
[0005] To solve the above problems, the present disclosure protects the local data of the participating parties based on differential privacy of the exponential mechanism. The participating parties can directly send the local model parameters to the federated server in plaintext, reducing the communication overhead and having no risk of leaking the training data.
[0006] Embodiments of the present disclosure provide a federated integrated learning method, apparatus, device, and computer-readable storage medium.
[0007] Embodiments of the present disclosure provide a federated ensemble learning method, including: selecting a first number of features from the feature sets of the participating parties according to a first probability distribution, where the first probability distribution is obtained based on the exponential mechanism for the feature sets of the participating parties; obtaining a plurality of logistic regression models based on at least some of the selected first number of features, and selecting a second number of logistic regression models from the plurality of logistic regression models according to a second probability distribution, where the second probability distribution is obtained based on the exponential mechanism for the plurality of logistic regression models; and sending at least some of the second number of logistic regression models to a fusion end to perform ensemble fusion based on the at least some of the logistic regression models and generate a federated ensemble model.
[0008] Embodiments of the present disclosure further provide a federated ensemble learning method, including: receiving at least one logistic regression model from each of a plurality of participating parties; performing deduplication processing on all the logistic regression models from the plurality of participating parties to remove duplicate logistic regression models; and performing ensemble fusion on the deduplicated logistic regression models to generate a federated ensemble model; where, for each of the plurality of participating parties, the at least one logistic regression model from the participating party includes at least some of the second number of logistic regression models as described in the above federated ensemble learning method.
[0009] Embodiments of the present disclosure provide a federated ensemble learning device, including: a feature selection module configured to select a first number of features from the feature sets of the participating parties according to a first probability distribution, where the first probability distribution is obtained based on the exponential mechanism for the feature sets of the participating parties; a model selection module configured to obtain a plurality of logistic regression models based on at least some of the selected first number of features, and select a second number of logistic regression models from the plurality of logistic regression models according to a second probability distribution, where the second probability distribution is obtained based on the exponential mechanism for the plurality of logistic regression models; and a model sending module configured to send at least some of the second number of logistic regression models to a fusion end to perform ensemble fusion based on the at least some of the logistic regression models and generate a federated ensemble model.
[0010] Embodiments of the present disclosure provide a federated ensemble learning device, including: one or more processors; and one or more memories, where computer-executable programs are stored in the one or more memories, and when the computer-executable programs are executed by the processors, the above-described federated ensemble learning method is performed.
[0011] Embodiments of the present disclosure provide a computer-readable storage medium, on which computer-executable instructions are stored, and the instructions are used to implement the federated integrated learning method as described above when executed by a processor.
[0012] Embodiments of the present disclosure provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the federated integrated learning method according to the embodiments of the present disclosure.
[0013] Compared with the traditional federated learning methods based on cryptography or secret sharing, the method provided by the embodiments of the present disclosure can send the model parameters obtained by training to the federated server in plaintext. There is only one message interaction between each participant and the federated server and the amount of transmitted data is small, significantly reducing the requirements for the communication network.
[0014] Compared with the traditional federated learning methods based on the random gradient descent model, the method provided by the embodiments of the present disclosure can train effective model parameters under the condition of unbalanced data distribution among participants, improving the stability of the federated learning method.
[0015] The method provided by the embodiments of the present disclosure performs feature selection and model selection based on differential privacy of the exponential mechanism locally at each participant, and sends the selected training model to the federated server for integrated fusion, thereby generating a federated integrated model with better performance. Through the method of the embodiments of the present disclosure, the parameters of the selected training model can be sent to the federated server in plaintext without using any cryptographic method, avoiding the problem of ciphertext expansion based on cryptographic methods, thereby realizing more efficient and low-communication-overhead federated learning while ensuring no risk of data leakage.
[0016] In addition, the method provided by the embodiments of the present disclosure can also support a scenario with only two participants by directly transmitting the training model between participants, and support direct communication and model fusion among multiple participants without a federated server. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some exemplary embodiments of the present disclosure. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0018] Figure 1 is a schematic diagram showing cryptography-based horizontal federated learning according to an embodiment of the present disclosure;
[0019] Figure 2A is a flowchart showing a federated ensemble learning method according to an embodiment of the present disclosure;
[0020] Figure 2B is a schematic diagram showing a federated ensemble learning method according to an embodiment of the present disclosure;
[0021] Figure 3A is a flowchart showing feature selection based on the exponential mechanism according to an embodiment of the present disclosure;
[0022] Figure 3B is a schematic diagram showing feature selection based on the exponential mechanism according to an embodiment of the present disclosure;
[0023] Figure 4A is a flowchart showing model construction according to an embodiment of the present disclosure;
[0024] Figure 4B is a flowchart showing model selection according to an embodiment of the present disclosure;
[0025] Figure 4C is a schematic diagram showing model construction and selection according to an embodiment of the present disclosure;
[0026] Figure 5 is a flowchart showing a federated ensemble learning method according to an embodiment of the present disclosure;
[0027] Figure 6A is a schematic diagram showing model fusion via a fusion center according to an embodiment of the present disclosure;
[0028] Figure 6B is a schematic diagram showing model fusion without a fusion center according to an embodiment of the present disclosure;
[0029] Figure 7 is a schematic diagram showing a federated ensemble learning device according to an embodiment of the present disclosure;
[0030] Figure 8 shows a schematic diagram of a federated ensemble learning device according to an embodiment of the present disclosure;
[0031] Figure 9 shows a schematic diagram of the architecture of an exemplary computing device according to an embodiment of the present disclosure; and
[0032] Figure 10 shows a schematic diagram of a storage medium according to an embodiment of the present disclosure. Detailed Description
[0033] To make the objectives, technical solutions, and advantages of the present disclosure more apparent, exemplary embodiments according to the present disclosure will be described in detail below with reference to the accompanying drawings. Apparently, the described embodiments are only a part rather than all of the embodiments of the present disclosure. It should be understood that the present disclosure is not limited by the exemplary embodiments described herein.
[0034] In this specification and the accompanying drawings, steps and elements that are substantially the same or similar are denoted by the same or similar reference numerals, and repeated descriptions of these steps and elements will be omitted. Meanwhile, in the description of the present disclosure, terms such as "first" and "second" are only used for differential description and should not be construed as indicating or implying relative importance or order.
[0035] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure pertains. The terms used herein are only for the purpose of describing the embodiments of the present invention and are not intended to limit the present invention.
[0036] To facilitate the description of the present disclosure, the following concepts related to the present disclosure are introduced.
[0037] The federated integrated learning method of the present disclosure may be based on artificial intelligence (AI). Artificial intelligence is the theory, method, technology, and application system that uses digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use knowledge to obtain the best results. In other words, artificial intelligence is a comprehensive technology in computer science that attempts to understand the essence of intelligence and produce a new intelligent machine that can respond in a way similar to human intelligence. For example, for the federated integrated learning method based on artificial intelligence, it can centrally fuse partial models trained based on partial features of each participant in a way similar to how humans collect data from multiple parties and comprehensively analyze these data to make specific decision judgments to form a better training model. By studying the design principles and implementation methods of various intelligent machines, artificial intelligence enables the federated integrated learning method of the present disclosure to have the function of automatically and real-time selecting partial features of each participant for model training based on the exponential mechanism of differential privacy and selecting partial models for central fusion.
[0038] The federated integrated learning method of the present disclosure can be based on federated learning technology. According to the distribution of data among different participants, federated learning can be divided into horizontal federated learning (HFL), vertical federated learning (VFL), and federated transfer learning (FTL). Among them, the essence of horizontal federated learning is the union of samples, which is applicable to scenarios where the business types among participants are the same but the reached users are different, that is, the feature overlap is large while the sample overlap is small. For example, among banks in different regions, their businesses are similar (features are similar), but the users are different (samples are different). The essence of vertical federated learning is the union of features, which is applicable to scenarios where the sample overlap is large while the feature overlap is small. For example, between a supermarket and a bank in the same region, the reached users are all residents of this region (samples are the same), but their businesses are different (features are different). When the feature and sample overlaps among participants are both small, federated transfer learning can be used to apply the model learned in the source domain to the target domain by leveraging the similarity among data, tasks, or models. For example, the collaboration between a bank and a supermarket in different regions. The embodiments of the present disclosure perform horizontal federated learning for scenarios where the data sets owned by each participant have the same feature space and different sample spaces. Its advantage lies in that it can increase the amount of data participating in training and only interact with the learning model, thereby protecting the data privacy and security of the participants to a certain extent.
[0039] In addition, the federated integrated learning method of the present disclosure can also be based on the differential privacy method. Differential Privacy (DP) is a mechanism to prevent differential attacks to protect user data privacy. Its purpose is to make the probability of obtaining the same result through model inference for two datasets that differ by only one record very close, removing individual characteristics while retaining statistical characteristics to protect user privacy. In federated learning, the model training algorithm does not distinguish between general characteristics and individual characteristics. Therefore, the trained model may inadvertently reveal the individual characteristics in the training set, and malicious attackers may obtain user privacy information from the model. Therefore, it is very necessary to use differential privacy technology to protect the training model. In applying differential privacy for privacy protection, the data to be processed can be mainly divided into two categories: numerical and non-numerical. Among them, for numerical data, the Laplace or Gaussian mechanism is generally used, and random noise is added to the obtained numerical result to achieve differential privacy; for non-numerical data, the exponential mechanism is generally used and a scoring function is introduced to obtain a score for each possible output, and after normalization, it is used as the probability value of the query return. It does not deterministically output a specific result after receiving the query, but returns the result with a certain probability value, thereby achieving differential privacy. Among them, this probability value is determined by the scoring function. The higher the score, the higher the probability of the result being output, and the lower the score, the lower the probability of being output. The federated integrated learning method of the present disclosure can implement differential privacy based on the exponential mechanism to protect the data privacy of each participating party while effectively training a better model.
[0040] In summary, the solutions provided by the embodiments of the present disclosure involve technologies such as artificial intelligence, federated learning, and differential privacy. The embodiments of the present disclosure will be further described below with reference to the accompanying drawings.
[0041] Figure 1 It is a schematic diagram showing horizontal federated learning based on cryptography according to an embodiment of the present disclosure.
[0042] Optionally, Figure 1 Both the shown horizontal federated learning system and the horizontal federated learning system of the present application can include K participating parties ( Figure 1 shown as participating party 0 to participating party (K - 1)). Among them, each participating party can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. Each participating party can have the same feature space and different sample spaces.
[0043] The core idea of traditional horizontal federated learning is that each participant trains a model locally using its own data, and then obtains a better global model through secure model aggregation based on cryptography (such as secure model parameter averaging, also known as federated averaging), or performs secure model aggregation based on secret sharing (i.e., encrypting using masks), or perturbs the generative model using the random gradient descent method based on the Gaussian mechanism to protect data security and privacy.
[0044] As Figure 1 shown, taking cryptography-based horizontal federated learning as an example, after each participant completes local model training, it can send encrypted gradients to the federated server through step ① for secure model aggregation at the federated server in step ②. Next, in step ③, the federated server can distribute the aggregated encrypted gradients back to each participant for each participant to decrypt the gradient and update the model locally (step ④), so as to train a better global model without revealing the data privacy of each participant.
[0045] However, in the model aggregation method based on cryptography (or secret sharing), the federated server and each participant cannot obtain the local models (or model parameters) of the participants in plaintext form, but can only obtain the encrypted model parameters. However, for the encrypted local model parameters, the federated server can usually only use the secure aggregation algorithm (i.e., the secure federated averaging algorithm) for model aggregation. In the secure aggregation algorithm, each participant needs to interact with the federated server multiple times to transfer model parameters, which has high requirements for network communication bandwidth and stability, and consumes a large amount of computing and communication overhead. In addition, the model aggregation method based on secure federated averaging usually requires at least three or more participants and does not support scenarios with only two participants.
[0046] In addition, in the existing random gradient descent model scheme based on the Gaussian mechanism, in the case of unbalanced data distribution among participants, it is difficult to learn the feature distribution of global data and data preprocessing methods such as upsampling or downsampling cannot be used, resulting in difficulty in generating effective model parameters. In this case, due to the poor model effects provided by each participant, the performance of the model aggregated by the federated server is difficult to guarantee.
[0047] Therefore, in view of the above problems, the present disclosure provides a federated ensemble learning method, which protects the local data of participants based on differential privacy of the exponential mechanism, eliminates the risk of data leakage and avoids the above problems.
[0048] Compared with traditional federated learning methods based on cryptography or secret sharing, the method provided by the embodiments of the present disclosure can send the trained model parameters to the federated server in plaintext. There is only one message interaction between each participant and the federated server, and the amount of transmitted data is small, significantly reducing the requirements for the communication network.
[0049] Compared with traditional federated learning methods based on the stochastic gradient descent model, the method provided by the embodiments of the present disclosure can train effective model parameters in the case of unbalanced participant data distributions, improving the stability of the federated learning method.
[0050] The method provided by the embodiments of the present disclosure performs feature selection and model selection based on differential privacy of the exponential mechanism locally at each participant, and sends the selected training model to the federated server for integrated fusion, thereby generating a federated integrated model with better performance. Through the method of the embodiments of the present disclosure, the parameters of the selected training model can be sent to the federated server in plaintext without using any cryptographic method, avoiding the problem of ciphertext expansion based on cryptographic methods, and thus achieving more efficient and low-communication-overhead federated learning while ensuring no risk of data leakage.
[0051] In addition, the method provided by the embodiments of the present disclosure can also support scenarios with only two participants by directly transmitting the training model between the participants, and support direct communication and model fusion between multiple participants without a federated server.
[0052] Figure 2A is a flowchart showing a federated integrated learning method 200 according to an embodiment of the present disclosure. Figure 2B is a schematic diagram showing a federated integrated learning method 200 according to an embodiment of the present disclosure.
[0053] As Figure 2A shown, in step 201, a first number of features can be selected from the feature set of the participant according to a first probability distribution, and the first probability distribution can be obtained based on the exponential mechanism for the feature set of the participant.
[0054] Optionally, each participant can perform feature selection locally using the training data it locally owns. Among them, the feature set of the participant can be a discrete data set after one-hot encoding, and each feature in the feature set can be a data taking values from discrete data of 0 and 1.
[0055] Optionally, taking a binary classification task (i.e., classification into {0, 1}) as an example, for each sample involved in the training data owned by each participating party locally, it can have a class label of 0 or 1, and this class label can have a certain correlation with the respective features of this sample.
[0056] For example, for a task of classifying by gender, the classification results can include male (class label 0) and female (class label 1), and for the feature data included in each sample (here each user or individual), such as the feature data regarding hair length and height, the values thereof may affect the classification results, and this influence can reflect the magnitude of the correlation between this feature and the classification results (i.e., the determined class label). Among them, this correlation can be divided into positive correlation and negative correlation. For example, for the feature of hair length, assuming that this feature takes the value of 0 within a predetermined smaller length value interval and takes the value of 1 within a predetermined larger length value interval, so in the case of the assumption that the hair length value of the sample with the class label 0 (male) is usually less than the hair length value of the sample with the class label 1 (female), it can be considered that the feature of hair length is positively correlated with the class label 0 and negatively correlated with the class label 1.
[0057] It should be understood that the above binary classification task is only used as an example rather than a limitation in the present disclosure. For a multi-classification task, it can be achieved by converting the multi-classification task into a binary classification task. For example, in the case where there are three classifications {A, B, C} in a multi-classification task, it can be converted into three "binary classification" tasks to achieve, that is: ① the binary classification task of A and {B, C}, ② the binary classification task of B and {A, C}, and ③ the binary classification task of C and {A, B}. The present disclosure does not limit the specific number of classes of the classification task.
[0058] Optionally, when each participating party performs feature selection locally, it can use differential privacy implemented based on the exponential mechanism to protect the training data. Specifically, step 201 can include the steps as Figure 3A shown.
[0059] Figure 3A is a flowchart showing feature selection based on the exponential mechanism according to an embodiment of the present disclosure. Figure 3B is a schematic diagram showing feature selection based on the exponential mechanism according to an embodiment of the present disclosure.
[0060] As Figure 3A shown, in step 2011, for each feature in the feature set of the participating party, determine the feature score of this feature, and the feature score is determined based on the correlation between this feature and the class labels corresponding to multiple samples of the participating party.
[0061] According to an embodiment of the present disclosure, each of the multiple samples of the participating parties may correspond to one of two class labels. For each of the features, the relevance between the feature and the two class labels may be indicated by the flip flag of the feature. That is, the relevance between the feature and the class label as described above may be indicated by this flip flag.
[0062] According to an embodiment of the present disclosure, the case where the flip flag is the first value indicates that the feature is co-directionally relevant to the first class label among the two class labels, and the feature score of the feature includes a first feature score determined based on the co-directional relevance between the feature and the first class label.
[0063] According to an embodiment of the present disclosure, the case where the flip flag is the second value indicates that the feature is anti-directionally relevant to the first class label among the two class labels, and the feature score of the feature includes a second feature score determined based on the anti-directional relevance between the feature and the first class label.
[0064] Similar to the above, the flip flag can indicate the co-directional relevance or anti-directional relevance between the feature and the class label of the sample, and the magnitude of the co-directional relevance or anti-directional relevance can be determined based on the sum of the distances between the class labels of the multiple samples of the participating parties and the value of the feature.
[0065] For example, assume the feature set of the participating party The class label set is {0, 1}. For the nth feature among them, assume its flip flag is q n (q n = 0 or 1), where q n = 0 indicates that the nth feature is co-directionally relevant to the first class label 0 among the two class labels (i.e., anti-directionally relevant to the second class label 1 among the two class labels), while q n = 1 indicates that the nth feature is anti-directionally relevant to the first class label 0 among the two class labels (i.e., co-directionally relevant to the second class label 1 among the two class labels). Therefore, a scoring function for implementing differential privacy based on the exponential mechanism can be constructed using it. That is, in the case of q n = 0, the first feature score is And in the case of q n = 1, the second feature score is Where X m,n represents the value of the nth feature of the mth data (the mth sample), y m represents the class label (0 or 1) of the mth data (the mth sample), represents outputting 1 when X m,n = y m and outputting 0 otherwise, and represents when 1 - Xm,n = y m Output 1 when it is satisfied, otherwise output 0. |I| represents the number of features included in the feature set.
[0066] As Figure 3B shown, after determining the feature scores of each feature in the feature set of the participating party, the probability of each feature being selected can be determined based on all the feature scores to form a complete feature selection probability distribution.
[0067] In step 2012, based on the exponential mechanism, the first probability distribution is determined according to the feature scores of each feature in the feature set of the participating party, and the first probability distribution includes the probability of each feature in the feature set of the participating party being selected.
[0068] According to an embodiment of the present disclosure, step 2012 may include: based on the first feature scores and the second feature scores of all features in the feature set of the participating party, for each feature in the feature set of the participating party, the probability of the feature being selected is determined based on the exponential mechanism, and the probability includes a co-directional probability associated with the first feature score and an anti-directional probability associated with the second feature score. The first probability distribution includes the co-directional probability and the anti-directional probability of each feature in the feature set of the participating party being selected.
[0069] As described above, the exponential mechanism uses a scoring function to introduce a certain degree of randomness when answering queries, so that the outputs have the same probability to a certain extent, thereby ensuring ε-differential privacy. Among them, the scoring function u maps the dataset output pair (x, r) to a score u(x, r) (x is the input dataset, and r is the output result). The exponential mechanism will output each possible output r with a certain probability, and the probability is proportional to exp(εu(x, r) / Δu), where Δu represents the sensitivity, and ε represents the privacy loss metric. Then the probability that the output o = r can be expressed as
[0070]
[0071] where Therefore, for any adjacent datasets x and x′, they have a certain probability of outputting the same result, and this probability is less than or equal to e ε . Therefore, it is very difficult for an observer to detect the subtle changes in the dataset by observing the output parameters, and it is impossible to reverse-infer a specific training data from the observed output parameters. In this way, the purpose of protecting data privacy can be achieved.
[0072] Optionally, a participant may select a first number of features (e.g., L features) from its feature set based on the exponential mechanism. Assuming the privacy budget for feature selection is ε1, the privacy cost for each feature selection can be an average allocation of the privacy budget, that is, each feature selection will consume of the privacy cost. In addition, the privacy cost for each feature selection may also vary depending on the specific feature. For example, features that have a greater impact on the model output results or are more likely to disclose user individual privacy (such as individual features) may consume a higher privacy cost when selected for model training. Therefore, a higher proportion of the privacy budget can be allocated to them. In the present disclosure, only the average privacy budget allocation is described as an example rather than a limitation, and other forms of representing privacy costs can also be applicable.
[0073] Therefore, according to the definition of the exponential mechanism above, the probability of a specific feature being selected can be calculated based on the feature scores of each feature in the participant's feature set. This probability includes the same-direction probability associated with the first feature score (i.e., when q n = 0) and the reverse probability associated with the second feature score (i.e., when q n = 1). Optionally, taking the nth feature as an example, assuming the sensitivity Δu = 1, its same-direction probability θ n (i.e., the probability of selecting the nth feature) can be expressed as:
[0074]
[0075] And its reverse probability θ |I|+n (i.e., the probability of selecting the flipped feature corresponding to the nth feature, where the flipped feature corresponding to the nth feature can be a feature of the same feature type as the nth feature but with the opposite correlation to the class label, which can be indicated by a flip flag) can be expressed as:
[0076]
[0077] Therefore, the same-direction probability and reverse probability of each feature in the participant's feature set being selected can be calculated as described above to form a first probability distribution.
[0078] In step 2013, a feature can be selected from the participant's feature set according to the first probability distribution.
[0079] According to an embodiment of the present disclosure, step 2013 may include: selecting a feature from the participant's feature set according to the first probability distribution and determining the flip flag of the feature.
[0080] As described above, the sum of the probabilities of all features in the feature set of the participating party being selected is 1. The probability of each feature being selected includes its same-direction probability and reverse-direction probability, and the same-direction probability and reverse-direction probability respectively correspond to different feature scores (the first feature score and the second feature score) of the same feature and the flip flag (q n = 0 and q n = 1). Therefore, when selecting a feature from the feature set based on the first probability distribution, not only the index of the feature is determined, but also the flip flag of the feature is determined.
[0081] In step 2014, the selected feature can be removed from the feature set of the participating party to update the feature set of the participating party and the first probability distribution, and continue to select features based on the updated feature set and the first probability distribution until the total number of the selected features reaches the first number.
[0082] Optionally, after each feature is selected, it can be removed from the feature set I, and the first probability distribution can also be updated as shown in formulas (2) and (3). Therefore, new features can be continuously selected based on the updated feature set I and the first probability distribution as described in step 2013 until the number of the selected features reaches the first number, so as to select the first number of features.
[0083] Therefore, as Figure 3B shown, after each feature selection is completed, it can be determined whether the number of the selected features reaches the predetermined first number. If the number of the selected features is insufficient, the feature set I and the first probability distribution can be updated to continue the feature selection, otherwise the first number of the selected features is obtained for the following local model training.
[0084] Next, returning to Figure 2A , in step 202, a plurality of logistic regression models can be obtained based on at least some of the first number of selected features, and a second number of logistic regression models can be selected from the plurality of logistic regression models according to the second probability distribution, where the second probability distribution is obtained based on the exponential mechanism for the plurality of logistic regression models.
[0085] Obtaining a plurality of logistic regression models based on at least some of the first number of selected features in step 202 can include the steps as Figure 4A shown. Selecting a second number of logistic regression models from the plurality of logistic regression models according to the second probability distribution in step 202 can include the steps as Figure 4B shown.
[0086] Figure 4A is a flowchart showing the model construction according to an embodiment of the present disclosure. Figure 4BIt is a flowchart showing model selection according to an embodiment of the present disclosure. Figure 4C It is a schematic diagram showing model construction and selection according to an embodiment of the present disclosure.
[0087] As Figure 4A shown, in step 20211, the optimal feature among the first number of features and its corresponding one-dimensional logistic regression model can be determined, and the feature score of the optimal feature is not less than the feature scores of other features among the first number of features.
[0088] Optionally, among the selected first number of features, the optimal feature may have the highest feature score, and a one-dimensional logistic regression model can be independently trained based on this optimal feature.
[0089] As Figure 4C shown, after completing the optimal feature selection and one-dimensional logistic regression model construction, single-group model selection can be entered. During this single-group model selection process, one round of model construction and selection can be performed to obtain a second number of logistic regression models.
[0090] In step 20212, a predetermined number of features can be randomly selected from the other features among the first number of features.
[0091] Optionally, randomly selecting a predetermined number of features from the other features among the first number of features can be, for example, a selection based on equal probability, that is, the probability of each of these other features being selected is equal. In addition, the above random selection can also be based on other selection methods, and the present disclosure does not limit this.
[0092] Optionally, after completing the local feature selection, the participating party can use the first number of features (such as L features) selected locally by it to construct a training model. For example, for a logistic regression model with dimension D + 1, several ( Figure 4C shown as D in
[0093] Note that the determination and use of the above optimal feature f can be optional, and the method of the present disclosure can also be carried out without using the optimal feature f.
[0094] In step 20213, for the optimal feature and the selected predetermined number of features, multiple logistic regression models may be constructed based on a predetermined weight value space, and the number of the multiple logistic regression models is related to the predetermined number and the number of weight values included in the predetermined weight value space.
[0095] Optionally, for the logistic regression model of the present disclosure, its output result is expressed as a weighted output of weights and features, where the weight values are discrete values and the dimension cannot be too large (preferably, less than or equal to 6). Therefore, assuming that the weight value space of the logistic regression model of the present disclosure is V = {0, 0.1, 0.2, 0.3, 0.4, 0.5}, for a set of feature sets (such as the above D + 1 features), T = |V| D+1 logistic regression models can be obtained, where |V| represents the number of weights included in the weight value space V (in this example, |V| = 6). For the sake of convenience of expression, the set of the optimal feature and the selected predetermined number of features (i.e., the above D + 1 features) is hereinafter represented by S.
[0096] Optionally, the weight corresponding to the optimal feature f may be set to a fixed value (for example, 0.5). Therefore, by the method of enumeration, for a set of feature sets (such as the above (D + 1) features), T = |V| D logistic regression models can be obtained.
[0097] Therefore, after the model construction in the single-group model selection as shown in Figure 4C is implemented based on the above steps 20212 and 20213, the model selection process in the single-group model selection can be entered.
[0098] As described above, the step of selecting the second number of logistic regression models from the multiple logistic regression models according to the second probability distribution in step 202 may include the steps as shown in Figure 4B .
[0099] As shown in Figure 4B , in step 20221, based on the prediction results of the class labels corresponding to the multiple samples of the participating party by the multiple logistic regression models, the model score of each logistic regression model in the multiple logistic regression models may be determined.
[0100] Optionally, the model score of each logistic regression model in the multiple logistic regression models may be determined according to the degree of consistency between the prediction results of each logistic regression model for all the samples of the participating party and the class labels corresponding to these samples respectively.
[0101] For example, the prediction result of the classification of the m-th sample of the participating party by the i-th logistic regression model among the multiple constructed logistic regression models Can be expressed as:
[0102]
[0103] Where z i,m Represents the output value of the i-th logistic regression model for the m-th sample, and Where S[d] represents the index of the selected feature, q S[a] Represents the flip flag of the selected feature, w i,d ∈V represents the weight of the logistic regression model. Therefore, Represents when z i,m ≥0.5, That is, it is determined that the class label of the m-th sample is 1, otherwise That is, it is determined that the class label of the m-th sample is 0.
[0104] Therefore, optionally, the model score (scoring function) H of the i-th logistic regression model i Can be calculated through the prediction result That is
[0105]
[0106] Where Represents that when the prediction result of the i-th logistic regression model for the m-th sample of the participating party Is the same (consistent) as the class label corresponding to this sample, the output is 1, otherwise the output is 0. Therefore, Reflects the degree of consistency between the prediction results of the i-th logistic regression model for all samples (M samples) of the participating party and the class labels corresponding to these samples respectively.
[0107] In step 20222, according to the model scores of each of the multiple logistic regression models, the second probability distribution is determined based on the exponential mechanism, and the second probability distribution includes the probability that each of the multiple logistic regression models is selected.
[0108] Optionally, similar to that described above with reference to step 2012 during feature selection, the probability that the i-th logistic regression model is selected determined based on the exponential mechanism can be expressed as:
[0109]
[0110] Where J = {1, 2,..., T}, ε2 represents the privacy budget for model selection, Indicates the privacy cost consumed by the current model selection. Here, G represents repeating the operation of randomly selecting D features from the first quantity of features and selecting the second quantity (S) of LR models G times to obtain G*S + 1 logistic regression models (including the one-dimensional logistic regression model corresponding to the above optimal features), specifically as described in the following reference step 20225.
[0111] In step 20223, according to the second probability distribution, select one logistic regression model from the multiple logistic regression models.
[0112] As described above, the probability of each logistic regression model in the determined multiple logistic regression models being selected forms a second probability distribution. Based on this second probability distribution, model selection can be performed from the multiple logistic regression models with a specific probability, thus adding randomness to the model selection result and model training.
[0113] In step 20224, the selected logistic regression model can be removed from the multiple logistic regression models to update the multiple logistic regression models and the second probability distribution, and continue to select logistic regression models based on the updated multiple logistic regression models and second probability distribution until the total number of the selected logistic regression models reaches the second quantity.
[0114] Therefore, as Figure 4C shown, after each model selection is completed, it can be judged whether the number of the selected models reaches the predetermined second quantity. If the number of the selected models is insufficient, the multiple logistic regression models and the second probability distribution can be updated to continue the model selection; otherwise, the second quantity of logistic regression models is determined.
[0115] In addition, step 2022 may further include step 20225: randomly selecting a predetermined quantity of features from the other features among the first quantity of features a predetermined number of times; and obtaining multiple logistic regression models based on the predetermined quantity of features selected each time in the predetermined number of times, and selecting the second quantity of logistic regression models from the multiple logistic regression models according to the second probability distribution to obtain a third quantity of logistic regression models, where the third quantity is the product of the second quantity and the predetermined number of times, and the third quantity of logistic regression models may include multiple groups of the second quantity of logistic regression models with the predetermined number of times as the number of groups.
[0116] Optionally, in order to obtain a better fusion model, more logistic regression models can be generated based on the selected first quantity of features. Therefore, the operations such as Figure 4CThe single-group model selection shown selects different features from the first quantity of features multiple times for model construction, so as to obtain multiple groups of the second quantity of logistic regression models, that is, the third quantity of logistic regression models. For example, by repeating the random selection of D features and S logistic regression models G times, G*S logistic regression models with a dimension of (D+1) can be obtained.
[0117] Therefore, the participating party can obtain the third quantity of logistic regression models and a one-dimensional logistic regression model based on the selected first quantity of features, as Figure 4C shown.
[0118] Next, in step 203, the participating party can send at least a part of the second quantity of logistic regression models to the fusion end to perform integrated fusion based on the at least a part of the logistic regression models and generate a federated integrated model.
[0119] According to an embodiment of the present disclosure, sending at least a part of the second quantity of logistic regression models to the fusion end may include: for each group of the second quantity of logistic regression models in the third quantity of logistic regression models, determining a logistic regression model that is better than the determined one-dimensional logistic regression model from the second quantity of logistic regression models as the at least a part of the logistic regression models; and sending the at least a part of the logistic regression models and the one-dimensional logistic regression model in each group of the second quantity of logistic regression models in the third quantity of logistic regression models to the fusion end in plain text.
[0120] Among them, according to an embodiment of the present disclosure, sending a logistic regression model may include sending the feature index, flip flag, and model weight parameters of the features corresponding to the logistic regression model.
[0121] Optionally, sending a logistic regression model may further include a feature name. For horizontal federated learning, the feature spaces of each participating party are the same, but the sample spaces are different. Among them, each feature of the participating party has an index or a name, and the features of each participating party are aligned. Therefore, each participating party can determine the feature name based on the determined feature index.
[0122] Optionally, after each participating party selects the third quantity (G*S) of logistic regression models, it will only retain the models that are better than the one-dimensional logistic regression model. Therefore, the number of models actually sent by each participating party to the fusion end may be less than (G*S+1), but at least equal to 1, that is, at least send the one-dimensional logistic regression model corresponding to the optimal feature f.
[0123] According to an embodiment of the present disclosure, determining a logistic regression model that is better than the determined one-dimensional logistic regression model from the second number of logistic regression models is based on a comparison between the model scores of the second number of logistic regression models and the model score of the one-dimensional logistic regression model, where the model score of the one-dimensional logistic regression model is determined based on the prediction results of the one-dimensional logistic regression model for the class labels corresponding to multiple samples of the participating party.
[0124] Optionally, in the case where the model score of the logistic regression model is greater than the model score of the one-dimensional logistic regression model, it can be considered that this logistic regression model is better than the one-dimensional logistic regression model, and thus it is more conducive to the training of the federated integrated model.
[0125] Therefore, as described above, the federated integrated learning method 200 of the present disclosure can perform local feature selection and model selection based on the exponential mechanism to achieve differential privacy, thereby realizing the protection of the data privacy of the participating parties. Figure 2B Schematically sorts out the main steps in the federated integrated learning method 200 and the target selection involved.
[0126] As Figure 2B shown, first, for the input feature set of the participating party (shown as N-dimensional features), feature selection can be performed based on the exponential mechanism to achieve differential privacy, so as to select L features from the N-dimensional features and determine the optimal feature f and its corresponding one-dimensional logistic regression model.
[0127] Then, randomly select D features (D < L) for model construction from the selected L features, and use the selected D features together with the feature f to construct a logistic regression model, thereby forming a set of logistic regression models (T = |V| D+1 logistic regression models), where the weights of the models take discrete values, for example, take values from V = {0, 0.1, 0.2, 0.3, 0.4, 0.5}.
[0128] Next, model selection is performed for this set of logistic regression models based on the exponential mechanism to achieve differential privacy, so as to select K logistic regression models from this set of logistic regression models.
[0129] The above model construction and model selection processes can be repeated G times (i.e., generate G sets of models) to select G*K models based on the exponential mechanism.
[0130] Therefore, (G*K + 1) logistic regression models can be generated based on the exponential mechanism, and at least a part of these models can be selectively sent to the fusion end.
[0131] It should be understood that the fusion end in the present disclosure can be a fusion center for centralized fusion of models of all participating parties, such as a federated server, or a fusion end based on distributed fusion, such as another participating party. Although the present disclosure is mainly described in the manner of centralized fusion based on a fusion end such as a federated server, in the absence of a fusion center, the federated integrated learning method of the present disclosure is equally applicable.
[0132] Figure 5 is a flowchart showing a federated integrated learning method 500 according to an embodiment of the present disclosure.
[0133] As Figure 5 shown, in step 501, at least one logistic regression model can be received from multiple participating parties respectively.
[0134] According to an embodiment of the present disclosure, for each of the multiple participating parties, the at least one logistic regression model from the participating party may include at least a part of each set of the second number of logistic regression models among the third number of logistic regression models described above and the one-dimensional logistic regression model.
[0135] Optionally, after receiving the local models (or model parameters) sent by two (or more or all) participating parties, the fusion end (such as a federated server) can perform integrated fusion on the received local models.
[0136] In step 502, all logistic regression models from the multiple participating parties can be de-duplicated to remove duplicate logistic regression models.
[0137] Optionally, due to the existence of multiple participating parties and the use of the exponential mechanism to achieve differential privacy, there may be a situation where the selected models of each participating party are repeated. The repeated models may indicate that there are identical parts in the multiple models generated by multiple participating parties, and the identical parts have the same effect on the classification task. Therefore, only one of the identical parts needs to be retained during fusion.
[0138] Therefore, before the fusion end performs fusion, it is necessary to de-duplicate the repeated models, that is, only one of the two or more repeated models is retained.
[0139] In step 503, the de-duplicated logistic regression models can be integrated and fused to generate a federated integrated model.
[0140] According to an embodiment of the present disclosure, integrating and fusing the de-duplicated logistic regression models includes performing voting integrated fusion on the de-duplicated logistic regression models, and the prediction result of the generated federated integrated model is based on the average of the prediction results of the de-duplicated logistic regression models and the one-dimensional logistic regression model.
[0141] Optionally, a fusion end such as a federated server can perform Federated Voting on the received local models from the participating parties. This Federated Voting method can be used for the fusion of classification models. For example, for a binary classification model (positive class and negative class), the classification result of the federated integrated model can be determined by the average value of the classification results of the local models of the participating parties. For example, for a piece of data to be classified, if the average value of the classification results of the local models of the participating parties is greater than 0.5, the federated integrated model can determine its classification result as the "positive class"; conversely, if the average value of the classification results of the local models of the participating parties is less than 0.5, the federated integrated model can determine its classification result as the "negative class", and when the average value of the classification results of the local models of the participating parties is equal to 0.5, the classification result can be simply determined by random selection.
[0142] Figure 6A is a schematic diagram showing model fusion via a fusion center according to an embodiment of the present disclosure. Figure 6B is a schematic diagram showing model fusion without passing through a fusion center according to an embodiment of the present disclosure.
[0143] Figure 6A The scenario shown is the above-mentioned reference Figure 5 described federated integrated learning scenario, where the scenario includes K participating parties. After each participating party completes the above-mentioned feature selection and model selection based on the exponential mechanism locally, the global model is updated through the federated server. There is only one model (or model parameters) transmission from each participating party to the federated server, and it is a message transmission in plaintext for the above-mentioned (G*K + 1) models. Similarly, there is also only one model (or model parameters) transmission in plaintext from the federated server to each participating party to transmit the federated integrated model obtained by centralized fusion.
[0144] However, as Figure 6B shown, in the case of only two participating parties (for example, participating party 1 and participating party 2), the two participating parties can communicate directly without performing model fusion through the federated server.
[0145] In addition, further, when there are more (for example, K) participating parties, these participating parties can also communicate through a ring topology or a mesh topology (P2P) to perform distributed fusion to generate a global model without relying on the federated server.
[0146] Therefore, it should be understood that the fusion end in the present disclosure can be a fusion center for centralized fusion of models of all participating parties, such as a federated server, or a fusion end based on distributed fusion, such as another participating party. Although the present disclosure is mainly described in the manner of centralized fusion based on a fusion end such as a federated server, in the absence of a fusion center, the federated integrated learning method of the present disclosure is equally applicable.
[0147] Figure 7 FIG. 4 is a schematic diagram showing a federated integrated learning device 700 according to an embodiment of the present disclosure.
[0148] The federated integrated learning device 700 may include a feature selection module 701, a model selection module 702, and a model sending module 703.
[0149] According to an embodiment of the present disclosure, the feature selection module 701 may be configured to select a first number of features from the feature set of the participating parties according to a first probability distribution, where the first probability distribution is obtained based on the exponential mechanism for the feature set of the participating parties.
[0150] According to an embodiment of the present disclosure, the feature selection module 701 selecting a first number of features from the feature set of the participating parties according to a first probability distribution may include operations as described in Figure 3A wherein when each participating party performs feature selection locally, differential privacy implemented based on the exponential mechanism can be used to protect the training data.
[0151] The model selection module 702 may be configured to obtain a plurality of logistic regression models based on at least a part of the selected first number of features, and select a second number of logistic regression models from the plurality of logistic regression models according to a second probability distribution, where the second probability distribution is obtained based on the exponential mechanism for the plurality of logistic regression models.
[0152] According to an embodiment of the present disclosure, the model selection module 702 obtaining a plurality of logistic regression models based on at least a part of the selected first number of features includes operations as described in Figure 4A the reference.
[0153] Optionally, after completing local feature selection, the participating party may use the first number of features (e.g., L features) selected locally by it to construct a training model. For example, for a logistic regression model with a dimension of D + 1, several ( Figure 4C shown as D in FIG. 4) features may be randomly selected from the first number of features and jointly constructed with the optimal feature f among the first number of features to form a logistic regression model with discrete space weights.
[0154] The model selection module 702 may select a second number of logistic regression models from the multiple logistic regression models according to the second probability distribution, which may include operations as described in reference Figure 4B .
[0155] Optionally, the model score of each logistic regression model in the multiple logistic regression models may be determined according to the degree of consistency between the prediction results of all samples of the participating party by each logistic regression model and the class labels corresponding to these samples respectively.
[0156] Optionally, according to the model scores of each logistic regression model in the multiple logistic regression models, the probability of each logistic regression model being selected in the multiple logistic regression models may be determined based on the exponential mechanism, so as to form a second probability distribution. Based on this second probability distribution, model selection may be performed from the multiple logistic regression models with a specific probability, thereby adding randomness to the model selection result and model training.
[0157] Optionally, in order to obtain a better fusion model, more logistic regression models may be generated based on the selected first number of features. Therefore, according to an embodiment of the present disclosure, the model selection module 702 may also be configured to perform operations as described in reference step 20225, that is, repeat the single-group model selection as shown in Figure 4C to select different features from the first number of features for model construction multiple times, so as to obtain multiple groups of the second number of logistic regression models, that is, the third number of logistic regression models.
[0158] Therefore, the participating party may obtain the third number of logistic regression models and a one-dimensional logistic regression model based on the selected first number of features.
[0159] The model sending module 703 may be configured to send at least a part of the second number of logistic regression models to the fusion side to perform integrated fusion based on the at least a part of the logistic regression models and generate a federated integrated model.
[0160] The model sending module 703 sending at least a part of the second number of logistic regression models to the fusion side may include operations as described in reference step 203.
[0161] Wherein, according to an embodiment of the present disclosure, sending a logistic regression model may include sending the feature index, flip flag, and model weight parameters of the features corresponding to the logistic regression model. Optionally, sending a logistic regression model may also include the feature name.
[0162] According to another aspect of the present disclosure, a federated integrated learning device is further provided. Figure 8Shows a schematic diagram of a federated integrated learning device 2000 according to an embodiment of the present disclosure.
[0163] As Figure 8 shown, the federated integrated learning device 2000 may include one or more processors 2010 and one or more memories 2020. Among them, computer-readable code is stored in the memory 2020, and when the computer-readable code is run by the one or more processors 2010, the federated integrated learning method described above can be executed.
[0164] The processor in the embodiment of the present disclosure may be an integrated circuit chip with signal processing capabilities. The above-mentioned processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components. It can implement or execute the various methods, steps and logic block diagrams disclosed in the embodiments of the present application. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., and may be of the X86 architecture or the ARM architecture.
[0165] Generally speaking, various example embodiments of the present disclosure may be implemented in hardware or dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects may be implemented in hardware, while other aspects may be implemented in firmware or software that can be executed by a controller, a microprocessor, or other computing devices. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, devices, systems, technologies, or methods described herein may be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general-purpose hardware or controllers or other computing devices, or some combination thereof.
[0166] For example, the method or device according to an embodiment of the present disclosure may also be implemented by means of Figure 9 the architecture of the computing device 3000 shown. As Figure 9 shown, the computing device 3000 may include a bus 3010, one or more CPUs 3020, a read-only memory (ROM) 3030, a random access memory (RAM) 3040, a communication port 3050 connected to the network, an input / output component 3060, a hard disk 3070, etc. The storage device in the computing device 3000, such as the ROM 3030 or the hard disk 3070, may store various data or files used for the processing and / or communication of the federated integrated learning method provided by the present disclosure and the program instructions executed by the CPU. The computing device 3000 may also include a user interface 3080. Of course, Figure 8The architecture shown is only exemplary. When implementing different devices, one or more components in the computing device shown may be omitted according to actual needs. Figure 9 in the computing device shown.
[0167] According to another aspect of the present disclosure, a computer-readable storage medium is also provided. Figure 10 FIG. 4000 shows a schematic diagram of the storage medium according to the present disclosure.
[0168] As Figure 10 shown, computer-readable instructions 4010 are stored on the computer storage medium 4020. When the computer-readable instructions 4010 are run by a processor, the federated integrated learning method according to the embodiments of the present disclosure described with reference to the above drawings can be executed. The computer-readable storage medium in the embodiments of the present disclosure may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDR SDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous link dynamic random access memory (SLDRAM), and direct memory bus random access memory (DRRAM). It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories. It should be noted that the memories of the methods described herein are intended to include, but are not limited to, these and any other suitable types of memories.
[0169] Embodiments of the present disclosure also provide a computer program product or a computer program. The computer program product or the computer program includes computer instructions, and the computer instructions are stored in a computer-readable storage medium. The processor of the computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the federated integrated learning method according to the embodiments of the present disclosure.
[0170] Embodiments of the present disclosure provide a federated integrated learning method, apparatus, device, and computer-readable storage medium.
[0171] The method provided by the embodiments of the present disclosure, compared with the traditional federated learning methods based on cryptography or secret sharing, can send the trained model parameters to the federated server in plaintext. There is only one message interaction between each participating party and the federated server, and the amount of transmitted data is small, significantly reducing the requirements for the communication network.
[0172] The method provided by the embodiments of the present disclosure, compared with the traditional federated learning methods based on the stochastic gradient descent model, can train effective model parameters under the condition of unbalanced data distribution among participating parties, improving the stability of the federated learning method.
[0173] The method provided by the embodiments of the present disclosure performs feature selection and model selection based on differential privacy of the exponential mechanism locally at each participating party, and sends the selected training model to the federated server for integrated fusion, thereby generating a federated integrated model with better performance. Through the method of the embodiments of the present disclosure, the parameters of the selected training model can be sent to the federated server in plaintext without using any cryptographic methods, avoiding the problem of ciphertext expansion based on cryptographic methods, and thus realizing more efficient and low-communication-overhead federated learning while ensuring no data leakage risk. In addition, the method provided by the embodiments of the present disclosure can also support the scenario with only two participating parties by directly transmitting the training models between the participating parties, and support direct communication and model fusion among multiple participating parties without a federated server.
[0174] It should be noted that the flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains at least one executable instruction for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0175] In general, the various example embodiments of the present disclosure can be implemented in hardware or in dedicated circuits, software, firmware, logic, or any combination thereof. Some aspects can be implemented in hardware, while other aspects can be implemented in firmware or software that can be executed by a controller, microprocessor, or other computing device. When aspects of the embodiments of the present disclosure are illustrated or described as block diagrams, flowcharts, or using some other graphical representation, it will be understood that the blocks, apparatus, systems, techniques, or methods described herein can be implemented as non-limiting examples in hardware, software, firmware, dedicated circuits or logic, general hardware or controllers or other computing devices, or some combination thereof.
[0176] The example embodiments of the present disclosure described in detail above are merely illustrative and not restrictive. Those skilled in the art should understand that various modifications and combinations can be made to these embodiments or their features without departing from the principles and spirit of the present disclosure, and such modifications should fall within the scope of the present disclosure.
Claims
1. A federated integrated learning method, comprising: Select a first number of features from the feature set of the participating party according to a first probability distribution, where the first probability distribution is determined based on the exponential mechanism according to the correlation between each feature in the feature set of the participating party and the class labels corresponding to multiple samples of the participating party; Based on at least some of the selected first number of features, obtain a plurality of logistic regression models, and select a second number of logistic regression models from the plurality of logistic regression models according to a second probability distribution, where the second probability distribution is determined based on the exponential mechanism according to the model scores of each logistic regression model in the plurality of logistic regression models; And Send at least some of the second number of logistic regression models to the fusion end to perform integrated fusion based on the at least some of the logistic regression models and generate a federated integrated model.
2. The method according to claim 1, wherein, Selecting a first number of features from the feature set of the participating party according to the first probability distribution includes: For each feature in the feature set of the participating party, determine a feature score of the feature, where the feature score is determined based on the correlation between the feature and the class labels corresponding to multiple samples of the participating party; Based on the feature scores of each feature in the feature set of the participating party, determine the first probability distribution based on the exponential mechanism, where the first probability distribution includes the probability that each feature in the feature set of the participating party is selected; According to the first probability distribution, select one feature from the feature set of the participating party; and Remove the selected feature from the feature set of the participating party to update the feature set of the participating party and the first probability distribution, and continue to select features based on the updated feature set and first probability distribution until the total number of selected features reaches the first number.
3. The method according to claim 2, wherein, Each sample in the multiple samples of the participating party corresponds to one of two class labels. For each of the features, the correlation between the feature and the two class labels is indicated by a flip flag of the feature, where the case where the flip flag is a first value indicates that the feature is co-directionally correlated with the first class label among the two class labels, and the feature score of the feature includes a first feature score determined based on the co-directional correlation between the feature and the first class label; The case where the flip flag is a second value indicates that the feature is inversely correlated with the first class label among the two class labels, and the feature score of the feature includes a second feature score determined based on the inverse correlation between the feature and the first class label.
4. The method according to claim 3, wherein, Determining the first probability distribution based on the exponential mechanism according to the feature scores of each feature in the feature set of the participating party includes: According to the first feature scores and the second feature scores of all features in the feature set of the participating party, for each feature in the feature set of the participating party, determine the probability that the feature is selected based on the exponential mechanism, where the probability includes a co-directional probability associated with the first feature score and an inverse probability associated with the second feature score; Among them, selecting a feature from the feature set of the participant according to the first probability distribution includes: Selecting a feature from the feature set of the participant according to the first probability distribution, and determining a flip flag of the feature, where the first probability distribution includes a same-direction probability and a reverse probability of each feature in the feature set of the participant being selected.
5. The method according to claim 2, wherein, Obtaining a plurality of logistic regression models based on at least a part of the first quantity of selected features includes: Determining an optimal feature among the first quantity of features and its corresponding one-dimensional logistic regression model, where a feature score of the optimal feature is not less than feature scores of other features among the first quantity of features; Randomly selecting a predetermined quantity of features from the other features among the first quantity of features; and For the optimal feature and the selected predetermined quantity of features, constructing a plurality of logistic regression models based on a predetermined weight value space, where a quantity of the plurality of logistic regression models is related to the predetermined quantity and a quantity of weight values included in the predetermined weight value space.
6. The method according to claim 5, wherein, Selecting a second quantity of logistic regression models from the plurality of logistic regression models according to a second probability distribution includes: Based on prediction results of class labels corresponding to a plurality of samples of the participant by the plurality of logistic regression models, determining a model score of each logistic regression model among the plurality of logistic regression models; Based on the model scores of each logistic regression model among the plurality of logistic regression models, determining the second probability distribution based on the exponential mechanism, where the second probability distribution includes a probability of each logistic regression model among the plurality of logistic regression models being selected; and Selecting a logistic regression model from the plurality of logistic regression models according to the second probability distribution; and Removing the selected logistic regression model from the plurality of logistic regression models to update the plurality of logistic regression models and the second probability distribution, and continuing to select logistic regression models based on the updated plurality of logistic regression models and the second probability distribution until a total quantity of the selected logistic regression models reaches the second quantity.
7. The method according to claim 6, further comprising: Randomly selecting a predetermined quantity of features from the other features among the first quantity of features for a predetermined number of times; And Obtaining a plurality of logistic regression models based on the predetermined quantity of features selected each time among the predetermined number of times, and selecting a second quantity of logistic regression models from the plurality of logistic regression models according to the second probability distribution to obtain a third quantity of logistic regression models, where the third quantity is a product of the second quantity and the predetermined number of times, and the third quantity of logistic regression models includes multiple groups of the second quantity of logistic regression models with the predetermined number of times as groups.
8. The method according to claim 7, wherein, Sending at least a part of the second quantity of logistic regression models to a fusion end includes: For each group of the second quantity of logistic regression models among the third quantity of logistic regression models, determining a logistic regression model that is superior to the determined one-dimensional logistic regression model from the second quantity of logistic regression models as the at least a part of the logistic regression models; and Send each set of at least a portion of the second quantity of logistic regression models and the one-dimensional logistic regression model in the third quantity of logistic regression models to the fusion end in plaintext form. Among them, sending the logistic regression model includes sending the feature index, flip flag, and model weight parameters of the features corresponding to the logistic regression model.
9. The method according to claim 8, wherein, Determining a logistic regression model that is better than the determined one-dimensional logistic regression model from the second quantity of logistic regression models is based on the comparison of the model scores of the second quantity of logistic regression models and the model score of the one-dimensional logistic regression model. Among them, the model score of the one-dimensional logistic regression model is determined based on the prediction results of the one-dimensional logistic regression model for the class labels corresponding to multiple samples of the participating party.
10. A federated integrated learning method, comprising: Receive at least one logistic regression model from multiple participating parties respectively. Perform deduplication processing on all the logistic regression models from the multiple participating parties to remove duplicate logistic regression models. And Perform integrated fusion on the deduplicated logistic regression models to generate a federated integrated model. Among them, for each participating party among the multiple participating parties, the at least one logistic regression model from the participating party includes at least a portion of the second quantity of logistic regression models as described in claims 1-9.
11. The method according to claim 10, wherein, Performing integrated fusion on the deduplicated logistic regression models includes performing voting integrated fusion on the deduplicated logistic regression models, and the prediction result of the generated federated integrated model is based on the average of the prediction results of the deduplicated logistic regression models and the one-dimensional logistic regression model.
12. A federated integrated learning device, comprising: A feature selection module, configured to select a first quantity of features from the feature set of the participating party according to a first probability distribution, where the first probability distribution is determined based on the exponential mechanism according to the correlation between each feature in the feature set of the participating party and the class label corresponding to multiple samples of the participating party. A model selection module, configured to obtain multiple logistic regression models based on at least a portion of the selected first quantity of features, and select a second quantity of logistic regression models from the multiple logistic regression models according to a second probability distribution, where the second probability distribution is determined based on the exponential mechanism according to the model score of each logistic regression model in the multiple logistic regression models. And A model sending module, configured to send at least a portion of the second quantity of logistic regression models to the fusion end to perform integrated fusion based on the at least a portion of the logistic regression models and generate a federated integrated model.
13. The device according to claim 12, wherein, The feature selection module selecting a first quantity of features from the feature set of the participating party according to the first probability distribution includes: For each feature in the feature set of the participating party, determine the feature score of the feature, where the feature score is determined based on the correlation between the feature and the class label corresponding to multiple samples of the participating party. Based on the feature scores of each feature in the feature set of the participant, determine the first probability distribution based on the exponential mechanism, where the first probability distribution includes the probability of each feature in the feature set of the participant being selected; Based on the first probability distribution, select a feature from the feature set of the participant; and Remove the selected feature from the feature set of the participant to update the feature set of the participant and the first probability distribution, and continue to select features based on the updated feature set and first probability distribution until the total number of selected features reaches the first number.
14. The device according to claim 13, wherein, The model selection module obtains multiple logistic regression models based on at least a part of the first number of selected features, including: Determine the optimal feature among the first number of features and its corresponding one-dimensional logistic regression model, where the feature score of the optimal feature is not less than the feature scores of other features among the first number of features; Randomly select a predetermined number of features from the other features among the first number of features; and For the optimal feature and the selected predetermined number of features, construct multiple logistic regression models based on a predetermined weight value space, where the number of the multiple logistic regression models is related to the predetermined number and the number of weight values included in the predetermined weight value space.
15. The device according to claim 14, wherein, The model selection module selects a second number of logistic regression models from the multiple logistic regression models according to the second probability distribution, including: Based on the prediction results of the class labels corresponding to multiple samples of the participant by the multiple logistic regression models, determine the model score of each logistic regression model among the multiple logistic regression models; According to the model scores of each logistic regression model among the multiple logistic regression models, determine the second probability distribution based on the exponential mechanism, where the second probability distribution includes the probability of each logistic regression model among the multiple logistic regression models being selected; and According to the second probability distribution, select a logistic regression model from the multiple logistic regression models; and Remove the selected logistic regression model from the multiple logistic regression models to update the multiple logistic regression models and the second probability distribution, and continue to select logistic regression models based on the updated multiple logistic regression models and second probability distribution until the total number of selected logistic regression models reaches the second number.
16. The device according to claim 15, wherein, The model selection module is further configured to: Randomly select a predetermined number of features from the other features among the first number of features a predetermined number of times; and Obtain multiple logistic regression models based on the predetermined number of features selected each time in the predetermined number of times, and select a second number of logistic regression models from the multiple logistic regression models according to the second probability distribution to obtain a third number of logistic regression models, where the third number is the product of the second number and the predetermined number of times, and the third number of logistic regression models includes multiple groups of the second number of logistic regression models with the predetermined number of times as the number of groups.
17. The device according to claim 16, wherein, The model sending module sends at least a part of the second number of logistic regression models to the fusion end, including: For each group of the second number of logistic regression models among the third number of logistic regression models, determine a logistic regression model that is better than the determined one-dimensional logistic regression model from the second number of logistic regression models as the at least a part of the logistic regression models; and Send the at least a part of the logistic regression models and the one-dimensional logistic regression model in each group of the second number of logistic regression models among the third number of logistic regression models to the fusion end in plain text form, wherein, sending the logistic regression model includes sending the feature index, flip flag, and model weight parameters of the features corresponding to the logistic regression model.
18. A federated integrated learning device, comprising: One or more processors; And One or more memories, which store computer-executable programs, and when the computer-executable programs are executed by the processors, the methods described in any one of claims 1-11 are executed.
19. A computer program product, the computer program product comprising computer instructions which, when run by a processor, cause a computer device to perform the method according to any one of claims 1-11.
20. A computer-readable storage medium having stored thereon computer-executable instructions which, when executed by a processor, are used to implement the method according to any one of claims 1-11.
Citation Information
Patent Citations
Method and apparatus for generating a combined isolation forest model for detecting anomalies in data
US20210049517A1