Federal learning method and device

By protecting local model parameters on the client nodes, generating non-real privacy model parameters, and aggregating model parameters on the server nodes, the problem of client data leakage in traditional federated learning solutions is solved, and data privacy and security guarantees are achieved.

CN120217419APending Publication Date: 2025-06-27HUAWEI TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202311837234.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-27
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In traditional federated learning solutions, the communication between the server and the client is likely to lead to local data leakage of the client, which cannot effectively protect the privacy and security of the data.

Method used

By protecting the local model parameters on the client node, non-real privacy model parameters are generated, and model parameters aggregation is performed on the server node, ensuring that the server or client of the attack behavior cannot deduce the local data of other client nodes.

Benefits of technology

It effectively avoids the local data leakage of the client, ensures the privacy and security of the data, and prevents the server or client from deducing the local data of other client nodes through non-real model parameters.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120217419A_ABST
    Figure CN120217419A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning method and device, and the method comprises the steps: a server node transmits model training messages to N client nodes which execute federated learning tasks, and sending a model training message to a client node, so that the client node performs model training according to the federated learning initial model parameter and the federated learning configuration information in the model training message to obtain a real local model parameter, and then performs privacy protection processing on the local model parameter according to the privacy protection configuration information in the model training message to obtain a false privacy model parameter. And then the client node sends the false privacy model parameters to the server node, so that the server node executes the subsequent operation of the federation learning task according to the false privacy model parameters, thereby preventing a certain client node from reversely deducing the local data of other client nodes through the real model parameters. And thus, the privacy and security of the data of the user are ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technologies, and in particular, to a federated learning method and apparatus. Background Art

[0002] Federated Learning refers to a privacy-preserving distributed machine learning method that includes a server and multiple clients.

[0003] In the traditional solution, the server generates initial model parameters based on a service request and sends the initial model parameters to multiple clients participating in model training. Each client uses its local data as training data, trains its local model according to the training data and the initial model parameters, and sends the trained model parameters to the server. The server aggregates the model parameters from multiple clients to obtain the model parameters for the next round of iteration, and sends the model parameters to multiple clients for the next round of model training, thereby realizing the iterative training of federated learning.

[0004] During the communication process between the server and multiple clients, it is easy to cause the leakage of the local data of the clients, and the privacy and security of the local data cannot be guaranteed. Summary of the Invention

[0005] Embodiments of this application provide a federated learning method and apparatus to avoid the leakage of local data of clients and ensure the privacy and security of local data.

[0006] This application is applicable to a federated learning system that includes a server node and client nodes; among them, after obtaining the real local model parameters, the client nodes perform privacy protection processing on them to obtain non-real privacy model parameters. The server node is used to aggregate the non-real privacy model parameters corresponding to the client nodes, implement the federated learning task with non-real model parameters, so as to ensure that a server node or a client node with an attack behavior cannot deduce the local data of other client nodes through the non-real model parameters, avoid the leakage of the local data of other client nodes, and ensure the privacy and security of the local data.

[0007] In a first aspect, this application proposes a federated learning method, which is applied to a server node in a federated learning system. The following is an explanation with the server node as the execution entity. The method includes:

[0008] The server node sends a first model training message to the first client node among the N client nodes; where the N client nodes are used to perform a federated learning task, the first client node is any one of the N client nodes, and the first model training message is the model training message corresponding to the first client node, including the initial federated learning model parameters, the federated learning configuration information, and the privacy protection configuration information. Further, the initial federated learning model parameters and the federated learning configuration information are used by the first client node for model training, and the privacy protection configuration information is used by the first client node for privacy protection processing of the local model parameters. That is to say, after receiving the first model training message, the first client node performs model training based on the initial federated learning model parameters and the federated learning configuration information to obtain local model parameters, and then performs privacy protection processing on the local model parameters according to the privacy protection configuration information to obtain privacy model parameters. Based on this, the server node receives the privacy model parameters from the first client node, and then the server node aggregates the privacy model parameters of the N client nodes to obtain model aggregation parameters for federated learning iteration, so as to implement subsequent federated learning iterations until the federated learning task is completed.

[0009] Through this solution, the server node sends a first model training message to the first client node to instruct the first client node to perform model training on its own local model based on the initial federated learning model parameters and the federated learning configuration information through its own local data to obtain real local model parameters; then perform privacy protection processing on the real local model parameters according to the privacy protection configuration information to obtain unreal local model parameters (i.e., privacy model parameters). Based on this, the model parameters received by the server node are unreal local model parameters. Even if the server node is a server node with an attack behavior, this server node cannot deduce the local data of the client node based on the unreal local model parameters. Moreover, after receiving the unreal local model parameters, the server node performs subsequent aggregation operations of the federated learning task according to the unreal model parameters, and the obtained model aggregation parameters are unreal. The client node with an attack behavior cannot reverse-deduce the local data of other client nodes through the unreal model aggregation parameters, thereby avoiding the leakage of the local data of other client nodes and ensuring the privacy and security of the local data.

[0010] In a possible design, the privacy protection configuration information includes privacy protection processing type information and privacy protection processing factors; where the privacy protection processing type information indicates the privacy protection processing method used by the first client node, and the privacy protection processing factors indicate the parameters used by the first client node for privacy protection processing of the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

[0011] In this way, the first client node can perform privacy protection processing on the locally trained model parameters according to the parameters indicated by the privacy protection processing factor and in accordance with the privacy protection processing method indicated by the privacy protection processing type information.

[0012] In a possible design, the server node determines the privacy protection processing factor corresponding to each of the N client nodes according to the privacy protection budget information supported by each client node among the N client nodes. Or rather, for the first client node among the N client nodes, the server node determines the privacy protection processing factor corresponding to the first client node according to the privacy protection budget information supported by the first client node. Among them, the privacy protection budget information supported by the first client node can represent the degree of decrease in model accuracy caused by performing privacy protection processing, that is, it represents the degree of decrease in the model accuracy corresponding to the privacy model parameters compared to the model accuracy corresponding to the local model parameters. In addition, the privacy protection budget information supported by the first client node can also represent the degree of protection of the local model parameters by the first client node when performing privacy protection processing. It can be understood that the degree of protection of the local model parameters by privacy protection processing is positively correlated with the degree of decrease in model accuracy caused by privacy protection processing; or rather, the greater the degree of protection of the local model parameters by privacy protection processing, the greater the degree of decrease in model accuracy caused by privacy protection processing.

[0013] Through this design, the privacy protection processing factor corresponding to each client node not only satisfies the privacy protection processing of each client node's own local model parameters, but also ensures that the degree of decrease in model accuracy caused by each client node performing privacy protection processing is within the allowable accuracy decrease range, avoiding the unavailability of the model parameters aggregated due to a large degree of decrease in model accuracy. Or rather, it makes the degree of decrease in model accuracy caused by each client node performing privacy protection processing smaller, avoiding an excessive accuracy impact on federated learning caused by the degree of decrease in model accuracy corresponding to each client node, and ensuring the availability of the model parameters after model aggregation.

[0014] Optionally, due to the difference in the privacy protection budget information supported by any two client nodes, the corresponding privacy protection processing factors are also different, thereby ensuring the rationality of the privacy protection processing factor corresponding to each client node.

[0015] In a possible design, the privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method; the privacy protection processing factor indicates the noise distribution parameters added to the model parameters when performing the first processing method on the local model parameters in each of the K rounds.

[0016] With this design, the first client node can add noise to the locally trained model parameters according to the noise distribution parameters indicated by the privacy protection processing factor, and then obtain the private model parameters, thereby achieving privacy protection for the locally trained model parameters.

[0017] Optionally, the first processing method includes one of the following noise addition methods: differential privacy noise addition method based on the Gaussian mechanism, differential privacy noise addition method based on the Laplace mechanism.

[0018] Optionally, the noise distribution parameters include one of the following parameters: differential privacy budget information corresponding to each of the K rounds, noise distribution parameters corresponding to each of the K rounds.

[0019] In a possible design, the privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method; the privacy protection processing factor indicates a threshold or a compression factor or a Top-k value used when performing the second processing method on the locally trained model parameters in each of the K rounds.

[0020] With this design, the first client node can select some model parameters (or model gradients) from the locally trained model parameters as private model parameters according to the threshold or compression factor or Top-k value indicated by the privacy protection processing factor, thereby achieving privacy protection for the locally trained model parameters. Among them, the threshold indicates the minimum value of the difference in model gradients between the current round and the previous round of locally trained model parameters; the compression factor indicates the number of model gradients selected from the locally trained model parameters; the Top-k value indicates the proportion of model gradients selected from the locally trained model parameters. Based on this, the second processing method includes a gradient compression processing method based on the Top-K threshold method.

[0021] In a possible design, the server node receives a federated learning service request message; wherein, the federated learning service request message includes privacy protection indication information, and the privacy protection indication information indicates that the client nodes executing the federated learning task need to perform privacy protection processing on the locally trained model parameters, and then the server node can determine N client nodes executing the federated learning task according to the privacy protection indication information.

[0022] With this design, it is ensured that the N client nodes are all client nodes capable of performing privacy protection processing on the locally trained model parameters; or rather, the N client nodes are all client nodes that can execute the federated learning task.

[0023] In a possible design, the federal service request message further includes model accuracy degradation range information; such that the determined N client nodes for performing the federated learning task not only have the privacy protection processing capability, but also the degree of model accuracy degradation caused by each client node performing privacy protection processing meets the model accuracy degradation range information required by the federated learning task.

[0024] In this design, the model accuracy degradation range information in the federal service request message indicates the allowable model accuracy degradation range during the federated learning task, ensuring that the degree of model accuracy degradation caused by the N client nodes performing privacy protection processing meets the requirements of the federated learning task, and thus ensuring the availability of the local models participating in model aggregation.

[0025] In a possible design, the server node obtains the configuration information of client nodes that meet the first condition by sending a first request message, where the first condition at least includes having the privacy protection processing capability and the model accuracy degradation range information indicating that the degree of model accuracy degradation caused by performing privacy protection processing meets the requirements of the federated learning task. Based on this, the server node receives the configuration information of N client nodes that meet the first condition.

[0026] In this design, the server node can receive the configuration information of N client nodes that meet the first condition in any one of the following two ways:

[0027] The first way is that the server node can send a first request message to a first device; where the first device is used to store the configuration information of client nodes, and thus the server node receives the configuration information of N client nodes that meet the first condition from the first device; where the configuration information corresponding to the N client nodes includes privacy protection processing capability information.

[0028] The second way is that the server node can separately send a first request message to at least two client nodes, and thus receive the configuration information from the N client nodes that meet the first condition among at least two client nodes.

[0029] Optionally, the first request message further includes one or more of the following information:

[0030] Privacy protection processing type information, which indicates the privacy protection processing method required by the federated learning task;

[0031] Threshold of the user privacy feature ratio, which indicates the minimum ratio of privacy data in the training data used by the client nodes performing the federated learning task;

[0032] A threshold of the dataset size, where the threshold of the dataset size indicates the minimum amount of training data used by the client nodes executing the federated learning task;

[0033] A threshold of the model training duration, where the threshold of the model training duration indicates the maximum model training duration per round for the client nodes executing the federated learning task;

[0034] A threshold of the privacy protection processing duration, where the threshold of the privacy protection processing duration indicates the maximum privacy protection processing duration per round for the client nodes executing the federated learning task.

[0035] In this design, based on the first client node among the N client nodes having the privacy protection processing ability and the information on the range of model accuracy degradation that meets the requirements of the federated learning task, the configuration information of the first client node further satisfies one or more of the following conditions:

[0036] The privacy protection processing type information of the first client node meets the privacy protection processing type information of the federated learning task;

[0037] The proportion of user privacy features of the first client node is greater than or equal to the threshold of the proportion of user privacy features;

[0038] The dataset size of the first client node is greater than or equal to the threshold of the dataset size;

[0039] The model training duration of the first client node is less than or equal to the threshold of the model training duration;

[0040] The privacy protection processing duration of the first client node is less than or equal to the threshold of the privacy protection processing duration;

[0041] In this way, the server node can achieve an optimal selection of the client nodes executing the federated learning task.

[0042] In a possible design, the federated learning configuration information further includes the number of learning rounds of the federated learning task; the number of learning rounds is determined by the server node according to the model training duration required by the federated learning task, and the model training duration and privacy protection processing duration of each client node among the N client nodes.

[0043] In this design, based on factors such as the amount of training data and hardware processing performance of each client node, the model training duration and privacy protection processing duration of each client node may be different. Based on this, the server node can group N client nodes as a set of client nodes, then calculate the total duration of the model training duration and privacy protection processing duration of each client node, and then determine the learning rounds corresponding to this set of client nodes according to the maximum total duration among the N total durations and the model training duration required by the federated learning task. That is to say, the learning rounds of the N client nodes are the same.

[0044] Optionally, the server node can also group the N client nodes according to the model training duration and privacy protection processing duration corresponding to the N client nodes to determine at least two groups of client nodes. Then, for any group of client nodes, the server node determines the learning rounds of this group of client nodes according to the model training duration required by the federated learning task and the model training duration and privacy protection processing duration corresponding to each client node in this group. That is to say, the learning rounds of this group of client nodes are the same, but the learning rounds of different groups of client nodes are different. In the scenario of at least two groups of client nodes, the server node can define the association relationship between the learning rounds of different groups of client nodes, and this association relationship indicates the timing when at least two groups of client nodes jointly participate in the model parameter aggregation operation. For example, the server node defines that the first group of client nodes participates in the model parameter aggregation operation together with the first group of client nodes in the 2n-th round and the first group of client nodes in the n-th round. Based on this, the server node can perform more rounds of model training for client nodes with shorter model training duration and privacy protection processing duration, and perform fewer rounds of model training for client nodes with longer model training duration and privacy protection processing duration, reducing the model training waiting time of client nodes with shorter model training duration and privacy protection processing duration, realizing flexible indication of client nodes to execute the federated learning task, and improving the efficiency of client nodes to execute the federated learning task.

[0045] In a possible design, after the server node receives the privacy model parameters from M client nodes among the N client nodes, it aggregates the privacy model parameters corresponding to the M client nodes to obtain model aggregation parameters for the next round of model training; then subtracts the first error from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error; and finally sends the model aggregation parameters after subtracting the first error to the M client nodes or the N client nodes respectively.

[0046] Reduce the difference between the non - real model aggregation parameters and the real model aggregation parameters through the first error in this design, increase the availability of the non - real model aggregation parameters, thereby reducing the degree of decline in model accuracy of the client node in the next round of model training, and ensuring the availability of the model parameters after model aggregation.

[0047] Optionally, the first error is determined according to the weights of the M client nodes and the noise estimation values of the M client nodes. The noise estimation value of the first client node among the M client nodes is determined by the server node according to the privacy protection budget information required by the federated learning task. That is to say, the first error can be determined based on the privacy protection processing factors corresponding to each client node. For example, if the parameters used by each client node for privacy protection processing are selected based on a Gaussian distribution interval, the first error can be the expectation of the Gaussian distribution interval. Based on this, the server node denoises the non - real model aggregation parameters through the first error based on the noise - adding method of the client nodes, thereby further reducing the difference between the non - real model aggregation parameters and the real model aggregation parameters.

[0048] In a possible design, after the server node receives the privacy model parameters from M client nodes among the N client nodes, it aggregates the privacy model parameters corresponding to the M client nodes (which can also be called model aggregation according to the privacy model parameters corresponding to the M client nodes) to obtain the model aggregation parameters for the next round; then it sends the model aggregation parameters to the M client nodes or the N client nodes respectively; where M is an integer less than or equal to N.

[0049] Through this design, the server node aggregates the non - real local model parameters (i.e., the privacy model parameters corresponding to the client nodes) corresponding to the M client nodes, thereby obtaining non - real model aggregation parameters, and then sends the non - real model parameters to the client nodes. Based on this, after a client node with an attack behavior receives the non - real model aggregation parameters from the server node, it cannot deduce the local data of other client nodes, thereby avoiding the leakage of the local data of other client nodes and ensuring the privacy and security of the local data.

[0050] In a possible design, after the server node receives the privacy model parameters from M client nodes among the N client nodes, it aggregates the privacy model parameters corresponding to the M client nodes to obtain the model aggregation parameters for the next round; then it performs privacy protection processing on the model aggregation parameters to obtain the privacy - protected model aggregation parameters; finally, it sends the privacy - protected model aggregation parameters to the M client nodes or the N client nodes respectively.

[0051] With this design, due to the privacy protection process for the non-genuine model aggregation parameters, the difference between the non-genuine model aggregation parameters and the genuine model aggregation parameters is further increased, and it is further prevented that the client nodes or server nodes with attack behaviors infer the local data of other client nodes.

[0052] In a possible design, the server node aggregates the privacy model parameters corresponding to the M client nodes according to the privacy model parameters corresponding to the M client nodes and the weights of the M client nodes.

[0053] In a possible design, the server node determines H client nodes from the M client nodes, and then aggregates the privacy model parameters corresponding to the H client nodes according to the privacy model parameters corresponding to the H client nodes and the weights of the H client nodes to obtain model aggregation parameters; wherein, the degree of decrease in the model accuracy of any client node among the H client nodes is less than or equal to the model accuracy decrease threshold, and H is a positive integer less than or equal to M.

[0054] In this design, by selecting client nodes with a relatively small degree of decrease in model accuracy for model aggregation, the difference between the non-genuine model aggregation parameters and the genuine model aggregation parameters after model aggregation is reduced, the usability of the non-genuine model aggregation parameters is increased, and the degree of decrease in the model accuracy of the client nodes in the next round of model training is reduced.

[0055] In a possible design, the server node receives model accuracy information from the M client nodes, and then determines the H client nodes among the M client nodes whose degree of decrease in model accuracy is less than or equal to the model accuracy decrease threshold; wherein, the model accuracy information indicates the degree of decrease in model accuracy caused by the privacy protection process of the client nodes for the locally trained model parameters.

[0056] In a possible design, the server node receives data difference information from the N client nodes, and then determines the weights of the N client nodes according to the data difference information corresponding to the N client nodes; wherein, the data difference information indicates the degree of difference between the training data in the client nodes and the training data corresponding to the initial model parameters of the federated learning; the data difference information is positively correlated with the weights, and the sum of the weights of the N client nodes is 1.

[0057] Through this design, client nodes with training data similar to the training data corresponding to the initial model parameters of federated learning account for a greater weight. Consequently, when model aggregation occurs, the private model parameters of these similar client nodes have a greater impact on the model aggregation parameters after model aggregation, improving the adaptability of the model aggregation parameters to the federated learning task.

[0058] In a second aspect, an embodiment of the present application proposes a federated learning method, which is applied to a client node in a federated learning system. The following will be described taking the first client node among at least two client nodes in the federated learning system as the execution subject. The method includes:

[0059] The first client node receives a first model training message from the server node; then, based on the federated learning initial model parameters and the federated learning configuration information in the first model training message, it performs model training to obtain local model parameters; then, it performs privacy protection processing on the local model parameters according to the privacy protection configuration information in the first model training message to obtain private model parameters; and finally, it sends the private model parameters to the server node.

[0060] In a possible design, the privacy protection configuration information includes the privacy protection processing type information corresponding to the first client node and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method that the first client node needs to use, and the privacy protection processing factor indicates the parameter used by the first client node to perform privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information; based on this, the first client node can perform privacy protection processing on the local model parameters according to the parameter indicated by the privacy protection processing factor based on the privacy protection processing method indicated by the privacy protection processing type information.

[0061] In a possible design, when the privacy protection processing type information indicates a first processing method, the first client node adds the noise of the current round to the local model parameters based on the noise addition method corresponding to the first processing method and according to the noise distribution parameters of the current round indicated by the privacy protection processing factor.

[0062] In a possible design, when the privacy protection processing type information indicates a second processing method, the first client node performs gradient compression on the local model parameters based on the gradient compression method corresponding to the second processing method and according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor.

[0063] In a possible design, the first client node sends model accuracy information to the server node, and notifies the server node of the degree of model accuracy degradation caused by the first client node's privacy protection processing of the local model parameters through the model accuracy information.

[0064] In a possible design, if the first client node receives a first request message from the server node, it feeds back its own configuration information to the server node.

[0065] Optionally, the configuration information of the first client node includes one or more of the following information:

[0066] The privacy protection processing type information of the first client node, indicating the method of privacy protection processing of the local model parameters supported by the first client node;

[0067] The privacy protection budget information of the first client node; indicating the degree of model accuracy degradation caused by the first client node's execution of privacy protection processing on the local model parameters;

[0068] The user privacy feature ratio of the first client node, indicating the proportion of private data in the training data used by the first client node to execute the federated learning task;

[0069] The dataset scale of the first client node, indicating the data volume of the training data used by the first client node to execute the federated learning task;

[0070] The model training duration of the first client node, indicating the model training duration of each round in the federated learning task executed by the first client node;

[0071] The privacy protection processing duration of the first client node, indicating the privacy protection processing duration of each round in the federated learning task indicated by the first client node;

[0072] In a possible design, the first client node registers its own configuration information with the first device; so that the server node can obtain the configuration information of the first client node from the first device.

[0073] Optionally, the configuration information of the first client node includes one or more of the following information:

[0074] The privacy protection processing type information of the first client node, indicating the method of privacy protection processing of the local model parameters supported by the first client node;

[0075] The privacy protection budget information of the first client node indicates the degree of model accuracy degradation caused by the first client node performing privacy protection processing on local model parameters;

[0076] The user privacy feature ratio of the first client node indicates the proportion of private data in the training data used by the first client node to perform the federated learning task;

[0077] The dataset scale of the first client node indicates the amount of training data used by the first client node to perform the federated learning task;

[0078] The model training duration of the first client node indicates the model training duration of each round in the federated learning task performed by the first client node;

[0079] The privacy protection processing duration of the first client node indicates the privacy protection processing duration of each round in the federated learning task indicated by the first client node.

[0080] In a third aspect, the present application provides a federated learning device, including units for performing each step in the above first aspect or second aspect. Optionally, the federated learning device may include a communication unit and a processing unit; the communication unit is used to receive and send data, and the processing unit is used to execute the method provided in the above first aspect or second aspect.

[0081] In a fourth aspect, an embodiment of the present application provides a federated learning system, including: a server node for performing the above first aspect and a first client node for performing the above second aspect.

[0082] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program runs on a computer, the computer is caused to execute the method described in any one of the above first aspect or second aspect.

[0083] In a sixth aspect, an embodiment of the present application provides a computer program product. When the computer program product runs on a computer, the computer is caused to execute the method described in any one of the above first aspect or second aspect.

[0084] In a seventh aspect, an embodiment of the present application provides a chip, the chip includes a processor and a data interface, and the processor reads instructions stored on a memory through the data interface and executes the method described in any one of the above first aspect or second aspect.

[0085] In a possible design, the chip may further include a memory, where instructions are stored, and the processor is configured to execute the instructions stored on the memory. When the instructions are executed, the processor is configured to perform the method described in any one of the first aspect or the second aspect above.

[0086] Based on the implementations provided in the above aspects, the embodiments of the present application can be further combined to provide more implementations.

[0087] The technical effects achievable in any one of the third aspect to the seventh aspect above can be correspondingly described with reference to the technical effects achievable in the first aspect and / or the second aspect above. Repeated descriptions will not be elaborated. BRIEF DESCRIPTION OF THE DRAWINGS

[0088] Figure 1 A schematic diagram of a federated learning scenario applicable to the embodiments of the present application;

[0089] Figure 2 A schematic diagram of a general process of federated learning applicable to the embodiments of the present application;

[0090] Figures 3A - 3C A schematic diagram of the system architecture applicable to the embodiments of the present application;

[0091] Figure 4 A method flow diagram for determining a client node for executing a federated learning task provided by the embodiments of the present application;

[0092] Figure 5 Another method flow diagram for determining a client node for executing a federated learning task provided by the embodiments of the present application;

[0093] Figure 6 A flow schematic diagram of a federated learning method provided by the embodiments of the present application;

[0094] Figure 7 A flow schematic diagram of a federated learning method provided by the embodiments of the present application;

[0095] Figure 8 A flow schematic diagram of a federated learning method provided by the embodiments of the present application;

[0096] Figure 9 A flow schematic diagram of a federated learning method provided by the embodiments of the present application;

[0097] Figure 10 A structural diagram of a federated learning device provided by the embodiments of the present application;

[0098] Figure 11 A structural diagram of a federated learning device provided by the embodiments of the present application. Detailed implementation manners

[0099] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. The terms used in the following embodiments are only for the purpose of describing specific embodiments, and are not intended to limit the present application. As used in the specification and claims of the present application, the singular forms "a", "an", "the", "above", "said", "this" are also intended to include the plural forms such as "one or more", unless the context clearly indicates otherwise. It should also be understood that in the embodiments of the present application, "one or more" means one, two or more than two; " / ", which describes the association relationship of associated objects, indicates that three relationships may exist; for example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after.

[0100] The reference to "one embodiment" or "some embodiments" etc. described in this specification means that a specific feature, structure or characteristic described in connection with the embodiment is included in one or more embodiments of the present application. Thus, the statements "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments" etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "comprise", "include", "have" and their variants all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0101] The "multiple" involved in the embodiments of the present application means greater than or equal to two. It should be noted that in the description of the embodiments of the present application, the terms "first", "second", etc. are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order.

[0102] For ease of understanding, first in combination with Figure 1 and Figure 2 , an exemplary description of the scenario and process of federated learning is given.

[0103] Referring to Figure 1 , federated learning includes a server node (Server) 110 and at least two client nodes (Client). Figure 1Taking client nodes 120a, 120b, and 120c as examples. As participants in federated learning, client nodes (120a, 120b, 120c) use their own local datasets as training data (or training samples) to perform local model training, obtain local model parameters, and send the obtained local model parameters to server node 110.

[0104] Server node 110, as the aggregator or coordinator of federated learning, is used to aggregate the model parameters of client nodes (120a, 120b, 120c). It can also be understood as aggregating the local models corresponding to client nodes (120a, 120b, 120c) based on the model parameters of client nodes (120a, 120b, 120c). Therefore, the process of server node 110 aggregating parameters can also be called model aggregation.

[0105] In some scenarios, the server node and client nodes have a model training logical function (ModelTraining Logical Function, MTLF). Therefore, the server node can also be called the central MTLF (or central MTLF), and the client node can also be called the local MTLF (or local MTLF). Optionally, in the federated learning framework in the scenario of the 3rd Generation Partnership Project (3GPP), the network data analytics function (Network Data Analytics Function, NWDAF) that supports federated learning includes central MTLF and local MTLF. Therefore, the server node can also be called the server network element (Server NWDAF), and the client node can also be called the client network element (Client NWDAF).

[0106] In the embodiments of this application, the model parameters obtained after training the client node are called local model parameters, or partial model parameters; the model parameters obtained after the server node performs model aggregation are called model aggregation parameters, or global model parameters.

[0107] Based on Figure 1 the architecture shown, the server node can be Server NWDAF, and the client node can be Client NWDAF. Figure 2 shows a method of federated learning. The general process of federated learning will be introduced below in combination with Figure 2 , where the network function (Network Function, NF), as the consumer that initiates the federated learning service request, triggers the federated learning process. This NF can be called the consumer NF, or the consumer NF for short.

[0108] S201: The consumer NF sends a federated learning service request to the Server NWDAF.

[0109] The federated learning service request is used to indicate the federated learning task to be executed. The type of the federated learning task can be horizontal federated learning, vertical federated learning, or federated transfer learning.

[0110] S202: The Server NWDAF sends the initial federated learning model parameters to the Client NWDAFs participating in model training. Based on Figure 1 the architecture shown, the number of Client NWDAFs participating in model training can be multiple.

[0111] The initial federated learning model parameters can be preset values based on experience, or obtained by the Server NWDAF through model training using training data. In a possible implementation, the training data used by the Server NWDAF can be preset or obtained based on the federated learning service request. The federated learning model used by the Server NWDAF can be a general machine learning model or a specific machine learning model constructed according to the requirements of the federated learning service request. Taking an image recognition task as an example, the federated learning model can be a convolutional neural network (CNN) model. For the convenience of understanding and distinction, the federated learning model used by the Server NWDAF can be called the global model.

[0112] S203: The Client NWDAF receives the initial federated learning model parameters from the Server NWDAF, and performs model training based on the initial federated learning model parameters to obtain local model parameters.

[0113] The federated learning model used by the Client NWDAF can be the same as the federated learning model used by the Server NWDAF, or the federated learning model used by the Client NWDAF is a partial model of the federated learning model used by the Server NWDAF. Based on this, the federated learning model used by the Client NWDAF can be called the local model or the partial model.

[0114] S204: The Client NWDAF sends the local model parameters to the Server NWDAF.

[0115] S205: The Server NWDAF receives the corresponding local model parameters from the Client NWDAF, and performs model aggregation based on the local model parameters of at least two Client NWDAFs to obtain model aggregation parameters.

[0116] S206: The Server NWDAF sends the status information of the federated learning to the consumer NF. Among them, the status information of the federated learning includes the current iteration round of the federated learning, the model aggregation parameters, etc.

[0117] S207: The server node sends model aggregation parameters to the client node so that the client node can perform the next round of iterative training based on the model aggregation parameters.

[0118] S208: The client node and the server node perform a round of training process until the federated learning ends. The training process of each round can refer to S203 to S207. Among them, the end of the federated learning can be indicated by the consumer NF, or the local model of the client node converges, or the global model of the server node converges.

[0119] In the above federated learning process, the client node and the server node directly interact with the local model parameters and the model aggregation parameters. Or rather, the server node receives the local model parameters from the client node, and each client node receives the model aggregation parameters from the server node. The model aggregation parameters received by each client node are the same. Based on this, the attacking nodes in the server node or each client node can deduce the local data of other client nodes according to the following attack process. Therefore, there is a risk of local data leakage of the client node in the above federated learning, and the privacy and security of the user's local data cannot be guaranteed.

[0120] Taking the server node as an attacker as an example, since the model of the server node is used as the global model, the server node can thus have the local models corresponding to the client nodes. The server node sends the model aggregation parameter W1 corresponding to the global model to the client node, and then receives the local model parameter W2 obtained by the client node through local model training according to the model aggregation parameter W1. Based on this, the server node can initialize a pair of input samples x` (or rather, training samples) and the label y` corresponding to the input samples, and then input the input samples x` and the label y` into the local model corresponding to the client node (at this time, the parameter of the local model is the model aggregation parameter W1), and perform model training to obtain the model parameter W3 after model training. Then the server node calculates the distance (or rather, similarity) between the model parameter W3 and the local model parameter W2. It can be understood that the greater the distance (or rather, similarity), the closer the input samples x` and the label y` are to the local data (training data) of this client node. Therefore, the server node can determine the input samples x` and the label y` corresponding to when the distance between the model parameter W3 and the local model parameter W2 satisfies the preset condition, and then regard the input samples x` and the label y` at this time as the local data of this client node. Or rather, the input samples x` and the label y` at this time are equivalent to the similar data of the local data of this client node, and thus are equivalent to deducing the local data of this client node.

[0121] Similarly, taking the client node as the attacker as an example, for the sake of easy understanding, the client node that executes the attack process is called the attacking client node. The attack process includes: the attacking client node initializes a pair of input samples x` (or training samples) and the labels y` corresponding to the input samples. Then the attacking client node receives the model aggregation parameter W4 from the server node. Then, the input samples x` and the labels y` are input into its own local model (at this time, the parameters of the local model are the model aggregation parameter W4) for model training to obtain the local model parameter W5. The attacking client node can receive the model aggregation parameter W6 corresponding to the next round of iteration from the server node again. Furthermore, the attacking client node can calculate the distance between the local model parameter W5 obtained by its own training and the model aggregation parameter W6. It can be understood that the smaller the distance between the local model parameter W5 and the model aggregation parameter W6, the closer the input samples x` and the labels y` are to the local data of other client nodes. The attacker adjusts the input samples x` and the labels y` to determine the input samples x` and the labels y` when the distance between the local model parameter W5 obtained by its own training and the model aggregation parameter W6 meets the preset condition. Then, the input samples x` and the labels y` at this time are used as the local data of other client nodes. Or rather, the input samples x` and the labels y` at this time are equivalent to the similar data of the local data of other client nodes, and thus are equivalent to deducing the local data of other client nodes. In a possible implementation manner, the input sample x` can be the normalized RSRP (Reference Signal Receiving Power) of 16 beams.

[0122] To solve the above problems, the embodiments of the present application provide a federated learning method for preventing the server node or the client node acting as an attacker from deducing the local data of other client nodes, thereby avoiding the leakage of the local data of other client nodes and ensuring the privacy and security of the local data of users. It can be understood that in this method, the method and the device are based on the same technical concept. Since the principles of the method and the device for solving problems are similar, the implementation of the device and the method can be referred to each other, and the repeated parts will not be described again.

[0123] This method can be applied to a federated learning system including a server node and at least two client nodes. Before explaining the embodiments of the present application in detail, the system architecture involved in the embodiments of the present application will be introduced first.

[0124] In a possible implementation manner, refer to Figure 3A, the embodiments of the present application can be applied to a federated learning system architecture under the core network domain of 3GPP. In this scenario, multiple NWDAF nodes implement federated learning. The NWDAF nodes can be core network elements, or NWDAF functions are deployed on the core network elements. Among them, the core network element serving as the server node is called Server NWDAF, and the core network element serving as the client node is called Client NWDAF.

[0125] In a possible implementation, refer to Figure 3B , the embodiments of the present application can be applied to a federated learning system architecture under an access network domain of 3GPP. In this scenario, the server node can be an element management system (EMS) with a machine learning training function (MLTF), and this server node is called Server EMS. The client node can be a base station, or a radio access network training logic function (RTLF) within the base station. Here, RAN is the English abbreviation for radio access network, that is, the radio access network. This client node is called Client RTLF, or Client gNB.

[0126] In a possible implementation, refer to Figure 3C , the embodiments of the present application can be applied to another federated learning system architecture under the access network domain of 3GPP. In this scenario, both the server node and the client node can be base stations, or RTLFs within the base stations. The server node is called Server RTLF, or Server gNB, and the client node is called Client RTLF, or Client gNB. Multiple RTLFs execute federated learning through Server RTLF and Client RTLF.

[0127] In the federated learning system architectures shown above in Figure 3A 、 Figure 3B 、 Figure 3C , the server node has the ability to aggregate models for federated learning, which is used to aggregate the local model parameters of the client nodes, or to perform model aggregation, and thus obtain model aggregation parameters; the client node has the ability to participate in federated learning, which is used to perform local model training based on local data to obtain local model parameters. It should be noted that the types of federated learning in the embodiments of the present application include but are not limited to: horizontal federated learning, vertical federated learning, or federated transfer learning. It should be understood that the above Figure 3A 、 Figure 3B 、Figure 3C The number of client nodes shown is only illustrative, and the number of client nodes is not limited herein.

[0128] In the embodiments of the present application, the client node has corresponding configuration information, and the configuration information may include one or more of the following information:

[0129] Privacy protection processing information, which is used to indicate whether the client node has the privacy protection processing ability, or in other words, whether the client node supports privacy protection processing for local model parameters. Exemplarily, the privacy protection processing information may be an indication information. When the value of the indication information is 1, it means that the client node has the privacy protection processing ability; when the value of the indication information is 0, it means that the client node does not have the privacy protection processing ability.

[0130] Privacy protection processing type information, which is used to indicate the way of privacy protection processing for local model parameters supported by the client node. Exemplarily, the ways of privacy protection processing for local model parameters include the first processing way and / or the second processing way; among them, the first processing way may be a noise addition way, such as a differential privacy noise addition way based on the Gaussian mechanism, a differential privacy noise addition way based on the Laplace mechanism, etc.; the second processing way may be a gradient compression way, including a gradient compression way based on the Top-K threshold method, etc. The present application does not limit the privacy protection processing ways supported by the client node.

[0131] Privacy protection budget information, which is used to indicate the degree of model accuracy degradation caused by the client node performing privacy protection processing, or in other words, to indicate the degree of protection of local model parameters by the client node performing privacy protection processing. Exemplarily, based on the first processing way, the privacy protection budget information of the client node may be the value range of the differential privacy budget ε in the differential privacy rule allowed by the client node; based on the second processing way, the privacy protection budget of the client node may be the value range of the threshold in the Top-K threshold method allowed by the client node, or the value range of Top-K, or the value range of the compression factor. It can be understood that the smaller the differential privacy budget ε, or the larger the threshold in the Top-K threshold method, or the smaller the Top-K value, or the smaller the compression factor, the greater the degree of model accuracy degradation caused by the client node performing privacy protection processing, or in other words, the higher the degree of protection of local model parameters by the client node performing privacy protection processing.

[0132] User privacy feature ratio, which is used to indicate the proportion of private data in the training data used by the client node to perform the federated learning task.

[0133] Dataset scale, which is used to indicate the data volume of the training data used by the client node to perform the federated learning task.

[0134] The model training duration is used to indicate the model training duration for each round in the client node's execution of the federated learning task. This model training duration is related to the hardware capabilities of the client node, the scale of the local dataset of the client node, etc.

[0135] The privacy protection processing duration is used to indicate the privacy protection processing duration for each round in the client node's execution of the federated learning task. This privacy protection processing duration is related to the hardware capabilities of the client node, the privacy protection processing methods supported by the client node, etc.

[0136] The hardware capabilities of the client node are used to indicate the performance of the client node in executing the federated learning task. Exemplarily, the hardware capabilities of the client node may include the processing capabilities of the CPU, the size of the memory, etc.

[0137] The identifier (Analytics ID) of the model analysis use case of the client node is used to indicate the local model use cases (or local models) supported by the client node. For example, the Analytics ID of the client node is used to indicate that the local model use cases supported by the client node include business experience analysis, end - user behavior analysis, terminal movement analysis, etc.

[0138] The service area of the client node is used to indicate the coverage range of the local model of the client node. For example, the service area of the client node is the cell ID, etc.

[0139] The federated learning type of the client node is used to indicate the federated learning types supported by the client node. For example, the federated learning types supported by the client node include horizontal federated learning, vertical federated learning, and / or federated transfer learning.

[0140] In a possible implementation, the above - mentioned configuration information is stored in the client node.

[0141] In another possible implementation, the client node can register the configuration information with a first device; where the first device can be a Network Repository Function (NRF), or other functional entities or network elements with data storage capabilities. Optionally, the client node sends a registration message to the first device, and this registration message includes the configuration information of the client node, and then the first device stores the configuration information of the client node.

[0142] In the embodiments of this application, before performing federated learning, the server node first selects the client nodes participating in the federated learning.

[0143] Refer to Figure 4 , which is a method flow for determining the client nodes executing the federated learning task provided by the embodiments of this application, and includes:

[0144] S401: The consumer NF sends a federated learning service request message to the server node, and the federated learning service request message includes relevant information for performing the federated learning task.

[0145] Optionally, the relevant information for performing the federated learning task may include requirement information corresponding to the federated learning task to indicate the requirements met by the federated learning task. The requirement information may include one or more of the following information (1) to (10):

[0146] (1) Privacy protection indication information, used to indicate that privacy protection processing is required when performing the federated learning task.

[0147] Optionally, the privacy protection indication information may be used to indicate that the client node performing the federated learning task needs to perform privacy protection processing on the local model parameters.

[0148] Optionally, the privacy protection indication information may also be used to indicate that the server node performing the federated learning task needs to perform privacy protection processing on the model aggregation parameters.

[0149] (2) Privacy protection processing type information, used to indicate the privacy protection processing method required by the federated learning task.

[0150] Optionally, the privacy protection processing type information may be used to indicate the method for the client node performing the federated learning task to perform privacy protection processing on the local model parameters.

[0151] Optionally, the privacy protection processing type information may also be used to indicate the method for the server node performing the federated learning task to perform privacy protection processing on the local model parameters.

[0152] (3) Model accuracy degradation range information, used to indicate the model accuracy degradation range required by the federated learning task. It can be understood as indicating the degree of model accuracy degradation allowed for performing the federated learning task.

[0153] Furthermore, the model accuracy degradation range information may include a threshold corresponding to the privacy protection processing method, indicating the maximum degree of model accuracy degradation caused by the privacy protection processing required by the federated learning task, or indicating the maximum degree of protection for the model parameters by the privacy protection processing required by the federated learning task.

[0154] (4) Threshold of the proportion of user privacy features, used to indicate the minimum proportion of privacy data in the training data used by the client node performing the federated learning task. Correspondingly, the proportion of user privacy features of the client nodes participating in the federated learning is greater than or equal to the threshold of the proportion of user privacy features.

[0155] (5) Threshold of the dataset size, which is used to indicate the minimum amount of training data used by the client nodes executing the federated learning task. Correspondingly, the amount of training data used by the client nodes participating in the federated learning is greater than or equal to the threshold of the dataset size.

[0156] (6) Threshold of the model training duration, which is used to indicate the maximum model training duration of each round of the client nodes executing the federated learning task. Correspondingly, the model training duration of each round of the client nodes participating in the federated learning is less than or equal to the threshold of the model training duration.

[0157] (7) Threshold of the privacy protection processing duration, which is used to indicate the maximum privacy protection processing duration of each round of the client nodes executing the federated learning task. Correspondingly, the duration of performing privacy protection processing on the local model parameters of the client nodes participating in the federated learning is less than or equal to the threshold of the privacy protection processing duration.

[0158] (8) Identification (Analytics ID) of the model analysis use case of the federated learning task, which is used to indicate that the client nodes executing the federated learning task need to support the Analytics ID of the federated learning task.

[0159] (9) Service area of the federated learning task, which is used to indicate that the service area of the client nodes executing the federated learning task is within the service area of the federated learning task.

[0160] (10) Federated learning type of the federated learning task, which is used to indicate that the client nodes executing the federated learning task support the federated learning type of the federated learning task.

[0161] S402: The server node sends a first request message to the first device, and the first request message is used to request to obtain the information of the client nodes executing the federated learning task. Among them, the information of the client nodes may include the identification information of the client nodes, and may also include all or part of the information in the configuration information of the client nodes. In this process, the first device is taken as an example of the NRF.

[0162] Optionally, the partial information may include the model training duration, the privacy protection processing duration, the privacy protection budget information, etc.

[0163] Optionally, the first request message may include the relevant information for executing the federated learning task. Optionally, the first request message may include one or more of the above information (1) to (10).

[0164] Optionally, the server node may also send information predefined for the federated learning task to the first device. For example, the information predefined for the federated learning task may include indication information of the service area of the server node.

[0165] S403: The first device sends a first response message to the server node; the first response message includes information of N client nodes for performing the federated learning task.

[0166] Optionally, the information of the N client nodes may include identification information corresponding to the N client nodes, and may also include all or part of the configuration information corresponding to the N client nodes. Optionally, the partial information may include model training duration, privacy protection processing duration, privacy protection budget information, etc.

[0167] After receiving the first request message, the first device selects client nodes that meet the corresponding requirements indicated by the requirement information based on the configuration information of the client nodes stored in the first device.

[0168] Exemplarily, the client nodes that meet the corresponding requirements indicated by the requirement information meet one or more of the following conditions:

[0169] (a) The client node has privacy protection processing capabilities; Exemplarily, the privacy protection processing information of the client node indicates that it has privacy protection processing capabilities.

[0170] (b) The way the client node performs privacy protection processing supports the privacy protection processing method required by the federated learning task indicated by the first request message.

[0171] (c) The privacy protection budget information of the client node meets the range of model accuracy degradation required by the federated learning task indicated by the first request message; Optionally, the degree of model accuracy degradation indicated by the privacy protection budget information of the client node is less than the maximum value of the model accuracy degradation range indicated by the first request message, or the degree of model accuracy degradation indicated by the privacy protection budget information of the client node is within the model accuracy degradation range indicated by the first request message. For example, if the range of the degree of model accuracy degradation indicated in the first request message is 0 - 5%, then the degree of model accuracy degradation indicated by the privacy protection budget information of the client node needs to be less than or equal to 5%.

[0172] (d) The proportion of user privacy features in the training data of the client node is greater than or equal to the threshold of the proportion of user privacy features indicated by the first request message.

[0173] (e) The model training duration of each round of the client node is less than or equal to the threshold of the model training duration indicated by the first request message.

[0174] (f) The duration for the client node to perform privacy protection processing on the local model parameters in each round is less than or equal to the threshold of the privacy protection processing duration indicated in the first request message.

[0175] (g) The Analytics ID supported by the client node includes the Analytics ID indicated in the first request message.

[0176] (h) The service area of the client node is within the service area of the federated learning task indicated in the first request message.

[0177] (i) The federated learning type supported by the client node supports the federated learning type of the federated learning task indicated in the first request message.

[0178] It can be understood that one or more of the above conditions (a) to (i) that the client node performing the federated learning task needs to meet correspond to the relevant information for performing the federated learning task indicated in the first request message. For example, if the first request message includes information (1), (2), (3), (4), then the client node performing the federated learning task needs to meet the above conditions (a), (b), (c), (d).

[0179] Based on the above description, the first device can determine N client nodes that meet the first request message, and then send the configuration information of the N client nodes to the server node.

[0180] Optionally, Figure 4 The described process can be applied to Figure 3A the system architecture shown.

[0181] Optionally, the privacy protection processing methods corresponding to the N client nodes are the same to facilitate subsequent aggregation calculations by the server node.

[0182] In a possible implementation, referring to Figure 5 , another method flow for determining the client node performing the federated learning task includes:

[0183] S501: The consumer NF sends a federated learning service request message to the server node; the federated learning service request message includes the relevant information for performing the federated learning task.

[0184] In this process, the relevant information for performing the federated learning task can refer to the above S401 and will not be elaborated here.

[0185] S502: The server node sends a first request message to multiple client nodes respectively; the first request message is used to request information of the client nodes for performing the federated learning task. Wherein, the information of the client nodes may include the identification information of the client nodes, and may also include all or part of the configuration information of the client nodes.

[0186] In this process, referring to the above S402, the first request message may include one or more of the above information (1) to (10), which will not be elaborated here. The process of multiple client nodes responding to the first request message will be specifically described in S503 below.

[0187] S503: The client node sends a first response message to the server node; the first response message includes the information of the client node.

[0188] For any one of the multiple client nodes, after receiving the first request message, the client node determines whether it meets the corresponding requirements indicated by the first request message based on its own configuration information. If the client node determines that its configuration information meets the conditions corresponding to the corresponding requirements for performing the federated learning task indicated by the first request message, it means that the client node can perform the federated learning task corresponding to the first request message. Based on this, the client node can send its own information to the server node, thereby indicating that the client node can perform the confederated learning task, that is, the client node is a client node for performing the confederated learning task. Optionally, the information of the client node itself may include the identification information of the client node, and may also include all or part of the configuration information of the client node.

[0189] Optionally, Figure 5 The described process can be applied to Figure 3B or Figure 3C the system architectures shown.

[0190] Based on the above description, after determining N client nodes for performing the federated learning task, refer to Figure 6 , and the N client nodes perform the federated learning task. For the convenience of description, the principle of the federated learning method of the first client node among the N client nodes will be introduced below, and the processing operations of other client nodes and the interaction process with the server node can refer to the first client node. Figure 6 As

[0191] shown in Figure 6 , the process includes:

[0192] S601: The server node sends a first model training message to the first client node; wherein, the first model training message includes federated learning initial model parameters, federated learning configuration information, and privacy protection configuration information.

[0193] Optionally, the federated learning initial model parameters in the first model training message can be obtained by the server node after inputting the baseline training data into the global model for model training. Among them, the baseline training data can be indicated by the federated learning service request message, or obtained by the server node according to the federated learning task indicated by the federated learning service request message. The baseline training data is not limited herein.

[0194] Optionally, the federated learning configuration information in the first model training message can include the Analytics ID of the federated learning task, the service area of the federated learning task, the type of federated learning of the federated learning task, etc.

[0195] In the embodiments of the present application, the federated learning configuration information further includes the number of learning rounds of the federated learning task. The number of learning rounds is determined by the server node according to the model training duration required by the federated learning task, as well as the model training duration and privacy protection processing duration of each of the N client nodes.

[0196] In a possible implementation, the server node calculates the total duration of the model training duration and privacy protection processing duration of each of the N client nodes, and then determines the number of learning rounds corresponding to the N client nodes according to the maximum total duration among the N total durations and the model training duration required by the federated learning task. For example, N = 3, that is, it includes client node C1, client node C2, and client node C3; assuming that the model training duration required by the federated learning task is 10h, the total duration of the model training duration and privacy protection processing duration of client node C1 is 0.5h, the total duration of the model training duration and privacy protection processing duration of client node C2 is 0.3h, and the total duration of the model training duration and privacy protection processing duration of client node C3 is 1h. Then, the number of learning rounds of the N client nodes is determined to be 10 times. It can be understood that in this implementation, the number of learning rounds of the N client nodes is the same.

[0197] In another possible implementation, the server node can divide the N client nodes into at least two groups of client nodes. Then, for any group of client nodes, calculate the total duration of the model training duration and the privacy protection processing duration of each client node in the group, and then determine the number of learning rounds corresponding to the group of client nodes based on the maximum total duration in the group and the model training duration required by the federated learning task. In this implementation, the server node can group the N client nodes according to the hardware performance, model training duration, and / or privacy protection processing duration corresponding to the N client nodes. Taking the above method as an example, the server node can use client node C1 and client node C2 as the first group of client nodes, and client node C3 as the second group of client nodes. Furthermore, it can be obtained that the number of learning rounds for the first group of client nodes is 20 times, and the number of learning rounds for the second group of client nodes is 10 times. It can be understood that the number of learning rounds for the same group of client nodes is the same, but the number of learning rounds corresponding to any two groups of client nodes may be different.

[0198] In this implementation, in order to ensure that the N client nodes participate in iterative training together at least once during the federated learning process, the server node can define the association relationship of the number of learning rounds between different groups of client nodes, and this association relationship indicates the timing when at least two groups of client nodes participate in the federated learning iteration together. For example, the number of learning rounds for the first group of client nodes above is twice that of the second group of client nodes. Therefore, the server node can define that the first group of client nodes participate in the federated learning iteration together with the second group of client nodes at the 2n-th learning round. Or rather, the server node only aggregates the models of the first group of client nodes at the 0.5h moment to implement the federated learning iteration of the first group of client nodes; at the 1h moment, it aggregates the models of the first group of client nodes and the second group of client nodes to implement the federated learning iteration of the first group of client nodes and the second group of client nodes; and so on. At the 1.5h moment, only the federated learning iteration of the first group of client nodes is implemented, and at the 2h moment, the federated learning iteration of the first group of client nodes and the second group of client nodes is implemented;... It can be understood that the above example exemplarily describes the timing when at least two groups of client nodes participate in the federated learning iteration according to the number of learning rounds of the two groups of client nodes. In addition, the server node can also define the timing when at least two groups of client nodes participate in the federated learning iteration based on the iteration duration of each round corresponding to at least two groups of client nodes, and the specific definition method is not specifically limited here.

[0199] In the above implementation, the server node can group the N client nodes according to the total duration of the model training duration and the privacy protection processing duration of each client node among the N client nodes, so as to distinguish at least two groups of client nodes with shorter total duration and longer total duration. The grouping method of the N client nodes is not specifically limited here. Based on this, the server node can perform more rounds of model training for the client nodes with shorter total duration, reduce the model training waiting time of the client nodes with shorter total duration, realize flexible selection of client nodes to execute the federated learning task, and improve the efficiency of the client nodes to execute the federated learning task.

[0200] Optionally, the privacy protection configuration information in the first model training message includes the privacy protection processing type information and the privacy protection processing factor corresponding to the first client node; wherein, the privacy protection processing type information is used to indicate the privacy protection processing method performed by the first client node on the local model parameters; the privacy protection processing factor is used to indicate the parameters used by the first client node to perform privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

[0201] In a possible implementation, if the privacy protection processing type information indicates the first processing method, the privacy protection processing factor indicates adding noise distribution parameters to the model parameters when performing the first processing method on the local model parameters in each of the K rounds. Exemplarily, the first processing method is the differential privacy noise addition method based on the Gaussian mechanism, and the privacy protection processing factor can be one of the following parameters: the total differential privacy budget ε for the K rounds, the differential privacy budget ε for each of the K rounds t , the Gaussian distribution interval (μ, σ t 2 ) for each of the K rounds, and the noise value N; where, the K rounds represent the learning rounds of the client node, and the sum of the differential privacy budgets ε for each of the K rounds t is equal to the total differential privacy budget ε for the K rounds, and the noise value N follows (μ, σ t 2 ).

[0202] In another possible implementation, if the privacy protection processing type information indicates the second processing method, the privacy protection processing factor indicates the threshold, compression factor, or Top-k value used when the second processing method is performed on the local model parameters in each of the K rounds. Exemplarily, the second processing method is a gradient compression method based on the Top-K threshold method, and the privacy protection processing factor can be one of the following parameters: threshold, Top-K value, compression factor. Among them, the threshold represents the threshold of the gradient change degree; the Top-K value represents the number of gradients to be selected; the compression factor represents the ratio of the gradients to be selected. For the specific implementation of privacy protection processing using the privacy protection processing factor, please refer to the description in S603 below.

[0203] In the embodiments of the present application, since the privacy protection budget information supported by each of the N client nodes may be different, the server node can determine the corresponding privacy protection processing factor for each client node. Taking the first client node as an example, assuming that the differential privacy budget ε supported by the first client node is 10, the server node can determine the corresponding differential privacy budget ε for the first client node to be 11, or the differential privacy budget ε in each of the K (e.g., K = 10) rounds, t or the Gaussian distribution interval (0, 1) in each of the K rounds. Alternatively, the server node determines a unified privacy protection processing factor for the N client nodes, that is, the privacy protection processing factors used by the N client nodes are the same. For example, the server node takes the average value (or maximum value, minimum value, etc.) of the privacy protection budget information supported by each of the N client nodes as the corresponding privacy protection processing factor for each client node.

[0204] S602: The first client node performs model training based on the federated learning initial model parameters and the federated learning configuration information to obtain local model parameters.

[0205] In this process, after the first client node receives the first model training message, the first client node uses the federated learning initial model parameters in the first model training message described above as the model parameters of the local model, and then performs model training according to its own local data and the federated learning configuration parameters in the first model training message, thereby obtaining local model parameters. In the embodiments of the present application, the model parameters can be model gradients or model weights, and the local data includes training data.

[0206] S603: The first client node performs privacy protection processing on the local model parameters according to the privacy protection configuration information to obtain privacy model parameters.

[0207] In this process, after the first client node trains to obtain the local model parameters, the first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information in the first model training message and the parameters indicated by the privacy protection processing factor, and then obtains the privacy model parameters.

[0208] In a possible implementation, the privacy protection processing type information indicates a first processing method. Assuming that the first processing method is the differential privacy noise addition method based on the Gaussian mechanism, the first client node adds the noise of the current round indicated by the privacy protection processing factor to the local model parameters according to the noise distribution parameters of the current round, and then obtains the privacy model parameters. Exemplarily, as shown in the following formula (1):

[0209]

[0210] Wherein, is the privacy model parameter of the t-th round of the first client node, W t is the local model parameter of the t-th round of the first client node, N t is the noise value of the t-th round of the first client node, N t obeys (μ t , σ t 2 ), (μ t , σ t 2 ) is the Gaussian distribution interval of the t-th round of the first client node. Based on the above description, it can be understood that t = 1 in this step.

[0211] For example, in the t-th round, the first client node selects the number of noise values corresponding to the local model parameters or at least one noise value (such as 0.5) from (μ t , σ t 2 ); then the first client node adds 0.5 to each gradient in the local model parameters, and then obtains the privacy model parameters corresponding to the local model parameters.

[0212] In another possible implementation, the privacy protection processing type information indicates a second processing method. Assuming that the second processing method is the gradient compression method based on the Top-K threshold method, the first client node performs gradient compression on the local model parameters according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor, and then obtains the privacy model parameters. Exemplarily, as shown in the following formula (2):

[0213]

[0214] Wherein, is the privacy model parameter of the first client node in the t-th round, W t is the local model parameter of the first client node in the t-th round. In this step, t = 1; thr represents the threshold or compression factor or Top-k value of the current round. |W t | > thr means that the degree of gradient change in W t is greater than the threshold, or the gradient value within the compression factor or Top-k value range; or rather, the threshold represents the threshold of the degree of gradient change. The first client node takes the gradients in the local model parameter that are greater than this threshold as the privacy model parameter; or the first client node sorts the gradients in the local model parameter according to the degree of gradient change, and then takes the gradients within the compression factor or Top-k value range as the privacy model parameter in the order from large to small according to the degree of gradient change; that is to say, the first client node selects some gradients from the local model parameter as the privacy model parameter according to the threshold or compression factor or Top-k value of the current round. In this way, when t = 1, the degree of gradient change in the t-th round can be the difference degree between the gradients in the local model parameter of the t-th round and the initial model parameter of the federated learning; when t > 1, the degree of gradient change in the t-th round can be the difference degree between the gradients in the local model parameter of the t-th round and the gradients in the local model parameter of the t - 1-th round; or it can be the difference degree between the gradients in the local model parameter of the t-th round and the gradients that have not been selected as the privacy model parameter in the historical rounds.

[0215] For example, in the t-th round, the first client node determines 70% of the gradients from the local model parameter according to the threshold, and then takes these 70% of the gradients as the privacy model parameter corresponding to the local model parameter; or the first client node selects the top 100 gradients with the degree of gradient change from the local model parameter according to the Top-k value, and then takes these 100 gradients as the privacy model parameter corresponding to the local model parameter; or the first client node selects the top 80% of the gradients with the degree of gradient change from the local model parameter according to the compression factor, and then takes these 80% of the gradients as the privacy model parameter corresponding to the local model parameter.

[0216] S604: The server node receives the privacy model parameter from the first client node and continues to execute the federated learning task.

[0217] In this process, the server node receives the privacy model parameter from the first client node, and then executes the subsequent process of the federated learning task according to the privacy model parameter.

[0218] In the embodiments of the present application, due to factors such as the hardware performance, communication performance, model training duration, and privacy protection processing duration of each client node being different, the server node may not be able to receive the privacy model parameters corresponding to all client nodes among the N client nodes, and can only receive the privacy model parameters from M client nodes among the N client nodes, and then perform model aggregation based on the privacy model parameters corresponding to the M client nodes, that is, M is a positive integer less than or equal to N. For example, in S601, the server node divides the N client nodes into two groups, then when performing federated learning iteration only for one of the groups, the server node can only receive the privacy model parameters corresponding to this group of client nodes.

[0219] Optionally, after the server node receives the privacy model parameters from the M client nodes, it can aggregate the privacy model parameters corresponding to the M client nodes to obtain the model aggregation parameters for the next round; further, the server node can determine the model aggregation parameters according to the privacy model parameters corresponding to the M client nodes and the weights of the M client nodes. Exemplarily, as shown in the following formula (3):

[0220]

[0221] Among them, QW t+1 is the model aggregation parameter for the (t + 1)-th round, QW t is the model aggregation parameter for the t-th round, P i is the weight of the i-th client node among the M client nodes, is the local model parameter of the i-th client node in the t-th round.

[0222] In a possible implementation manner, the weight of each client node can be determined by the server node according to the data volume of the training data of the N client nodes; among them, the weight is positively correlated with the data volume of the training data, and the sum of the weights of the N client nodes is 1.

[0223] In another possible implementation, the weight of each client node may be determined by the server node based on the data difference information of N client nodes; wherein the data difference information indicates the degree of difference between the training data of the client node and the training data corresponding to the initial model parameters of the federated learning (or the baseline training data used by the server node), the data difference information is negatively correlated with the weight (or the degree of difference between the training data of the client node and the baseline training data is positively correlated with the weight), and the sum of the weights of the N client nodes is 1. In this implementation, the data difference information of the client node may be represented by parameters such as the L1 norm, the L2 norm, the cosine distance, the Hamming distance, etc. The greater the degree of difference between the training data of the client node and the baseline training data, the greater the weight of the client node, indicating that the contribution of the training data of the client node is greater, thereby improving the training efficiency of federated learning. Taking the data difference information as the L2 norm as an example, the following formula (4) is exemplified:

[0224]

[0225] Among them, P i is the weight of the i-th client node among N client nodes, is the data difference information in the form of L2 norm of the i-th client node among the N client nodes.

[0226] Optionally, after receiving the privacy model parameters from the M client nodes, the server node can also determine the model aggregation parameters according to H client nodes among the M client nodes and the weights of the H client nodes; wherein the degree of decrease in model accuracy of any client node among the H client nodes is less than or equal to the model accuracy decrease threshold. It can be understood that the model accuracy decrease threshold can be set according to the privacy protection budget information of the federated learning task, such as the model accuracy decrease threshold is 5%, and the model accuracy decrease threshold is not specifically limited here.

[0227] In a possible implementation, the degree of decrease in model accuracy (or model accuracy information) of any client node among the M client nodes can be calculated by the client node and sent to the server node. Taking the first client node as an example, in any round, the first model accuracy of the first client node and the second model accuracy of the first client node calculate the degree of decrease in model accuracy of the first client node. Among them, q represents the degree of decline in the model accuracy of the first client node; q1 is the first model accuracy of the first client node, which represents the accuracy of the model prediction made by the first client node using local model parameters; q2 is the second model accuracy of the first client node, which represents the accuracy of the model prediction made by the first client node using privacy model parameters.

[0228] In this implementation manner, the process for the first client node to determine the model accuracy may include: The first client node inputs the prediction data (or prediction samples) in its own local data into the local model to obtain an output result, and then compares the output result with the true result corresponding to the prediction data, thereby obtaining the model accuracy; Based on this, the first model accuracy is equivalent to that determined after performing the above process with the local model parameters without privacy protection processing as the parameters of the local model, and the second model accuracy is equivalent to that determined after performing the above process with the privacy model parameters with privacy protection processing as the parameters of the local model. It can be understood that the degree of decrease in the model accuracy of the first client node is caused by the first client node performing privacy protection processing on the locally trained local model parameters.

[0229] In another possible implementation manner, after the first client node determines the first model accuracy and the second model accuracy, it sends the first model accuracy and the second model accuracy to the server node as model accuracy information, and then the server node determines the degree of decrease in the model accuracy of the first client node based on the first model accuracy and the second model accuracy of the first client node, and then selects H client nodes for model aggregation from the M client nodes. The specific calculation process refers to the above description and will not be elaborated here.

[0230] Based on the above method, the server node selects client nodes with a smaller degree of decrease in model accuracy to participate in model aggregation, so as to reduce the difference between the non-real model aggregation parameters obtained after model aggregation and the real model aggregation parameters, increase the usability of the non-real model aggregation parameters, and reduce the error of the client nodes in the next round of model training.

[0231] In the embodiments of the present application, after the server node performs model aggregation, it obtains model aggregation parameters for implementing the next round of iteration of federated learning.

[0232] In a possible implementation manner, the server node subtracts the first error from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error, and then sends the model aggregation parameters after subtracting the first error to the M client nodes or the N client nodes, so that the M client nodes or the N client nodes perform the next round of iteration of federated learning based on the model aggregation parameters after subtracting the first error. Among them, the first error is determined according to the weights of the M client nodes and the noise estimation values of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node according to the privacy protection budget information required by the federated learning task; Exemplarily, referring to formula (1) in S603, the noise estimation value E of the first client node in the t-th round t is (μ t , σ t 2)'s expected value, i.e., E = μ t .

[0233] In this implementation, by subtracting the first error from the model aggregation parameter, the difference between the untrue model aggregation parameter and the true model aggregation parameter is reduced, thereby increasing the usability of the untrue model aggregation parameter and increasing the accuracy of the next round of federated learning of the client node.

[0234] In a possible implementation, the server node performs privacy protection processing on the model aggregation parameter to obtain the privacy-protected model aggregation parameter, and then sends the privacy-protected model aggregation parameter to M client nodes or N client nodes, so that the M client nodes or N client nodes perform the next round of iterative federated learning according to the privacy-protected model aggregation parameter. Among them, the method for the server node to perform privacy protection processing is the same as that for the client node to perform privacy protection processing in S603, which will not be elaborated here.

[0235] In this implementation, privacy protection processing is performed on the untrue model aggregation parameter, further increasing the difference between the untrue model aggregation parameter and the true model aggregation parameter, and further ensuring the privacy and security of the user's data.

[0236] In a possible implementation, the server node directly sends the model aggregation parameter to M client nodes or N client nodes, so that the M client nodes or N client nodes perform the next round of iterative federated learning according to the model aggregation parameter (refer to S602 and S603).

[0237] In this implementation, no processing is performed on the untrue model aggregation parameter. By directly sending the untrue model aggregation parameter to M client nodes or N client nodes, the efficiency of federated learning is improved.

[0238] Through this method, the server node performs the aggregation process of federated learning according to the privacy model parameter of the client node to obtain the untrue model aggregation parameter, thereby realizing the use of untrue model parameters for iteration in the federated learning process. Therefore, as an attacker, the server node or the client node cannot calculate the true difference between the input sample x` and the label y` and the true local data of the client node based on the untrue model parameter, and thus cannot deduce the local data of other client nodes, thereby avoiding the leakage of the local data of other client nodes and ensuring the privacy and security of the user's local data.

[0239] To better elaborate the above technical solutions, the federated learning method provided by the present application will be further elaborated below in combination with Implementation Modes 1 to 3.

[0240] Implementation Method 1:

[0241] Implementation Method 1 is described by taking Figure 3A the scenario as an example. Refer to Figure 7 , and the process includes:

[0242] S701: The server node (Server NWDAF) receives a federated learning service request message; the federated learning service request message includes information required for performing a federated learning task. Exemplarily, the federated learning service request message is sent by a consumer NF (consumer NF).

[0243] Referring to the description of S401 above, the federated learning service request message includes the conditions that should be met for performing a federated learning task. After the server node responds to the federated learning service request message, it first determines whether its own configuration information can meet this condition. If it can, it performs the subsequent steps; otherwise, it rejects the request.

[0244] Optionally, the configuration information of the server node includes one or more of the following information:

[0245] The privacy protection processing type information of the server node, which is used to indicate the way of performing privacy protection processing on the model aggregation parameters supported by the server node, and can also indicate that the server node has the privacy protection processing ability.

[0246] The privacy protection budget information of the server node; it is used to indicate the degree of decrease in the model accuracy of the global model corresponding to the server node, and can also indicate the degree of protection of the model aggregation parameters by the server node when performing privacy protection processing.

[0247] The identifier (Analytics ID) of the model analysis use case of the server node, which is used to indicate the model use case (or the global model use case) supported by the server node.

[0248] The service area of the server node, which is used to indicate the coverage range of the global model of the server node.

[0249] The federated learning type of the server node, which is used to indicate the federated learning type supported by the server node.

[0250] Optionally, if the server node has the privacy protection ability and the privacy protection ability of the server node supports the privacy protection processing method indicated by the federated learning service request message, the server node performs the subsequent steps.

[0251] Optionally, if the degree of decrease in the model accuracy indicated by the privacy protection budget of the server node meets the range of the model accuracy decrease required for the federated learning task, the server node performs the subsequent steps.

[0252] Optionally, if the identifier of the model analysis use case of the server node contains the identifier of the model analysis use case of the federated learning task, and the service area of the server node is within the service area of the federated learning task, the server node performs the subsequent steps.

[0253] S702: In response to the federated learning service request message, the server node determines N client nodes (Client NWDAF) that execute the federated learning task through the first device.

[0254] In this process, the federated learning service request message may include: privacy protection indication information, privacy protection processing type information (such as the indicated privacy protection processing method is the first processing method), and the model accuracy degradation range information is: differential privacy budget ε = 15 (indicating that the model accuracy degradation degree is 0-5%). Optionally, the federated learning service request message further includes one or more pieces of information in (4) to (10) in S401.

[0255] Refer to the above Figure 4 According to the process description of S403 above, the server node sends a first request message to the first device, so that the first device determines N client nodes that execute the federated learning task according to the first request message, and the configuration information of the N client nodes meets the corresponding requirements for executing the federated learning task indicated by the federated learning service request message.

[0256] Taking the first client node among the N client nodes as an example, in this process, in the configuration information of the first client node, the privacy protection processing information of the first client node indicates that it has the privacy protection processing ability; and, the privacy protection processing method supported by the first client node is the differential privacy noise addition method based on the Gaussian mechanism; and, the privacy protection budget of the first client node is: differential privacy budget ε = 16 (indicating that the model accuracy degradation degree of the local model is 4%, that is, the model accuracy degradation degree of the first client node meets the above model accuracy degradation range information). Optionally, other configuration information of the first client node all meets the corresponding requirements for executing the federated learning task indicated by the federated learning service request message.

[0257] Optionally, the N client nodes perform privacy protection processing in the same way, so as to facilitate the aggregation calculation of subsequent federated learning tasks and improve the efficiency of federated learning.

[0258] Optionally, after the server node determines N client nodes that execute the federated learning task, it sends a response message to the consumer NF; the response message may include a feedback on the successful establishment of the federated learning task.

[0259] S703: The server node determines the number of learning rounds of the federated learning task, and determines the noise distribution parameters added to the model parameters when the N client nodes execute the first processing method according to the number of learning rounds.

[0260] In this process, the number of learning rounds of the N client nodes is the same, or in any round, the N client nodes jointly participate in model aggregation.

[0261] Referring to the description of the number of learning rounds in S601 above, taking the first client node as an example, the server node determines that the number of learning rounds K of the first client node is 20, determines the differential privacy budget ε for the first client node to be 40, and then determines the differential privacy budgets ε1, ε2,..., ε of each round based on a preset rule (which can be an average rule, a rule from small to large, etc.). t 、……ε 10 Taking the current round t as 1 as an example, assuming ε t=1 is 2. Then the server node determines the Gaussian distribution interval of the current round Finally, the Gaussian distribution interval (0, 0.25) of the current round is used as the noise distribution parameters added to the model parameters when the first client node executes the first processing method.

[0262] S704: The server node sends model training messages to the N client nodes respectively; the model training messages include the initial model parameters of federated learning, the federated learning configuration information corresponding to the client nodes, and the privacy protection configuration information of the client nodes.

[0263] Taking the first client node as an example, in this process, the federated learning configuration information corresponding to the first client node includes the number of learning rounds of the first client node; the privacy protection configuration information of the first client node includes information indicating that the privacy protection processing method is the differential privacy noise addition method based on the Gaussian mechanism, and the Gaussian distribution interval (0, 0.25) of the current round. The initial model parameters of federated learning can also be understood as the model parameters used by the first client node in the first-round model training.

[0264] The first model training message of the first client node may also include the baseline data of the server node. Furthermore, the first client node determines the data difference information between the baseline data set and its own training data, and this data difference information is used to determine the weight of the client node.

[0265] S705: The client node performs model training according to the model training message and performs privacy protection processing on the locally trained model parameters.

[0266] Based on S704, taking the first client node as an example, the first client node uses the federated learning initial model parameters as the model parameters to be used in the first round of model training for the local model. Then, it obtains its corresponding training data from the network file of the data provider, and then inputs its training data into the local model. After the local model converges or the training ends, the local model parameters after model training are obtained. In this process, there is a one-to-one correspondence between the client node and the data provider.

[0267] After the first client node determines the local model parameters, the first client node selects at least one noise value N (such as 0.1) from (0, 0.25); then the first client node adds 0.1 to each gradient in the local model parameters to obtain the privacy model parameters corresponding to the local model parameters.

[0268] S706: The client node predicts the model accuracy information; the model accuracy information represents the degree of decrease in model accuracy caused by the privacy protection process.

[0269] Based on S705, taking the first client node as an example, the first client node uses the local model parameters as the gradients of the local model, then obtains its corresponding prediction samples from its own network file, and then inputs its prediction samples into the local model to obtain the first model accuracy of 80% corresponding to the local model parameters; similarly, the first client node uses the privacy model parameters after privacy protection processing as the gradients of the local model, and then inputs its prediction samples into the local model to obtain the second model accuracy of 78% corresponding to the privacy model parameters.

[0270] S707: The client node sends the privacy model parameters and the model accuracy information.

[0271] Taking the first client node as an example, the first client node sends its corresponding first model accuracy, second model accuracy, and privacy model parameters to the server node.

[0272] Based on S704, the first client node also sends its corresponding data difference information to the server node.

[0273] S708: The server node determines H client nodes participating in model aggregation according to the model accuracy information of the client nodes, and performs model aggregation based on the local model parameters of the H client nodes to obtain the model aggregation parameters.

[0274] In this process, among the N client nodes, some client nodes do not send their local model parameters to the server node within the model aggregation period of the current round (or the reporting period of the current round, or the learning duration of the current round in the federated learning task) due to factors such as incomplete model training. Therefore, the server node can only receive the private model parameters corresponding to M client nodes among the N client nodes. Based on this, the client nodes that do not send their local model parameters to the server node within the model aggregation period of the current round will send model training progress messages to the server node. Furthermore, after receiving the model training progress messages, the server node can determine whether the client nodes will participate in the subsequent federated learning process.

[0275] After the server node receives the private model parameters corresponding to the M client nodes, it determines H client nodes participating in model aggregation from the M client nodes according to the model accuracy information corresponding to the M client nodes. Assume that the model accuracy degradation threshold is 5%. Taking the i-th client node among the M client nodes as an example, the degree of model accuracy degradation of the i-th client node is Therefore, the i-th client node is one of the H client nodes participating in model aggregation.

[0276] Based on S707, the server node can receive the data difference information corresponding to the M client nodes, and then determine the weights corresponding to the M client nodes according to the data difference information corresponding to the M client nodes, or determine the weights corresponding to the H client nodes according to the data difference information corresponding to the H client nodes. Therefore, after the server node determines the H client nodes, it can also determine the weights corresponding to the H client nodes, and then perform model aggregation according to the weights corresponding to the H client nodes and the private model parameters of the H client nodes to obtain model aggregation parameters (refer to S603 above).

[0277] S709: The server node subtracts the first error from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error.

[0278] Referring to the description of formula (3) above, in this process, the model aggregation parameters after subtracting the first error are calculated according to the following formula (5):

[0279]

[0280] Wherein, is the model aggregation parameter after subtracting the first error for model training in the (t + 1)-th round, QW t+1 is the model aggregation parameter in the (t + 1)-th round, P i is the weight of the i-th client node among the H client nodes, E iis the expectation of the Gaussian distribution interval corresponding to the i-th client node; in this process, taking the Gaussian distribution interval ((0, 0.25)) as an example, i.e., E i = 0.

[0281] S710: The server node sends model training messages for the next round to N client nodes respectively; the model training messages for the next round include the model aggregation parameters after subtracting the first error for the next round of model training; the model training messages for the next round also include the federated learning configuration information corresponding to the client node and the privacy protection configuration information of the client node.

[0282] Taking the first client node as an example, the privacy protection configuration information corresponding to the first client node in this process includes information indicating that the privacy protection processing method is the differential privacy noise addition method based on the Gaussian mechanism and the Gaussian distribution interval for the next round

[0283] S711: Repeat S705 - S710 until the server node determines that the federated learning end condition is met, and end the federated learning. The federated learning end condition includes reaching the model training duration required by the federated learning task, or the end indication message sent by the consumer NF, or the global model convergence of the server node, or the local model convergence of any client node.

[0284] In this process, the server node ends the federated learning according to the federated learning end condition. Optionally, the server node ends the federated learning when the learning duration of the federated learning reaches the model training duration required by the federated learning task; optionally, the server node ends the federated learning based on the end indication message sent by the consumer NF; optionally, the server node ends the federated learning after the global model converges; optionally, the server node ends the federated learning after the local model converges. After the federated learning ends, the server node sends the model aggregation parameters of the last round of the federated learning and the federated learning status to the consumer NF and N client nodes respectively.

[0285] In Implementation 1, after the client node receives the model training message, it performs model training to obtain the real local model parameters (or actual local model parameters), and then adds noise to the local model parameters to obtain non-real local model parameters (or false local model parameters), so as to prevent the server node from deducing the training data of the client node based on the real local model parameters. Since the privacy protection budget used by the client node for privacy protection processing meets the privacy protection budget required by the federated learning task, the non-real local model parameters are similar to the real local model parameters, or the non-real local model parameters meet the model aggregation requirements of the federated learning task. Therefore, the server node can continue to execute the subsequent federated learning process based on the non-real local model parameters on the basis of meeting the requirements of the federated learning task. Since the server node performs model aggregation based on the non-real local model parameters, the model aggregation parameters obtained after the server node performs model aggregation are also false (or non-real model aggregation parameters) are obtained, thereby preventing attackers in multiple client nodes from deducing the local data of other client nodes based on the real model aggregation parameters. Based on this, as an attacker, the server node or the client node cannot deduce the local data of other client nodes through the untrue model parameters, thus avoiding the leakage of the local data of other client nodes and ensuring the privacy and security of the user's local data.

[0286] Implementation 2:

[0287] Implementation 2 is described by taking Figure 3B the scenario as an example. It is described by applying it to different scenarios on the basis of Implementation 1. For the similarities, please refer to the relevant description in combination with Figure 7 and will not be elaborated below.

[0288] Implementation 2 can be applied to Figure 3B the scenario. Refer to Figure 8 and the process includes:

[0289] S801: The server node (Server EMS) receives the federated learning service request message; the federated learning service request message includes the information required to perform the federated learning task in the first processing manner.

[0290] Referring to the description of S701 above, in this process, the configuration information of the server node itself meets the information required to perform the federated learning task, and the subsequent steps are executed.

[0291] S802: The server node sends a first request message to at least two client nodes respectively; the first request message includes the federated learning service request message, which is used to instruct the client nodes that execute the federated learning task in the first processing manner to return a response message, and the response message includes the configuration information of the client nodes.

[0292] Referring to the description of S702 above, assuming that the configuration information of any one of the at least two client nodes (taking the first client node as an example in this process) satisfies the federated learning service request message, then all the at least two client nodes return response messages. That is to say, the at least two client nodes are equivalent to N client nodes (Client RTLF) that execute the federated learning task.

[0293] S803: The server node determines the learning rounds of the federated learning task, and determines the noise distribution parameters added to the model parameters when the N client nodes execute the first processing manner according to the learning rounds.

[0294] In this process, in order to improve the efficiency of the client nodes in executing the federated learning task, the server node divides the N client nodes into multiple groups. For the convenience of description, in this process, the N client nodes are exemplarily divided into two groups, but the number of groups is not limited. Among them, the sum of the model training duration and the privacy protection processing duration of any one of the client nodes in the first group of client nodes is less than the sum of the model training duration and the privacy protection processing duration of any one of the client nodes in the second group of client nodes.

[0295] Referring to the description of the learning rounds in S601 above, assuming that the learning rounds K1 of the first group of client nodes is 20 and the learning rounds K2 of the second group of client nodes is 10. Based on this, for any one of the client nodes in the first group of client nodes (taking the first client node as an example in this process), the server node determines the differential privacy budget ε for the first client node to be 40, and then determines the differential privacy budget ε1, ε2,... ε t 、……ε 10 。Taking the current round t as 1 as an example, assuming ε t=1 is 2. Then the server node determines the Gaussian distribution interval of the current round Finally, the Gaussian distribution interval (0, 0.25) of the current round is used as the noise distribution parameters added to the model parameters when the first client node executes the first processing manner.

[0296] For any client node in the second group of client nodes (this process takes the second client node as an example), the server node determines the differential privacy budget ε for the second client node to be 50, and then determines the differential privacy budgets ε1, ε2, ……, ε for each round of the second client node. t ……ε 20 . Taking the current round t as 1 as an example, assume ε t=1 is 5. Then the server node determines the Gaussian distribution interval for the current round Finally, the Gaussian distribution interval (0, 0.04) for the current round is used as the noise distribution parameter added to the model parameters when the second client node executes the first processing method.

[0297] S804: The server node sends model training messages to N client nodes respectively; the model training messages include the initial model parameters of federated learning, the federated learning configuration information corresponding to the client node, and the privacy protection configuration information for the client node to execute the first processing method.

[0298] Based on the above S803, the federated learning configuration information corresponding to the first client node includes the learning rounds of the first client node; the privacy protection configuration information of the first client node includes information indicating that the privacy protection processing method is the differential privacy noise addition method based on the Gaussian mechanism and the Gaussian distribution interval (0, 0.25) for the current round.

[0299] The federated learning configuration information corresponding to the second client node includes the learning rounds of the second client node; the privacy protection configuration information of the second client node includes information indicating that the privacy protection processing method is the differential privacy noise addition method based on the Gaussian mechanism and the Gaussian distribution interval (0, 0.04) for the current round.

[0300] S805: The N client nodes respectively perform model training according to the model training messages, and perform privacy protection processing on the locally trained model parameters in the first processing method.

[0301] Based on S804, the first client node uses the initial model parameters of federated learning as the model parameters used in the first-round model training of the local model, then inputs its own training data into the local model, and after the local model converges or the training ends, obtains the locally trained model parameters after model training. After the first client node determines its own local model parameters, the first client node selects at least one noise value N1 (such as 0.1) from (0, 0.25); then the first client node adds 0.1 to each gradient in the local model parameters, thereby obtaining the privacy model parameters corresponding to the local model parameters.

[0302] The second client node uses the federated learning initial model parameters as the model parameters for the local model in the first round of model training, and then inputs its own training data into the local model. After the local model converges or the training ends, the second local model parameters after model training are obtained. After the second client node determines its own second local model parameters, the first client node selects at least one noise value N2 (such as 0.02) from (0, 0.04); then the second client node adds 0.02 to each gradient in the second local model parameters, thereby obtaining the second privacy model parameters corresponding to the second local model parameters.

[0303] S806: N client nodes respectively predict model accuracy information; the model accuracy information represents the degree of model accuracy degradation caused by privacy protection processing in the first processing manner.

[0304] Referring to S706, any client node in the first group of client nodes and the second group of client nodes performs model prediction in the same way, which will not be elaborated here.

[0305] S807: N client nodes send privacy model parameters and model accuracy information.

[0306] The first client node sends the first model accuracy of its corresponding local model parameters, the second model accuracy of the privacy model parameters, and the privacy model parameters to the server node.

[0307] Similarly, the second client node sends the first model accuracy of its corresponding local model parameters, the second model accuracy of the second privacy model parameters, and the privacy model parameters to the server node.

[0308] S808: The server node determines the client nodes participating in model aggregation according to the model accuracy information of the client nodes, and performs model aggregation according to the local model parameters of the client nodes participating in model aggregation to obtain model aggregation parameters.

[0309] In this process, since the learning rounds of the first group of client nodes are greater than those of the second group of client nodes, in the 1st, 3rd, 5th,... rounds of the first group of client nodes, the server node only determines the client nodes participating in model aggregation from the first group of client nodes, and then performs model aggregation according to the local model parameters of the client nodes participating in model aggregation at this time to obtain the model aggregation parameters used by the group of client nodes in the 2nd, 4th, 6th,... rounds.

[0310] At the 2nd, 4th, 6th,... rounds of the first group of client nodes (or at the 1st, 2nd, 3rd,... rounds of the second group of client nodes), the server node determines the client nodes participating in model aggregation from the first group of client nodes and the second group of client nodes (equivalent to N client nodes), and then performs model aggregation based on the local model parameters of the client nodes participating in model aggregation at this time, to obtain the model aggregation parameters used by the first group of client nodes at the 3rd, 5th, 7th,... rounds, or to obtain the model aggregation parameters used by the second group of client nodes at the 2nd, 3rd, 4th,... rounds.

[0311] S809: The server node subtracts the first error from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error.

[0312] Based on the above S808, at the 1st, 3rd, 5th,... rounds of the first group of client nodes, the server node subtracts the first error from the model aggregation parameters used by the first group of client nodes at the 2nd, 4th, 6th,... rounds to obtain the model aggregation parameters used by the first group of client nodes at the 2nd, 4th, 6th,... rounds after subtracting the first error. The first error at this time is determined according to the expectation of the Gaussian distribution interval corresponding to the first group of client nodes, and the specific implementation refers to the above formula (5).

[0313] At the 2nd, 4th, 6th,... rounds of the first group of client nodes, the server node subtracts the first error from the model aggregation parameters used by the first group of client nodes at the 3rd, 5th, 7th,... rounds to obtain the model aggregation parameters used by the first group of client nodes at the 3rd, 5th, 7th,... rounds after subtracting the first error. The first error at this time is determined according to the expectations of the Gaussian distribution intervals corresponding to the first group of client nodes and the second group of client nodes.

[0314] S810: The server node sends the model training messages for the next round to N client nodes respectively; the model training messages for the next round include the model aggregation parameters after subtracting the first error for the next round of model training, the federated learning configuration information corresponding to the client nodes, and the privacy protection configuration information of the client nodes.

[0315] Referring to S710, the specific implementation of this step will not be elaborated here.

[0316] S811: Repeat S805 - S810 until the server node determines that the federated learning end condition is met, and end the federated learning. Referring to the federated learning end condition and the process of ending the federated learning described in the above S711, it will not be elaborated here.

[0317] In Implementation 2, on the basis of ensuring the privacy and security of the user's local data, the server node can group N client nodes to distinguish at least two groups of client nodes with different total durations. Based on this, the server node can perform more rounds of model training for the first group of client nodes with a shorter total duration, reduce the model training waiting time of the first group of client nodes, flexibly select client nodes to execute the federated learning task, and improve the efficiency of the client nodes in executing the federated learning task.

[0318] Implementation 3:

[0319] Implementation 3 is described by taking Figure 3C as an example. It is a description of the way of applying different privacy protection processes on the basis of Implementation 2. For the similarities, please refer to the relevant descriptions in combination with Figure 7 and / or Figure 8 which will not be elaborated below.

[0320] Implementation 3 can be applied to the Figure 3C scenario. Refer to Figure 9 and the process includes:

[0321] S901: The server node (Server RTLF) receives a federated learning service request message; the federated learning service request message includes the information required to execute the federated learning task in the second processing mode.

[0322] Referring to the description of S701 above, in this process, the configuration information of Server RTLF itself meets the information required to execute the federated learning task, and the subsequent steps are executed.

[0323] S902: The server node sends a first request message to at least two client nodes respectively; the first request message includes the federated learning service request message, which is used to instruct the client nodes that execute the federated learning task in the second processing mode to return a response message, and the response message includes the configuration information of the client nodes.

[0324] Referring to the description of S802 above, assume that at least two client nodes are equivalent to N client nodes (Client RTLF) that execute the federated learning task.

[0325] S903: The server node determines the number of learning rounds of the federated learning task, and determines the threshold or compression factor or Top-k value used when the N client nodes execute the second processing mode according to the number of learning rounds.

[0326] Referring to the description of S802 above, the server node divides the N client nodes into two groups. Among them, the learning round K1 of the first group of client nodes is 20, and the learning round K2 of the second group of client nodes is 10.

[0327] Referring to the description of S401 above, the server node can determine the corresponding threshold or compression factor or Top-k value for each client node according to the privacy protection budget required by the federated learning task; that is, the thresholds or compression factors or Top-k values of any two client nodes are different. Alternatively, the server node can determine the corresponding threshold or compression factor or Top-k value for each group of client nodes according to the privacy protection budget required by the federated learning task; that is, the thresholds or compression factors or Top-k values of any two groups of client nodes are different, but the thresholds or compression factors or Top-k values of the client nodes in the same group are the same. Alternatively, the server node can determine a unified threshold or compression factor or Top-k value for the N client nodes according to the privacy protection budget required by the federated learning task.

[0328] For the sake of convenience of description, this process takes a unified threshold or compression factor or Top-k value as an example. Suppose the threshold is 0.01, or the compression factor is 80%, or the Top-k value is 100.

[0329] S904: The server node sends model training messages to the N client nodes respectively; the model training messages include the initial model parameters of federated learning, the federated learning configuration information corresponding to the client nodes, and the privacy protection configuration information for the client nodes to execute the second processing method.

[0330] Based on S903 above, in the first model training message corresponding to the first client node, the federated learning configuration information includes the learning round of the first client node, and the privacy protection configuration information includes the determined threshold (0.01) or compression factor (80%) or Top-k value (100).

[0331] In the second model training message corresponding to the second client node, the federated learning configuration information includes the learning round of the second client node, and the privacy protection configuration information includes the determined threshold (0.01) or compression factor (80%) or Top-k value (100).

[0332] S905: The N client nodes respectively perform model training according to the model training messages, and perform privacy protection processing on the locally trained model parameters in the second processing method.

[0333] Based on S904, after the first client node performs model training according to the initial model parameters of federated learning, the federated learning configuration information, and its own training data, it obtains local model parameters. Then, the first client node determines the degree of change between each gradient in the local model parameters and each corresponding gradient in the initial model parameters of federated learning, and uses the gradients in the local model parameters with a degree of change greater than 0.01 (i.e., the threshold) as the privacy model parameters corresponding to the local model parameters. Alternatively, based on the order of the degree of change from large to small, 80% (i.e., the compression factor) of the gradients are selected from the local model parameters as the privacy model parameters corresponding to the local model parameters. Alternatively, based on the order of the degree of change from large to small, the first 100 (i.e., the Top-k value) gradients are selected from the local model parameters as the privacy model parameters corresponding to the local model parameters.

[0334] S906: The N client nodes respectively predict model accuracy information; the model accuracy information represents the degree of decrease in model accuracy caused by performing privacy protection processing in the second processing manner.

[0335] Referring to S706, any client node in the first group of client nodes and the second group of client nodes performs model prediction in the same way, which will not be elaborated here.

[0336] S907: The client node sends the privacy model parameters and the model accuracy information.

[0337] S908: The server node determines the client nodes participating in model aggregation according to the model accuracy information of the client nodes, and performs model aggregation according to the local model parameters of the client nodes participating in model aggregation to obtain model aggregation parameters.

[0338] Referring to S807 and S808 above, the model aggregation method in this process is the same as the implementation method described in S807 and S808 above, which will not be elaborated here.

[0339] S909: The server node respectively sends model training messages for the next round to the N client nodes; the model training messages for the next round include model aggregation parameters, the federated learning configuration information corresponding to the client nodes, and the privacy protection configuration information of the client nodes.

[0340] Referring to S710, the specific implementation of this step will not be elaborated here.

[0341] S910: Repeat S905 - S909 until the server node determines that the federated learning end condition is met, and end the federated learning.

[0342] Referring to the federated learning end condition and the process of ending the federated learning described in S711 above, which will not be elaborated here.

[0343] In Implementation 3, after the client node receives the model training message, it performs model training to obtain the real local model parameters, and then selects some gradients from the real local model parameters according to the threshold, compression factor, or Top-k value, and further uses these gradients as the local model parameters (or privacy model parameters) after privacy protection processing. Since the privacy model parameters are only some gradients in the real local model parameters, the privacy model parameters are incomplete local model parameters, which is equivalent to non-real local model parameters, so as to prevent the server node from deducing the local data of the client node based on the real local model parameters. Similarly, the model aggregation parameters obtained by the server node after aggregating according to the privacy model parameters are also non-real, thus preventing attackers in multiple client nodes from deducing the local data of other client nodes based on the real model aggregation parameters, thereby avoiding the leakage of the local data of other client nodes and ensuring the privacy and security of the user's local data.

[0344] It should be noted that each step involved in the above embodiments or examples can be executed by the corresponding device, or by components such as modules, chips, processors, or chip systems within the device. The embodiments of the present application do not limit this. The above embodiments are only described by taking the execution by the corresponding device as an example. In addition, the specific implementation methods or examples in the above embodiments do not limit the solutions provided by the embodiments of the present application.

[0345] Optionally, in each of the above embodiments, some steps can be selected for implementation, and the order of the steps in the figure can also be adjusted for implementation. The present application does not limit this. It should be understood that implementing some steps in the figure, adjusting the order of the steps, or combining them with each other for specific implementation all fall within the protection scope of the present application.

[0346] It can be understood that, in order to implement the functions in the above embodiments, each device involved in the above embodiments includes the corresponding hardware structure and / or software module for executing each function. Those skilled in the art should easily realize that, combining the units and method steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application scenario and design constraint conditions of the technical solution.

[0347] It can be understood that the above-described network architecture and application scenarios described in the embodiments of the present application are for more clearly explaining the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Those of ordinary skill in the art know that with the evolution of the network architecture and the emergence of new services, the technical solutions provided by the embodiments of the present invention are equally applicable to similar technical problems.

[0348] It should be noted that the "steps" in the embodiments of this application are only for illustration, which is a way of expression for better understanding of the embodiments and does not constitute a substantial limitation on the execution of the solution of this application. For example, the "steps" can also be understood as "features". In addition, these steps do not impose any limitation on the execution order of the solution of this application. Any operation such as changing the step order, merging steps or splitting steps that does not affect the realization of the overall solution and forms a new technical solution is also within the scope disclosed in this application.

[0349] Based on the same technical concept, this application also provides a federated learning device, and the federated learning device can be applied to a federated learning system as Figures 3A - 3C shown. The federated learning device is used to implement the method provided in the above embodiments, and the federated learning device can be applied to the server node or client node involved in the above embodiments. Referring to Figure 10 as shown, the federated learning device 1000 includes a communication unit 1001 and a processing unit 1002.

[0350] The communication unit 1001 is used to receive and send data and support the communication between the federated learning device 1000 and other devices.

[0351] The processing unit 1002 is used to control and manage the actions of the federated learning device 1000 and execute the steps performed by the server node or client node in the federated learning methods provided in the above various embodiments or examples.

[0352] Optionally, the federated learning device 1000 further includes a storage unit for storing the program code and / or data of the federated learning device 1000.

[0353] The communication unit 1001 can be referred to as an input / output unit, a transceiver unit, etc. The communication unit 1001 can be a transceiver or a communication interface; the processing unit 1002 can be a processor. When the federated learning device 1000 is a module (such as a chip) in a communication device, the communication unit 1001 can be an input / output interface, an input / output circuit or an input / output pin, etc., and can also be referred to as an interface, a communication interface or an interface circuit, etc.; the processing unit 1002 can be a processor, a processing circuit or a logic circuit, etc.

[0354] In one implementation, the federated learning device 1000 can be applied to Figure 6 the server node in the embodiment shown. The processing unit 1002 is used for:

[0355] Determine N client nodes for executing the federated learning task, where N is an integer greater than 1;

[0356] According to the communication unit 1001, a first model training message is sent to the first client node among the N client nodes; the first model training message includes initial federated learning model parameters, federated learning configuration information, and privacy protection configuration information; the initial federated learning model parameters and the federated learning configuration information are used for model training, and the privacy protection configuration information is used for privacy protection processing; according to the communication unit 1001, the privacy model parameters are received from the first client node; the privacy model parameters are obtained by the first client node based on model training and privacy protection processing.

[0357] Optionally, the privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method used by the first client node, and the privacy protection processing factor indicates the parameter used by the first client node to perform privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

[0358] Optionally, the processing unit 1002 is specifically configured to:

[0359] According to the privacy protection budget information supported by each client node among the N client nodes, determine the privacy protection processing factor of each client node among the N client nodes; the privacy protection budget information supported by the first client node among the N client nodes represents the degree of model accuracy degradation caused by the first client node performing privacy protection processing.

[0360] Optionally, the privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method; the privacy protection processing factor indicates the noise distribution parameter added to the model parameters when the first processing method is performed on the local model parameters in each of the K rounds.

[0361] Optionally, the privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method; the privacy protection processing factor indicates the threshold or compression factor or Top-k value used when the second processing method is performed on the local model parameters in each of the K rounds.

[0362] Optionally, the processing unit 1002 is specifically configured to:

[0363] According to the communication unit 1001, a federated learning service request message is received. The federated learning service request message includes privacy protection indication information, and the privacy protection indication information indicates that the client nodes performing the federated learning task need to perform privacy protection processing on local model parameters; N client nodes performing the federated learning task are determined according to the privacy protection indication information, and the N client nodes have the privacy protection processing ability, and the privacy protection processing ability is the ability to perform privacy protection processing on local model parameters.

[0364] Optionally, the federated service request message further includes model accuracy degradation range information; the processing unit 1002 is specifically configured to:

[0365] According to the model accuracy degradation range information, N client nodes performing the federated learning task are determined, and the degree of model accuracy degradation caused by the N client nodes performing privacy protection processing satisfies the model accuracy degradation range information required by the federated learning task.

[0366] Optionally, the processing unit 1002 is specifically configured to:

[0367] Send a first request message to a first device according to the communication unit 1001; the first request message is used to obtain configuration information of client nodes that meet a first condition, and the first condition at least includes having the privacy protection processing ability and the degree of model accuracy degradation caused by performing privacy protection processing satisfying the model accuracy degradation range information required by the federated learning task; the first device is used to store the configuration information of client nodes.

[0368] Receive a first response message from the first device according to the communication unit 1001, and the first response message includes configuration information of N client nodes that meet the first condition, and the configuration information corresponding to the N client nodes includes privacy protection processing ability information.

[0369] Optionally, the processing unit 1002 is specifically configured to:

[0370] Send first request messages to at least two client nodes respectively according to the communication unit 1001; receive corresponding first response messages from N client nodes that meet the first condition, and the first response message corresponding to each client node among the N client nodes includes the configuration information corresponding to the client node.

[0371] Optionally, the first request message includes one or more of the following information:

[0372] Privacy protection processing type information, where the privacy protection processing type information indicates the privacy protection processing method required by the federated learning task; a threshold for the proportion of user privacy features, where the threshold for the proportion of user privacy features indicates the minimum proportion of private data in the training data used by the client nodes executing the federated learning task; a threshold for the dataset size, where the threshold for the dataset size indicates the minimum amount of data in the training data used by the client nodes executing the federated learning task; a threshold for the model training duration, where the threshold for the model training duration indicates the maximum model training duration per round for the client nodes executing the federated learning task; a threshold for the privacy protection processing duration, where the threshold for the privacy protection processing duration indicates the maximum privacy protection processing duration per round for the client nodes executing the federated learning task.

[0373] Optionally, the processing unit 1002 is further configured to:

[0374] Receive privacy model parameters from M client nodes among the N client nodes according to the communication unit 1001, where M is a positive integer less than or equal to N; aggregate the privacy model parameters corresponding to the M client nodes to obtain model aggregation parameters; subtract a first error from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error; the first error is determined according to the weights of the M client nodes and the noise estimation values of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node according to the privacy protection budget information required by the federated learning task; send the model aggregation parameters after subtracting the first error to the M client nodes or the N client nodes respectively according to the communication unit 1001.

[0375] Optionally, the processing unit 1002 is specifically configured to:

[0376] Determine H client nodes from the M client nodes, where the degree of decrease in model accuracy of any client node among the H client nodes is less than or equal to a model accuracy decrease threshold, and H is a positive integer less than or equal to M; aggregate the privacy model parameters corresponding to the H client nodes according to the privacy model parameters corresponding to the H client nodes and the weights of the H client nodes.

[0377] Optionally, the processing unit 1002 is specifically configured to:

[0378] Receiving, by the communication unit 1001, model accuracy information from the M client nodes, where the model accuracy information indicates the degree of model accuracy degradation caused by the privacy protection processing of the locally trained model parameters by the client nodes; determining, from the M client nodes, the H client nodes whose model accuracy degradation degree is less than or equal to the model accuracy degradation threshold.

[0379] In one implementation, the federated learning device 1000 can be applied to Figure 6 the first client node in the illustrated embodiment. The processing unit 1002 is configured to:

[0380] Receiving, by the communication unit 1001, a first model training message from the server node; where the first model training message includes federated learning initial model parameters, federated learning configuration information, and privacy protection configuration information; performing model training according to the federated learning initial model parameters and the federated learning configuration information to obtain locally trained model parameters; performing privacy protection processing on the locally trained model parameters according to the privacy protection configuration information to obtain privacy model parameters; and sending, by the communication unit 1001, the privacy model parameters to the server node.

[0381] Optionally, the privacy protection configuration information includes privacy protection processing type information corresponding to the first client node and a privacy protection processing factor; where the privacy protection processing type information indicates the privacy protection processing method to be used, and the privacy protection processing factor indicates the parameter used by the first client node to perform privacy protection processing on the locally trained model parameters according to the privacy protection processing method indicated by the privacy protection processing type information;

[0382] The processing unit 1002 is specifically configured to:

[0383] Performing privacy protection processing on the locally trained model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameter indicated by the privacy protection processing factor.

[0384] Optionally, the privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method;

[0385] The processing unit 1002 is specifically configured to:

[0386] Adding the noise of the current round to the locally trained model parameters according to the noise distribution parameters of the current round indicated by the privacy protection processing factor.

[0387] Optionally, the privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method;

[0388] The processing unit 1002 is specifically configured to:

[0389] Perform gradient compression on the local model parameters according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor.

[0390] Optionally, the processing unit 1002 is further configured to:

[0391] Send model accuracy information to the server node through the communication unit 1001, where the model accuracy information indicates the degree of model accuracy degradation caused by the privacy protection processing of the local model parameters by the first client node.

[0392] It should be noted that the division of modules in the embodiments of the present application is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present application, each functional unit may be integrated in a processing unit, may exist separately physically, or two or more units may be integrated in one unit. The above integrated unit may be implemented in the form of hardware or in the form of a software functional unit.

[0393] If the above integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) or a processor to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: USB flash drive, mobile hard disk, read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disc, etc., which can store program codes.

[0394] Based on the above embodiments, the embodiments of the present application further provide a federated learning device, and the federated learning device may be a server node or a client node in the federated learning system as shown in Figures 3A - 3C The federated learning device can implement the methods in the above embodiments and has the functions of the federated learning device 1000. Refer to Figure 11As shown, the federated learning device 1100 includes a transceiver 1101, a processor 1102, and a memory 1103. Among them, the transceiver 1101, the processor 1102, and the memory 1103 are interconnected with each other.

[0395] Optionally, the transceiver 1101, the processor 1102, and the memory 1103 are interconnected with each other through a bus 1104. The bus 1104 can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus, etc. The bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 11 only a thick line is used to represent it in the figure, but it does not mean that there is only one bus or one type of bus.

[0396] The transceiver 1101 is used to receive and send signals to implement communication with other devices.

[0397] The function of the processor 1102 can refer to the description in the above embodiments and will not be elaborated here.

[0398] Among them, the processor 1102 can be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP, etc. The processor 1102 can further include a hardware chip. The above hardware chip can be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The above PLD can be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. When implementing the above functions, the processor 1102 can be implemented by hardware, and of course, it can also implement the corresponding software through the hardware. The steps of the method disclosed in the above embodiments of the present application can be directly reflected as being completed by the execution of the processor 1102, or being completed by the combination of the hardware and software modules in the processor 1102.

[0399] The memory 1103 is used to store program instructions, data, etc. Specifically, the program instructions may include program code, and the program code includes computer operation instructions. The memory 1103 may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as at least one disk memory, hard disk drive (HDD), or solid state drive (SSD). The memory 1103 may also be any other medium that can be used to carry or store program code in the form of instructions or data structures and can be accessed by a computer, and the present application does not limit this. The processor 1102 executes the program instructions stored in the memory 1103 to implement the above functions, thereby implementing the method provided in the above embodiments.

[0400] Based on the above embodiments, an embodiment of the present application further provides a federated learning system, which includes a server node and a client node. The server node is used to implement the steps executed by the server node in the method provided in the above embodiments, and the client node is used to implement the steps executed by the client node in the method provided in the above embodiments.

[0401] Based on the above embodiments, an embodiment of the present application further provides a computer program product, which includes a computer program; when the computer program runs on a computer, it causes the computer to execute the method provided in the above embodiments.

[0402] Based on the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a computer, it causes the computer to execute the method provided in the above embodiments.

[0403] Optionally, the above computer may but is not limited to including communication devices such as terminal devices and network devices.

[0404] Among them, the storage medium may be any available medium that can be accessed by a computer. Taking this as an example but not limited to: the computer-readable medium may include RAM, ROM, EEPROM, CD-ROM or other optical disc storage, magnetic disk storage medium or other magnetic storage devices, or any other medium that can be used to carry or store the desired program code in the form of instructions or data structures and can be accessed by a computer.

[0405] Based on the above embodiments, the embodiments of the present application further provide a chip, which is used to read the computer program stored in the memory and implement the method provided in the above embodiments. Optionally, the chip may include a processor, which is coupled to the memory and is used to read the computer program stored in the memory and implement the method provided in the above embodiments. Optionally, the chip may further include components such as a memory, a communication interface, and a power supply module. The memory is used to store the computer program; the communication interface is used to receive and send data; and the power supply unit is used to supply power to the processor.

[0406] Based on the above embodiments, the embodiments of the present application provide a chip system, which includes a processor and is used to support a computer device to implement the functions involved in the federated learning system in the above embodiments. In a possible design, the chip system further includes a memory, which is used to store the necessary programs and data of the computer device. The chip system may be composed of chips or may include chips and other discrete devices.

[0407] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, optical storage, etc.) containing computer-usable program code.

[0408] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the present application. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0409] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device, and the instruction device implements the functions specified in Figure 1 one or more of the processes or multiple processes and / or blocks Figure 1 one or more of the blocks or multiple blocks.

[0410] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to generate a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or multiple processes and / or blocks Figure 1 one process or multiple processes and / or blocks Figure 1 or steps for implementing the functions specified in multiple blocks.

[0411] Obviously, those skilled in the art can make various changes and modifications to this application without departing from the scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalent technologies, this application is also intended to include these changes and modifications.

Claims

1. A federated learning method, characterized in that, Applied to a federated learning system including a server node and at least two client nodes, the method includes: The server node determines N client nodes for performing a federated learning task, where N is an integer greater than 1; the N client nodes are some or all of the at least two client nodes; The server node sends a first model training message to a first client node among the N client nodes; the first model training message includes initial federated learning model parameters, federated learning configuration information, and privacy protection configuration information; the initial federated learning model parameters and the federated learning configuration information are used for model training, and the privacy protection configuration information is used for privacy protection processing; The server node receives the privacy model parameters from the first client node; the privacy model parameters are obtained by the first client node based on model training and privacy protection processing.

2. The method according to claim 1, wherein The privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method used by the first client node, and the privacy protection processing factor indicates the parameter used by the first client node to perform privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

3. The method according to claim 2, wherein The method further includes: The server node determines the privacy protection processing factor for each of the N client nodes according to the privacy protection budget information supported by each client node among the N client nodes; the privacy protection budget information supported by the first client node among the N client nodes represents the degree of model accuracy degradation caused by the first client node performing privacy protection processing.

4. The method according to claim 2 or 3, characterized in that, The privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method; The privacy protection processing factor indicates the noise distribution parameter added to the model parameters when performing the first processing method on the local model parameters in each of the K rounds.

5. The method according to claim 2 or 3, characterized in that, The privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method; The privacy protection processing factor indicates the threshold or compression factor or Top-k value used when performing the second processing method on the local model parameters in each of the K rounds.

6. The method according to any one of claims 1-5, characterized in that, The server node determines N client nodes for performing a federated learning task, including: The server node receives a federated learning service request message, and the federated learning service request message includes privacy protection indication information, and the privacy protection indication information indicates that the client node performing the federated learning task needs to perform privacy protection processing on the local model parameters; The server node determines N client nodes for performing the federated learning task according to the privacy protection indication information, and the N client nodes have the privacy protection processing ability, and the privacy protection processing ability is the ability to perform privacy protection processing on the local model parameters; The federated service request message further includes model accuracy degradation range information; The server node determines N client nodes to execute the federated learning task, including: The server node determines N client nodes to execute the federated learning task according to the model accuracy degradation range information, and the degree of model accuracy degradation caused by the privacy protection processing of the N client nodes meets the model accuracy degradation range information required by the federated learning task.

7. The method according to claim 6, characterized in that, Determining N client nodes to execute the federated learning task, including: The server node sends a first request message to the first device; the first request message is used to obtain the configuration information of client nodes that meet the first condition, and the first condition includes at least having the privacy protection processing ability and the degree of model accuracy degradation caused by the privacy protection processing meeting the model accuracy degradation range information required by the federated learning task; the first device is used to store the configuration information of client nodes. The server node receives a first response message from the first device, and the first response message includes the configuration information of N client nodes that meet the first condition, and the configuration information corresponding to the N client nodes includes privacy protection processing ability information. Alternatively, the server node separately sends a first request message to at least two client nodes; The server node separately receives corresponding first response messages from N client nodes that meet the first condition, and the first response message corresponding to each of the N client nodes includes the configuration information corresponding to the client node.

8. The method according to claim 7, characterized in that The method further includes: The first request message includes one or more of the following information: Privacy protection processing type information, which indicates the privacy protection processing method required by the federated learning task; Threshold of the user privacy feature ratio, which indicates the minimum ratio of privacy data in the training data used by the client nodes executing the federated learning task; Threshold of the dataset size, which indicates the minimum data volume of the training data used by the client nodes executing the federated learning task; Threshold of the model training duration, which indicates the maximum model training duration of each round of the client nodes executing the federated learning task; Threshold of the privacy protection processing duration, which indicates the maximum privacy protection processing duration of each round of the client nodes executing the federated learning task.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: The server node receives privacy model parameters from M client nodes among the N client nodes, where M is a positive integer less than or equal to N; The server node aggregates the privacy model parameters corresponding to the M client nodes to obtain model aggregation parameters. The server node subtracts the first error from the model aggregation parameter to obtain the model aggregation parameter after subtracting the first error; the first error is determined according to the weights of the M client nodes and the noise estimation values of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node according to the privacy protection budget information required by the federated learning task; The server node sends the model aggregation parameter after subtracting the first error to the M client nodes or the N client nodes respectively.

10. The method according to claim 9, characterized in that, The server node aggregates the privacy model parameters corresponding to the M client nodes to obtain a model aggregation parameter, including: The server node determines H client nodes from the M client nodes, and the degree of decrease in model accuracy of any client node among the H client nodes is less than or equal to the model accuracy decrease threshold, where H is a positive integer less than or equal to M; The server node aggregates the privacy model parameters corresponding to the H client nodes according to the privacy model parameters corresponding to the H client nodes and the weights of the H client nodes.

11. The method according to claim 10, wherein The server node determines H client nodes from the M client nodes, including: The server node receives model accuracy information from the M client nodes, and the model accuracy information indicates the degree of decrease in model accuracy caused by the client node performing privacy protection processing on the locally trained model parameters; The server node determines the H client nodes from the M client nodes whose model accuracy decrease degree is less than or equal to the model accuracy decrease threshold.

12. A federated learning method, characterized in that, Applied to a federated learning system including a server node and at least two client nodes, where at least two client nodes include a first client node, the method includes: The first client node receives a first model training message from the server node; wherein, the first model training message includes federated learning initial model parameters, federated learning configuration information, and privacy protection configuration information; The first client node performs model training according to the federated learning initial model parameters and the federated learning configuration information to obtain locally trained model parameters; The first client node performs privacy protection processing on the locally trained model parameters according to the privacy protection configuration information to obtain privacy model parameters; The first client node sends the privacy model parameters to the server node.

13. The method according to claim 12, wherein The privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method to be used, and the privacy protection processing factor indicates the parameter used by the client node to perform privacy protection processing on the locally trained model parameters according to the privacy protection processing method indicated by the privacy protection processing type information; The first client node performs privacy protection processing on the locally trained model parameters according to the privacy protection configuration information, including: The first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor.

14. The method according to claim 13, wherein The privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method; The first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor, including: The first client node adds the noise of the current round to the local model parameters according to the noise distribution parameters of the current round indicated by the privacy protection processing factor.

15. The method according to claim 13, characterized in that, The privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method; The first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor, including: The first client node performs gradient compression on the local model parameters according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor.

16. The method according to any one of claims 12-15, characterized in that, The method further includes: The first client node sends model accuracy information to the server node, and the model accuracy information indicates the degree of decrease in model accuracy caused by the first client node performing privacy protection processing on the local model parameters.

17. A federated learning device, characterized in that, The device includes: A communication unit, configured to receive and send data; A processing unit, configured to execute the method according to any one of claims 1-11 or 12-16.

18. A federated learning system, characterized in that, Including: A server node for executing any one of claims 1-11, and a first client node for executing any one of claims 12-16.

19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is caused to execute the method according to any one of claims 1-11 or 12-16.

20. A chip, characterized in that, The chip is coupled to the memory and is configured to read and execute program instructions stored in the memory to implement the method according to any one of claims 1-11 or 12-16.

Citation Information

Cited By

  • Privacy protection-oriented user behavior data distribution modeling method and terminal equipment

    CN121480277A