Federated learning method and apparatus

By privacy protection of the local model parameters of the client node in the federated learning system and non-real model parameters are generated, the problem of data leakage in traditional federated learning is solved, and data privacy protection and security protection are achieved.

WO2025139501A1PCT designated stage expired Publication Date: 2025-07-03HUAWEI TECH CO LTD

Patent Information

Application Number
PCT/CN2024/133635
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-27
Filing Date
2024-11-21
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

In traditional federated learning, the communication between the server and multiple clients is likely to lead to the leakage of local data on the client, and the privacy and security of local data cannot be effectively guaranteed.

Method used

By privacy protection of real local model parameters on the client node, non-real privacy model parameters are generated, and model aggregation is performed on the server node, ensuring that the attacker cannot deduce local data of other client nodes through non-real model parameters.

Benefits of technology

It effectively prevents the server or client from deriving local data from other client nodes as attackers, ensuring the privacy and security of local data. At the same time, while meeting the model accuracy requirements, it reduces the reduction in model accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024133635_03072025_PF_FP_ABST
    Figure CN2024133635_03072025_PF_FP_ABST
Patent Text Reader

Abstract

A federated learning method and apparatus. The method comprises: a server node respectively sends a model training message to N client nodes that execute a federated learning task, so that the client node performs model training on the basis of federated learning initial model parameters and federated learning configuration information in the model training message, to obtain real local model parameters, and then performs privacy protection processing on the local model parameters on the basis of privacy protection configuration information in the model training message, to obtain false privacy model parameters. Then, the client node sends the false privacy model parameters to the server node, so that the server node executes subsequent operations of the federated learning task on the basis of the false privacy model parameters, so that a certain client node is prevented from inferring local data of other client nodes by means of real model parameters, thereby ensuring the privacy and security of data of a user.
Need to check novelty before this filing date? Find Prior Art

Description

A federated learning method and device

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the State Intellectual Property Office of the People's Republic of China on December 27, 2023, with application number 202311837234.7 and application name "A Federated Learning Method and Device", the entire contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of communication technology, and in particular to a federated learning method and apparatus. Background Art

[0004] Federated Learning refers to a privacy-preserving distributed machine learning method that includes a server and multiple clients.

[0005] In traditional solutions, the server generates initial model parameters based on service requests and sends them to multiple clients participating in model training. Each client uses its local data as training data, trains its local model based on the training data and initial model parameters, and sends the trained model parameters to the server. The server aggregates the model parameters from multiple clients to obtain the model parameters for the next iteration. It then sends these model parameters to multiple clients for the next round of model training, thus achieving iterative training in federated learning.

[0006] During the communication between the server and multiple clients, the client's local data is easily leaked, and the privacy and security of the local data cannot be guaranteed. Summary of the Invention

[0007] The embodiments of the present application provide a federated learning method and apparatus to prevent leakage of local data of a client and to ensure the privacy and security of local data.

[0008] This application is applicable to a federated learning system, which includes server nodes and client nodes. After obtaining real local model parameters, the client nodes perform privacy protection processing on them, thereby obtaining non-real privacy model parameters. The server nodes are used to aggregate the non-real privacy model parameters corresponding to the client nodes, thereby executing federated learning tasks using non-real model parameters. This ensures that server nodes or client nodes with attacking behavior cannot deduce the local data of other client nodes using non-real model parameters, thereby preventing the leakage of local data of other client nodes and ensuring the privacy and security of local data.

[0009] First, this application proposes a federated learning method, which is applied to a server node in a federated learning system. The following description uses the server node as the execution subject. The method includes:

[0010] The server node sends a first model training message to a first client node among N client nodes; wherein the N client nodes are used to perform a federated learning task, the first client node is any client node among the N client nodes, and the first model training message is a model training message corresponding to the first client node, including initial federated learning model parameters, federated learning configuration information, and privacy protection configuration information. Furthermore, the initial federated learning model parameters and the federated learning configuration information are used by the first client node for model training, and the privacy protection configuration information is used by the first client node to perform privacy protection processing on local model parameters. That is, after receiving the first model training message, the first client node performs model training based on the initial federated learning model parameters and the federated learning configuration information to obtain local model parameters, and then performs privacy protection processing on the local model parameters based on the privacy protection configuration information to obtain private model parameters. Based on this, the server node receives the private model parameters from the first client node, and then performs model aggregation based on the private model parameters of the N client nodes to obtain model aggregation parameters for federated learning iterations, thereby enabling subsequent federated learning iterations until the federated learning task is completed.

[0011] Through this solution, the server node sends a first model training message to the first client node to instruct the first client node to train its own local model using its own local data based on the federated learning initial model parameters and federated learning configuration information to obtain true local model parameters. The true local model parameters are then privacy-protected according to the privacy protection configuration information to obtain non-true local model parameters (i.e., privacy model parameters). Based on this, the model parameters received by the server node are non-true local model parameters. Even if the server node is an attacking server node, the server node cannot infer the local data of the client node based on the non-true local model parameters. Moreover, after receiving the non-true local model parameters, the server node performs subsequent aggregation operations of the federated learning task based on the non-true model parameters, obtaining non-true model aggregation parameters. The attacking client node cannot infer the local data of other client nodes using the non-true model aggregation parameters, thereby preventing the leakage of local data of other client nodes and ensuring the privacy and security of local data.

[0012] In one possible design, the privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein the privacy protection processing type information indicates the privacy protection processing method used by the first client node, and the privacy protection processing factor indicates the parameters used by the first client node when performing privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

[0013] In this way, the first client node can perform privacy protection processing on the trained local model parameters based on the privacy protection processing method indicated by the privacy protection processing type information and according to the parameters indicated by the privacy protection processing factor.

[0014] In one possible design, the server node determines the privacy protection processing factor corresponding to each of the N client nodes based on the privacy protection budget information supported by each of the N client nodes. In other words, the server node determines the privacy protection processing factor corresponding to the first client node among the N client nodes based on the privacy protection budget information supported by the first client node. The privacy protection budget information supported by the first client node can characterize the degree of decrease in model accuracy caused by performing privacy protection processing, that is, the degree of decrease in model accuracy corresponding to the privacy model parameters compared to the model accuracy corresponding to the local model parameters. In addition, the privacy protection budget information supported by the first client node can also characterize the degree of protection of the local model parameters by the first client node when performing privacy protection processing. It can be understood that the degree of protection of the local model parameters by the privacy protection processing is positively correlated with the degree of decrease in model accuracy caused by the privacy protection processing; in other words, the greater the degree of protection of the local model parameters by the privacy protection processing, the greater the degree of decrease in model accuracy caused by the privacy protection processing.

[0015] This design ensures that the privacy protection factor corresponding to each client node not only satisfies the privacy protection of each client node's local model parameters, but also ensures that the degree of model accuracy degradation caused by each client node's privacy protection processing is within the allowable accuracy degradation range, thus preventing the unavailability of model parameters after model aggregation due to a large degree of model accuracy degradation. In other words, the degree of model accuracy degradation caused by each client node's privacy protection processing is minimized, preventing the accuracy degradation of each client node's model from having a significant impact on the accuracy of federated learning, and ensuring the availability of model parameters after model aggregation.

[0016] Optionally, any two client nodes may have different corresponding privacy protection processing factors due to the different privacy protection budget information they support, thereby ensuring the rationality of the privacy protection processing factor corresponding to each client node.

[0017] In one possible design, the privacy protection processing type information indicates a first processing method, which is a noise addition method; the privacy protection processing factor indicates that a noise distribution parameter is added to the model parameter when the first processing method is performed on the local model parameter in each round in K rounds.

[0018] Through this design, the first client node can add noise to the trained local model parameters according to the noise distribution parameters indicated by the privacy protection processing factor, and then obtain the privacy model parameters, thereby achieving privacy protection processing of the local model parameters.

[0019] Optionally, the first processing method includes one of the following noisy methods: a differential privacy noisy method based on a Gaussian mechanism, and a differential privacy noisy method based on a Laplace mechanism.

[0020] Optionally, the noise distribution parameter includes one of the following parameters: differential privacy budget information corresponding to each of the K rounds, and a noise distribution parameter corresponding to each of the K rounds.

[0021] In one possible design, the privacy protection processing type information indicates a second processing method, which is a gradient compression method; the privacy protection processing factor indicates a threshold or compression factor or Top-k value used when performing the second processing method on the local model parameters in each of K rounds.

[0022] Through this design, the first client node can select some model parameters (or model gradients) from the trained local model parameters as privacy model parameters according to the threshold, compression factor, or Top-k value indicated by the privacy protection processing factor, thereby achieving privacy protection processing of the local model parameters. Among them, the threshold indicates the minimum value of the model gradient difference between the local model parameters in the current round and the previous round; the compression factor indicates the number of model gradients in the selected local model parameters; and the Top-k value indicates the proportion of model gradients in the selected local model parameters. Based on this, the second processing method includes a gradient compression processing method based on the Top-K threshold method.

[0023] In one possible design, the server node receives a federated learning service request message; wherein, the federated learning service request message includes privacy protection indication information, and the privacy protection indication information indicates that the client node executing the federated learning task needs to perform privacy protection processing on the local model parameters, and then the server node can determine the N client nodes executing the federated learning task based on the privacy protection indication information.

[0024] Through this design, it is ensured that the N client nodes are all client nodes capable of performing privacy protection processing on local model parameters; in other words, the N client nodes are all client nodes that can perform federated learning tasks.

[0025] In one possible design, the federated service request message also includes model accuracy degradation range information; so that the N client nodes determined to execute the federated learning task not only have privacy protection processing capabilities, but also the degree of model accuracy degradation caused by each client node executing privacy protection processing meets the model accuracy degradation range information required by the federated learning task.

[0026] In this design, the model accuracy degradation range information in the federated service request message indicates the allowed model accuracy degradation range during the federated learning task, ensuring that the degree of model accuracy degradation caused by the privacy protection processing performed by N client nodes meets the requirements of the federated learning task, thereby ensuring the availability of local models participating in model aggregation.

[0027] In one possible design, the server node obtains configuration information of client nodes that meet a first condition by sending a first request message, where the first condition includes at least privacy-preserving processing capabilities and information about a model accuracy degradation range that satisfies the requirements of the federated learning task due to the privacy-preserving processing. Based on this, the server node receives configuration information of N client nodes that meet the first condition.

[0028] In this design, the server node can receive the configuration information of N client nodes that meet the first condition in either of the following two ways:

[0029] In the first method, the server node can send a first request message to the first device; wherein, the first device is used to store the configuration information of the client node, and then the server node receives the configuration information of N client nodes that meet the first condition from the first device; wherein, the configuration information corresponding to the N client nodes includes privacy protection processing capability information.

[0030] In a second manner, the server node may send a first request message to at least two client nodes respectively, thereby receiving configuration information from N client nodes that meet the first condition among the at least two client nodes.

[0031] Optionally, the first request message further includes one or more of the following information:

[0032] Privacy protection processing type information, where the privacy protection processing type information indicates a privacy protection processing method required by the federated learning task;

[0033] a threshold value of a user privacy feature ratio, where the threshold value indicates a minimum value of a ratio of private data to training data used by a client node executing the federated learning task;

[0034] A data set size threshold, where the data set size threshold indicates a minimum amount of training data used by a client node executing the federated learning task;

[0035] A threshold value of model training duration, where the threshold value of model training duration indicates a maximum value of model training duration in each round of a client node executing the federated learning task;

[0036] A privacy protection processing time threshold indicates a maximum value of the privacy protection processing time of each round of the client node executing the federated learning task.

[0037] In this design, the first client node among the N client nodes has privacy-preserving processing capabilities and information about the model accuracy degradation range that meets the requirements of the federated learning task. The configuration information of the first client node also satisfies one or more of the following conditions:

[0038] The privacy protection processing type information of the first client node meets the privacy protection processing type information of the federated learning task;

[0039] The user privacy feature ratio of the first client node is greater than or equal to a threshold value of the user privacy feature ratio;

[0040] The data set size of the first client node is greater than or equal to a data set size threshold;

[0041] The model training duration of the first client node is less than or equal to a threshold of the model training duration;

[0042] The privacy protection processing duration of the first client node is less than or equal to a threshold of the privacy protection processing duration;

[0043] In this way, the server node can select the best client node to perform the federated learning task.

[0044] In one possible design, the federated learning configuration information also includes a learning round of the federated learning task; the learning round is determined by the server node based on the model training duration required by the federated learning task, and the model training duration and privacy protection processing duration of each client node in the N client nodes.

[0045] In this design, the model training duration and privacy protection processing duration for each client node may vary based on factors such as the amount of training data and hardware processing performance. Therefore, the server node can group N client nodes together and calculate the sum of the model training duration and privacy protection processing duration for each client node. The server node then determines the learning round for this group of client nodes based on the maximum sum of the N durations and the model training duration required by the federated learning task. In other words, the learning rounds for the N client nodes are the same.

[0046] Optionally, the server node can also group the N client nodes according to the model training duration and privacy protection processing duration corresponding to the N client nodes to determine at least two groups of client nodes. Then, for any group of client nodes, the server node determines the learning rounds of this group of client nodes based on the model training duration required by the federated learning task and the model training duration and privacy protection processing duration corresponding to each client node in this group. In other words, the learning rounds of this group of client nodes are the same, but the learning rounds of client nodes in different groups are different. In a scenario with at least two groups of client nodes, the server node can define an association relationship between the learning rounds of different groups of client nodes, and the association relationship indicates the timing for at least two groups of client nodes to jointly participate in the model parameter aggregation operation. For example, the server node defines that the first group of client nodes in the 2nth round and the first group of client nodes in the nth round jointly participate in the model parameter aggregation operation. Based on this, the server node can perform more rounds of model training for client nodes with shorter model training time and privacy protection processing time, and perform fewer rounds of model training for client nodes with longer model training time and privacy protection processing time, thereby reducing the model training waiting time of client nodes with shorter model training time and privacy protection processing time, and flexibly instructing client nodes to perform federated learning tasks, thereby improving the efficiency of client nodes in performing federated learning tasks.

[0047] In one possible design, after the server node receives the privacy model parameters from M of the N client nodes, it aggregates the privacy model parameters corresponding to the M client nodes to obtain model aggregation parameters for the next round of model training; then it subtracts the first error from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error; and finally, it sends the model aggregation parameters after subtracting the first error to the M client nodes or the N client nodes, respectively.

[0048] Through the first error in this design, the difference between the non-real model aggregation parameters and the real model aggregation parameters is reduced, and the availability of the non-real model aggregation parameters is increased, thereby reducing the degree of decline in model accuracy of the client node in the next round of model training and ensuring the availability of the model parameters after model aggregation.

[0049] Optionally, the first error is determined based on the weights of the M client nodes and the noise estimation values ​​of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node based on the privacy protection budget information required by the federated learning task. In other words, the first error can be determined based on the privacy protection processing factor corresponding to each client node. For example, if the parameters used by each client node when performing privacy protection processing are selected based on the Gaussian distribution interval, the first error can be the expectation of the Gaussian distribution interval. Based on this, the server node reduces the noise of the non-real model aggregation parameters through the first error based on the noise addition method of the client node, thereby further reducing the difference between the non-real model aggregation parameters and the real model aggregation parameters.

[0050] In one possible design, after the server node receives privacy model parameters from M of the N client nodes, it aggregates the privacy model parameters corresponding to the M client nodes (which can also be called performing model aggregation based on the privacy model parameters corresponding to the M client nodes) to obtain model aggregation parameters for the next round; then, it sends the model aggregation parameters to the M client nodes or the N client nodes respectively; where M is an integer less than or equal to N.

[0051] Through this design, the server node aggregates the non-authentic local model parameters corresponding to the M client nodes (i.e., the privacy model parameters corresponding to the client nodes) to obtain non-authentic model aggregate parameters, which are then sent to the client nodes. This prevents an attacking client node from inferring the local data of other client nodes after receiving the non-authentic model aggregate parameters from the server node, thereby preventing the leakage of local data of other client nodes and ensuring the privacy and security of local data.

[0052] In one possible design, after the server node receives privacy model parameters from M of the N client nodes, it aggregates the privacy model parameters corresponding to the M client nodes to obtain model aggregation parameters for the next round; then, it performs privacy protection processing on the model aggregation parameters to obtain model aggregation parameters after privacy protection processing; and finally, it sends the model aggregation parameters after privacy protection processing to the M client nodes or the N client nodes, respectively.

[0053] Through this design, due to the privacy protection processing of the non-real model aggregation parameters, the difference between the non-real model aggregation parameters and the real model aggregation parameters is further increased, further preventing client nodes or server nodes with attacking behavior from pushing out local data of other client nodes.

[0054] In one possible design, the server node aggregates the privacy model parameters corresponding to the M client nodes based on the privacy model parameters corresponding to the M client nodes and the weights of the M client nodes.

[0055] In one possible design, the server node determines H client nodes from the M client nodes, and then aggregates the privacy model parameters corresponding to the H client nodes based on the privacy model parameters corresponding to the H client nodes and the weights of the H client nodes to obtain model aggregation parameters; wherein the degree of model accuracy degradation of any client node among the H client nodes is less than or equal to the model accuracy degradation threshold, and H is a positive integer less than or equal to M.

[0056] In this design, by selecting client nodes with less model accuracy degradation for model aggregation, the difference between the non-real model aggregation parameters and the real model aggregation parameters after model aggregation is reduced, the availability of the non-real model aggregation parameters is increased, and the degree of model accuracy degradation of the client nodes in the next round of model training is reduced.

[0057] In one possible design, the server node receives model accuracy information from the M client nodes, and then determines the H client nodes from the M client nodes whose model accuracy degradation is less than or equal to the model accuracy degradation threshold; wherein the model accuracy information indicates the degree of model accuracy degradation caused by the client node performing privacy protection processing on the trained local model parameters.

[0058] In one possible design, the server node receives data difference information from the N client nodes, and then determines the weights of the N client nodes based on the data difference information corresponding to the N client nodes; wherein the data difference information indicates the degree of difference between the training data in the client node and the training data corresponding to the initial model parameters of the federated learning; the data difference information is positively correlated with the weight, and the sum of the weights of the N client nodes is 1.

[0059] Through this design, client nodes whose training data is similar to the training data corresponding to the initial model parameters of federated learning are given a greater weight. As a result, when the model is aggregated, the privacy model parameters of these similar client nodes have a greater influence on the model aggregation parameters after model aggregation, thereby improving the adaptability of the model aggregation parameters to the federated learning task.

[0060] In a second aspect, embodiments of the present application provide a federated learning method, which is applied to a client node in a federated learning system. The following description uses the first client node among at least two client nodes in the federated learning system as the execution subject. The method includes:

[0061] The first client node receives a first model training message from the server node; performs model training according to the federated learning initial model parameters and the federated learning configuration information in the first model training message to obtain local model parameters; performs privacy protection processing on the local model parameters according to the privacy protection configuration information in the first model training message to obtain private model parameters; and sends the private model parameters to the server node;

[0062] In one possible design, the privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor corresponding to the first client node; wherein the privacy protection processing type information indicates the privacy protection processing method that the first client node needs to use, and the privacy protection processing factor indicates the parameters used by the first client node when performing privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information; based on this, the first client node can perform privacy protection processing on the local model parameters based on the privacy protection processing method indicated by the privacy protection processing type information and according to the parameters indicated by the privacy protection processing factor.

[0063] In one possible design, when the privacy protection processing type information indicates a first processing method, the first client node adds the noise of the current round to the local model parameters based on the noise addition method corresponding to the first processing method and according to the noise distribution parameters of the current round indicated by the privacy protection processing factor.

[0064] In one possible design, when the privacy protection processing type information indicates the second processing method, the first client node performs gradient compression on the local model parameters based on the gradient compression method corresponding to the second processing method and according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor.

[0065] In one possible design, the first client node sends model accuracy information to the server node, and notifies the server node through the model accuracy information of the degree of model accuracy degradation caused by the first client node performing privacy protection processing on the local model parameters.

[0066] In one possible design, if the first client node receives a first request message from the server node, it feeds back its own configuration information to the server node.

[0067] Optionally, the configuration information of the first client node includes one or more of the following information:

[0068] The privacy protection processing type information of the first client node indicates a privacy protection processing method for local model parameters supported by the first client node;

[0069] Privacy protection budget information of the first client node; indicating the degree of degradation of model accuracy caused by the first client node performing privacy protection processing on local model parameters;

[0070] a user privacy feature ratio of the first client node, indicating a ratio of private data in the training data used by the first client node to perform the federated learning task;

[0071] The dataset size of the first client node indicates the amount of training data used by the first client node to perform the federated learning task;

[0072] The model training duration of the first client node indicates the model training duration of each round of the federated learning task performed by the first client node;

[0073] The privacy protection processing duration of the first client node indicates the privacy protection processing duration of each round of the federated learning task that the first client node instructs to execute.

[0074] In one possible design, the first client node registers its own configuration information with the first device, so that the service node can obtain the configuration information of the first client node from the first device.

[0075] Optionally, the configuration information of the first client node includes one or more of the following information:

[0076] The privacy protection processing type information of the first client node indicates a privacy protection processing method for local model parameters supported by the first client node;

[0077] privacy protection budget information of the first client node, indicating a degree of degradation of model accuracy caused by the first client node performing privacy protection processing on local model parameters;

[0078] a user privacy feature ratio of the first client node, indicating a ratio of private data in the training data used by the first client node to perform the federated learning task;

[0079] The dataset size of the first client node indicates the amount of training data used by the first client node to perform the federated learning task;

[0080] The model training duration of the first client node indicates the model training duration of each round of the federated learning task performed by the first client node;

[0081] The privacy protection processing duration of the first client node indicates the privacy protection processing duration of each round of the federated learning task that the first client node instructs to execute.

[0082] In a third aspect, the present application provides a federated learning device, comprising a unit for executing each step of the first or second aspect. Optionally, the federated learning device may include a communication unit and a processing unit; the communication unit is configured to receive and transmit data, and the processing unit is configured to execute the method provided in the first or second aspect.

[0083] In a fourth aspect, an embodiment of the present application provides a federated learning system, comprising: a server node for executing the above-mentioned first aspect and a first client node for executing the above-mentioned second aspect.

[0084] In a fifth aspect, an embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run on a computer, the computer executes any one of the methods in the first or second aspects above.

[0085] In a sixth aspect, an embodiment of the present application provides a computer program product, which, when executed on a computer, enables the computer to execute any one of the methods in the first or second aspects above.

[0086] In the seventh aspect, an embodiment of the present application provides a chip, which includes a processor and a data interface. The processor reads instructions stored in a memory through the data interface and executes any one of the methods described in the first or second aspect above.

[0087] In one possible design, the chip may further include a memory storing instructions, and the processor is used to execute the instructions stored in the memory. When the instructions are executed, the processor is used to perform the method described in any one of the first or second aspects above.

[0088] Based on the implementations provided in the above aspects, the embodiments of the present application can be further combined to provide more implementations.

[0089] The technical effects that can be achieved in any of the third to seventh aspects mentioned above can be referred to the description of the technical effects that can be achieved in the first and / or second aspects mentioned above, and the repetitions will not be discussed here. BRIEF DESCRIPTION OF THE DRAWINGS

[0090] FIG1 is a schematic diagram of a federated learning scenario applicable to an embodiment of the present application;

[0091] FIG2 is a schematic diagram of a general process of federated learning applicable to an embodiment of the present application;

[0092] 3A-3C are schematic diagrams of system architectures applicable to embodiments of the present application;

[0093] FIG4 is a method flow diagram of a method for determining a client node that performs a federated learning task provided by an embodiment of the present application;

[0094] FIG5 is a flow chart of another method for determining a client node that performs a federated learning task provided by an embodiment of the present application;

[0095] FIG6 is a flow chart of a federated learning method provided in an embodiment of the present application;

[0096] FIG7 is a flow chart of a federated learning method provided in an embodiment of the present application;

[0097] FIG8 is a flow chart of a federated learning method provided in an embodiment of the present application;

[0098] FIG9 is a flow chart of a federated learning method provided in an embodiment of the present application;

[0099] FIG10 is a structural diagram of a federated learning device provided in an embodiment of the present application;

[0100] FIG11 is a structural diagram of a federated learning device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0101] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. The terms used in the following embodiments are only for the purpose of describing specific embodiments and are not intended to be limiting of the present application. As used in the specification and appended claims of the present application, the singular expressions "one", "a", "said", "above", "the" and "this" are intended to also include expressions such as "one or more", unless there is a clear contrary indication in the context. It should also be understood that in the embodiments of the present application, "one or more" refers to one, two or more; "and / or" describes the association relationship of associated objects, indicating that three relationships may exist; for example, A and / or B can mean: A exists alone, A and B exist at the same time, and B exists alone, where A and B can be singular or plural. The character " / " generally indicates that the related objects before and after are in an "or" relationship.

[0102] References to "one embodiment" or "some embodiments" in this specification mean that a particular feature, structure, or characteristic described in conjunction with that embodiment is included in one or more embodiments of the present application. Thus, phrases such as "in one embodiment," "in some embodiments," "in other embodiments," and "in yet other embodiments" appearing in various places in this specification do not necessarily refer to the same embodiment, but rather mean "one or more but not all embodiments," unless otherwise specifically emphasized. The terms "including," "comprising," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0103] The "multiple" involved in the embodiments of the present application means greater than or equal to two. It should be noted that in the description of the embodiments of the present application, the words "first" and "second" are only used for the purpose of distinguishing the description and cannot be understood as indicating or implying relative importance or order.

[0104] To facilitate understanding, we first use Figures 1 and 2 to illustrate the scenarios and processes of federated learning.

[0105] Referring to Figure 1 , federated learning involves a server node (Server) 110 and at least two client nodes (Client), with client nodes 120a, 120b, and 120c used as an example. As participants in federated learning, client nodes (120a, 120b, 120c) use their local datasets as training data (or training samples) to train local models, obtain local model parameters, and then send the trained local model parameters to server node 110.

[0106] The server node 110 serves as the aggregator or coordinator of federated learning and is used to aggregate the model parameters of the client nodes (120a, 120b, 120c). This can also be understood as performing model aggregation on the local models corresponding to the client nodes (120a, 120b, 120c) based on the model parameters of the client nodes (120a, 120b, 120c). Therefore, the process of parameter aggregation by the server node 110 can also be called model aggregation.

[0107] In some scenarios, the server node and the client node have a Model Training Logical Function (MTLF), so the server node can also be called the central MTLF (or central MTLF), and the client node can also be called the local MTLF (or local MTLF). Optionally, in the federated learning framework in the 3rd Generation Partnership Project (3GPP) scenario, the network data analytics function (NWDAF) that supports federated learning includes the central MTLF and the local MTLF, so the server node can also be called the central network element (Server NWDAF) and the client node can also be called the local network element (Client NWDAF).

[0108] In an embodiment of the present application, the model parameters obtained after training the client node are referred to as local model parameters, or local model parameters; the model parameters obtained after model aggregation of the server node are referred to as model aggregation parameters, or global model parameters.

[0109] Based on the architecture shown in Figure 1, the server node can be a Server NWDAF and the client node can be a Client NWDAF. Figure 2 illustrates a federated learning method. The following describes the general process of federated learning, using Figure 2 as a reference. A network function (NF), acting as a consumer that initiates a federated learning service request, triggers the federated learning process. This NF can be referred to as a consumer NF.

[0110] S201: The consumer NF sends a federated learning service request to the server node (Server NWDAF).

[0111] The federated learning service request is used to indicate the federated learning task that needs to be executed. The type of federated learning task can be horizontal federated learning, vertical federated learning, or federated transfer learning.

[0112] S202: The server node sends the federated learning initial model parameters to the client node (Client NWDAF) participating in the model training. Based on the architecture shown in FIG1 , there can be multiple Client NWDAFs participating in the model training.

[0113] The initial model parameters for federated learning can be preset values ​​based on experience, or obtained by the server node through model training based on training data. In one possible implementation, the training data used by the server node can be preset or obtained based on a federated learning service request. The federated learning model used by the server node can be a general machine learning model or a specific machine learning model built based on the requirements of the federated learning service request. Taking the image recognition task as an example, the federated learning model can be a convolutional neural network (CNN) model. For ease of understanding and distinction, the federated learning model used by the server node can be referred to as the global model.

[0114] S203: The client node receives the federated learning initial model parameters from the server node, performs model training according to the federated learning initial model parameters, and obtains local model parameters.

[0115] The federated learning model used by the client node can be the same as the federated learning model used by the server node, or the federated learning model used by the client node can be a partial model of the federated learning model used by the server node. Based on this, the federated learning model used by the client node can be called a local model or a partial model.

[0116] S204: The client node sends local model parameters to the server node.

[0117] S205: The server node receives corresponding local model parameters from the client node, performs model aggregation according to the local model parameters of at least two client nodes, and obtains model aggregation parameters.

[0118] S206: The server node sends the federated learning status information to the consumer NF. The federated learning status information includes the current iteration of the federated learning, model aggregation parameters, etc.

[0119] S207: The server node sends the model aggregation parameters to the client node, so that the client node performs the next round of iterative training according to the model aggregation parameters.

[0120] S208: The client node and the server node perform a round of training until federated learning ends. Each round of training can refer to S203 to S207. The end of federated learning can be indicated by the consumer NF, or when the local model of the client node converges, or when the global model of the server node converges.

[0121] In the aforementioned federated learning process, client nodes directly exchange local model parameters and model aggregate parameters with server nodes. Specifically, the server node receives local model parameters from the client node, and each client node receives model aggregate parameters from the server node. Each client node receives the same model aggregate parameters. Based on this, an attacking node within the server node or within each client node can deduce the local data of other client nodes using the following attack process. Therefore, this federated learning process carries the risk of client node local data leakage and cannot guarantee the privacy and security of users' local data.

[0122] Taking the server node as an example of an attacker, since the server node's model serves as a global model, the server node can then possess a local model corresponding to the client node. The server node sends the model aggregation parameter W1 corresponding to the global model to the client node, and then receives the local model parameter W2 obtained by the client node through local model training based on the model aggregation parameter W1. Based on this, the server node can initialize a pair of input samples x' (or training samples) and the label y' corresponding to the input samples, and then input the input samples x' and label y' into the local model corresponding to the client node (the parameters of the local model at this time are the model aggregation parameter W1), and perform model training to obtain the model parameters W3 after model training. The server node then calculates the distance (or similarity) between the model parameters W3 and the local model parameters W2. It can be understood that the larger the distance (or similarity), the closer the input sample x' and label y' are to the local data (training data) of the client node. Therefore, the server node can determine the input sample x' and label y' corresponding to the time when the distance between model parameter W3 and local model parameter W2 meets the preset conditions, and then use the input sample x' and label y' at this time as the local data of the client node. In other words, the input sample x' and label y' at this time are equivalent to similar data of the local data of the client node, and thus are equivalent to inferring the local data of the client node.

[0123] Similarly, taking the client node as an example of an attacker, for ease of understanding, the client node that executes the attack process is referred to as an attack client node. The attack process includes: the attack client node initializes a pair of input samples x` (or training samples) and the label y` corresponding to the input sample, and then the attack client node receives the model aggregation parameter W4 from the server node, and then inputs the input sample x` and label y` into its own local model (at this time, the parameter of the local model is the model aggregation parameter W4), performs model training, and obtains the local model parameter W5. The attack client node can receive the model aggregation parameter W6 corresponding to the next round of iteration from the server node again, and then the attack client node can calculate the distance between the local model parameter W5 obtained by its own training and the model aggregation parameter W6. It can be understood that the smaller the distance between the local model parameter W5 and the model aggregation parameter W6, the closer the input sample x` and the label y` are to the local data of other client nodes. By adjusting the input sample x` and label y`, the attacker determines the input sample x` and label y` corresponding to the time when the distance between the local model parameter W5 obtained through self-training and the model aggregation parameter W6 meets the preset conditions, and then uses the input sample x` and label y` at this time as the local data of other client nodes. In other words, the input sample x` and label y` at this time are equivalent to similar data of the local data of other client nodes, which is equivalent to deducing the local data of other client nodes. In one possible implementation, the input sample x` can be the normalized 16-beam RSRP (Reference Signal Receiving Power).

[0124] To address the above-mentioned issues, an embodiment of the present application provides a federated learning method for preventing a server node or client node acting as an attacker from deducing the local data of other client nodes, thereby avoiding leakage of local data of other client nodes and protecting the privacy and security of the user's local data. It is understood that in this method, the method and apparatus are based on the same technical concept. Since the principles of the method and apparatus for solving the problem are similar, the implementation of the apparatus and method can refer to each other, and the repeated parts will not be repeated.

[0125] The method can be applied to a federated learning system including a server node and at least two client nodes. Before explaining the embodiment of the present application in detail, the system architecture involved in the embodiment of the present application is first introduced.

[0126] In one possible implementation, referring to Figure 3A , embodiments of the present application can be applied to a federated learning system architecture within the 3GPP core network domain. In this scenario, federated learning is implemented by multiple NWDAF nodes. NWDAF nodes can be core network elements, or in other words, core network elements with NWDAF functionality deployed on them. A core network element that serves as a server node is called a Server NWDAF, and a core network element that serves as a client node is called a Client NWDAF.

[0127] In one possible implementation, referring to FIG3B , an embodiment of the present application can be applied to a federated learning system architecture in a 3GPP access network domain. In this scenario, the server node can be an element management system (EMS) with a machine learning training function (MLTF), and the server node is referred to as a Server EMS. The client node can be a base station, or a radio access network model training logic function (RAN Training Logic Function, RTLF) within the base station, and the client node is referred to as a Client RTLF, or a Client gNB.

[0128] In one possible implementation, referring to Figure 3C , embodiments of the present application can be applied to a federated learning system architecture within another 3GPP access network domain. In this scenario, both the server node and the client node can be a base station or a RTLF within a base station. The server node is referred to as a Server RTLF, or Server gNB, and the client node is referred to as a Client RTLF, or Client gNB. Federated learning can be performed across multiple RTLFs using the Server RTLF and Client RTLF.

[0129] In the federated learning system architecture shown in Figures 3A, 3B, and 3C above, the server node has the model aggregation capability of federated learning, which is used to aggregate the local model parameters of the client node, or for model aggregation, and then obtain model aggregation parameters; the client node has the ability to participate in federated learning, which is used to perform local model training based on local data and obtain local model parameters. It should be noted that the types of federated learning in the embodiments of the present application include but are not limited to: horizontal federated learning, vertical federated learning, or federated transfer learning. It should be understood that the number of client nodes shown in Figures 3A, 3B, and 3C above is only for reference, and the number of client nodes is not limited here.

[0130] In the embodiment of the present application, the client node has corresponding configuration information, which may include one or more of the following information:

[0131] The privacy protection processing information is used to indicate whether the client node has privacy protection processing capabilities, or whether the client node supports privacy protection processing of local model parameters. Exemplarily, the privacy protection processing information may be indication information, where a value of 1 indicates that the client node has privacy protection processing capabilities; and a value of 0 indicates that the client node does not have privacy protection processing capabilities.

[0132] Privacy protection processing type information is used to indicate the privacy protection processing methods supported by the client node for local model parameters. Exemplarily, the privacy protection processing methods for local model parameters include a first processing method and / or a second processing method; wherein the first processing method can be a noise addition method, such as a differential privacy noise addition method based on a Gaussian mechanism, a differential privacy noise addition method based on a Laplace mechanism, etc.; the second processing method can be a gradient compression method, such as a gradient compression method based on a Top-K threshold method, etc. This application does not limit the privacy protection processing methods supported by the client node.

[0133] Privacy protection budget information is used to indicate the degree of degradation in model accuracy caused by the client node's privacy protection processing, or in other words, the degree of protection of local model parameters by the client node's privacy protection processing. Exemplarily, based on the first processing approach, the client node's privacy protection budget information can be the range of differential privacy budget ε in the differential privacy rule allowed by the client node; based on the second processing approach, the client node's privacy protection budget can be the range of threshold values ​​in the Top-K thresholding method, the range of Top-K values, or the range of compression factor values ​​allowed by the client node. It will be understood that the smaller the differential privacy budget ε, or the larger the threshold value in the Top-K thresholding method, or the smaller the Top-K value, or the smaller the compression factor, the greater the degree of degradation in model accuracy caused by the client node's privacy protection processing, or the higher the degree of protection of local model parameters by the client node's privacy protection processing.

[0134] The user privacy feature ratio is used to indicate the proportion of private data in the training data used by the client node to perform federated learning tasks.

[0135] Dataset size, which indicates the amount of training data used by the client node to perform federated learning tasks.

[0136] Model training duration indicates the duration of each round of model training for the client node in the federated learning task. This duration is related to the client node's hardware capabilities and the size of the local dataset.

[0137] The privacy protection processing duration is used to indicate the duration of privacy protection processing for each round of the federated learning task executed by the client node. This privacy protection processing duration is related to the hardware capabilities of the client node and the privacy protection processing methods supported by the client node.

[0138] The hardware capabilities of the client node are used to indicate the performance of the client node in executing the federated learning task. For example, the hardware capabilities of the client node may include the processing power of the CPU, the size of the memory, etc.

[0139] The client node's model analysis use case identifier (Analytics ID) is used to indicate the local model use cases (or local models) supported by the client node. For example, the client node's Analytics ID is used to indicate that the local model use cases supported by the client node include business experience analysis, end-user behavior analysis, and end-mobility analysis.

[0140] The service area of ​​the client node is used to indicate the coverage of the local model of the client node. For example, the service area of ​​the client node is a cell ID.

[0141] The client node's federated learning type indicates the type of federated learning supported by the client node. For example, the supported federated learning types include horizontal federated learning, vertical federated learning, and / or federated transfer learning.

[0142] In a possible implementation, the above configuration information is stored in the client node.

[0143] In another possible implementation, the client node may register the configuration information with a first device; wherein the first device may be a Network Repository Function (NRF) or other functional entity or network element with data storage capabilities. Optionally, the client node sends a registration message to the first device, where the registration message includes the configuration information of the client node, and the first device then stores the configuration information of the client node.

[0144] In an embodiment of the present application, before executing federated learning, the server node first selects a client node to participate in the federated learning.

[0145] Referring to FIG. 4 , a method for determining a client node that performs a federated learning task provided in an embodiment of the present application includes:

[0146] S401: The consumer NF sends a federated learning service request message to the server node. The federated learning service request message includes relevant information for executing the federated learning task.

[0147] Optionally, the relevant information for executing the federated learning task may include requirement information corresponding to the federated learning task, to indicate the requirements that the federated learning task meets. The requirement information may include one or more of the following information (1) to (10):

[0148] (1) Privacy protection indication information, which is used to indicate that privacy protection processing is required when executing federated learning tasks.

[0149] Optionally, the privacy protection indication information can be used to indicate that the client node executing the federated learning task needs to perform privacy protection processing on the local model parameters.

[0150] Optionally, the privacy protection indication information can also be used to indicate that the server node executing the federated learning task needs to perform privacy protection processing on the model aggregation parameters.

[0151] (2) Privacy protection processing type information, which is used to indicate the privacy protection processing method required by the federated learning task.

[0152] Optionally, the privacy protection processing type information can be used to indicate how the client node executing the federated learning task performs privacy protection processing on the local model parameters.

[0153] Optionally, the privacy protection processing type information can also be used to indicate how the server node executing the federated learning task performs privacy protection processing on local model parameters.

[0154] (3) Model accuracy degradation range information, which is used to indicate the model accuracy degradation range required by the federated learning task. It can be understood as indicating the degree of model accuracy degradation allowed for executing the federated learning task.

[0155] Furthermore, the model accuracy degradation range information may include a threshold value corresponding to the privacy protection processing method, indicating the maximum degree of degradation of model accuracy caused by the privacy protection processing required by the federated learning task, or indicating the maximum degree of protection of model parameters caused by the privacy protection processing required by the federated learning task.

[0156] (4) A threshold for the proportion of user privacy features, which indicates the minimum proportion of private data in the training data used by the client node performing the federated learning task. Accordingly, the proportion of user privacy features of the client node participating in the federated learning task is greater than or equal to the threshold for the proportion of user privacy features.

[0157] (5) A dataset size threshold, which indicates the minimum amount of training data used by the client nodes that execute the federated learning task. Accordingly, the amount of training data used by the client nodes participating in the federated learning task is greater than or equal to the dataset size threshold.

[0158] (6) A model training duration threshold, which indicates the maximum model training duration for each round of the client node executing the federated learning task. Accordingly, the model training duration for each round of the client node participating in the federated learning task is less than or equal to the model training duration threshold.

[0159] (7) A privacy protection processing duration threshold, which indicates the maximum privacy protection processing duration of each round of the client node performing the federated learning task. Accordingly, the duration of the privacy protection processing of the local model parameters by the client node participating in the federated learning task in each round is less than or equal to the privacy protection processing duration threshold.

[0160] (8) The identifier of the model analysis use case of the federated learning task (Analytics ID), which is used to indicate that the client node executing the federated learning task needs to support the Analytics ID of the federated learning task.

[0161] (9) The service area of ​​the federated learning task is used to indicate that the service area of ​​the client node executing the federated learning task is within the service area of ​​the federated learning task.

[0162] (10) The federated learning type of the federated learning task, which is used to indicate that the client node executing the federated learning task supports the federated learning type of the federated learning task.

[0163] S402: The server node sends a first request message to the first device, requesting information about the client node that will perform the federated learning task. The client node information may include identification information of the client node and all or part of the client node's configuration information. In this process, the first device is an NRF.

[0164] Optionally, the partial information may include model training time, privacy protection processing time, privacy protection budget information, etc.

[0165] Optionally, the first request message may include relevant information for executing the federated learning task. Optionally, the first request message may include one or more of the above information (1) to (10).

[0166] Optionally, the server node may further send predefined information related to the federated learning task to the first device. For example, the predefined information related to the federated learning task may include indication information of a service area of ​​the server node.

[0167] S403: The first device sends a first response message to the server node, where the first response message includes information of the N client nodes that perform the federated learning task.

[0168] Optionally, the information of the N client nodes may include identification information corresponding to the N client nodes, and may also include all or part of the configuration information corresponding to the N client nodes. Optionally, the partial information may include model training duration, privacy protection processing duration, privacy protection budget information, etc.

[0169] After receiving the first request message, the first device selects a client node that meets the corresponding requirements indicated by the requirement information based on the configuration information of the client nodes stored in the first device.

[0170] Exemplarily, a client node that meets the corresponding requirements indicated by the requirement information satisfies one or more of the following conditions:

[0171] (a) The client node has privacy protection processing capability; illustratively, the privacy protection processing information of the client node indicates that it has privacy protection processing capability.

[0172] (b) The client node performs privacy protection processing in a manner that supports the privacy protection processing method required by the federated learning task indicated by the first request message.

[0173] (c) The client node's privacy protection budget information meets the model accuracy degradation range required by the federated learning task indicated in the first request message; optionally, the model accuracy degradation indicated by the client node's privacy protection budget information is less than the maximum value of the model accuracy degradation range indicated in the first request message, or the model accuracy degradation indicated by the client node's privacy protection budget information is within the model accuracy degradation range indicated in the first request message. For example, if the model accuracy degradation range indicated in the first request message is 0-5%, the model accuracy degradation indicated by the client node's privacy protection budget information must be less than or equal to 5%.

[0174] (d) The proportion of user privacy features of the training data of the client node is greater than or equal to a threshold of the proportion of user privacy features indicated by the first request message.

[0175] (e) The model training duration of each round of the client node is less than or equal to the threshold of the model training duration indicated by the first request message.

[0176] (f) the duration of each round of privacy protection processing performed by the client node on the local model parameters is less than or equal to the threshold of the privacy protection processing duration indicated in the first request message;

[0177] (g) The Analytics IDs supported by the client node include the Analytics ID indicated in the first request message.

[0178] (h) The service area of ​​the client node is in the service area of ​​the federated learning task indicated in the first request message.

[0179] (i) The federated learning type of the client node supports the federated learning type of the federated learning task indicated in the first request message.

[0180] It can be understood that one or more of the above conditions (a) to (i) that the client node that executes the federated learning task needs to meet corresponds to the relevant information for executing the federated learning task indicated in the first request message. For example, if the first request message includes information (1), (2), (3), and (4), the client node that executes the federated learning task needs to meet the above conditions (a), (b), (c), and (d).

[0181] Based on the above description, the first device may determine N client nodes that satisfy the first request message, and then send configuration information of the N client nodes to the server node.

[0182] Optionally, the process described in FIG. 4 may be applied to the system architecture shown in FIG. 3A .

[0183] Optionally, the privacy protection processing methods corresponding to the N client nodes are the same to facilitate subsequent aggregation calculations by the server node.

[0184] In one possible implementation, referring to FIG. 5 , another method flow for determining a client node that performs a federated learning task includes:

[0185] S501: The consumer NF sends a federated learning service request message to the server node. The federated learning service request message includes relevant information for executing the federated learning task.

[0186] In this process, the relevant information for executing the federated learning task can be referred to the above S401, which will not be described in detail.

[0187] S502: The server node sends a first request message to each of the plurality of client nodes, wherein the first request message is used to request information about the client nodes that perform the federated learning task. The client node information may include identification information of the client node and may also include all or part of configuration information of the client node.

[0188] In this process, referring to S402 above, the first request message may include one or more of the above information (1) to (10), which will not be described in detail here. The following S503 specifically describes the process of multiple client nodes responding to the first request message.

[0189] S503: The client node sends a first response message to the server node, where the first response message includes information about the client node.

[0190] For any client node among multiple client nodes, after the client node receives the first request message, the client node determines whether it meets the corresponding requirements indicated by the first request message based on its own configuration information. If the client node determines that its own configuration information meets the conditions corresponding to the corresponding requirements for executing the federated learning task indicated by the first request message, it means that the client node can execute the federated learning task corresponding to the first request message. Based on this, the client node can send its own information to the server node to indicate that the client node can execute the federated learning task, that is, the client node is a client node that executes the federated learning task. Optionally, the information of the client node itself may include the identification information of the client node, and may also include all or part of the information in the configuration information of the client node.

[0191] Optionally, the process described in FIG. 5 may be applied to the system architecture shown in FIG. 3B or FIG. 3C .

[0192] Based on the above description, after determining the N client nodes that will perform the federated learning task, as shown in Figure 6 , the N client nodes will perform the federated learning task. For ease of description, the following describes the principles of the federated learning method in Figure 6 using the first client node among the N client nodes. The processing operations of the other client nodes and the interaction process with the server node can be referenced to the first client node.

[0193] As shown in Figure 6, the process includes:

[0194] S601: The server node sends a first model training message to the first client node. The first model training message includes federated learning initial model parameters, federated learning configuration information, and privacy protection configuration information.

[0195] Optionally, the federated learning initial model parameters in the first model training message may be obtained by the server node inputting baseline training data into the global model for model training. The baseline training data may be indicated by the federated learning service request message, or obtained by the server node according to the federated learning task indicated by the federated learning service request message. The baseline training data is not limited here.

[0196] Optionally, the federated learning configuration information in the first model training message may include the Analytics ID of the federated learning task, the service area of ​​the federated learning task, the federated learning type of the federated learning task, etc.

[0197] In an embodiment of the present application, the federated learning configuration information also includes the learning rounds of the federated learning task. The learning rounds are determined by the server node based on the model training duration required by the federated learning task, as well as the model training duration and privacy protection processing duration of each of the N client nodes.

[0198] In one possible implementation, the server node calculates the sum of the model training duration and privacy protection processing duration for each of the N client nodes, and then determines the learning rounds corresponding to the N client nodes based on the maximum sum of the N duration sums and the model training duration required by the federated learning task. For example, N = 3, which includes client nodes C1, client node C2, and client node C3; assuming that the model training duration required by the federated learning task is 10 hours, the sum of the model training duration and privacy protection processing duration of client node C1 is 0.5 hours, the sum of the model training duration and privacy protection processing duration of client node C2 is 0.3 hours, and the sum of the model training duration and privacy protection processing duration of client node C3 is 1 hour, thereby determining that the learning rounds of the N client nodes are 10 times. It can be understood that in this implementation, the learning rounds of the N client nodes are the same.

[0199] In another possible implementation, the server node can divide N client nodes into at least two groups of client nodes, and then for any group of client nodes, calculate the sum of the model training time and privacy protection processing time of each client node in the group, and then determine the learning round corresponding to the group of client nodes based on the maximum total time in the group and the model training time required by the federated learning task. In this implementation, the server node can group the N client nodes according to the hardware performance, model training time and / or privacy protection processing time corresponding to the N client nodes. Based on the above method, the server node can use client node C1 and client node C2 as the first group of client nodes, and client node C3 as the second group of client nodes. It can be concluded that the learning rounds of the first group of client nodes are 20 times, and the learning rounds of the second group of client nodes are 10 times. It can be understood that the learning rounds of the same group of client nodes are the same, but the learning rounds corresponding to any two groups of client nodes may be different.

[0200] In this implementation, to ensure that N client nodes jointly participate in at least one iterative training iteration during federated learning, the server node can define an association between learning rounds between different groups of client nodes. This association indicates when at least two groups of client nodes jointly participate in a federated learning iteration. For example, if the learning rounds of the first group of client nodes are twice the learning rounds of the second group of client nodes, the server node can define that the first group of client nodes jointly participate in a federated learning iteration in the 2nth learning round and the second group of client nodes jointly participate in the nth learning round. In other words, at time 0.5 hours, the server node only aggregates the model for the first group of client nodes, achieving a federated learning iteration for the first group of client nodes; at time 1 hour, it aggregates the model for the first and second groups of client nodes, achieving a federated learning iteration for the first and second groups of client nodes; and so on. At time 1.5 hours, only the first group of client nodes undergoes a federated learning iteration, and at time 2 hours, both the first and second groups of client nodes undergo a federated learning iteration, and so on. It can be understood that the above example describes the timing when at least two groups of client nodes jointly participate in the federated learning iteration based on the learning rounds of the two groups of client nodes. In addition, the server node can also define the timing when at least two groups of client nodes jointly participate in the federated learning iteration based on the iteration duration of each round corresponding to the at least two groups of client nodes. The definition method is not specifically limited here.

[0201] In the above implementation, the server node can group the N client nodes based on the sum of the model training duration and privacy protection processing duration of each client node, thereby distinguishing at least two groups of client nodes with a shorter sum of durations and a longer sum of durations. The grouping method of the N client nodes is not specifically limited here. Based on this, the server node can perform more rounds of model training on client nodes with shorter sums of durations, reducing the model training wait time for client nodes with shorter sums of durations, enabling flexible selection of client nodes to execute federated learning tasks, and improving the efficiency of client nodes in executing federated learning tasks.

[0202] Optionally, the privacy protection configuration information in the first model training message includes privacy protection processing type information and privacy protection processing factor corresponding to the first client node; wherein, the privacy protection processing type information is used to indicate the privacy protection processing method performed by the first client node on the local model parameters; the privacy protection processing factor is used to indicate the parameters used by the first client node when performing privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

[0203] In one possible implementation, if the privacy-preserving processing type information indicates the first processing method, the privacy-preserving processing factor indicates the noise distribution parameter to be added to the model parameters when the first processing method is performed on the local model parameters in each of the K rounds. For example, the first processing method is a differential privacy noise addition method based on the Gaussian mechanism, and the privacy-preserving processing factor can be one of the following parameters: the total differential privacy budget ε for the K rounds, the differential privacy budget ε for each of the K rounds t , the Gaussian distribution interval (μ, σ) of each round in K rounds t 2 ), noise value N; where K rounds represent the learning rounds of the client node, and the differential privacy budget ε of each round in K rounds is t The sum of is equal to the total differential privacy budget ε of K rounds, and the noise value N obeys (μ, σ t 2 ),

[0204] In another possible implementation, if the privacy protection processing type information indicates the second processing method, then the privacy protection processing factor indicates the threshold, compression factor, or Top-k value used when executing the second processing method on the local model parameters in each of the K rounds. Exemplarily, the second processing method is a gradient compression method based on the Top-K threshold method, and the privacy protection processing factor can be one of the following parameters: threshold, Top-K value, and compression factor. The threshold represents the threshold value of the degree of gradient change; the Top-K value represents the number of gradients to be selected; and the compression factor represents the ratio of gradients to be selected. For the specific implementation of privacy protection processing using the privacy protection processing factor, please refer to the description in S603 below.

[0205] In an embodiment of the present application, since the privacy protection budget information supported by each of the N client nodes may be different, the server node can determine a corresponding privacy protection processing factor for each client node. Taking the first client node as an example, assuming that the differential privacy budget ε supported by the first client node is 10, the server node can determine the corresponding differential privacy budget ε for the first client node as 11, or the differential privacy budget ε for each of K (such as K = 10) rounds. t is 1, or a Gaussian distribution interval (0, 1) for each of the K rounds. Alternatively, the server node determines a unified privacy protection processing factor for the N client nodes, i.e., the N client nodes use the same privacy protection processing factor. For example, the server node uses the average (or maximum, minimum, etc.) of the privacy protection budget information supported by each of the N client nodes as the privacy protection processing factor corresponding to each client node.

[0206] S602: The first client node performs model training according to the federated learning initial model parameters and the federated learning configuration information to obtain local model parameters.

[0207] In this process, after the first client node receives the first model training message, it uses the federated learning initial model parameters in the first model training message as the model parameters of the local model. It then performs model training based on its own local data and the federated learning configuration parameters in the first model training message to obtain local model parameters. In this embodiment of the present application, the model parameters can be model gradients or model weights, and the local data includes training data.

[0208] S603: The first client node performs privacy protection processing on the local model parameters according to the privacy protection configuration information to obtain private model parameters.

[0209] In this process, after the first client node obtains the local model parameters through training, the first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information in the first model training message and the parameters indicated by the privacy protection processing factor, thereby obtaining the privacy model parameters.

[0210] In one possible implementation, the privacy protection processing type information indicates a first processing method. Assuming that the first processing method is a differential privacy noise addition method based on a Gaussian mechanism, the first client node adds the current round noise to the local model parameters based on the current round noise distribution parameters indicated by the privacy protection processing factor, thereby obtaining the privacy model parameters. An example is shown in the following formula (1):

[0211] in, is the privacy model parameter of the first client node in the tth round, W t is the local model parameter of the first client node in the tth round, N t is the noise value of the first client node in the tth round, N t Obey (μ t , σ t 2 ), (μ t , σ t 2 ) is the Gaussian distribution interval of the first client node in the tth round. Based on the above description, it can be understood that in this step, t=1.

[0212] For example, in the tth round, the first client node starts from (μ t , σ t 2) Select a number of noise values ​​corresponding to the local model parameters or at least one noise value (such as 0.5); then the first client node increases each gradient in the local model parameters by 0.5, thereby obtaining the privacy model parameters corresponding to the local model parameters.

[0213] In another possible implementation, the privacy protection processing type information indicates a second processing method. Assuming that the second processing method is a gradient compression method based on the Top-K threshold method, the first client node performs gradient compression on the local model parameters according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor, thereby obtaining the privacy model parameters. An example is shown in the following formula (2):

[0214] in, is the privacy model parameter of the first client node in the tth round, W t is the local model parameter of the first client node in the tth round, where t=1 in this step; thr represents the threshold or compression factor or Top-k value of the current round, |W t |>thr means W t In this approach, the first client node selects gradients from the local model parameters that are greater than a threshold or fall within the compression factor or Top-k value range. Alternatively, the threshold represents a threshold for the gradient change, and the first client node uses the gradients in the local model parameters that are greater than the threshold as the privacy model parameters. Alternatively, the first client node sorts the gradients in the local model parameters according to their gradient change, and then selects the gradients that fall within the compression factor or Top-k value range in descending order of gradient change as the privacy model parameters. In other words, the first client node selects a portion of the gradients from the local model parameters as the privacy model parameters based on the threshold, compression factor, or Top-k value of the current round. In this approach, when t = 1, the gradient change in round t can be the difference between the gradients in the local model parameters in round t and the initial model parameters of federated learning. When t is greater than 1, the gradient change in round t can be the difference between the gradients in the local model parameters in round t and the gradients in the local model parameters in round t-1, or the difference between the gradients in the local model parameters in round t and gradients in previous rounds that were not selected as privacy model parameters.

[0215] For example, in the tth round, the first client node determines 70% of the gradients from the local model parameters based on the threshold, and then uses these 70% of the gradients as the privacy model parameters corresponding to the local model parameters; or the first client node selects the top 100 gradients in terms of gradient change from the local model parameters based on the Top-k value, and then uses these 100 gradients as the privacy model parameters corresponding to the local model parameters; or the first client node selects the top 80% of the gradients in terms of gradient change from the local model parameters based on the compression factor, and then uses these 80% of the gradients as the privacy model parameters corresponding to the local model parameters.

[0216] S604: The server node receives the privacy model parameters from the first client node to continue executing the federated learning task.

[0217] In this process, the server node receives the privacy model parameters from the first client node, and then performs the subsequent process of the federated learning task based on the privacy model parameters.

[0218] In the embodiments of the present application, due to differences in hardware performance, communication performance, model training duration, privacy protection processing duration, and other factors of each client node, the server node may not be able to receive privacy model parameters corresponding to all N client nodes, and can only receive privacy model parameters from M client nodes out of the N client nodes, and then perform model aggregation based on the privacy model parameters corresponding to the M client nodes, where M is a positive integer less than or equal to N. For example, in S601, the server node divides the N client nodes into two groups. Then, when implementing federated learning iterations for only one of the groups, the server node can only receive the privacy model parameters corresponding to this group of client nodes.

[0219] Optionally, after receiving the privacy model parameters from the M client nodes, the server node may aggregate the privacy model parameters corresponding to the M client nodes to obtain the model aggregation parameters for the next round. Furthermore, the server node may determine the model aggregation parameters based on the privacy model parameters corresponding to the M client nodes and the weights of the M client nodes. An example is shown in the following formula (3):

[0220] Among them, QW t+1 is the model aggregation parameter of the t+1th round, QW t is the model aggregation parameter of the tth round, P i is the weight of the i-th client node among the M client nodes, is the local model parameter of the i-th client node in the t-th round.

[0221] In one possible implementation, the weight of each client node may be determined by the server node based on the amount of training data of the N client nodes; wherein the weight is positively correlated with the amount of training data, and the sum of the weights of the N client nodes is 1.

[0222] In another possible implementation, the weight of each client node can be determined by the server node based on the data difference information of N client nodes; wherein the data difference information indicates the degree of difference between the training data of the client node and the training data corresponding to the initial model parameters of the federated learning (or the baseline training data used by the server node), the data difference information is negatively correlated with the weight (or the degree of difference between the training data of the client node and the baseline training data is positively correlated with the weight), and the sum of the weights of the N client nodes is 1. In this implementation, the data difference information of the client node can be represented by parameters such as L1 norm, L2 norm, cosine distance, and Hamming distance. The greater the degree of difference between the training data of the client node and the baseline training data, the greater the weight of the client node, indicating that the contribution of the training data of the client node is greater, thereby improving the training efficiency of federated learning. Taking the data difference information as the L2 norm as an example, the following formula (4) is shown as an example:

[0223] Among them, P i is the weight of the i-th client node among N client nodes, is the data difference information in the form of L2 norm of the i-th client node among the N client nodes.

[0224] Optionally, after receiving the privacy model parameters from the M client nodes, the server node may further determine the model aggregation parameter based on H client nodes among the M client nodes and the weights of the H client nodes; wherein the degree of model accuracy degradation of any of the H client nodes is less than or equal to a model accuracy degradation threshold. It is understood that the model accuracy degradation threshold can be set based on the privacy protection budget information of the federated learning task, such as a model accuracy degradation threshold of 5%. The model accuracy degradation threshold is not specifically limited herein.

[0225] In one possible implementation, the degree of model accuracy degradation (or model accuracy information) of any client node among the M client nodes can be calculated by the client node and sent to the server node. Taking the first client node as an example, in any round, the first model accuracy and the second model accuracy of the first client node are used to calculate the degree of model accuracy degradation of the first client node. Wherein, q represents the degree of decline in the model accuracy of the first client node; q1 is the first model accuracy of the first client node, which represents the accuracy of the model prediction made by the first client node using the local model parameters; q2 is the second model accuracy of the first client node, which represents the accuracy of the model prediction made by the first client node using the privacy model parameters.

[0226] In another possible implementation, after the first client node determines the first model accuracy and the second model accuracy, it sends the first model accuracy and the second model accuracy as model accuracy information to the server node. The server node then determines the degree of decline in model accuracy of the first client node based on the first model accuracy and the second model accuracy of the first client node, and then selects H client nodes that perform model aggregation from the M client nodes. The specific calculation process is described above and will not be repeated here.

[0227] Based on the above method, the server node selects client nodes with less model accuracy degradation to participate in model aggregation, so as to reduce the difference between the non-real model aggregation parameters and the real model aggregation parameters obtained after model aggregation, increase the availability of non-real model aggregation parameters, and reduce the error of client nodes in the next round of model training.

[0228] In an embodiment of the present application, after performing model aggregation, the server node obtains model aggregation parameters for implementing the next round of iteration of federated learning.

[0229] In one possible implementation, the server node subtracts the first error from the model aggregation parameter to obtain the model aggregation parameter after subtracting the first error, and then sends the model aggregation parameter after subtracting the first error to M client nodes or N client nodes, so that the M client nodes or N client nodes perform the next round of federated learning based on the model aggregation parameter after subtracting the first error. The first error is determined based on the weights of the M client nodes and the noise estimation values ​​of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node based on the privacy protection budget information required by the federated learning task; for example, referring to formula (1) in S603, the noise estimation value E of the first client node in the tth round is t is (μ t , σ t 2 )’s expected value, such as E t =μ t .

[0230] In this implementation, by subtracting the first error from the model aggregation parameter, the difference between the non-real model aggregation parameter and the real model aggregation parameter is reduced, thereby increasing the availability of the non-real model aggregation parameter and increasing the accuracy of the client node in the next round of federated learning.

[0231] In one possible implementation, the server node performs privacy protection processing on the model aggregate parameters to obtain the privacy-protected model aggregate parameters. The server node then sends the privacy-protected model aggregate parameters to M or N client nodes, so that the M or N client nodes perform the next round of federated learning based on the privacy-protected model aggregate parameters. The server node performs privacy protection processing in the same manner as the client node performs privacy protection processing in S603 and is not further described here.

[0232] In this implementation, privacy protection processing is performed on the non-real model aggregation parameters, further increasing the difference between the non-real model aggregation parameters and the real model aggregation parameters, and further ensuring the privacy and security of user data.

[0233] In one possible implementation, the server node directly sends the model aggregation parameters to M client nodes or N client nodes, so that the M client nodes or N client nodes perform the next round of federated learning based on the model aggregation parameters (refer to S602 and S603).

[0234] In this implementation, no processing is performed on the non-real model aggregation parameters, and the efficiency of federated learning is improved by directly sending the non-real model aggregation parameters to M client nodes or N client nodes.

[0235] Through this method, the server node performs the federated learning aggregation process based on the client node's private model parameters, obtaining non-real model aggregation parameters. This allows iteration using these non-real model parameters during the federated learning process. Therefore, an attacker, either the server or client node, cannot calculate the true difference between the input sample x' and label y' and the client node's actual local data based on the non-real model parameters, and thus cannot deduce the local data of other client nodes. This prevents the leakage of other client nodes' local data and ensures the privacy and security of users' local data.

[0236] In order to better illustrate the above technical solution, the federated learning method provided in this application is further explained below in combination with implementation methods 1 to 3.

[0237] Implementation method 1:

[0238] Implementation 1 is described using the scenario in FIG. 3A as an example. Referring to FIG. 7 , the process includes:

[0239] S701: The server node (Server NWDAF) receives a federated learning service request message, which includes information required to perform a federated learning task. Exemplarily, the federated learning service request message is sent by a consumer NF.

[0240] Referring to the description of S401 above, the federated learning service request message includes the conditions that must be met to execute the federated learning task. After responding to the federated learning service request message, the server node first determines whether its own configuration information can meet the conditions. If so, it executes the subsequent steps; otherwise, it rejects the request.

[0241] Optionally, the configuration information of the server node includes one or more of the following information:

[0242] The privacy protection processing type information of the server node is used to indicate the privacy protection processing method of the model aggregation parameters supported by the server node, and can also indicate that the server node has privacy protection processing capabilities.

[0243] The privacy protection budget information of the server node is used to indicate the degree of degradation of the model accuracy of the global model corresponding to the server node, and can also indicate the degree of protection of the model aggregation parameters by the privacy protection processing performed by the server node.

[0244] The identifier of the model analysis use case of the server node (Analytics ID) is used to indicate the model use cases (or global model use cases) supported by the server node.

[0245] The service area of ​​the server node is used to indicate the coverage of the global model of the server node.

[0246] The federated learning type of the server node is used to indicate the federated learning type supported by the server node.

[0247] Optionally, if the server node has privacy protection capabilities and the privacy protection capabilities of the server node support the privacy protection processing method indicated by the federated learning service request message, the server node executes subsequent steps.

[0248] Optionally, if the degree of model accuracy degradation indicated by the privacy protection budget of the server node meets the model accuracy degradation range required by the federated learning task, the server node executes subsequent steps.

[0249] Optionally, if the identifier of the model analysis use case of the server node includes the identifier of the model analysis use case of the federated learning task, and the service area of ​​the server node is within the service area of ​​the federated learning task, the server node executes subsequent steps.

[0250] S702: The server node determines N client nodes (Client NWDAF) that perform the federated learning task through the first device in response to the federated learning service request message.

[0251] In this process, the federated learning service request message may include: privacy protection indication information, privacy protection processing type information (e.g., the indicated privacy protection processing method is the first processing method), and model accuracy degradation range information: differential privacy budget ε = 15 (indicating that the model accuracy degradation range is 0-5%). Optionally, the federated learning service request message also includes one or more of the information in (4) to (10) in S401.

[0252] Referring to the process description of S403 in Figure 4 above, the service node sends a first request message to the first device, so that the first device determines N client nodes that perform the federated learning task based on the first request message, and the configuration information of the N client nodes meets the corresponding requirements for performing the federated learning task indicated by the federated learning service request message.

[0253] The following example uses the first client node among N client nodes. In this process, in the configuration information of the first client node, the privacy protection processing information of the first client node indicates that it has privacy protection processing capabilities; the privacy protection processing method supported by the first client node is the differential privacy noising method based on the Gaussian mechanism; and the privacy protection budget of the first client node is: differential privacy budget ε = 16 (indicating that the model accuracy of the local model has decreased by 4%, that is, the model accuracy of the first client node has decreased by 4%, that is, the model accuracy of the first client node meets the above-mentioned model accuracy decrease range information). Optionally, other configuration information of the first client node meets the corresponding requirements for executing the federated learning task indicated by the federated learning service request message.

[0254] Optionally, the N client nodes perform privacy protection processing in the same way to facilitate aggregated computing of subsequent federated learning tasks and improve the efficiency of federated learning.

[0255] Optionally, after the server node determines the N client nodes that execute the federated learning task, it feeds back a response message to the consumer NF; the response message may include feedback that the federated learning task is successfully established.

[0256] S703: The server node determines a learning round of the federated learning task, and determines, according to the learning round, noise distribution parameters to be added to the model parameters when the N client nodes execute the first processing method.

[0257] In this process, the learning rounds of N client nodes are the same, or in any round, N client nodes participate in model aggregation together.

[0258] Referring to the description of the learning round in S601 above, taking the first client node as an example, the server node determines that the learning round K of the first client node is 20, determines the differential privacy budget ε for the first client node to be 40, and then determines the differential privacy budget ε1, ε2, ..., ε1 of each round based on the preset rules (which can be the average rule, the small to large rule, etc.). t 、……、ε 10 . Take the current round t as 1 for example, assuming ε t=1 is 2. Then the server node determines the Gaussian distribution interval of the current round Finally, the Gaussian distribution interval (0, 0.25) of the current round is used as the noise distribution parameter added to the model parameters when the first client node executes the first processing method.

[0259] S704: The server node sends a model training message to each of the N client nodes. The model training message includes the initial model parameters of the federated learning, the federated learning configuration information corresponding to the client node, and the privacy protection configuration information of the client node.

[0260] Taking the first client node as an example, in this process, the federated learning configuration information corresponding to the first client node includes the first client node's learning round; the first client node's privacy protection configuration information includes information indicating that the privacy protection processing method is differential privacy noise based on the Gaussian mechanism, and the Gaussian distribution interval of the current round (0, 0.25). The federated learning initial model parameters can also be understood as the model parameters used by the first client node during the first round of model training.

[0261] The first model training message of the first client node may also include baseline data of the server node, so that the first client node determines data difference information between the baseline data set and its own training data, and the data difference information is used to determine the weight of the client node.

[0262] S705: The client node performs model training according to the model training message, and performs privacy protection processing on the local model parameters obtained through training.

[0263] Based on S704, taking the first client node as an example, the first client node uses the initial federated learning model parameters as the model parameters for the local model during the first round of model training. It then obtains its own training data from the data provider's network file and inputs its own training data into the local model. After the local model converges or training is completed, the trained local model parameters are obtained. In this process, there is a one-to-one correspondence between client nodes and data providers.

[0264] After the first client node determines the local model parameters, the first client node selects at least one noise value N (such as 0.1) from (0, 0.25); then the first client node increases each gradient in the local model parameters by 0.1, thereby obtaining the privacy model parameters corresponding to the local model parameters.

[0265] S706: The client node predicts model accuracy information, where the model accuracy information indicates the degree of model accuracy degradation caused by the privacy protection process.

[0266] Based on S705, taking the first client node as an example, the first client node uses the local model parameters as the gradient of the local model, then obtains the corresponding prediction samples of its own from its own network file, and then inputs its own prediction samples into the local model, obtaining a first model accuracy of 80% corresponding to the local model parameters; similarly, the first client node uses the privacy model parameters after privacy protection processing as the gradient of the local model, and then inputs its own prediction samples into the local model, obtaining a second model accuracy of 78% corresponding to the privacy model parameters.

[0267] S707: The client node sends privacy model parameters and model accuracy information.

[0268] Taking the first client node as an example, the first client node sends its corresponding first model accuracy, second model accuracy and privacy model parameters to the server node.

[0269] Based on S704 , the first client node also sends its corresponding data difference information to the server node.

[0270] S708: The server node determines H client nodes that participate in model aggregation based on the model accuracy information of the client node, and performs model aggregation based on the local model parameters of the H client nodes to obtain model aggregation parameters.

[0271] In this process, because some of the N client nodes do not send their local model parameters to the server node within the current round of model aggregation (or the current round of reporting period, or the current round of learning duration in the federated learning task) due to factors such as incomplete model training, the server node can only receive the private model parameters corresponding to M client nodes out of the N client nodes. Based on this, client nodes that do not send their local model parameters to the server node within the current round of model aggregation will send model training progress messages to the server node. After receiving the model training progress messages, the server node can then determine whether the client node will participate in the subsequent federated learning process.

[0272] After receiving the privacy model parameters corresponding to the M client nodes, the server node determines H client nodes from the M client nodes to participate in the model aggregation based on the model accuracy information corresponding to the M client nodes. Assuming that the model accuracy drop threshold is 5%, taking the i-th client node among the M client nodes as an example, the model accuracy drop degree of the i-th client node is Therefore, the i-th client node serves as one of the H client nodes participating in model aggregation.

[0273] Based on S707, the server node may receive data difference information corresponding to the M client nodes, and then determine the weights corresponding to the M client nodes based on the data difference information corresponding to the M client nodes, or determine the weights corresponding to the H client nodes based on the data difference information corresponding to the H client nodes. Therefore, after the server node determines the H client nodes, it may also determine the weights corresponding to the H client nodes, and then perform model aggregation based on the weights corresponding to the H client nodes and the privacy model parameters of the H client nodes to obtain the model aggregation parameters (see S603 above).

[0274] S709: The server node subtracts the first error from the model aggregation parameter to obtain the model aggregation parameter after subtracting the first error.

[0275] Referring to the description of the above formula (3), in this process, the model aggregation parameter after subtracting the first error is calculated according to the following formula (5):

[0276] in, is the model aggregation parameter after subtracting the first error for the t+1th round of model training, QW t+1 is the model aggregation parameter of the t+1th round, P i is the weight of the i-th client node among the H client nodes, E iis the expectation of the Gaussian distribution interval corresponding to the i-th client node; in this process, the Gaussian distribution interval (0, 0.25) is taken as an example, such as E i =0.

[0277] S710: The server node sends the next round of model training messages to N client nodes respectively. The next round of model training messages includes the model aggregation parameters used for the next round of model training after subtracting the first error; the next round of model training messages also includes the federated learning configuration information corresponding to the client node and the privacy protection configuration information of the client node.

[0278] Taking the first client node as an example, the privacy protection configuration information corresponding to the first client node in this process includes information indicating that the privacy protection processing method is the differential privacy noise method based on the Gaussian mechanism and the Gaussian distribution interval of the next round.

[0279] S711: Repeat S705-S710 until the server node determines that the federated learning termination conditions are met and terminates the federated learning. The federated learning termination conditions include reaching the model training duration required by the federated learning task, or receiving an end indication message from the consumer NF, or the global model convergence of the server node, or the local model convergence of any client node.

[0280] In this process, the server node terminates federated learning based on the federated learning termination conditions. Optionally, the server node terminates federated learning when the learning duration reaches the model training duration required by the federated learning task; optionally, the server node terminates federated learning based on a termination indication message sent by the consumer NF; optionally, the server node terminates federated learning after the global model converges; optionally, the server node terminates federated learning after the local model converges. After the federated learning process concludes, the server node sends the model aggregation parameters of the last round of federated learning to the consumer NF and the N client nodes, respectively.

[0281] In implementation method 1, after receiving the model training message, the client node performs model training to obtain true local model parameters (or actual local model parameters), then adds noise to the local model parameters to obtain false local model parameters (or false local model parameters). This prevents the server node from deducing the client node's training data based on the true local model parameters. Because the privacy protection budget used by the client node for privacy protection processing meets the privacy protection budget required by the federated learning task, the false local model parameters are similar to the true local model parameters, or in other words, the false local model parameters meet the model aggregation requirements of the federated learning task. Therefore, the server node can continue to perform subsequent federated learning processes based on the false local model parameters on the basis of meeting the requirements of the federated learning task. Since the server node performs model aggregation based on the false local model parameters, the model aggregation parameters obtained after the server node model aggregation are also false (or false model aggregation parameters are obtained), thereby preventing attackers in multiple client nodes from deducing the local data of other client nodes based on the true model aggregation parameters. Based on this, the server node or client node acting as an attacker cannot deduce the local data of other client nodes through unrealistic model parameters, thereby avoiding the leakage of local data of other client nodes and ensuring the privacy and security of users' local data.

[0282] Implementation 2:

[0283] Implementation method 2 is described using the scenario of FIG3B as an example. It is a description based on implementation method 1 applied to a different scenario. For similarities, please refer to the relevant description in conjunction with FIG7 , which will not be repeated below.

[0284] Implementation 2 can be applied to the scenario in FIG3B . Referring to FIG8 , the process includes:

[0285] S801: The server node (Server EMS) receives a federated learning service request message, where the federated learning service request message includes information required to execute a federated learning task in a first processing mode.

[0286] Referring to the description of S701 above, in this process, if the configuration information of the server node itself meets the information required to perform the federated learning task, the subsequent steps are executed.

[0287] S802: The server node sends a first request message to at least two client nodes respectively, where the first request message includes the federated learning service request message, and is used to instruct the client node that performs the federated learning task in the first processing mode to return a response message, where the response message includes the configuration information of the client node.

[0288] Referring to the description of S702 above, assuming that the configuration information of any of the at least two client nodes (using the first client node in this process as an example) satisfies the federated learning service request message, then both client nodes return a response message. In other words, the at least two client nodes are equivalent to the N client nodes (Client RTLF) executing the federated learning task.

[0289] S803: The server node determines a learning round of the federated learning task, and determines, according to the learning round, noise distribution parameters to be added to the model parameters when the N client nodes execute the first processing method.

[0290] In this process, to improve the efficiency of client nodes in executing federated learning tasks, the server node divides N client nodes into multiple groups. For ease of description, this process exemplifies the division of N client nodes into two groups, but the number of groups is not limited. The sum of the model training duration and the privacy protection processing duration for any client node in the first group is less than the sum of the model training duration and the privacy protection processing duration for any client node in the second group.

[0291] Referring to the description of learning rounds in S601 above, it is assumed that the learning rounds K1 of the first group of client nodes is 20, and the learning rounds K2 of the second group of client nodes is 10. Based on this, for any client node in the first group of client nodes (this process takes the first client node as an example), the server node determines the differential privacy budget ε for the first client node as 40, and then determines the differential privacy budgets ε1, ε2, ..., ε1 of each round of the first client node. t 、……、ε 10 . Take the current round t as 1 for example, assuming ε t=1 is 2. Then the server node determines the Gaussian distribution interval of the current round Finally, the Gaussian distribution interval (0, 0.25) of the current round is used as the noise distribution parameter added to the model parameters when the first client node executes the first processing method.

[0292] For any client node in the second group of client nodes (this process takes the second client node as an example), the server node determines the differential privacy budget ε for the second client node as 50, and then determines the differential privacy budget ε1, ε2, ..., ε t 、……、ε 20 . Take the current round t as 1 for example, assuming ε t=1 is 5. Then the server node determines the Gaussian distribution interval of the current round Finally, the Gaussian distribution interval (0, 0.04) of the current round is used as the noise distribution parameter added to the model parameters when the second client node executes the first processing method.

[0293] S804: The server node sends a model training message to each of the N client nodes. The model training message includes the initial model parameters of federated learning, the federated learning configuration information corresponding to the client node, and the privacy protection configuration information for the client node to execute the first processing method.

[0294] Based on the above S803, the federated learning configuration information corresponding to the first client node includes the learning round of the first client node; the privacy protection configuration information of the first client node includes information indicating that the privacy protection processing method is a differential privacy noisy method based on the Gaussian mechanism and the Gaussian distribution interval (0, 0.25) of the current round.

[0295] The federated learning configuration information corresponding to the second client node includes the learning round of the second client node; the privacy protection configuration information of the second client node includes information indicating that the privacy protection processing method is a differential privacy noising method based on the Gaussian mechanism and the Gaussian distribution interval of the current round (0, 0.04).

[0296] S805: N client nodes perform model training respectively according to the model training message, and perform privacy protection processing on the local model parameters obtained through training using the first processing method.

[0297] Based on S804, the first client node uses the initial federated learning model parameters as the model parameters for the local model during the first round of model training. It then inputs its own training data into the local model. After the local model converges or training is completed, the trained local model parameters are obtained. After the first client node determines its own local model parameters, it selects at least one noise value N1 (e.g., 0.1) from (0, 0.25). The first client node then increases each gradient in the local model parameters by 0.1 to obtain the privacy model parameters corresponding to the local model parameters.

[0298] The second client node uses the initial federated learning model parameters as the model parameters for the local model during the first round of model training. It then inputs its own training data into the local model. After the local model converges or training completes, it obtains the second local model parameters after model training. After the second client node determines its own second local model parameters, the first client node selects at least one noise value N2 (e.g., 0.02) from (0, 0.04); the second client node then increases each gradient in the second local model parameters by 0.02 to obtain the second privacy model parameters corresponding to the second local model parameters.

[0299] S806: N client nodes respectively predict model accuracy information, where the model accuracy information indicates a degree of degradation of the model accuracy caused by the privacy protection processing performed in the first processing manner.

[0300] Referring to S706 , the manner in which any client node in the first group of client nodes and any client node in the second group of client nodes performs model prediction is the same, which will not be described in detail here.

[0301] S807: N client nodes send privacy model parameters and model accuracy information.

[0302] The first client node sends the first model accuracy of the local model parameters, the second model accuracy of the privacy model parameters, and the privacy model parameters corresponding to itself to the server node.

[0303] Similarly, the second client node sends the first model accuracy of its own corresponding local model parameters, the second model accuracy of the second privacy model parameters, and the privacy model parameters to the server node.

[0304] S808: The server node determines the client nodes participating in model aggregation according to the model accuracy information of the client nodes, and performs model aggregation according to the local model parameters of the client nodes participating in model aggregation to obtain model aggregation parameters.

[0305] In this process, since the learning round of the first group of client nodes is greater than the learning round of the second group of client nodes, in the 1st, 3rd, 5th,... rounds of the first group of client nodes, the server node only determines the client nodes participating in the model aggregation from the first group of client nodes, and then performs model aggregation based on the local model parameters of the client nodes participating in the model aggregation at this time, and obtains the model aggregation parameters used by the group client nodes in the 2nd, 4th, 6th,... rounds.

[0306] During the 2nd, 4th, 6th, ... rounds of the first group of client nodes (or during the 1st, 2nd, 3rd, ... rounds of the second group of client nodes), the server node determines the client nodes participating in model aggregation from the first group of client nodes and the second group of client nodes (equivalent to N client nodes), and then performs model aggregation based on the local model parameters of the client nodes participating in the model aggregation at this time, to obtain the model aggregation parameters used by the first group of client nodes in the 3rd, 5th, 7th, ... rounds, or to obtain the model aggregation parameters used by the second group of client nodes in the 2nd, 3rd, 4th, ... rounds.

[0307] S809: The server node subtracts the first error from the model aggregation parameter to obtain the model aggregation parameter after subtracting the first error.

[0308] Based on S808 above, during the first, third, fifth, ... rounds of the first group of client nodes, the server node subtracts the first error from the model aggregation parameters used by the first group of client nodes during the second, fourth, sixth, ... rounds, to obtain the model aggregation parameters used by the first group of client nodes during the second, fourth, sixth, ... rounds after subtracting the first error. The first error at this time is determined based on the expectation of the Gaussian distribution interval corresponding to the first group of client nodes. For specific implementation, refer to the above formula (5).

[0309] During the second, fourth, sixth, ... rounds of the first group of client nodes, the server node subtracts the first error from the model aggregation parameters used by the first group of client nodes in the third, fifth, seventh, ... rounds, to obtain the model aggregation parameters used by the first group of client nodes in the third, fifth, seventh, ... rounds after subtracting the first error. The first error is determined based on the expectation of the Gaussian distribution intervals corresponding to the first and second groups of client nodes.

[0310] S810: The server node sends the next round of model training messages to N client nodes respectively. The next round of model training messages includes the model aggregation parameters used for the next round of model training after subtracting the first error, the federated learning configuration information corresponding to the client node, and the privacy protection configuration information of the client node.

[0311] With reference to S710 , the specific implementation of this step will not be described in detail.

[0312] S811: Repeat S805-S810 until the server node determines that the federated learning termination condition is met and terminates the federated learning. Refer to the federated learning termination conditions and the process of terminating the federated learning described in S711 above, which will not be repeated here.

[0313] In Implementation 2, while ensuring the privacy and security of the user's local data, the server node can group the N client nodes to identify at least two groups of client nodes with different total durations. Based on this, the server node can perform more rounds of model training on the first group of client nodes with the shorter total durations, reducing the model training wait time for the first group of client nodes. This allows for flexible selection of client nodes for federated learning tasks and improves the efficiency of client nodes in executing federated learning tasks.

[0314] Implementation 3:

[0315] Implementation method 3 is described using Figure 3C as an example. It is a description of the way of applying different privacy protection processing based on implementation method 2. For similarities, please refer to the relevant description in combination with Figure 7 and / or Figure 8, which will not be repeated below.

[0316] Implementation 3 can be applied to the scenario in FIG3C . Referring to FIG9 , the process includes:

[0317] S901: The server node (Server RTLF) receives a federated learning service request message, where the federated learning service request message includes information required to execute a federated learning task in a second processing mode.

[0318] Referring to the description of S701 above, in this process, the configuration information of Server RTLF itself meets the information required to execute the federated learning task, and the subsequent steps are executed.

[0319] S902: The server node sends a first request message to at least two client nodes respectively, where the first request message includes the federated learning service request message, and is used to instruct the client node that performs the federated learning task in the second processing mode to return a response message, where the response message includes the configuration information of the client node.

[0320] With reference to the description of S802 above, it is assumed that at least two client nodes are equivalent to N client nodes (Client RTLF) that execute the federated learning task.

[0321] S903: The server node determines a learning round of the federated learning task, and determines, according to the learning round, a threshold value or a compression factor or a Top-k value used by the N client nodes when executing the second processing method.

[0322] Referring to the description of S802 above, the server node divides N client nodes into two groups, wherein the learning rounds K1 of the first group of client nodes are 20, and the learning rounds K2 of the second group of client nodes are 10.

[0323] With reference to the description of S401 above, the server node may determine a corresponding threshold value, compression factor, or Top-k value for each client node based on the privacy protection budget required by the federated learning task; that is, the threshold value, compression factor, or Top-k value for any two client nodes is different. Alternatively, the server node may determine a corresponding threshold value, compression factor, or Top-k value for each group of client nodes based on the privacy protection budget required by the federated learning task; that is, the threshold value, compression factor, or Top-k value for any two groups of client nodes is different, but the threshold value, compression factor, or Top-k value for the client nodes in the same group is the same. Alternatively, the server node may determine a unified threshold value, compression factor, or Top-k value for N client nodes based on the privacy protection budget required by the federated learning task.

[0324] For ease of description, this process takes a unified threshold, compression factor, or Top-k value as an example, assuming that the threshold is 0.01, or the compression factor is 80%, or the Top-k value is 100.

[0325] S904: The server node sends a model training message to each of the N client nodes. The model training message includes the initial model parameters of federated learning, the federated learning configuration information corresponding to the client node, and the privacy protection configuration information for the client node to execute the second processing method.

[0326] Based on the above S903, in the first model training message corresponding to the first client node, the federated learning configuration information includes the learning round of the first client node, and the privacy protection configuration information includes the determined threshold (0.01) or compression factor (80%) or Top-k value (100).

[0327] In the second model training message corresponding to the second client node, the federated learning configuration information includes the learning round of the second client node, and the privacy protection configuration information includes a determined threshold (0.01) or a compression factor (80%) or a Top-k value (100).

[0328] S905: N client nodes perform model training respectively according to the model training message, and perform privacy protection processing on the local model parameters obtained through training using the second processing method.

[0329] Based on S904, the first client node performs model training based on the federated learning initial model parameters, the federated learning configuration information, and its own training data to obtain local model parameters. The first client node then determines the degree of change between each gradient in the local model parameters and each gradient corresponding to the federated learning initial model parameters, and uses the gradient in the local model parameters with a degree of change greater than 0.01 (i.e., the threshold) as the privacy model parameter corresponding to the local model parameter. Alternatively, based on the order of the degree of change from large to small, 80% (i.e., the compression factor) of the gradient is selected from the local model parameters as the privacy model parameter corresponding to the local model parameter. Alternatively, based on the order of the degree of change from large to small, the top 100 (i.e., Top-k values) of the gradients are selected from the local model parameters as the privacy model parameters corresponding to the local model parameters.

[0330] S906: N client nodes respectively predict model accuracy information, where the model accuracy information indicates a degree of degradation of the model accuracy caused by the privacy protection processing performed in the second processing manner.

[0331] Referring to S706 , the manner in which any client node in the first group of client nodes and any client node in the second group of client nodes performs model prediction is the same, which will not be described in detail here.

[0332] S907: The client node sends privacy model parameters and model accuracy information.

[0333] S908: The server node determines the client nodes participating in model aggregation according to the model accuracy information of the client nodes, and performs model aggregation according to the local model parameters of the client nodes participating in the model aggregation to obtain model aggregation parameters.

[0334] With reference to S807 and S808 above, the model aggregation method in this process is the same as the implementation method described in S807 and S808 above, and will not be described in detail here.

[0335] S909: The server node sends the next round of model training messages to N client nodes respectively. The next round of model training messages includes model aggregation parameters, federated learning configuration information corresponding to the client node, and privacy protection configuration information of the client node.

[0336] With reference to S710 , the specific implementation of this step will not be described in detail.

[0337] S910: Repeat S905-S909 until the server node determines that the federated learning end condition is met and ends the federated learning.

[0338] Refer to the federated learning end section and the process of ending federated learning described in S711 above, which will not be repeated here.

[0339] In implementation method 3, after receiving the model training message, the client node performs model training to obtain the real local model parameters, and then selects part of the gradient from the real local model parameters according to the threshold or compression factor or Top-k value, and then uses this part of the gradient as the local model parameter after privacy protection (or privacy model parameter). Since the privacy model parameter is only a part of the gradient in the real local model parameter, the privacy model parameter is an incomplete local model parameter, which is equivalent to a non-real local model parameter, thereby preventing the server node from deducing the local data of the client node based on the real local model parameter. Similarly, the model aggregation parameter obtained by the server node after model aggregation based on the privacy model parameter is also non-real, thereby preventing attackers in multiple client nodes from deducing the local data of other client nodes based on the real model aggregation parameter, thereby avoiding the leakage of local data of other client nodes and ensuring the privacy and security of users' local data.

[0340] It should be noted that each step involved in the above embodiments or examples can be performed by a corresponding device, or by a module, chip, processor, or chip system within the device, and the embodiments of the present application do not limit this. The above embodiments are only described as being performed by a corresponding device. In addition, the specific implementation methods or examples in the above embodiments do not limit the solutions provided in the embodiments of the present application.

[0341] Optionally, in each of the above embodiments, some steps may be selected for implementation, and the order of the steps in the diagrams may be adjusted for implementation, and this application does not limit this. It should be understood that executing some of the steps in the diagrams, adjusting the order of the steps, or combining them for specific implementation all fall within the scope of protection of this application.

[0342] It is understandable that in order to implement the functions in the above embodiments, the various devices involved in the above embodiments include hardware structures and / or software modules corresponding to the execution of each function. It should be readily apparent to those skilled in the art that, in combination with the units and method steps of the various examples described in the embodiments disclosed in this application, this application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a function is executed in the form of hardware or computer software driving hardware depends on the specific application scenario and design constraints of the technical solution.

[0343] It can be understood that the above-mentioned network architecture and application scenarios described in the embodiments of the present application are intended to more clearly illustrate the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided by the embodiments of the present invention. Ordinary technicians in this field can know that with the evolution of network architecture and the emergence of new services, the technical solutions provided by the embodiments of the present invention are also applicable to similar technical problems.

[0344] It should be noted that the "steps" in the embodiments of this application are merely illustrative and serve as a method of expression for a better understanding of the embodiments. They do not constitute a substantial limitation on the execution of the solutions of this application. For example, the "steps" can also be understood as "features." Furthermore, the steps do not constitute any limitation on the execution order of the solutions of this application. Any changes to the order of steps, or any operations such as step merging or step splitting that do not affect the implementation of the overall solution, resulting in new technical solutions, are also within the scope of this application.

[0345] Based on the same technical concept, this application also provides a federated learning device, which can be applied to the federated learning system shown in Figures 3A-3C. The federated learning device is used to implement the methods provided in the above embodiments and can be applied to the server nodes or client nodes involved in the above embodiments. Referring to Figure 10, the federated learning device 1000 includes a communication unit 1001 and a processing unit 1002.

[0346] The communication unit 1001 is used to receive and send data, and supports the federated learning device 1000 to communicate with other devices.

[0347] The processing unit 1002 is used to control and manage the actions of the federated learning device 1000 and execute the steps performed by the server node or the client node in the federated learning method provided in each of the above embodiments or examples.

[0348] Optionally, the federated learning device 1000 further includes a storage unit for storing program code and / or data of the federated learning device 1000 .

[0349] The communication unit 1001 can be referred to as an input / output unit, a transceiver unit, etc., and can be a transceiver or a communication interface; the processing unit 1002 can be a processor. When the federated learning device 1000 is a module (e.g., a chip) in a communication device, the communication unit 1001 can be an input / output interface, an input / output circuit, or an input / output pin, etc., and can also be referred to as an interface, a communication interface, or an interface circuit; the processing unit 1002 can be a processor, a processing circuit, or a logic circuit, etc.

[0350] In one embodiment, the federated learning device 1000 may be applied to the server node of the embodiment shown in Figure 6. The processing unit 1002 is configured to:

[0351] Determine N client nodes that execute the federated learning task, where N is an integer greater than 1;

[0352] According to the communication unit 1001, a first model training message is sent to the first client node among the N client nodes; the first model training message includes federated learning initial model parameters, federated learning configuration information and privacy protection configuration information; the federated learning initial model parameters and the federated learning configuration information are used for model training, and the privacy protection configuration information is used for privacy protection processing; according to the communication unit 1001, the privacy model parameters are received from the first client node; the privacy model parameters are obtained by the first client node based on model training and privacy protection processing.

[0353] Optionally, the privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method used by the first client node, and the privacy protection processing factor indicates the parameters used by the first client node when performing privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

[0354] Optionally, the processing unit 1002 is specifically configured to:

[0355] Determine a privacy protection processing factor for each of the N client nodes based on privacy protection budget information supported by each of the N client nodes; the privacy protection budget information supported by a first client node among the N client nodes represents a degree of decrease in model accuracy caused by the first client node performing privacy protection processing.

[0356] Optionally, the privacy protection processing type information indicates a first processing method, which is a noise addition method; the privacy protection processing factor indicates that a noise distribution parameter is added to the model parameter when the first processing method is performed on the local model parameter in each round in K rounds.

[0357] Optionally, the privacy protection processing type information indicates a second processing method, which is a gradient compression method; the privacy protection processing factor indicates a threshold or compression factor or Top-k value used when performing the second processing method on the local model parameters in each of the K rounds.

[0358] Optionally, the processing unit 1002 is specifically configured to:

[0359] A federated learning service request message is received according to the communication unit 1001, where the federated learning service request message includes privacy protection indication information, where the privacy protection indication information indicates that the client node executing the federated learning task needs to perform privacy protection processing on the local model parameters; N client nodes executing the federated learning task are determined according to the privacy protection indication information, where the N client nodes have a privacy protection processing capability, where the privacy protection processing capability is the ability to perform privacy protection processing on the local model parameters.

[0360] Optionally, the federated service request message further includes model accuracy degradation range information; the processing unit 1002 is specifically configured to:

[0361] According to the model accuracy degradation range information, N client nodes that execute the federated learning task are determined, and the degree of model accuracy degradation caused by the N client nodes performing privacy protection processing meets the model accuracy degradation range information required by the federated learning task.

[0362] Optionally, the processing unit 1002 is specifically configured to:

[0363] Sending a first request message to the first device according to the communication unit 1001; the first request message is used to obtain configuration information of a client node that meets a first condition, the first condition at least including having a privacy protection processing capability and a degree of model accuracy degradation caused by performing the privacy protection processing that meets the model accuracy degradation range required by the federated learning task; the first device is used to store the configuration information of the client node;

[0364] According to the communication unit 1001, a first response message is received from the first device, where the first response message includes configuration information of N client nodes that meet the first condition, and the configuration information corresponding to the N client nodes includes privacy protection processing capability information.

[0365] Optionally, the processing unit 1002 is specifically configured to:

[0366] According to the communication unit 1001, a first request message is sent to at least two client nodes respectively; corresponding first response messages are received from N client nodes that meet the first condition respectively, and the first response message corresponding to each client node in the N client nodes includes configuration information corresponding to the client node.

[0367] Optionally, the first request message includes one or more of the following information:

[0368] Privacy protection processing type information, the privacy protection processing type information indicates the privacy protection processing method required by the federated learning task; the threshold value of the user privacy feature ratio, the threshold value of the user privacy feature ratio indicates the minimum ratio of private data in the training data used by the client node that executes the federated learning task; the threshold value of the data set size, the threshold value of the data set size indicates the minimum amount of data of the training data used by the client node that executes the federated learning task; the threshold value of the model training time, the threshold value of the model training time indicates the maximum value of the model training time of each round of the client node that executes the federated learning task; the threshold value of the privacy protection processing time, the threshold value of the privacy protection processing time indicates the maximum value of the privacy protection processing time of each round of the client node that executes the federated learning task.

[0369] Optionally, the processing unit 1002 is further configured to:

[0370] According to the communication unit 1001, privacy model parameters are received from M client nodes among the N client nodes, where M is a positive integer less than or equal to N; the privacy model parameters corresponding to the M client nodes are aggregated to obtain model aggregation parameters; a first error is subtracted from the model aggregation parameters to obtain the model aggregation parameters after subtracting the first error; the first error is determined based on the weights of the M client nodes and the noise estimation values ​​of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node based on the privacy protection budget information required by the federated learning task; according to the communication unit 1001, the model aggregation parameters after subtracting the first error are sent to the M client nodes or the N client nodes respectively.

[0371] Optionally, the processing unit 1002 is specifically configured to:

[0372] Determine H client nodes from the M client nodes, where a degree of model accuracy degradation of any client node among the H client nodes is less than or equal to a model accuracy degradation threshold, where H is a positive integer less than or equal to M; and aggregate the privacy model parameters corresponding to the H client nodes according to the privacy model parameters corresponding to the H client nodes and the weights of the H client nodes.

[0373] Optionally, the processing unit 1002 is specifically configured to:

[0374] According to the communication unit 1001, model accuracy information is received from the M client nodes, where the model accuracy information indicates a degree of model accuracy degradation caused by the client nodes performing privacy protection processing on the trained local model parameters; and the H client nodes whose degree of model accuracy degradation is less than or equal to the model accuracy degradation threshold are determined from the M client nodes.

[0375] In one embodiment, the federated learning apparatus 1000 may be applied to the first client node in the embodiment shown in Figure 6. The processing unit 1002 is configured to:

[0376] A first model training message is received from the server node through the communication unit 1001; wherein the first model training message includes federated learning initial model parameters, federated learning configuration information and privacy protection configuration information; model training is performed according to the federated learning initial model parameters and the federated learning configuration information to obtain local model parameters; privacy protection processing is performed on the local model parameters according to the privacy protection configuration information to obtain privacy model parameters; and the privacy model parameters are sent to the server node through the communication unit 1001.

[0377] Optionally, the privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor corresponding to the first client node; wherein the privacy protection processing type information indicates a privacy protection processing method to be used, and the privacy protection processing factor indicates a parameter used by the first client node when performing privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information;

[0378] The processing unit 1002 is specifically configured to:

[0379] Privacy protection processing is performed on the local model parameters according to the privacy protection processing mode indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor.

[0380] Optionally, the privacy protection processing type information indicates a first processing mode, and the first processing mode is a noise addition mode;

[0381] The processing unit 1002 is specifically configured to:

[0382] According to the noise distribution parameters of the current round indicated by the privacy protection processing factor, the noise of the current round is added to the local model parameters.

[0383] Optionally, the privacy protection processing type information indicates a second processing mode, and the second processing mode is a gradient compression mode;

[0384] The processing unit 1002 is specifically configured to:

[0385] Gradient compression is performed on the local model parameters according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor.

[0386] Optionally, the processing unit 1002 is further configured to:

[0387] Model accuracy information is sent to the server node through the communication unit 1001, where the model accuracy information indicates a degree of degradation of the model accuracy caused by the first client node performing privacy protection processing on the local model parameters.

[0388] It should be noted that the division of modules in the embodiments of the present application is illustrative and is merely a logical functional division. In actual implementation, other division methods may be used. Furthermore, the functional units in the various embodiments of the present application may be integrated into a single processing unit, or may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.

[0389] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) or a processor to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0390] Based on the above embodiments, embodiments of the present application further provide a federated learning device, which may be a server node or a client node in the federated learning system shown in Figures 3A-3C. The federated learning device can implement the methods in the above embodiments and has the functions of a federated learning apparatus 1000. Referring to Figure 11, the federated learning device 1100 includes a transceiver 1101, a processor 1102, and a memory 1103. The transceiver 1101, the processor 1102, and the memory 1103 are interconnected.

[0391] Optionally, the transceiver 1101, the processor 1102, and the memory 1103 are interconnected via a bus 1104. The bus 1104 may be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus may be classified as an address bus, a data bus, a control bus, etc. For ease of illustration, FIG11 shows only one thick line, but this does not mean that there is only one bus or only one type of bus.

[0392] The transceiver 1101 is used to receive and send signals to achieve communication with other devices.

[0393] The functions of the processor 1102 can be referred to the description in the above embodiments and will not be repeated here.

[0394] The processor 1102 may be a central processing unit (CPU), a network processor (NP), or a combination of a CPU and an NP. The processor 1102 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof. When implementing the above functions, the processor 1102 may be implemented through hardware, or may also execute corresponding software implementations through hardware. The steps of the method disclosed in the above embodiments of the present application may be directly executed by the processor 1102, or executed by a combination of hardware and software modules in the processor 1102.

[0395] The memory 1103 is used to store program instructions, data, etc. Specifically, the program instructions may include program code, which includes computer operation instructions. The memory 1103 may include volatile memory, such as random access memory (RAM); it may also include non-volatile memory, such as at least one disk memory, hard disk drive (HDD), or solid state drive (SSD). The memory 1103 can also be any other medium that can be used to carry or store program code in the form of instructions or data structures and can be accessed by a computer, and this application does not limit this. The processor 1102 executes the program instructions stored in the memory 1103 to implement the above functions, thereby implementing the method provided in the above embodiment.

[0396] Based on the above embodiments, the embodiments of the present application also provide a federated learning system, which includes a server node and a client node. The server node is used to implement the steps performed by the server node in the method provided in the above embodiments, and the client node is used to implement the steps performed by the client node in the method provided in the above embodiments.

[0397] Based on the above embodiments, the embodiments of the present application further provide a computer program product, which includes a computer program; when the computer program runs on a computer, the computer executes the method provided in the above embodiments.

[0398] Based on the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a computer, the computer executes the method provided in the above embodiments.

[0399] Optionally, the above-mentioned computer may include, but is not limited to, communication devices such as terminal devices and network devices.

[0400] The storage medium may be any available medium that can be accessed by a computer. By way of example and not limitation, computer-readable media may include RAM, ROM, EEPROM, CD-ROM or other optical disk storage, magnetic disk storage media or other magnetic storage devices, or any other medium that can be used to carry or store desired program code in the form of instructions or data structures and can be accessed by a computer.

[0401] Based on the above embodiments, the embodiments of the present application further provide a chip, which is used to read a computer program stored in a memory to implement the method provided in the above embodiments. Optionally, the chip may include a processor, which is coupled to the memory and is used to read the computer program stored in the memory to implement the method provided in the above embodiments. Optionally, the chip may also include components such as a memory, a communication interface, and a power supply module. The memory is used to store computer programs; the communication interface is used to receive and send data; and the power supply unit is used to power the processor.

[0402] Based on the above embodiments, the present application provides a chip system that includes a processor for supporting a computer device in implementing the functions of the federated learning system described in the above embodiments. In one possible design, the chip system also includes a memory for storing the necessary programs and data for the computer device. The chip system can be composed of a single chip or include a chip and other discrete components.

[0403] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, optical storage, etc.) that contain computer-usable program code.

[0404] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the present application. It should be understood that each flow and / or box in the flow chart and / or block diagram, as well as the combination of the flow chart and / or box in the flow chart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for implementing the functions specified in one or more flow charts and / or one or more boxes in the block diagram.

[0405] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0406] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0407] Obviously, those skilled in the art may make various changes and modifications to the present application without departing from the scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is intended to include these modifications and variations.

Claims

1. A federated learning method, characterized in that, Applied to a federated learning system including a server node and at least two client nodes, the method includes: The server node determines N client nodes for performing the federated learning task, where N is an integer greater than 1; the N client nodes are some or all of the at least two client nodes; The server node sends a first model training message to a first client node among the N client nodes; the first model training message includes initial federated learning model parameters, federated learning configuration information, and privacy protection configuration information; the initial federated learning model parameters and the federated learning configuration information are used for model training, and the privacy protection configuration information is used for privacy protection processing; The server node receives the private model parameters from the first client node; the private model parameters are obtained by the first client node based on model training and privacy protection processing.

2. The method according to claim 1, wherein The privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method used by the first client node, and the privacy protection processing factor indicates the parameter used by the first client node to perform privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information.

3. The method according to claim 2, wherein The method further includes: The server node determines the privacy protection processing factor for each of the N client nodes according to the privacy protection budget information supported by each client node among the N client nodes; the privacy protection budget information supported by the first client node among the N client nodes represents the degree of model accuracy degradation caused by the first client node performing privacy protection processing.

4. The method according to claim 2 or 3, characterized in that, The privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method; The privacy protection processing factor indicates the noise distribution parameter added to the model parameters when performing the first processing method on the local model parameters in each of the K rounds.

5. The method according to claim 2 or 3, characterized in that, The privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method; The privacy protection processing factor indicates the threshold or compression factor or Top-k value used when performing the second processing method on the local model parameters in each of the K rounds.

6. The method according to any one of claims 1-5, characterized in that, The server node determines N client nodes for performing the federated learning task, including: The server node receives a federated learning service request message, and the federated learning service request message includes privacy protection indication information, and the privacy protection indication information indicates that the client nodes performing the federated learning task need to perform privacy protection processing on the local model parameters; The server node determines N client nodes for performing the federated learning task according to the privacy protection indication information, and the N client nodes have the privacy protection processing ability, and the privacy protection processing ability is the ability to perform privacy protection processing on the local model parameters; The federated service request message further includes model accuracy degradation range information; The server node determines N client nodes to execute the federated learning task, including: The server node determines N client nodes to execute the federated learning task according to the model accuracy degradation range information. The degree of model accuracy degradation caused by the N client nodes performing privacy protection processing meets the model accuracy degradation range information required by the federated learning task.

7. The method according to claim 6, characterized in that, Determining N client nodes to execute the federated learning task, including: The server node sends a first request message to the first device; the first request message is used to obtain the configuration information of client nodes that meet the first condition. The first condition includes at least having the privacy protection processing ability and the degree of model accuracy degradation caused by performing privacy protection processing meeting the model accuracy degradation range information required by the federated learning task; the first device is used to store the configuration information of client nodes. The server node receives a first response message from the first device. The first response message includes the configuration information of N client nodes that meet the first condition. The configuration information corresponding to the N client nodes includes privacy protection processing ability information. Alternatively, the server node separately sends a first request message to at least two client nodes; The server node separately receives corresponding first response messages from N client nodes that meet the first condition. The first response message corresponding to each of the N client nodes includes the configuration information corresponding to the client node.

8. The method according to claim 7, characterized in that, The method further includes: The first request message includes one or more of the following information: Privacy protection processing type information, which indicates the privacy protection processing method required by the federated learning task; Threshold of the user privacy feature ratio, which indicates the minimum ratio of private data in the training data used by the client nodes executing the federated learning task; Threshold of the dataset size, which indicates the minimum data volume of the training data used by the client nodes executing the federated learning task; Threshold of the model training duration, which indicates the maximum model training duration per round of the client nodes executing the federated learning task; Threshold of the privacy protection processing duration, which indicates the maximum privacy protection processing duration per round of the client nodes executing the federated learning task.

9. The method according to any one of claims 1-8, characterized in that, The method further includes: The server node receives privacy model parameters from M client nodes among the N client nodes, where M is a positive integer less than or equal to N; The server node aggregates the privacy model parameters corresponding to the M client nodes to obtain model aggregation parameters. The server node subtracts the first error from the model aggregation parameter to obtain the model aggregation parameter after subtracting the first error; the first error is determined according to the weights of the M client nodes and the noise estimation values of the M client nodes, and the noise estimation value of the first client node among the M client nodes is determined by the server node according to the privacy protection budget information required by the federated learning task; The server node sends the model aggregation parameter after subtracting the first error to the M client nodes or the N client nodes respectively.

10. The method according to claim 9, wherein The server node aggregates the privacy model parameters corresponding to the M client nodes to obtain a model aggregation parameter, including: The server node determines H client nodes from the M client nodes, and the model accuracy degradation degree of any client node among the H client nodes is less than or equal to the model accuracy degradation threshold, where H is a positive integer less than or equal to M; The server node aggregates the privacy model parameters corresponding to the H client nodes according to the privacy model parameters corresponding to the H client nodes and the weights of the H client nodes.

11. The method according to claim 10, wherein The server node determines H client nodes from the M client nodes, including: The server node receives model accuracy information from the M client nodes, and the model accuracy information indicates the degree of model accuracy degradation caused by the client node performing privacy protection processing on the locally trained model parameters; The server node determines the H client nodes from the M client nodes whose model accuracy degradation degree is less than or equal to the model accuracy degradation threshold.

12. A federated learning method, characterized in that, Applied to a federated learning system including a server node and at least two client nodes, where at least two client nodes include a first client node, the method includes: The first client node receives a first model training message from the server node; wherein, the first model training message includes federated learning initial model parameters, federated learning configuration information, and privacy protection configuration information; The first client node performs model training according to the federated learning initial model parameters and the federated learning configuration information to obtain locally trained model parameters; The first client node performs privacy protection processing on the locally trained model parameters according to the privacy protection configuration information to obtain privacy model parameters; The first client node sends the privacy model parameters to the server node.

13. The method according to claim 12, wherein The privacy protection configuration information includes privacy protection processing type information and a privacy protection processing factor; wherein, the privacy protection processing type information indicates the privacy protection processing method to be used, and the privacy protection processing factor indicates the parameter used by the client node to perform privacy protection processing on the locally trained model parameters according to the privacy protection processing method indicated by the privacy protection processing type information; The first client node performs privacy protection processing on the locally trained model parameters according to the privacy protection configuration information, including: The first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor.

14. The method according to claim 13, wherein The privacy protection processing type information indicates a first processing method, and the first processing method is a noise addition method. The first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor, including: The first client node adds the noise of the current round to the local model parameters according to the noise distribution parameters of the current round indicated by the privacy protection processing factor.

15. The method according to claim 13, wherein The privacy protection processing type information indicates a second processing method, and the second processing method is a gradient compression method. The first client node performs privacy protection processing on the local model parameters according to the privacy protection processing method indicated by the privacy protection processing type information and the parameters indicated by the privacy protection processing factor, including: The first client node performs gradient compression on the local model parameters according to the threshold or compression factor or Top-k value of the current round indicated by the privacy protection processing factor.

16. The method according to any one of claims 12 - 15, characterized in that, The method further includes: The first client node sends model accuracy information to the server node, and the model accuracy information indicates the degree of model accuracy degradation caused by the first client node performing privacy protection processing on the local model parameters.

17. A federated learning device, characterized in that, The device includes: A communication unit, configured to receive and send data; A processing unit, configured to execute the method according to any one of claims 1-11, or configured to execute the method according to any one of claims 12-16.

18. A federated learning system, characterized in that, Including: A server node for executing any one of claims 1-11, and a first client node for executing any one of claims 12-16.

19. A computer-readable storage medium, characterized in that, A computer program is stored in the computer-readable storage medium. When the computer program runs on a computer, the computer is caused to execute the method according to any one of claims 1-11, or the computer is caused to execute the method according to any one of claims 12-16.

20. A chip, characterized in that, The chip is coupled to the memory and is configured to read and execute program instructions stored in the memory to implement the method according to any one of claims 1-11, or to implement the method according to any one of claims 12-16.

Citation Information

Patent Citations

  • Model training method based on federated learning

    CN111046433A

  • Federal learning model training method based on conditional privacy set intersection

    CN114386069A

  • Node selection and aggregation optimization system and method for federated learning under micro-service architecture

    CN114418109A

  • Federal learning-based model training method and device

    CN116186772A

  • Federal learning-oriented privacy protection method

    CN117294469A

Cited By

  • Federal learning method and system based on differential privacy and zero knowledge proof

    CN121150970A

  • Data processing method and equipment for protecting end-side data and medium

    CN121234402A