Multi-party large model joint training method and device based on federal learning
Through the improvement of dual central server architecture and secret sharing, the security risks and communication costs of cross-domain collaborative training in federated learning are solved, and efficient parallel training of large models is achieved.
Patent Information
- Application Number
- CN202510484026.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-08-01
AI Technical Summary
The existing federated learning technology has a high threshold for implementing the cross-domain collaboration scenario of private domain data, complex system structure, poor robustness, high communication costs, and it is difficult for a single computing node to train large models.
Using a dual central server architecture, the large model parameters are evenly divided into computing nodes, random seeds are used to generate subsecrets shared by secrets, and the final update gradient is calculated through the dual central server, reducing communication between clients and achieving parallel model training.
Reduces security risks, improves system robustness, reduces communication costs, and supports training of large models, simplifying the secret sharing process.
Smart Images

Figure CN120415786A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular to a method and device for joint training of multiple large models based on federated learning. Background Art
[0002] Since its introduction, big model technology has demonstrated outstanding performance in many fields. The performance improvement of big models often relies on larger model parameters, more training data, and stronger computing resources. With the rapid depletion of public data, the utilization of private domain data becomes imminent. However, the current training method for storing data in centralized datasets cannot meet the high privacy protection requirements of private domain data.
[0003] As an important privacy computing technology, federated learning can train large models without leaving the domain, ensuring that the data is available but invisible. It is an effective way to achieve cross-domain collaborative joint modeling and fully tap the value of private domain data.
[0004] Taking the most common Fedavg 1 The algorithm is used as an example. The entire computing framework consists of a central server Server and n computing clients Worker1, Worker2, ...Worker n The computing client is the data holder. The specific training process is as follows:
[0005] a) Initialization: The central server initializes the global model parameters and sends them to all participating clients.
[0006] b) Local training: Each client trains the model using local data and updates the local model parameters.
[0007] c) Parameter upload: The client uploads the updated model parameters to the central server.
[0008] d) Parameter aggregation: The central server calculates the average of all uploaded parameters and broadcasts this average back to all clients.
[0009] e) Iteration: Repeat the above steps until the model converges or reaches the preset number of iterations
[0010] During the entire training process, the data holder only needs to pass the model parameters obtained from each round of training to the central server, and the private domain data is always kept locally and not transmitted externally.
[0011] But in some cases, the input data can be reconstructed through model parameters or training gradients 2 Even a small probability of partial reconstruction will pose a serious threat to the security of private domain data. The open source framework Federatedscope 3A secret sharing mechanism is introduced, and the specific process is as follows:
[0012] a) Initialization: Each client initializes the model parameters locally according to the pre-defined random seeds.
[0013] b) Local training: Each client uses the local data to train the model and obtains the updated gradient of the current round of the model.
[0014] c) Secret sharing: Each client splits the gradient into n sub-secrets, where n is the total number of clients, and the sum of the n sub-secrets is equal to the gradient. The client itself retains one sub-secret and shares one sub-secret with each of the other clients. After the sharing is completed, each client holds n sub-secrets from different clients, and sums up these n sub-secrets to obtain a new encrypted gradient.
[0015] d) Parameter upload: The client uploads the encrypted gradient to the central server.
[0016] e) Parameter aggregation: The central server calculates the average value of all the uploaded encrypted gradients and broadcasts this average value back to all clients. The clients use the received gradient to update the local model.
[0017] f) Iteration: Repeat the above steps until the model converges or reaches the preset number of iterations.
[0018] The following disadvantages exist in the above technology:
[0019] 1. The implementation threshold is relatively high in the actual cross-domain collaboration scenario of private data.
[0020] Existing technical solutions require that each client participating in the collaborative calculation can communicate with each other. However, in the actual scenario, for the protection of many highly sensitive private data, the holder requires minimizing the connection with the external system as much as possible. The original data system is often only used within the internal network. Enabling connections with all other clients means significantly relaxing the access restrictions, greatly increasing the security risks of the original system, and making it very difficult to implement the solution.
[0021] 2. The system structure is relatively complex and the robustness is poor.
[0022] Existing solutions require all clients to perform gradient secret sharing. During the secret sharing process, if any client crashes or has a communication timeout, it will cause the originally transmitted sub-secrets to be unable to correctly reconstruct the new gradient, and the remaining normal clients can only be forced to perform secret sharing again.
[0023] 3. The communication cost is relatively high.
[0024] Since all clients need to communicate pairwise during the secret sharing period, the time consumption is O(n 2), where n is the number of clients. As the number of clients participating in collaborative computing increases, the communication cost increases significantly.
[0025] 4. Each client's single computing node cannot meet the requirements for large model training.
[0026] Due to the huge number of parameters in large models, the video memory of a single GPU often cannot support large model training. Moreover, to accelerate the training cycle, a large number of GPUs are generally used for simultaneous training. There is a lack of application scenarios in the existing solutions where a client consists of multiple computing nodes. Summary of the Invention
[0027] The purpose of the present invention is to provide a multi-party large model joint training method and device based on federated learning in view of the deficiencies of the existing technology.
[0028] The purpose of the present invention is achieved through the following technical solutions: A multi-party large model joint training method based on federated learning includes the following steps:
[0029] (1) Evenly divide the large model to be trained according to the parameter size among each computing node of the computing client, set s as the random seed for all nodes, and initialize the model parameters;
[0030] (2) Each computing client uses the private data it holds to train the large model. After completing one round of training, combine the gradient parameters of each computing node into the complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub-secrets of secret sharing; calculate T i = G i - R i as another sub-secret;
[0031] (3) Send the sub-secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub-secrets and to obtain the final updated gradient W = R + T and send it to each computing client;
[0032] (4) Each computing client divides W into node update gradients and sends them to the corresponding nodes to complete model update;
[0033] (5) Repeat steps (2) to (4) until the large model converges or reaches the preset number of iterations.
[0034] Further, the random seed s is randomly generated by the first central server or the second central server.
[0035] Further, combine the gradient parameters of each computing node into a complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing; calculate T i = G i - R i as another sub-secret, including:
[0036] Each computing client selects a computing node as the master node, and non-master nodes send their gradient parameters to the master node; the master node combines the gradient parameters of each computing node into a complete gradient G i , and generates a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing, and calculates T i = G i - E i as another sub-secret.
[0037] Further, send the sub-secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub-secrets and to obtain the final updated gradient W = R + T and send it to each computing client, including:
[0038] Send the T calculated by the second central server to the first central server, and the first central server calculates the final updated gradient W and sends it to each computing client, and each computing client selects a computing node as the master node to receive.
[0039] Further, the master node of each computing client divides W into node update gradients and sends them to the corresponding nodes to complete the model update.
[0040] Further, based on the parameter division in step (1), the master node of each computing client divides W into corresponding node update gradients and sends them to the corresponding nodes.
[0041] The present invention also provides a multi-party large model joint training device based on federated learning, including:
[0042] A parameter division and initialization module, configured to evenly divide the large model to be trained into each computing node of the computing client according to the parameter size, and all nodes set s as the random seed to initialize the model parameters;
[0043] A model training module is used for each computing client to train a large model using its own private data. After one round of training, the gradient parameters of each computing node are combined into a complete gradient G i ; and generate a random number matrix E i with the same shape as G i as one of the sub-secrets for secret sharing; calculate T i = G i - R i as another sub-secret;
[0044] Send the sub-secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub-secrets and to obtain the final updated gradient W = R + T and send it to each computing client;
[0045] Each computing client divides W into node update gradients and sends them to the corresponding nodes to complete the model update;
[0046] Repeat the above process until the large model converges or reaches the preset number of iterations.
[0047] The present invention also provides an electronic device, including a memory and a processor, and the memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement the above-mentioned multi-party large model joint training method based on federated learning.
[0048] The present invention also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, it implements the above-mentioned multi-party large model joint training method based on federated learning.
[0049] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, it implements the above-mentioned multi-party large model joint training method based on federated learning.
[0050] Compared with the prior art, the beneficial effects of the present invention are:
[0051] 1. Using a dual central server communication architecture, all clients only need to communicate with the central server and do not need to communicate with other clients, avoiding security risks to the original data system.
[0052] 2. The secret sharing scheme is more concise. Only two sub-secrets need to be split and sent to the central server. Even if some clients are down or communication times out during the sharing process, there is no need to re-perform secret sharing. The current round of update is completed by aggregating the gradients of the current valid clients, and the robustness is stronger.
[0053] 3. In the prior art, the communication cost for secret sharing between clients is O(n 2 ), while the communication cost of the present invention is only O(n).
[0054] 4. The present invention realizes model parallel training under the federated learning framework and can support the training of large models with a large number of scale parameters. BRIEF DESCRIPTION OF THE DRAWINGS
[0055] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0056] Figure 1 It is a schematic flow chart of a method for joint training of multi-party large models based on federated learning provided by an embodiment of the present invention;
[0057] Figure 2 It is a schematic diagram of a method for joint training of multi-party large models based on federated learning provided by Embodiment 1 of the present invention;
[0058] Figure 3 It is a schematic diagram of a module of a device for joint training of multi-party large models based on federated learning provided by an embodiment of the present invention;
[0059] Figure 4 It is a schematic diagram of the structure of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0060] The present invention will be described in detail below with reference to the drawings. Without conflict, the features in the following embodiments and implementation manners can be combined with each other.
[0061] A method for joint training of multi-party large models based on federated learning according to the present invention is shown in Figure 1 and includes the following steps:
[0062] (1) Evenly divide the large model to be trained into each computing node of the computing client according to the parameter size. Set s as the random seed for all nodes and initialize the model parameters;
[0063] (2) Each computing client uses the private data it holds to train the large model. After completing one round of training, combine the gradient parameters of each computing node into a complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing; calculate Ti = G i - R i as another sub - secret;
[0064] (3) Send the sub - secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub - secrets and Get the final updated gradient W = R + T and send it to each computing client;
[0065] (4) Each computing client divides W into node - updated gradients and sends them to the corresponding nodes to complete the model update;
[0066] (5) Repeat steps (2) to (4) until the large - model converges or reaches the preset number of iterations.
[0067] In one embodiment, the random seed s is randomly generated by the first central server or the second central server.
[0068] In one embodiment, combine the gradient parameters of each computing node into a complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub - secrets for secret sharing; calculate T i = G i - R i as another sub - secret, including:
[0069] Each computing client selects a computing node as the master node, and the non - master nodes send their gradient parameters to the master node; the master node combines the gradient parameters of each computing node into a complete gradient G i , and generates a random number matrix R i with the same shape as G i as one of the sub - secrets for secret sharing, and calculate T i = G i - R i as another sub - secret.
[0070] In one embodiment, send the sub - secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub - secrets and Get the final updated gradient W = R + T and send it to each computing client, including:
[0071] Send the T calculated by the second central server to the first central server, and the first central server calculates the final updated gradient W and sends it to each computing client. Each computing client selects a computing node as the master node to receive it.
[0072] In one embodiment, the master node of each computing client divides W into node update gradients and sends them to the corresponding nodes to complete model update.
[0073] In one embodiment, based on the parameter division in step (1), the master node of each computing client divides W into corresponding node update gradients and sends them to the corresponding nodes.
[0074] Embodiment 1:
[0075] A multi-party large model joint training method based on federated learning, including two central servers Server1, Server2 and n computing clients Worker1, Worker2,... Worker n , and the number of computing nodes owned by each computing client is m1, m2... m n , and each computing client selects a computing node as the master node; see Figure 2 , and this training method includes:
[0076] (1) Server1 generates a random integer s as a seed, and sends s and the model type model_type jointly agreed by everyone to the master nodes of all computing clients;
[0077] (2) The master nodes of all computing clients formulate a model allocation plan according to the model type model_type. The allocation principle is to evenly divide the large model into each computing node according to the parameter size. The allocation plan of Worker i is m i sets of model parameter name collections The master node distributes the random seed s, the model type model_type and the model parameter name collection to the corresponding nodes. All nodes set s as the random seed, and then initialize the model parameters according to the model type model_type and the model parameter name collection to obtain the model parameters
[0078] (3) All computing clients use the private domain data they hold to train the large model. After completing one round of training, the m i computing nodes of Worker i obtain partial gradients of the model update The non-master nodes send their gradient parameters to the master node, and the master node will sort them according to the parameter names and combine them into a complete gradient G i;
[0079] (4) Worker i 's master node generates a random number matrix R i with the same shape as G i As one of the sub - secrets of secret sharing, calculate T i = G i - R i As another sub - secret, then send R i to Server1 and send T i to Server2;
[0080] (5) Server2 calculates the sum of the sub - secrets Send T to Server1; Server1 calculates the sum of the sub - secrets Get the final updated gradient W = R + T, and Server1 sends W to the master nodes of all computing clients;
[0081] (6) Worker i 's master node, after receiving the updated gradient W, divides the updated gradient W into node - updated gradients according to the set of model parameter names in step (2) and sends them to the corresponding nodes. After the nodes receive the updated gradients, they complete the model update model i_j = model i_j - w i_j ; j = 1, 2…, m i ;
[0082] (7) Repeat steps (3) to (6) until the large - model converges or reaches the preset number of iterations.
[0083] The present invention also provides a multi - party large - model joint training device based on federated learning. See Figure 3 , including:
[0084] A parameter division and initialization module, which is used to evenly divide the large - model to be trained into each computing node of the computing clients according to the parameter size, and all nodes set s as the random seed to initialize the model parameters;
[0085] A model training module, which is used for each computing client to train the large - model using its own private domain data. After completing one round of training, combine the gradient parameters of each computing node into a complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub - secrets of secret sharing; calculate T i = G i - R iAs another sub-secret;
[0086] Send the sub-secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub-secrets and obtain the final updated gradient W = R + T and send it to each computing client;
[0087] Each computing client divides W into node update gradients and sends them to the corresponding nodes to complete the model update;
[0088] Repeat the above process until the large model converges or reaches the preset number of iterations.
[0089] It should be noted that the device embodiments shown in this embodiment match the content of the above method embodiments. You can refer to the content of the above method embodiments and will not be elaborated here.
[0090] Figure 4 This is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Please refer to Figure 4 , the electronic device provided in this embodiment includes: a memory and a processor. Among them, the memory is used to store information including program instructions, and the processor is used to control the execution of the program instructions. When the program instructions are loaded and executed by the processor, a method for jointly training a multi-party large model based on federated learning of the present invention is implemented. It should be noted that in addition to Figure 4 the memory and processor shown, according to its actual functions, the electronic device may also include other hardware, which will not be elaborated here.
[0091] The present invention also provides a computer-readable storage medium, on which a computer program is stored. When the program is executed by a processor, the above-mentioned method for jointly training a multi-party large model based on federated learning is implemented.
[0092] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the above-mentioned method for jointly training a multi-party large model based on federated learning is implemented.
[0093] The computer-readable storage medium may be an internal storage unit of any data processing-capable device described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be any data processing-capable device, such as a plug-in hard disk, a Smart Media Card (SMC), an SD card, a Flash Card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any data processing-capable device and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any data processing-capable device, and may also be used to temporarily store data that has been output or is to be output.
[0094] Those skilled in the art should understand that the embodiments of the present application may be provided as a method, a system, or a computer program product. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0095] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the specified function in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks
[0096] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the specified function in Figure 1 one or more of the processes Figure 1 or multiple processes and / or blocks
[0097] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are performed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for realizing the functions specified in one process or a plurality of processes and / or blocks Figure 1 one process or a plurality of processes and / or blocks Figure 1 or steps for realizing the functions specified in one block or a plurality of blocks.
[0098] The above embodiments are only used to illustrate the design concept and features of the present invention, and the purpose is to enable those skilled in the art to understand the content of the present invention and implement it accordingly. The protection scope of the present invention is not limited to the above embodiments. Therefore, all equivalent changes or modifications made according to the principles and design concepts disclosed in the present invention are within the protection scope of the present invention.
Claims
1. A method for jointly training multi-party large models based on federated learning, characterized in that, Including the following steps: (1) Evenly divide the large model to be trained according to the parameter size among each computing node of the computing client, set s as the random seed for all nodes, and initialize the model parameters; (2) Each computing client trains a large model using the private data it holds. After completing one round of training, the gradient parameters of each computing node are combined into a complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing; calculate T i = G i - R i as another sub-secret; (3) Send the sub-secret R i and T i to the first central server and the second central server respectively, calculate the sum of the sub-secrets and obtain the final updated gradient W = R + T and send it to each computing client; (4) Each computing client divides W into node update gradients and sends them to the corresponding nodes to complete model update; (5) Repeat steps (2) to (4) until the large model converges or reaches the preset number of iterations.
2. The multi-party large model joint training method based on federated learning according to claim 1, characterized in that The random seed s is randomly generated by the first central server or the second central server.
3. A method for jointly training multi-party large models based on federated learning according to claim 1, characterized in that, Combine the gradient parameters of each computing node into the complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing; calculate T i = G i - R i as another sub-secret, including: Each computing client selects a computing node as the master node, and the non-master nodes send their gradient parameters to the master node; the master node combines the gradient parameters of each computing node into the complete gradient G i , and generates a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing, and calculates T i = G i - R i as another sub-secret.
4. A method for jointly training multi-party large models based on federated learning according to claim 1, characterized in that, Send the sub - secrets R i and T i to the first central server and the second central server respectively, calculate the sum of the sub - secrets and obtain the final updated gradient W = R + T and send it to each computing client, including: Send the T calculated by the second central server to the first central server, and the first central server calculates the final update gradient W and sends it to each computing client, and each computing client selects a computing node as the master node to receive.
5. A method for jointly training multi-party large models based on federated learning according to claim 4, characterized in that, The master node of each computing client divides W into node update gradients and sends them to the corresponding nodes to complete model update.
6. A method for jointly training multi-party large models based on federated learning according to claim 5, characterized in that, Based on the parameter division in step (1), the master node of each computing client divides W into corresponding node update gradients and sends them to the corresponding nodes.
7. A multi-party large model joint training device based on federated learning, characterized in that, Including: A parameter division and initialization module, which is used to evenly divide the large model to be trained according to the parameter size among each computing node of the computing client, set s as the random seed for all nodes, and initialize the model parameters; The model training module is used for each computing client to train the large model using its own private data. After one round of training, the gradient parameters of each computing node are combined into a complete gradient G i ; and generate a random number matrix R i with the same shape as G i as one of the sub-secrets for secret sharing; calculate T i = G i - R i as another sub-secret; Send the sub-secret R i and T i to the first central server and the second central server respectively, and calculate the sum of the sub-secrets and Obtain the final updated gradient W = R + T and send it to each computing client; Each computing client divides W into node update gradients and sends them to the corresponding nodes to complete model update; Repeat the above process until the large model converges or reaches the preset number of iterations.
8. An electronic device, comprising a memory and a processor, characterized in that, The memory is coupled to the processor; wherein, the memory is used to store program data, and the processor is used to execute the program data to implement a method for joint training of a multi-party large model based on federated learning according to any one of claims 1-6 above.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements a method for joint training of a multi-party large model based on federated learning according to any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements a method for joint training of a multi-party large model based on federated learning according to any one of claims 1-6.