Federal learning-based model training method, device and equipment

By identifying the target hidden layer and performing parameter aggregation and tuning in federated learning, the problem of inaccurate model training in the financial field is solved, achieving more efficient model training and reducing costs.

CN120822579AActive Publication Date: 2025-10-21ZHONGJINKE INFORMATION TECH CO LTD +1
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510678452.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-10-21
Estimated Expiration
2045-05-26

AI Technical Summary

Technical Problem

Existing federated learning methods have the problem of inaccurate model training in the financial field.

Method used

By determining the target hidden layer between the master and slave computing nodes, and aggregating and optimizing the parameters of these hidden layers, more accurate model parameters can be obtained, avoiding the computational power consumption and bandwidth occupation caused by full parameter updates.

Benefits of technology

This improves the accuracy of model training and reduces communication costs and computational costs for each participant in the federated learning process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120822579A_ABST
    Figure CN120822579A_ABST
Patent Text Reader

Abstract

The invention discloses a federated learning-based model training method, device and equipment, which are applied to a main computing node, and comprise the following steps: determining a plurality of first target hidden layers based on an initial first large language model obtained by training, and obtaining initial first sub-parameters corresponding to the first target hidden layers; receiving an initial second sub-parameter which is sent by each slave computing node and corresponds to each second target hidden layer in the initial second large language model; performing parameter aggregation at least based on each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregation parameter; and sending the initial aggregation parameter to each slave computing node to enable each slave computing node to carry out next round of model tuning based on the initial aggregation parameter, receiving a current second sub-parameter sent by each slave computing node, stopping model tuning until a predetermined tuning condition is met, taking the current aggregation parameter as a target aggregation parameter, and sending the target aggregation parameter to each slave computing node. And obtaining the target large language model. According to the method and the device, the final tuning result can be more accurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the financial field and the field of federated learning technology, and in particular to a model training method, apparatus and equipment based on federated learning. Background Art

[0002] Federated Learning is a distributed learning technology that combines traditional cryptography and machine learning. It aims to establish a federated learning model based on distributed data sets for multiple data providers who are unwilling or unable to expose their plaintext data.

[0003] However, in the financial field, existing federated learning methods suffer from inaccurate model training. Summary of the Invention

[0004] In view of this, the present invention provides a model training method, device and equipment based on federated learning, the main purpose of which is to solve the problem of inaccurate model training in the current federated learning method.

[0005] To solve the above problems, this application provides a model training method based on federated learning, which is applied to the main computing node, including: Determining a plurality of first target hidden layers based on the initial first language model obtained through training, and obtaining initial first sub-parameters corresponding to each first target hidden layer; receiving initial second sub-parameters corresponding to respective second target hidden layers in the initial second largest language model sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing node; Performing parameter aggregation based on at least each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregated parameter; The initial aggregation parameters are sent to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, the current aggregation parameters are used as the target aggregation parameters, and the target large language model is obtained.

[0006] To solve the above problems, this application provides a model training method based on federated learning, which is applied to slave computing nodes, including: Determining a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtaining initial second sub-parameters corresponding to each second target hidden layer; Sending each initial second sub-parameter to the main computing node, so that the main computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain an initial aggregated parameter; Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the main computing node until the predetermined tuning conditions are met. Then stop model tuning, use the received current aggregation parameters as the target aggregation parameters, and obtain the target large language model.

[0007] To solve the above problems, this application provides a model training device based on federated learning, including: a first determining module, configured to determine a plurality of first target hidden layers based on the initial first large language model obtained through training, and obtain initial first sub-parameters corresponding to the first target hidden layers; a first receiving module, configured to receive initial second sub-parameters corresponding to respective second target hidden layers in the initial second largest language model, sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing nodes; an aggregation module, configured to perform parameter aggregation based at least on each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregated parameter; The first sending module is used to send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, the current aggregation parameters are used as the target aggregation parameters, and the target large language model is obtained.

[0008] To solve the above problems, this application provides a model training device based on federated learning, including: a second determining module, configured to determine a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtain initial second sub-parameters corresponding to the respective second target hidden layers; A second sending module is configured to send each initial second sub-parameter to a master computing node, so that the master computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain an initial aggregated parameter; The second receiving module is used to receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the main computing node until the predetermined tuning conditions are met. The model tuning is stopped, and the received current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.

[0009] To solve the above problems, the present application provides an electronic device, which includes at least a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the computer program on the memory, it implements the steps of the model training method based on federated learning as described in any of the above claims.

[0010] The present application discloses a model training method, apparatus and device based on federated learning. In each round of training, the main computing node determines the first target hidden layer from the first language model obtained through training, and uses the parameters of each first target hidden layer as the parameters to be tuned. At the same time, the main computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the main node can aggregate the parameters of each first target hidden layer with the parameters of each second target hidden layer to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate next round of model training based on the aggregated parameters. In the present application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupancy of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0011] The above description is only an overview of the technical solution of the present invention. In order to more clearly understand the technical means of the present invention, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are specifically listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Various other advantages and benefits will become apparent to those skilled in the art upon reading the detailed description of the preferred embodiment below. The accompanying drawings are for illustration purposes only and are not to be considered as limiting the present invention. The same reference symbols are used throughout the drawings to represent the same components. In the drawings: Figure 1 This is a flowchart of a model training method based on federated learning according to an embodiment of the present application; Figure 2 This is a schematic diagram of a three-layer topology structure in another embodiment of the present application; Figure 3 This is a schematic diagram of a distributed machine learning process in another embodiment of the present application; Figure 4 This is a flowchart of a model training method based on federated learning according to another embodiment of the present application; Figure 5 This is a structural block diagram of a model training device based on federated learning in another embodiment of the present application; Figure 6 This is a structural block diagram of a model training device based on federated learning in another embodiment of the present application; Figure 7 This is a structural block diagram of an electronic device according to another embodiment of the present application. DETAILED DESCRIPTION

[0013] Various aspects and features of the present application are described herein with reference to the accompanying drawings.

[0014] It should be understood that various modifications may be made to the embodiments of the present application. Therefore, the above description should not be considered as limiting, but merely as an example of an embodiment. Other modifications within the scope and spirit of the present application will occur to those skilled in the art.

[0015] The accompanying drawings, which are incorporated in and constitute a part of the specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.

[0016] These and other characteristics of the present application will become apparent from the following description of a preferred form of embodiment given as a non-limiting example with reference to the accompanying drawings.

[0017] It should also be understood that although the present application has been described with reference to certain specific examples, those skilled in the art will readily be able to implement many other equivalent forms of the present application.

[0018] The above and other aspects, features and advantages of the present application will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings.

[0019] Specific embodiments of the present application will be described hereinafter with reference to the accompanying drawings; however, it should be understood that the embodiments described are merely examples of the present application and may be implemented in a variety of ways. Familiar and / or repetitive functions and structures are not described in detail to avoid obscuring the present application with unnecessary or redundant details. Therefore, the specific structural and functional details described herein are not intended to be limiting, but rather serve merely as a basis and representative basis for the claims to teach those skilled in the art to variously utilize the present application with substantially any suitable detailed structure.

[0020] This specification may use the phrases "in one embodiment," "in another embodiment," "in yet another embodiment," or "in other embodiments," which may all refer to one or more of the same or different embodiments according to the present application.

[0021] The present application embodiment provides a model training method based on federated learning, which can be applied to the leading party / main computing node of federated learning, such as Figure 1As shown, the method in this embodiment includes the following steps: Step S101, determining a plurality of first target hidden layers based on the initial first large language model obtained through training, and obtaining initial first sub-parameters corresponding to the respective first target hidden layers; In this step, the pre-trained original large language model can be fine-tuned based on the first conversation dataset to obtain an initial first large language model. The first conversation dataset is high-quality conversation data related to bond business locally stored on the primary computing node. Specifically, the first conversation data includes several question-and-answer data.

[0022] In this step, after the main computing node obtains the initial first large language model through training, it can determine the global score of each hidden layer in the initial large language model, and then determine several first target hidden layers to be adjusted based on the global score of each hidden layer, so as to facilitate the subsequent parameter aggregation processing based on the initial first sub-parameters of each first target hidden layer. In other words, if the initial large language model contains, for example, 10 hidden layers, the method in this application can be used to determine 3 hidden layers from these 10 hidden layers as the first target hidden layers, and then the parameters corresponding to each first target hidden layer in the initial first training parameters obtained through training are used as the initial first sub-parameters.

[0023] Step S102: receiving initial second sub-parameters corresponding to respective second target hidden layers in the initial second largest language model sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing nodes; In this step, after the slave computing node obtains the initial second-largest language model through training, it also determines the global score of each hidden layer in the initial second-largest language model, and then determines several second target hidden layers to be adjusted based on the global score of each hidden layer, so as to facilitate the subsequent sending of the initial second sub-parameters corresponding to each second target hidden layer to the master computing node. Thus, the master computing node can receive the initial second sub-parameters corresponding to each second target hidden layer sent by each slave computing node.

[0024] Step S103, performing parameter aggregation based on at least each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregated parameter; In this step, when performing parameter aggregation, the initial parameters of each hidden layer can be combined for parameter aggregation. That is, based on the initial parameters of each hidden layer, each initial first sub-parameter, and each initial second sub-parameter, parameter aggregation is performed to obtain initial aggregated parameters. In this step, after the number of training rounds exceeds one, the tuned parameters of each hidden layer obtained from the previous round of model tuning can also be used as initial parameters to perform parameter aggregation in combination with the initial parameters. That is, based on the tuned parameters of each hidden layer obtained from the previous round of model tuning, each current first sub-parameter, and each current second sub-parameter, parameter aggregation is performed to obtain current aggregated parameters.

[0025] For example, a large language model has eight layers, labeled 1 to 8, with initial parameters H1, H2, H3, ..., H8 corresponding to each layer. Each computation node (master and slave) selects several target hidden layers with the highest scores, labeled a, b, ...

[0026] For example, the main computing node a determines three first target hidden layers: a1, a2, and a3. The first target hidden layers a1, a2, and a3 correspond to the initial first sub-parameters A1, A2, and A3, respectively.

[0027] Three second target hidden layers are determined from the computing node b: b2, b3 and b4. The second target hidden layers b2, b3 and b4 correspond to the initial second sub-parameters B2, B3 and B4 respectively.

[0028] Three second target hidden layers are determined from the computing node c: c1, c3 and c5; the second target hidden layers c1, c3 and c5 correspond to the initial second sub-parameters C1, C3 and C5 respectively.

[0029] After collecting the sub-parameters of all target hidden layers, the master computation node can then perform the arithmetic averaging of the layers that need updating to obtain the optimized model parameters, labeled M1 through M8. Thus, M1 = (H1 + A1 + C1) ÷ 3. M2 = (H2 + A2 + B2) ÷ 3. M3 = (H3 + A3 + B3 + C3) ÷ ​​4. M4 = (H4 + B4) ÷ 2. M5 = (H5 + C5) ÷ 2. M6 = H6. M7 = H7. M8 = H8.

[0030] In step S104, the initial aggregation parameters are sent to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.

[0031] In this step, after completing parameter aggregation, the master computing node can send the aggregated parameters to each slave computing node, so that the slave computing node can perform the next round of model tuning based on the aggregated parameters until the predetermined model conditions are met. In this embodiment, the predetermined model tuning conditions can be that the number of tuning rounds is greater than the predetermined number of rounds or the error corresponding to the current aggregated parameters is less than a predetermined error threshold.

[0032] In this embodiment, a model training method based on federated learning is provided. In each round of training, the main computing node determines the first target hidden layer from the first language model obtained through training, and uses the parameters of each first target hidden layer as the parameters to be tuned. At the same time, the main computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the main node can aggregate the parameters of each first target hidden layer with the parameters of each second target hidden layer to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate next round of model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupancy of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0033] Based on the above embodiment, another embodiment of the present application provides a model training method based on federated learning, which can be specifically applied to a master computing node. The method in this embodiment includes the following steps: Step S201: Perform model training on a predetermined base model based on a target vocabulary to obtain initial first intermediate parameters; wherein the target vocabulary is obtained by merging a second original vocabulary of a slave computing node and a first original vocabulary locally on the master computing node; In this step, when the base model is trained, the following steps are specifically included: Step S201 - 1 , installation and deployment of a multimodal data conversion toolset.

[0034] In this step, the leading party / main computing node of the joint modeling can pre-install and deploy the multimodal data conversion tool set, and send the multimodal data conversion tool to each slave computing node, so that each slave computing node can process data of various modalities based on the multimodal data conversion tool set, thereby constructing a second original vocabulary.

[0035] Step S201-2: Installation and initialization of the federated learning collaborative network.

[0036] In this step, the master computing node and each slave computing node can also pre-deploy federated learning components. In this embodiment, the framework of the federated learning collaboration network can be implemented using other software products with the same functions such as FATE, SecretFlow, PaddleFL, etc.

[0037] After completing the deployment of the federated learning components, the master computing node will send the initial parameters of the base large language model / base model to each slave computing node through GRPC. In this embodiment, other large language models with the same functions such as Chinese-LLaMA and OpenChineseLLaMA can be used as base models, and the vocabulary expansion uses the SentencePiece tool. In the data preprocessing step, each computing node processes the local multimodal bond data and the deployed multimodal data conversion tool set into Chinese text data / Chinese corpus as the original data for training the federated multimodal large language model. The original data is not sent externally during the processing process to protect the privacy and security of the business data of each institution.

[0038] Step S201-3: construct a target vocabulary.

[0039] In this step, the master computing node can process the local multimodal bond data into Chinese text data / Chinese language training through the deployed multimodal data conversion tool set, and then train the first original vocabulary based on the Chinese text data / Chinese corpus. Similarly, after obtaining the local Chinese text data / Chinese corpus, the slave computing node can train the second original vocabulary based on the local Chinese text data / Chinese corpus. That is, the leading party / master computing node of the joint modeling trains the Chinese vocabulary / first original vocabulary chn_center.model on the bond data within the general Chinese corpus, and all participants / slave computing nodes of the joint modeling each train the Chinese vocabulary / second original vocabulary chn_nodei.model in the Chinese corpus composed of internal business data, and send it to the master computing node. The master computing node merges the above-mentioned first original vocabulary and the second original vocabulary of each slave computing node, and uses them as the target vocabulary for federal training of the base model after excluding duplicate tokens.

[0040] Step S201 - 4 : constructing a first target topology structure.

[0041] In this step, the main computing node may construct a first target topology structure corresponding to the main computing node based on the first server corresponding to the main computing node and the graphics processors deployed on each of the first servers.

[0042] Step S201-5: Perform model training on the predetermined base model.

[0043] In this step, model training can be performed based on the first target topology using the target vocabulary. That is, the main computing node uses a ring-type distributed learning method and the target vocabulary to train the predetermined base model for the first target topology to obtain the initial first intermediate parameters.

[0044] Build a ring-type distributed machine learning structure within each computing node. By default, the 3D ring structure of the 3D-Torus algorithm is used to form a three-layer topology structure in the federated learning node (inter-node > intra-node machine -> single machine with multiple cards), forming a three-dimensional decomposition, such as Figure 2 As shown. S1.X represents the slave computing node, and the bandwidth between it and the top-level master computing node S2.0 is w2 (usually forwarded through Nginx, etc., using the bandwidth of the financial dedicated line interconnection); S0.X is a single server, and the bandwidth between it and the gateway of the slave computing node is w1 (usually an inter-machine switch); A0, B0, etc. are single GPUs, and the bandwidth between them is w0 (such as a single machine with multiple cards). Distributed machine learning process is as follows Figure 3 As shown, in the first phase, four scatter-reduce operations (A0-A1-A2, B0-B1-B2, C0-C1-C2, and D0-D1-D2) are executed simultaneously within a node (using the same switch S0.X, with bandwidth w0). In the second phase, six scatter-reduce operations (A0-B0, A1-B1, A2-B2, C0-D0, C1-D1, and C2-D2) are executed simultaneously between nodes (using the same switch S1.X, with bandwidth w1). These operations occur simultaneously, sharing the S1.x bandwidth. In the third phase, six scatter-reduce operations (A0-C0, A1-C1, A2-C2, B0-D0, B1-D1, and B2-D2) are executed simultaneously between nodes, sharing the S1.x bandwidth on S2.0. After Step 3, the data on each GPU, S(ize) / 12, represents the complete data after all Reduce-Scatter reductions have been completed. At this point, all Scatter-Reduce operations are complete. The fourth all-gather phase is executed in the same order, but in reverse, ultimately completing the all-reduce phase. When applied to large language models with fewer parameters, or when computing nodes do not interact with multiple machines or GPUs, the 2D-Torus algorithm's 2D ring structure can be used to reduce one layer of data segmentation and interaction. This simplifies the training process to: intra-group scatter-reduce -> inter-group all-reduce -> intra-group all-gather. This will not be further described.

[0045] During the specific implementation of this embodiment, the LoRA method can be used to apply the above-mentioned 3D-Torus structure to pre-train a federated multimodal large language model. The master computing node generates a dimensionality reduction matrix A and a dimensionality increase matrix B as training objects, where matrix A is initialized to a random Gaussian distribution and matrix B is initialized to a zero matrix. The matrix parameters of the pre-trained model remain frozen. The input and output dimensions of the model remain unchanged during training, and the BA and pre-trained model parameters are superimposed on the output. The master computing node sends the A and B matrices to all slave computing nodes. All computing nodes involved in this task use local training data sets and perform local training on a batch (epoch) of data using a distributed machine learning structure of a 3D ring.

[0046] Step S202: Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, parameter aggregation is performed to obtain initial aggregated intermediate parameters, and the initial aggregated intermediate parameters are sent to the slave computing nodes for the slave computing nodes to perform the next round of model training based on the initial aggregated intermediate parameters, and the current second intermediate parameters sent by each slave computing node are received until the predetermined model training stop condition is met, and the current aggregated intermediate parameters are used as the target aggregated intermediate parameters to obtain the original large language model.

[0047] During the implementation of this step, after all slave compute nodes complete a batch of training locally, they send the trained intermediate parameters (i.e., matrices A and B) to the master compute node via GPRC. The master compute node reads the intermediate parameters from all slave compute nodes, aggregates them using the FedAvg algorithm, and then sends the aggregated parameters back to all slave compute nodes to begin training a new batch (next round). This process repeats until the model training stop condition is met, resulting in the original large language model M. The stopping condition can be the completion of a predetermined batch or number of rounds of training, or the attainment of a satisfactory loss or the occurrence of certain errors. The predetermined batch / number of rounds and loss conditions are preset hyperparameters. Errors that cause training to stop may include network communication issues, timeouts between federated learning nodes, abnormal node status, and other issues.

[0048] Step S203 , performing model tuning processing on the original large language model obtained through pre-training based on the first dialogue dataset to obtain an initial first large language model.

[0049] During the specific implementation of this step, after obtaining the original large language model M, it can be fine-tuned. This fine-tuning process is based on the high-quality conversation datasets local to each computing node (i.e., the first conversation dataset local to the master computing node and the second conversation dataset local to each slave computing node). Specifically, after completing federated training of the original large language model M, each node prepares its own high-quality conversation dataset related to bond business. The master computing node prepares the first conversation dataset, and each slave computing node prepares the second conversation dataset. These high-quality conversation datasets contain questions and answers likely to be asked in various bond business contexts. In this embodiment, the data in the high-quality conversation datasets is multimodal. For example, input questions may include scanned copies of business-related documents. This multimodal input data is converted to Chinese text using the aforementioned multimodal data conversion toolset, which will not be further detailed here.

[0050] Step S204: determining a plurality of first target hidden layers based on the initial first large language model obtained through training, and obtaining initial first sub-parameters corresponding to the respective first target hidden layers; During implementation, this step can determine the global score of each hidden layer in the initial first-largest language model, and determine a number of first target hidden layers to be adjusted based on the global score of each hidden layer. Specifically, for the initial first-largest language model, the correlation matrix between any two hidden layer parameters can be determined; the global score of each hidden layer can be determined based on the eigenvalues ​​of each correlation matrix corresponding to the same hidden layer; and the number of hidden layers with global scores greater than a predetermined score threshold can be determined as the first target hidden layers. Alternatively, a predetermined number of hidden layers with the highest global scores can be selected as the first target hidden layers.

[0051] That is, during the prompt tuning process of each batch of data, the score of each layer of the large language model is dynamically calculated through the correlation matrix of the hidden state, and the appropriate layer is selected for prompt tuning, avoiding the huge computing power consumption and bandwidth occupation of the traditional method of full parameter update, reducing the communication cost in the federated learning process and the local computing cost of each participant, while improving the accuracy of model training. Specifically, during each prompt tuning process, for any two hidden layers in the large model, the matrix K composed of the cosine similarity between the parameters of the two hidden layers can be used to represent their correlation. The eigenvalue λ of the matrix K is Kijrepresents the difference between two hidden layers i and j in this data set. Therefore, the sum of the eigenvalues ​​of the correlation matrix K between hidden layer i and all other layers represents the global score of hidden layer i. A higher score indicates greater similarity between that layer and the other layers. The global scores of all hidden layers are calculated, and a predetermined number of hidden layers with the highest scores are selected as the layers to be tuned. The predetermined number is a hyperparameter and can be predetermined based on the square root of the total number of hidden layers.

[0052] Step S205: receiving initial second sub-parameters corresponding to the second target hidden layers in the initial second largest language model sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing nodes; In this step, each slave computing node also fine-tunes its local original large language model based on the second conversation dataset local to the slave computing node, obtaining an initial second-largest language model. The node then determines the global score for each hidden layer in the initial second-largest language model, and based on the global score for each hidden layer, determines a number of second target hidden layers to be adjusted. The specific process for determining the second target hidden layer is similar to the process for determining the first target hidden layer in the above embodiment and will not be repeated here. After determining the second target hidden layer, each slave computing node can send the initial second sub-parameters corresponding to the second target hidden layer to the master computing node.

[0053] Step S206, performing parameter aggregation based on at least the initial first sub-parameters and the initial second sub-parameters to obtain initial aggregated parameters; When performing parameter aggregation, the initial parameters of each hidden layer may be combined for parameter aggregation, that is, based on the initial parameters of each hidden layer, each initial first sub-parameter and each initial second sub-parameter, parameter aggregation is performed to obtain initial aggregation parameters.

[0054] In step S207, the initial aggregation parameters are sent to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.

[0055] During the specific implementation of this step, after obtaining the original large language model M, it can be fine-tuned. The specific fine-tuning process is based on high-quality conversation datasets local to each computing node (i.e., a first conversation dataset local to the master computing node and a second conversation dataset local to each slave computing node). Specifically, after completing federated training of the original large language model M, each node prepares its own high-quality conversation data related to bond business (the master computing node prepares the first conversation dataset, and each slave computing node prepares the second conversation dataset). This data contains questions and answers likely to be asked in various bond business scenarios. Using the same distributed training process and intermediate parameter interaction process based on the federated learning framework as steps S201-S202, the model is fine-tuned, ultimately obtaining the target large language model M_SFT. In this embodiment, the input data is all multimodal data. For example, the input questions may include scanned copies of business-related documents. This multimodal input data is converted into Chinese text using the aforementioned multimodal data conversion toolkit, which will not be further described here.

[0056] In this step, after completing parameter aggregation, the master computing node can send the aggregated parameters to each slave computing node, so that the slave computing node can perform the next round of model tuning based on the aggregated parameters until the predetermined model conditions are met. In this embodiment, the predetermined model tuning conditions can be that the number of tuning rounds is greater than the predetermined number of rounds or the error corresponding to the current aggregated parameters is less than a predetermined error threshold.

[0057] Step S208: Based on the constructed first question and answer set, the initial reward model is trained as a federated model to obtain a target reward model; In this step, after obtaining the target large language model, in order to make the final target large language model more accurate and reliable, the target reward model can be further trained to facilitate subsequent federated reinforcement learning of the target large language model based on the target reward model.

[0058] In this step, the first question-answer set consists of several partial-order dialogue pairs. Each partial-order dialogue pair contains a sample question, several sample answers corresponding to the sample question, and the answer ranking of each sample answer. Each computing node (master and slave) prepares partial-order dialogue pairs within its business domain. These are manually annotated question-answer sample ranking sequences, including the ranking of prompt words and corresponding answers.

[0059] Specifically, the master computing node prepares a security dataset, and each slave computing node prepares an availability dataset. The partially ordered pairs of questions in the dataset are designed by each node based on its own business perspective. The answers are generated by multiple calls to the federated large language model, which has been optimized in step S207. The generated answers are manually ranked. The partially ordered pairs in the availability dataset are primarily used to improve the accuracy of the large model's output. For example: [Question] What materials are required to open an interbank bond market settlement account? [Best Answer] 1. Central Clearing Company Account Business Application Form 2. Central Clearing Company Market Participant Basic Information Form 3. Copy of Business License 4. Financial Business License (if applicable) 5. Signature Page of Bond Issuance, Registration, and Agency Redemption Service Agreement 6... [Inferior Answer] When preparing to open an interbank bond market settlement account, a series of relevant materials and documents are typically required to ensure smooth account opening and subsequent trading activities. These required materials may vary depending on the requirements of different banks and regulatory agencies, and typically include legal person identification, legal representative identification, financial reports, credit rating reports, trader information, account management agreements, other relevant documents, and other materials required by the bank.

[0060] During the specific implementation of this step, the master computing node can perform model training on the initial reward model based on the first question and answer set to obtain the initial first reward parameters; then, the master computing node can receive the initial second reward parameters sent from each computing node; based on the initial first intermediate parameters and each initial second reward parameter, perform parameter aggregation to obtain the initial aggregated reward parameters, and send the initial aggregated reward parameters to each slave computing node, so that the slave computing node can perform the next round of model training based on the initial aggregated reward parameters, and receive the current second reward parameters sent by each slave computing node, until the predetermined model reward stop condition is met, and use the current aggregated reward parameters as the target aggregated reward parameters to obtain the target reward model M_RM.

[0061] In this step, the loss function of the reward model is:

[0062] Where x is the input prompt (i.e. question), y is the answer, and y c Represents the sort ratio y r The sentence at the front is the one with better answer effect. m(r) represents the degree of difference between the two sentences. There are four selectable values: 3, 2, 1, and 0. For example, for two sentences answering y I and y J ,y I than y J When it is much better, the reward model gives the two answers a score that satisfies L(y I )>L(y J), and at this time m(r)=3.

[0063] Step S209: performing federated reinforcement learning on the target large language model based on the target reward model to obtain a reinforced target large language model; In this step, after training to obtain the target reward model M_RM, the master computing node can score the answers output by the target large language model based on the target reward model, thereby further strengthening the target large language model based on the answer scores. Specifically, the master computing node uses the local training dataset to generate results / answers / responses based on the target large language model M_SFT. It then uses the target reward model M_RM to calculate the reward (i.e., score) for the answers / responses. Based on the reward, it adjusts the model parameters to obtain initial first reinforcement parameters. Simultaneously, the master computing node receives initial second reinforcement parameters from each slave computing node, aggregates the parameters based on the initial first reinforcement parameters and each initial second reinforcement parameter, obtains initial aggregate reinforcement parameters, and sends these initial aggregate reinforcement parameters to each slave computing node for the next round of model reinforcement based on the initial aggregate reinforcement parameters. The master computing node then receives the current second reinforcement parameters from each slave computing node until a predetermined model reinforcement stop condition is met. At this point, the current aggregate reinforcement parameters are used as the target aggregate reinforcement parameters to obtain the enhanced target large language model M_PPO.

[0064] In this embodiment, after completing the training of the target large language model M_PPO, each computing node can save the trained M_PPO in each federated learning node, which can be directly used for reasoning. Therefore, the business customers of each institution can directly input multimodal data through the front-end developed by each node, and after converting it into Chinese text through the multimodal data conversion tool set locally at the node, call M_PPO to obtain the required output. At this time, there is no need for data interaction between nodes, which protects the data privacy of business customers from being leaked.

[0065] In the method of this embodiment, during each round of training, the main computing node determines the first target hidden layer from the first language model obtained from the training, and uses the parameters of each first target hidden layer as the parameters to be tuned. At the same time, the main computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the main node can aggregate the parameters of each first target hidden layer with the parameters of each second target hidden layer to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate next round of model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupancy of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0066] Another embodiment of the present application provides a model training method based on federated learning, which can be applied to each participant / slave computing node of federated learning, such as Figure 4 As shown, the method in this embodiment includes the following steps: Step S301, determining a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtaining initial second sub-parameters corresponding to the respective second target hidden layers; In this step, the slave computing node can first perform model tuning on the pre-trained original large language model based on the second conversation dataset to obtain an initial second large language model. The second conversation dataset is high-quality conversation data related to bond business locally generated by the slave computing node. Specifically, the second conversation data includes several question-and-answer data.

[0067] In this step, after the computing node obtains the initial second-largest language model through training, it is possible to determine the global score of each hidden layer in the initial large language model, and then determine several second target hidden layers to be adjusted based on the global score of each hidden layer, so as to facilitate subsequent parameter aggregation processing based on the initial second sub-parameters of each second target hidden layer. In other words, if the initial large language model contains, for example, 10 hidden layers, the method in this application can be used to determine 3 hidden layers from these 10 hidden layers as the first target hidden layers, and then the parameters corresponding to each second target hidden layer in the initial second training parameters obtained through training are used as the initial second sub-parameters.

[0068] Step S302: Send each initial second sub-parameter to the main computing node, so that the main computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain an initial aggregated parameter; In this step, after the slave computing node sends the initial second sub-parameters corresponding to each second target hidden layer to the master computing node, the master computing node can receive the initial second sub-parameters corresponding to each second target hidden layer sent by each slave computing node, and then the master computing node performs parameter aggregation based on the local first initial sub-parameters and the initial second sub-parameters of each slave computing node to obtain initial aggregated parameters. Among them, when performing parameter aggregation, the master computing node can perform parameter aggregation in combination with the initial parameters of each hidden layer, that is, based on the initial parameters of each hidden layer, each initial first sub-parameter and each initial second sub-parameter, perform parameter aggregation to obtain initial aggregated parameters. In this step, after the number of training rounds is greater than one round, the tuning parameters of each hidden layer obtained by the previous round of model tuning can also be used as the initial parameters to perform parameter aggregation in combination with the initial parameters. That is, based on the tuning parameters of each hidden layer obtained by the previous round of model tuning, each current first sub-parameter and each current second sub-parameter, parameter aggregation is performed to obtain the current aggregated parameters.

[0069] For example, a large language model has ten layers, labeled 1 to 8, with initial parameters H1, H2, H3, ..., H8 corresponding to each layer. Each computation node (master and slave) selects several target hidden layers with the highest scores, labeled a, b, ..., respectively.

[0070] For example, the main computing node a determines three first target hidden layers: a1, a2, and a3. The first target hidden layers a1, a2, and a3 correspond to the initial first sub-parameters A1, A2, and A3, respectively.

[0071] Three second target hidden layers are determined from the computing node b: b2, b3 and b4. The second target hidden layers b2, b3 and b4 correspond to the initial second sub-parameters B2, B3 and B4 respectively.

[0072] Three second target hidden layers are determined from the computing node c: c1, c3 and c5; the second target hidden layers c1, c3 and c5 correspond to the initial second sub-parameters C1, C3 and C5 respectively.

[0073] After collecting the sub-parameters of all target hidden layers, the master computation node can then perform the arithmetic averaging of the layers that need updating to obtain the optimized model parameters, labeled M1 through M8. Thus, M1 = (H1 + A1 + C1) ÷ 3. M2 = (H2 + A2 + B2) ÷ 3. M3 = (H3 + A3 + B3 + C3) ÷ ​​4. M4 = (H4 + B4) ÷ 2. M5 = (H5 + C5) ÷ 2. M6 = H6. M7 = H7. M8 = H8.

[0074] Step S303: Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the master computing node until the predetermined tuning conditions are met. Then, stop model tuning, use the received current aggregation parameters as the target aggregation parameters, and obtain the target large language model.

[0075] After receiving the initial aggregation parameters, the slave computing node can perform the next round of model tuning based on the initial aggregation parameters until the predetermined model conditions are met. In this embodiment, the predetermined model tuning conditions can be that the number of tuning rounds is greater than the predetermined number of rounds or the error corresponding to the current aggregation parameters is less than a predetermined error threshold.

[0076] In this embodiment, a model training method based on federated learning is used. In each round of training, the slave computing node determines the second target hidden layer from the second largest language model obtained from the training, and sends the parameters of each second target hidden layer as the parameters to be tuned to the main computing node. The main computing node can perform aggregation based on the parameters of each second target hidden layer and the parameters of the first target hidden layer determined locally to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate next round of model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupation of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0077] Based on the above embodiment, another embodiment of the present application provides a model training method based on federated learning, which can be specifically applied to each participant / slave computing node of federated learning. The method in this embodiment includes the following steps: Step S401: performing model training on a predetermined base model based on a second original vocabulary to obtain initial second intermediate parameters; In this step, when the computing node trains the base model, the following steps are specifically included: Step S401 - 1 , installation and deployment of a multimodal data conversion toolset.

[0078] In this step, the master computing node of the joint modeling project can pre-install a multimodal data conversion toolkit and send it to each slave computing node. Each slave computing node can then pre-install the multimodal data conversion toolkit and process data of various modalities based on the toolkit to construct a second original vocabulary.

[0079] Step S401-2: Installation and initialization of the federated learning collaborative network.

[0080] In this step, the master computing node can send the federated learning component to each slave computing node, so that each slave computing node can install the federated learning component. The framework of the federated learning collaborative network can be implemented using other software products with similar functions, such as FATE, SecretFlow, and PaddleFL.

[0081] After the federated learning components are deployed, the slave compute nodes can receive the base large language model / base model initial parameters sent by the master compute node. In this embodiment, other large language models with the same functionality, such as Chinese-LLaMA and OpenChineseLLaMA, can be used as the base model, and the vocabulary expansion tool SentencePiece is used.

[0082] Step S401-3: construct a second original vocabulary.

[0083] In this step, the slave computing node can process the local multimodal bond data into Chinese text data / Chinese language training using the deployed multimodal data conversion tool set. The slave computing node then trains the second original vocabulary, chn_nodei.model, based on the Chinese text data / Chinese corpus and sends it to the master computing node. The master computing node, which is then responsible for joint modeling, constructs a target vocabulary based on the first original vocabulary, chn_center.model, and each of the second original vocabulary, chn_nodei.model.

[0084] Step S401-4: construct a second target topology structure.

[0085] In this step, the slave computing node may construct a second target topology structure corresponding to the slave computing node based on the second server corresponding to the slave computing node and the graphics processors deployed on each second server.

[0086] Step S401-5: Perform model training on the predetermined base model.

[0087] In this step, each slave computing node may adopt a ring-type distributed learning method and utilize the second original vocabulary to perform model training on a predetermined base model for the second target topology structure.

[0088] Step S402: Send the initial second intermediate parameters to the master computing node, so that the master computing node performs parameter aggregation based on the initial second intermediate parameters and the initial first intermediate parameters of the master computing node to obtain initial aggregated intermediate parameters. Step S403: Receive the initial aggregated intermediate parameters sent by the main computing node, perform the next round of model training based on the initial aggregated intermediate parameters, and send the current second intermediate parameters obtained through training to the main computing node until the predetermined training conditions are met. Then, stop the model training, use the received current aggregated intermediate parameters as the target aggregated intermediate parameters, and obtain the original large language model.

[0089] During the implementation of this step, after all slave compute nodes complete a batch of training locally, they send the trained intermediate parameters (i.e., matrices A and B) to the master compute node via GPRC. The master compute node reads the intermediate parameters from all slave compute nodes, aggregates them using the FedAvg algorithm, and then sends the aggregated parameters back to all slave compute nodes to begin training a new batch (next round). This process repeats until the model training stop condition is met, resulting in the original large language model M. The stopping condition can be the completion of a predetermined batch or number of rounds of training, or the attainment of a satisfactory loss or the occurrence of certain errors. The predetermined batch / number of rounds and loss conditions are preset hyperparameters. Errors that cause training to stop may include network communication issues, timeouts between federated learning nodes, abnormal node status, and other issues.

[0090] Step S404: performing model tuning processing on the original large language model obtained through pre-training based on the second dialogue dataset to obtain an initial second large language model.

[0091] During the specific implementation of this step, after obtaining the original large language model M, it can be fine-tuned. This fine-tuning process is based on the high-quality conversation datasets local to each computing node (i.e., the first conversation dataset local to the master computing node and the second conversation dataset local to each slave computing node). Specifically, after completing federated training of the original large language model M, each node prepares its own high-quality conversation dataset related to bond business. The master computing node prepares the first conversation dataset, and each slave computing node prepares the second conversation dataset. These high-quality conversation datasets contain questions and answers likely to be asked in various bond business contexts. In this embodiment, the data in the high-quality conversation datasets is multimodal. For example, input questions may include scanned copies of business-related documents. This multimodal input data is converted to Chinese text using the aforementioned multimodal data conversion toolset, which will not be further detailed here.

[0092] Step S405: determining a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtaining initial second sub-parameters corresponding to the respective second target hidden layers; During implementation, this step can determine the global score of each hidden layer in the initial second-largest language model, and determine a number of second target hidden layers to be adjusted based on the global score of each hidden layer. Specifically, for the initial second-largest language model, the correlation matrix between any two hidden layer parameters can be determined; the global score of each hidden layer can be determined based on the eigenvalues ​​of each correlation matrix corresponding to the same hidden layer; and the number of hidden layers with global scores greater than a predetermined score threshold can be determined as the second target hidden layers. Alternatively, a predetermined number of hidden layers with the highest global scores can be selected as the second target hidden layers.

[0093] That is, during the prompt tuning process of each batch of data, the score of each layer of the large language model is dynamically calculated through the correlation matrix of the hidden state, and the appropriate layer is selected for prompt tuning, avoiding the huge computing power consumption and bandwidth occupation of the traditional method of full parameter update, reducing the communication cost in the federated learning process and the local computing cost of each participant, while improving the accuracy of model training. Specifically, during each prompt tuning process, for any two hidden layers in the large model, the matrix K composed of the cosine similarity between the parameters of the two hidden layers can be used to represent their correlation. The eigenvalue λ of the matrix K is Kij represents the difference between two hidden layers i and j in this data set. Therefore, the sum of the eigenvalues ​​of the correlation matrix K between hidden layer i and all other layers represents the global score of hidden layer i. A higher score indicates greater similarity between that layer and the other layers. The global scores of all hidden layers are calculated, and a predetermined number of hidden layers with the highest scores are selected as the layers to be tuned. The predetermined number is a hyperparameter and can be predetermined based on the square root of the total number of hidden layers.

[0094] Step S406: Send each initial second sub-parameter to the main computing node, so that the main computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain an initial aggregated parameter; In this step, after each slave computing node determines a number of second target hidden layers, it can send the initial second sub-parameters corresponding to the second target hidden layers to the master computing node. The master computing node can then perform parameter aggregation based on the initial parameters of each hidden layer, the initial first sub-parameters, and the initial second sub-parameters to obtain initial aggregated parameters. The parameters from the previous round, the initial first sub-parameters, and the initial second sub-parameters are aggregated to obtain initial aggregated parameters.

[0095] Step S407: Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained through tuning to the master computing node until the predetermined tuning conditions are met. Then, stop model tuning, use the received current aggregation parameters as target aggregation parameters, and obtain the target large language model. In this embodiment, the predetermined tuning condition may be that the number of tuning rounds is greater than a predetermined number of rounds or the error corresponding to the current aggregation parameter is less than a predetermined error threshold.

[0096] Step S408: Based on the constructed second question and answer set, the initial reward model is trained as a federated model to obtain a target reward model; In this step, after obtaining the target large language model, in order to make the final target large language model more accurate and reliable, the target reward model can be further trained from the computing node to facilitate subsequent federated reinforcement learning of the target large language model based on the target reward model.

[0097] In this step, the second question-and-answer set includes several partial-order dialogue pairs, each of which contains a sample question, several sample answers corresponding to the sample question, and the answer ranking of each sample answer. The process by which each slave computing node constructs the second question-and-answer set in this embodiment is similar to the process by which the master computing node constructs the first question-and-answer set in the above embodiment, and will not be further described here.

[0098] During the specific implementation of this step, the slave computing nodes can perform model training on the initial reward model based on the second question-answer set to obtain initial second reward parameters. These initial second reward parameters are then sent to the master computing node. The master computing node can then aggregate parameters based on the initial second reward parameters sent from each computing node and the initial first reward parameters obtained through local training to obtain initial aggregated reward parameters, which are then sent to each slave computing node. Each slave computing node then performs the next round of model training based on the initial aggregated reward parameters and sends the current second reward parameters obtained through training to the master computing node. Once the predetermined training conditions are met, model training is terminated and the received current aggregated reward parameters are used as the target aggregated reward parameters to obtain the target reward model M_RM.

[0099] In this step, the loss function of the reward model is:

[0100] Where x is the input prompt (i.e. question), y is the answer, and y c Represents the sort ratio y rThe sentence at the front is the one with better answer effect. m(r) represents the degree of difference between the two sentences. There are four selectable values: 3, 2, 1, and 0. For example, for two sentences answering y I and y J ,y I than y J When it is much better, the reward model gives the two answers a score that satisfies L(y I )>L(y J ), and at this time m(r)=3.

[0101] Step S409: performing federated reinforcement learning on the target large language model based on the target reward model to obtain a reinforced target large language model; In this step, after the target reward model M_RM is obtained through training, the slave computing node can score the answer output by the target large language model based on the target reward model, thereby further strengthening the learning of the target large language model according to the score of the answer. Specifically, the slave computing node uses the local training data set to generate results / answers / responses based on the target large language model M_SFT, and uses the target reward model M_RM to calculate the reward (i.e., score) of the answer / response. Based on the reward, the model parameters are adjusted to obtain the initial second reinforcement parameters, and the initial second reinforcement parameters are sent to the master computing node at the same time. Thus, the master node can perform parameter aggregation based on the initial second reinforcement parameters sent by each slave computing node and the initial first reinforcement parameters obtained through local training, obtain the initial aggregated reinforcement parameters, and send the initial aggregated reinforcement parameters to each slave computing node. Then, the slave computing node can receive the initial aggregation enhancement parameters sent by the master computing node, and perform the next round of model enhancement based on the initial aggregation enhancement parameters, and send the current second enhancement parameters obtained by the model enhancement to the master computing node until the predetermined enhancement conditions are met. Then, the model enhancement is stopped, and the received current aggregation enhancement parameters are used as the target aggregation enhancement parameters to obtain the enhanced target large language model M_PPO.

[0102] In this embodiment, after completing the training of the target large language model M_PPO, each computing node can save the trained M_PPO in each federated learning node, which can be directly used for reasoning. Therefore, the business customers of each institution can directly input multimodal data through the front-end developed by each node, and after converting it into Chinese text through the multimodal data conversion tool set locally at the node, call M_PPO to obtain the required output. At this time, there is no need for data interaction between nodes, which protects the data privacy of business customers from being leaked.

[0103] In the method of this embodiment, during each round of training, the slave computing node determines the second target hidden layer from the second largest language model obtained from the training, takes the parameters of each second target hidden layer as the parameters to be tuned, and sends them to the main computing node. Thus, the main computing node can aggregate based on the parameters of each second target hidden layer and the parameters of the first target hidden layer determined locally to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate training of the next round of model based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupation of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0104] Another embodiment of the present application provides a model training device based on federated learning, such as Figure 5 Shown, including: A first determining module 11 is configured to determine a plurality of first target hidden layers based on the initial first large language model obtained through training, and obtain initial first sub-parameters corresponding to the first target hidden layers; A first receiving module 12 is configured to receive initial second sub-parameters corresponding to respective second target hidden layers in the initial second largest language model, sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing nodes; an aggregation module 13, configured to perform parameter aggregation based on at least the initial first sub-parameters and the initial second sub-parameters to obtain initial aggregated parameters; The first sending module 14 is used to send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, the current aggregation parameters are used as the target aggregation parameters, and the target large language model is obtained.

[0105] In the specific implementation of this embodiment, the model training device based on federated learning also includes a first training module, which is used to perform model tuning processing on the original large language model obtained by pre-training based on the first dialogue data set before determining a number of first target hidden layers to obtain an initial first large language model.

[0106] In the specific implementation of this embodiment, the first determination module is specifically used to: for the initial first large language model, determine the correlation matrix between any two hidden layer parameters respectively; determine the global score of each hidden layer based on the eigenvalues ​​of each correlation matrix corresponding to the same hidden layer; and determine several hidden layers whose global scores are greater than a predetermined score threshold as the first target hidden layer.

[0107] In the specific implementation process of this embodiment, the model training device based on federated learning further includes a first pre-training module, which is used to: pre-train the original large language model obtained, specifically to: Performing model training on a predetermined base model based on a target vocabulary to obtain initial first intermediate parameters; wherein the target vocabulary is obtained by merging a second original vocabulary of the slave computing node and a first original vocabulary locally on the master computing node; Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, parameter aggregation is performed to obtain initial aggregated intermediate parameters, and the initial aggregated intermediate parameters are sent to the slave computing nodes so that the slave computing nodes can perform the next round of model training based on the initial aggregated intermediate parameters, and receive the current second intermediate parameters sent by each slave computing node until a predetermined model training stop condition is met, and use the current aggregated intermediate parameters as the target aggregated intermediate parameters to obtain the original large language model.

[0108] In a specific implementation of this embodiment, the model training device based on federated learning further includes a first construction module, which is configured to: construct a first target topology structure corresponding to the main computing node based on the first server corresponding to the main computing node and the graphics processor deployed on each first server; The first pre-training module is specifically used to: for the first target topology structure, adopt a ring-type distributed learning method and use a target vocabulary to perform model training on a predetermined base model.

[0109] In the specific implementation process of this embodiment, the model training device based on federated learning further includes a first reward model training module and a first reinforcement module; the first reward model training module is used to: perform federated model training on the initial reward model based on the constructed first question and answer set to obtain a target reward model; The first reinforcement module is used to perform federated reinforcement learning on the target large language model based on the target reward model to obtain a reinforced target large language model.

[0110] In the device of this embodiment, during each round of training, the main computing node determines the first target hidden layer from the first language model obtained from the training, and uses the parameters of each first target hidden layer as the parameters to be tuned. At the same time, the main computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the main node can aggregate the parameters of each first target hidden layer with the parameters of each second target hidden layer to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate next round of model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupancy of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0111] Another embodiment of the present application provides a model training device based on federated learning, such as Figure 6 Shown, including: A second determining module 21 is configured to determine a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtain initial second sub-parameters corresponding to the respective second target hidden layers; A second sending module 22 is configured to send each initial second sub-parameter to a master computing node, so that the master computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain an initial aggregated parameter; The second receiving module 23 is used to receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the master computing node until the predetermined tuning conditions are met. The model tuning is stopped, and the received current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.

[0112] During the specific implementation of this embodiment, the second determination module is specifically used to: for the initial second largest language model, determine the correlation matrix between any two hidden layer parameters respectively; determine the global score of each hidden layer based on the eigenvalues ​​of each correlation matrix corresponding to the same hidden layer; and determine several hidden layers whose global scores are greater than a predetermined score threshold as the second target hidden layers.

[0113] In the specific implementation process of this embodiment, the model training device based on federated learning further includes a second pre-training module, which is used to: pre-train the original large language model obtained, specifically to: The predetermined base model is trained based on the second original vocabulary to obtain initial second intermediate parameters; the initial second intermediate parameters are sent to the main computing node, so that the main computing node can perform parameter aggregation based on the initial second intermediate parameters and the initial first intermediate parameters of the main computing node to obtain initial aggregated intermediate parameters; the initial aggregated intermediate parameters sent by the main computing node are received, and the next round of model training is performed based on the initial aggregated intermediate parameters, and the current second intermediate parameters obtained by training are sent to the main computing node until the predetermined training conditions are met, then the model training is stopped, and the received current aggregated intermediate parameters are used as the target aggregated intermediate parameters to obtain the original large language model.

[0114] In the specific implementation process of this embodiment, the model training device based on federated learning further includes a second construction module; the second construction module is used to: construct a second target topology corresponding to the slave computing node based on the second server corresponding to the slave computing node and the graphics processor deployed on each second server; The second pre-training module is used to: for the second target topology structure, adopt a ring-distributed learning method and use the second original vocabulary to perform model training on a predetermined base model.

[0115] In the specific implementation process of this embodiment, the model training device based on federated learning further includes a second reward model training module and a second reinforcement module; The second reward model training module is used to: perform federated model training on the initial reward model based on the constructed second question and answer set to obtain a target reward model; The second reinforcement module is used to perform federated reinforcement learning on the target large language model based on the target reward model to obtain a reinforced target large language model.

[0116] The device in this embodiment, during each round of training, determines the second target hidden layer from the second largest language model obtained from the training from the computing node, and sends the parameters of each second target hidden layer as the parameters to be tuned to the main computing node, so that the main computing node can aggregate based on the parameters of each second target hidden layer and the parameters of the first target hidden layer determined locally to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate training of the next round of model based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then aggregating the parameters based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupation of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0117] Another embodiment of the present application provides an electronic device, such as Figure 7 As shown, it at least includes a memory 1 and a processor 2. The memory 1 stores a computer program. When the processor 2 executes the computer program on the memory 1, it implements the following method steps: Step 1: determining a plurality of first target hidden layers based on the initial first language model obtained through training, and obtaining initial first sub-parameters corresponding to each first target hidden layer; Step 2: receiving initial second sub-parameters corresponding to each second target hidden layer in the initial second largest language model sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing nodes; Step 3: Perform parameter aggregation based on at least each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregated parameter; Step 4: Send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, stop model tuning, use the current aggregation parameters as the target aggregation parameters, and obtain the target large language model.

[0118] Alternatively, implement the following method steps: Step 1: determining a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtaining initial second sub-parameters corresponding to each second target hidden layer; Step 2: Send each initial second sub-parameter to the main computing node, so that the main computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain an initial aggregated parameter; Step 3: Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the main computing node until the predetermined tuning conditions are met. Then, stop model tuning, use the received current aggregation parameters as the target aggregation parameters, and obtain the target large language model.

[0119] The specific implementation process of the above method steps can be found in any of the above-mentioned embodiments of the model training method based on federated learning, and this embodiment will not be repeated here.

[0120] For the electronic device in this embodiment, during each round of training, the main computing node determines the first target hidden layer from the first language model obtained through training, and uses the parameters of each first target hidden layer as the parameters to be tuned. At the same time, the main computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the main node can aggregate the parameters of each first target hidden layer with the parameters of each second target hidden layer to obtain aggregated parameters, so that the aggregated parameters are more reasonable and accurate, laying the foundation for the subsequent accurate next round of model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then performing parameter aggregation based on the target hidden layer to tune the target hidden layer, the final tuning result can be made more accurate, and at the same time, the huge computing power consumption and bandwidth occupancy of the full parameter update in the traditional method can be avoided, thereby reducing the communication cost in the federated learning process and the local computing cost of each participant.

[0121] The above embodiments are merely exemplary embodiments of the present application and are not intended to limit the scope of the present application. The scope of protection of the present application is defined by the claims. Those skilled in the art may make various modifications or equivalent substitutions to the present application within the essence and scope of protection of the present application, and such modifications or equivalent substitutions shall also be deemed to fall within the scope of protection of the present application.

Claims

1. A model training method based on federated learning, applied to a master computing node, characterized in that: include: Determining a plurality of first target hidden layers based on the initial first language model obtained through training, and obtaining initial first sub-parameters corresponding to each first target hidden layer; receiving initial second sub-parameters corresponding to respective second target hidden layers in the initial second largest language model sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing node; Performing parameter aggregation based on at least each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregated parameter; The initial aggregation parameters are sent to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, the current aggregation parameters are used as the target aggregation parameters, and the target large language model is obtained.

2. The method according to claim 1, wherein Before determining the plurality of first target hidden layers, the method further includes: Based on the first dialogue dataset, the original large language model obtained by pre-training is fine-tuned to obtain an initial first large language model.

3. The method according to claim 1, wherein The determining of a plurality of first target hidden layers based on the initial first language model obtained through training specifically includes: For the initial first language model, determine the correlation matrix between any two hidden layer parameters; Determine the global score of each hidden layer based on the eigenvalues ​​of the correlation matrices corresponding to the same hidden layer; Several hidden layers whose global scores are greater than a predetermined score threshold are determined as first target hidden layers.

4. The method according to claim 1, wherein The method further includes: pre-training an original large language model, specifically including: Performing model training on a predetermined base model based on a target vocabulary to obtain initial first intermediate parameters; wherein the target vocabulary is obtained by merging a second original vocabulary of the slave computing node and a first original vocabulary locally on the master computing node; Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, parameter aggregation is performed to obtain initial aggregated intermediate parameters, and the initial aggregated intermediate parameters are sent to the slave computing nodes so that the slave computing nodes can perform the next round of model training based on the initial aggregated intermediate parameters, and receive the current second intermediate parameters sent by each slave computing node until a predetermined model training stop condition is met, and use the current aggregated intermediate parameters as the target aggregated intermediate parameters to obtain the original large language model.

5. The method according to claim 4, wherein Before performing model training on a predetermined base model based on the target vocabulary, the method further includes: Building a first target topology structure corresponding to the main computing node based on the first server corresponding to the main computing node and the graphics processors deployed on each first server; The model training of the predetermined base model based on the target vocabulary specifically includes: For the first target topology, a ring-type distributed learning method is adopted and the target vocabulary is used to perform model training on the predetermined base model.

6. The method according to claim 1, wherein After obtaining the target large language model, the method further includes: Based on the constructed first question-answer set, the initial reward model is trained as a federated model to obtain the target reward model. Based on the target reward model, federated reinforcement learning is performed on the target large language model to obtain the reinforced target large language model.

7. A model training method based on federated learning, applied to a slave computing node, characterized in that: include: Determining a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtaining initial second sub-parameters corresponding to each second target hidden layer; Sending each initial second sub-parameter to the main computing node, so that the main computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain an initial aggregated parameter; Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the main computing node until the predetermined tuning conditions are met. Then stop model tuning, use the received current aggregation parameters as the target aggregation parameters, and obtain the target large language model.

8. A model training device based on federated learning, characterized in that: include: a first determining module, configured to determine a plurality of first target hidden layers based on the initial first large language model obtained through training, and obtain initial first sub-parameters corresponding to the first target hidden layers; a first receiving module, configured to receive initial second sub-parameters corresponding to respective second target hidden layers in the initial second largest language model, sent by each slave computing node; wherein the initial second largest language model is obtained by training the original large language model by the slave computing nodes; an aggregation module, configured to perform parameter aggregation based at least on each initial first sub-parameter and each initial second sub-parameter to obtain an initial aggregated parameter; The first sending module is used to send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, the current aggregation parameters are used as the target aggregation parameters, and the target large language model is obtained.

9. A model training device based on federated learning, characterized in that: include: a second determining module, configured to determine a plurality of second target hidden layers based on the initial second largest language model obtained through training, and obtain initial second sub-parameters corresponding to the respective second target hidden layers; A second sending module is configured to send each initial second sub-parameter to a master computing node, so that the master computing node performs parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain an initial aggregated parameter; The second receiving module is used to receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the main computing node until the predetermined tuning conditions are met. The model tuning is stopped, and the received current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.

10. An electronic device, characterized in that: The system comprises at least a memory and a processor, wherein a computer program is stored on the memory, and when the processor executes the computer program on the memory, the processor implements the steps of the model training method based on federated learning as described in any one of claims 1 to 6 or claim 7.

Citation Information

Patent Citations

  • Prediction method and system for error back propagation neural network and server

    CN105373830A

  • Data transmission method and device, equipment and storage medium

    CN115150228A

  • Model aggregation method, device and equipment, federal learning system and storage medium

    CN117808125A

  • Internet of vehicles federal dynamic sparse training method based on cluster expansion

    CN119906969A

  • Object recognition using a convolutional neural network trained by principal component analysis and repeated spectral clustering

    US20190164047A1