A model training method, apparatus, and device based on federated learning
Patent Information
- Application Number
- CN202510678452.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-26
- Publication Date
- 2026-09-01
- Estimated Expiration
- 2045-05-26
AI Technical Summary
[0004]有鉴于此,本发明提供了一种基于联邦学习的模型训练方法、装置及设备,主要目的在于解决目前联邦学习方法存在模型训练不准确的问题
[0010]本申请中的一种基于联邦学习的模型训练方法、装置及设备,在每轮训练时,主计算节点通过从训练获得的第一大语言模型中确定出第一目标隐含层,并将各第一目标隐含层的参数作为待调优的参数,与此同时主计算节点还会接收各从计算节点发送的各第二目标隐含层的参数,由此主节点就可以将各第一目标隐含层的参数与各第二目标隐含层的参数进行聚合,获得聚合参数,使得聚合获得的参数更加合理、准确,为后续基于聚合参数精准的进行下一轮模型训练奠定了基础。本申请中,通过从各隐含层中确定待调优的目标隐含层,然后基于目标隐含层进行参数聚合、以对目标隐含层进行调优,能够使得最终的调优结果更加精准,同时能够避免传统方法中进行全量参数更新的巨大算力消耗和带宽占用,降低了联邦学习过程中的通信成本和每个参与方本地的计算成本。
Smart Images

Figure CN120822579B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of finance and federated learning, and particularly to a model training method, apparatus and device based on federated learning. Background Technology
[0002] Federated learning is a distributed learning technology that combines traditional cryptography and machine learning. It aims to build a federated learning model based on a distributed dataset for multiple data providers who are unwilling or unable to expose their plaintext data.
[0003] However, in the financial field, existing federated learning methods suffer from inaccurate model training. Summary of the Invention
[0004] In view of this, the present invention provides a model training method, apparatus and device based on federated learning, the main purpose of which is to solve the problem of inaccurate model training in current federated learning methods.
[0005] To address the aforementioned issues, this application provides a model training method based on federated learning, applied to the master computing node, comprising: Based on the initial first large language model obtained from training, determine several first target hidden layers, and obtain the initial first sub-parameters corresponding to each first target hidden layer; Receive the initial second sub-parameters sent from each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; Based at least on each initial first sub-parameter and each initial second sub-parameter, parameter aggregation is performed to obtain initial aggregated parameters; The initial aggregation parameters are sent to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. The current second sub-parameters sent by each slave computing node are received until the predetermined tuning conditions are met. Then the model tuning is stopped, and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.
[0006] To address the aforementioned issues, this application provides a model training method based on federated learning, applicable to computational nodes, including: Based on the initial second-largest language model obtained through training, several second-target hidden layers are determined, and the initial second sub-parameters corresponding to each second-target hidden layer are obtained. Each initial second sub-parameter is sent to the master computing node so that the master computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain the initial aggregated parameters; The system receives the initial aggregation parameters sent by the master node, performs the next round of model tuning based on the initial aggregation parameters, and sends the current second sub-parameter obtained from the tuning to the master computing node. When the predetermined tuning conditions are met, the model tuning stops, and the received current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.
[0007] To address the aforementioned problems, this application provides a model training device based on federated learning, comprising: The first determining module is used to determine several first target hidden layers based on the initial first large language model obtained through training, and to obtain the initial first sub-parameters corresponding to each first target hidden layer; The first receiving module is used to receive the initial second sub-parameters sent by each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; The aggregation module is used to aggregate parameters based on at least each initial first sub-parameter and each initial second sub-parameter to obtain initial aggregated parameters; The first sending module is used to send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. It also receives the current second sub-parameters sent by each slave computing node until the predetermined tuning conditions are met, at which point the model tuning stops, and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.
[0008] To address the aforementioned problems, this application provides a model training device based on federated learning, comprising: The second determination module is used to determine several second target hidden layers based on the initial second large language model obtained through training, and to obtain the initial second sub-parameters corresponding to each second target hidden layer. The second sending module is used to send each initial second sub-parameter to the main computing node, so that the main computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain the initial aggregated parameters; The second receiving module is used to receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the master computing node until the predetermined tuning conditions are met. Then, the model tuning stops, and the received current aggregation parameters are used as the target aggregation parameters to obtain the target large language model.
[0009] To address the aforementioned problems, this application provides an electronic device, comprising at least a memory and a processor, wherein the memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the model training method based on federated learning as described in any of the preceding claims.
[0010] This application discloses a model training method, apparatus, and device based on federated learning. In each training round, the master computing node determines the first target hidden layer from the first large language model obtained through training, and uses the parameters of each first target hidden layer as parameters to be tuned. Simultaneously, the master computing node receives the parameters of each second target hidden layer sent by each slave computing node. The master node can then aggregate the parameters of each first target hidden layer and each second target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for subsequent, precise model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer and then aggregating parameters based on the target hidden layer for tuning, the final tuning result is more accurate. Simultaneously, it avoids the huge computational power consumption and bandwidth occupation of traditional methods involving full parameter updates, reducing communication costs and local computational costs for each participant in the federated learning process.
[0011] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0012] Various other advantages and benefits will become apparent to those skilled in the art upon reading the following detailed description of preferred embodiments. The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings: Figure 1 A flowchart illustrating a model training method based on federated learning, as described in an embodiment of this application; Figure 2 This is a schematic diagram of a 3-layer topology in another embodiment of this application; Figure 3 This is a schematic diagram of a distributed machine learning process in another embodiment of this application; Figure 4 A flowchart illustrating a model training method based on federated learning, as described in another embodiment of this application; Figure 5 This is a structural block diagram of a model training device based on federated learning, according to another embodiment of this application. Figure 6 This is a structural block diagram of a model training device based on federated learning, according to another embodiment of this application. Figure 7 This is a structural block diagram of an electronic device according to another embodiment of this application. Detailed Implementation
[0013] Various embodiments and features of this application are described herein with reference to the accompanying drawings.
[0014] It should be understood that various modifications can be made to the embodiments described herein. Therefore, the above description should not be considered as limiting, but merely as an example of embodiments. Other modifications within the scope and spirit of this application will be apparent to those skilled in the art.
[0015] The accompanying drawings, which are included in and form part of this specification, illustrate embodiments of the present application and, together with the general description of the present application given above and the detailed description of the embodiments given below, serve to explain the principles of the present application.
[0016] These and other features of this application will become apparent from the following description of preferred forms of embodiments given as non-limiting examples, with reference to the accompanying drawings.
[0017] It should also be understood that although this application has been described with reference to some specific examples, those skilled in the art can certainly implement many other equivalent forms of this application.
[0018] The above and other aspects, features and advantages of this application will become more apparent when taken in conjunction with the accompanying drawings and in view of the following detailed description.
[0019] Specific embodiments of this application are described thereafter with reference to the accompanying drawings; however, it should be understood that the claimed embodiments are merely examples of this application, which can be implemented in various ways. Well-known and / or repeated functions and structures are not described in detail to avoid unnecessary or redundant details that could obscure the application. Therefore, the specific structural and functional details claimed herein are not intended to be limiting, but merely serve as the basis and representative basis for the claims to teach those skilled in the art to use this application in a variety of substantially any suitable detailed structures.
[0020] This specification may use the phrases “in one embodiment,” “in another embodiment,” “in yet another embodiment,” or “in other embodiments,” all of which may refer to one or more of the same or different embodiments according to this application.
[0021] This application provides a model training method based on federated learning, which can be specifically applied to the leading party / master computing node in federated learning, such as... Figure 1As shown, the method in this embodiment includes the following steps: Step S101: Based on the initial first large language model obtained through training, determine several first target hidden layers and obtain the initial first sub-parameters corresponding to each first target hidden layer; In this step, the pre-trained original large language model can be fine-tuned based on the first dialogue dataset to obtain the initial first large language model. The first dialogue dataset consists of high-quality dialogue data related to bond business, located locally on the main computing node. Specifically, the first dialogue data contains several question-and-answer data points.
[0022] In this step, after the main computing node obtains the initial large language model through training, it can determine the global score of each hidden layer in the initial large language model, and then determine several first target hidden layers to be adjusted based on the global score of each hidden layer. This facilitates subsequent parameter aggregation processing based on the initial first sub-parameters of each first target hidden layer. That is, if the initial large language model contains, for example, 10 hidden layers, the method in this application can determine 3 hidden layers as first target hidden layers from these 10 hidden layers, and then the parameters corresponding to each first target hidden layer in the initial first training parameters obtained during training are used as initial first sub-parameters.
[0023] Step S102: Receive the initial second sub-parameters sent by each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; In this step, after the slave computing node obtains the initial second-largest language model through training, it also determines the global score of each hidden layer in the initial second-largest language model. Then, based on the global score of each hidden layer, it determines several second target hidden layers to be adjusted, so that the initial second sub-parameters corresponding to each second target hidden layer can be sent to the master computing node. In this way, the master computing node can receive the initial second sub-parameters corresponding to each second target hidden layer sent by each slave computing node.
[0024] Step S103: At least based on each initial first sub-parameter and each initial second sub-parameter, perform parameter aggregation to obtain initial aggregated parameters; In this step, parameter aggregation can be performed by combining the initial parameters of each hidden layer. That is, based on the initial parameters, initial first sub-parameters, and initial second sub-parameters of each hidden layer, parameter aggregation is performed to obtain the initial aggregated parameters. In this step, after more than one training epoch, the tuning parameters of each hidden layer obtained in the previous model tuning epoch can also be used as initial parameters for parameter aggregation. That is, based on the tuning parameters, current first sub-parameters, and current second sub-parameters of each hidden layer obtained in the previous model tuning epoch, parameter aggregation is performed to obtain the current aggregated parameters.
[0025] For example, a large language model has eight layers, labeled from 1 to 8, with initial parameters H1, H2, H3...H8 for each layer. Each computation node (master and slave nodes) selects several target hidden layers with the highest scores and labels them a, b... respectively.
[0026] For example, the main computing node a determines three first target hidden layers: a1, a2, and a3. The first target hidden layers a1, a2, and a3 correspond to the initial first sub-parameters A1, A2, and A3, respectively.
[0027] Three second target hidden layers are determined from the computing node b: b2, b3 and b4. The second target hidden layers b2, b3 and b4 correspond to the initial second sub-parameters B2, B3 and B4, respectively.
[0028] Three second target hidden layers are determined from the computing node c: c1, c3 and c5; the second target hidden layers c1, c3 and c5 correspond to the initial second sub-parameters C1, C3 and C5 respectively.
[0029] Therefore, after collecting the sub-parameters of all target hidden layers, the master compute node can obtain the optimized model parameters by performing an arithmetic average on the layers that need updating, labeled as M1 to M8. Thus, M1 = (H1 + A1 + C1) ÷ 3. M2 = (H2 + A2 + B2) ÷ 3. M3 = (H3 + A3 + B3 + C3) ÷ 4. M4 = (H4 + B4) ÷ 2. M5 = (H5 + C5) ÷ 2. M6 = H6. M7 = H7. M8 = H8.
[0030] Step S104: Send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. Receive the current second sub-parameter sent by each slave computing node until the predetermined tuning conditions are met. Stop the model tuning, use the current aggregation parameters as the target aggregation parameters, and obtain the target large language model.
[0031] In this step, after parameter aggregation is completed, the master compute node can send the aggregated parameters to each slave compute node, so that the slave compute nodes can perform the next round of model tuning based on the aggregated parameters until the predetermined model conditions are met. In this embodiment, the predetermined model tuning conditions can be that the number of tuning rounds is greater than a predetermined number of rounds or that the error corresponding to the current aggregated parameters is less than a predetermined error threshold.
[0032] This embodiment presents a federated learning-based model training method. In each training round, the master computing node determines the first target hidden layer from the first large language model obtained through training, and uses the parameters of each first target hidden layer as parameters to be tuned. Simultaneously, the master computing node receives the parameters of each second target hidden layer from each slave computing node. The master node can then aggregate the parameters of each first and second target hidden layer to obtain aggregated parameters, making the aggregated parameters more reasonable and accurate. This lays the foundation for subsequent, precise model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer and then aggregating parameters based on the target hidden layer for tuning, the final tuning result is more accurate. Simultaneously, it avoids the huge computational power consumption and bandwidth occupation of full parameter updates in traditional methods, reducing the communication cost and local computational cost of each participant in the federated learning process.
[0033] Based on the above embodiments, another embodiment of this application provides a model training method based on federated learning, which can be applied to the master computing node. The method in this embodiment includes the following steps: Step S201: Train the predetermined base model based on the target vocabulary to obtain the initial first intermediate parameters; wherein, the target vocabulary is obtained by merging the second original vocabulary from the computing node and the first original vocabulary local to the main computing node; In this step, the training of the base model specifically includes the following steps: Step S201-1, Installation and deployment of the multimodal data conversion toolset.
[0034] In this step, the lead / master computing node of the joint modeling can pre-install and deploy a multimodal data conversion toolkit, and send the multimodal data conversion toolkit to each slave computing node, so that each slave computing node can process data of various modalities based on the multimodal data conversion toolkit, thereby constructing a second original vocabulary.
[0035] Step S201-2, Installation and initialization of the federated learning collaboration network.
[0036] In this step, the master computing node and each slave computing node can also pre-deploy federated learning components. In this embodiment, the framework of the federated learning collaborative network can be implemented using other software products with similar functions, such as FATE, SecretFlow, and PaddleFL.
[0037] After the federated learning component is deployed, the master compute node will send the base large language model / base model initial parameters to each slave compute node via gRPC. In this embodiment, other large language models with similar functionality, such as Chinese-LLaMA and OpenChineseLLaMA, can be used as the base model, and the vocabulary expansion uses the SentencePiece tool. In the data preprocessing step, each compute node processes its local multimodal bond data into Chinese text data / Chinese corpus data using the deployed multimodal data conversion toolset, which serves as the raw data for training the federated multimodal large language model. During the processing, the raw data is not sent externally to protect the privacy and security of each institution's business data.
[0038] Step S201-3: Construct the target vocabulary.
[0039] In this step, the master compute node can process local multimodal bond data into Chinese text data / Chinese corpus using a deployed multimodal data conversion toolset, and then train a first original vocabulary based on the Chinese text data / Chinese corpus. Similarly, after obtaining local Chinese text data / Chinese corpus, slave compute nodes can train a second original vocabulary based on their local Chinese text data / Chinese corpus. That is, the leading party / master compute node in the joint modeling trains the Chinese vocabulary / first original vocabulary chn_center.model on the bond data within the general Chinese corpus, while all participants / slave compute nodes in the joint modeling train their own Chinese vocabulary / second original vocabulary chn_nodei.model on the Chinese corpus composed of their internal business data, and send them to the master compute node. The master compute node merges the first original vocabulary and the second original vocabulary of each slave compute node, excluding duplicate tokens, and uses this as the target vocabulary for federated training of the base model.
[0040] Step S201-4: Construct the first target topology.
[0041] In this step, the master computing node can construct a first target topology corresponding to the master computing node based on the first server corresponding to the master computing node and the graphics processors deployed on each first server.
[0042] Step S201-5: Train the predetermined base model.
[0043] In this step, model training can be performed based on the first target topology and the target vocabulary. That is, the main computing node, for the first target topology, uses a ring-distributed learning approach and the target vocabulary to train the predetermined base model to obtain the initial first intermediate parameters.
[0044] A ring-shaped distributed machine learning architecture is constructed within each computing node. The default 3D ring structure using the 3D-Torus algorithm forms a three-layer topology within the federated learning nodes (inter-node > intra-node machines -> single-machine multi-GPU), resulting in a three-dimensional decomposition, such as... Figure 2 As shown. S1.X represents a slave compute node, with a bandwidth of w2 between it and the top-level master compute node S2.0 (generally via forwarding through Nginx, using bandwidth from a dedicated financial line); S0.X is a single server, with a bandwidth of w1 between it and the gateway of the slave compute nodes (generally an inter-machine switch); A0, B0, etc., represent individual GPUs, with a bandwidth of w0 between them (e.g., a single machine with multiple GPUs). The distributed machine learning process is as follows: Figure 3 As shown, in the first phase, four Scatter-Reduce operations (A0-A1-A2, B0-B1-B2, C0-C1-C2, D0-D1-D2) are executed simultaneously within a node (under the same S0.X switch, bandwidth w0). In the second phase, six Scatter-Reduce operations (A0-B0, A1-B1, A2-B2, C0-D0, C1-D1, C2-D2) are executed simultaneously between nodes (under the same S1.X switch, bandwidth w1). These occur concurrently and share the S1.x bandwidth. In the third phase, six Scatter-Reduce operations (A0-C0, A1-C1, A2-C2, B0-D0, B1-D1, B2-D2) are executed simultaneously between nodes, while sharing the S1.x bandwidth on S2.0. After Step 3, each GPU has S(size) / 12 of data, which is the complete data after all Reduce-Scatter scattering reduction operations are completed. At this point, all Scatter-Reduce scattering reduction operations are finished. The fourth stage, all-gather, will be executed in the exact same but reversed order, ultimately completing all-reduce all-gather. When applied to large language models with fewer parameters, or when there is no multi-machine / multi-GPU interaction among computing nodes, the 2D ring structure of the 2D-Torus algorithm can be used to reduce one layer of data partitioning interaction. The training process is simplified to intra-group scatter-reduce -> inter-group all-reduce -> intra-group all-gather, which will not be elaborated further.
[0045] In this embodiment, the LoRA method can be used to pre-train the federated multimodal large language model using the aforementioned 3D-Torus structure. The master compute node generates a reduced-dimensional matrix A and an increased-dimensional matrix B as training objects. Matrix A is initialized with a random Gaussian distribution, and matrix B is initialized as a zero matrix. The matrix parameters of the pre-trained model remain frozen. During training, the input and output dimensions of the model remain unchanged, and the output is the superposition of BA and the parameters of the pre-trained model. The master compute node sends matrices A and B to all slave compute nodes. All compute nodes involved in this task use their local training datasets to perform local training on a batch (epoch) of data using a 3D ring distributed machine learning structure.
[0046] Step S202: Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, perform parameter aggregation to obtain initial aggregated intermediate parameters, and send the initial aggregated intermediate parameters to the slave computing nodes so that the slave computing nodes can perform the next round of model training based on the initial aggregated intermediate parameters. Receive the current second intermediate parameters sent by each slave computing node until the predetermined model training stopping condition is met, and use the current aggregated intermediate parameters as the target aggregated intermediate parameters to obtain the original large language model.
[0047] In this step, after all slave computing nodes perform a batch of training locally, they send the intermediate parameters (i.e., matrices A and B) to the master computing node via GPRC. The master computing node reads the intermediate parameters from all slave computing nodes, aggregates them using the FedAvg algorithm, and then sends the aggregated parameters back to all slave computing nodes to start a new batch (next round) of training. This process is repeated until the model training stopping condition is met, resulting in the original large language model M. The stopping condition can be the completion of a predetermined batch / predetermined number of training rounds, or the acquisition of a satisfactory loss or the occurrence of certain errors. The predetermined batch / predetermined number of rounds and the loss condition are preset hyperparameters, while errors causing training to stop may include network communication problems such as timeouts between federated learning nodes or abnormal node states.
[0048] Step S203: Based on the first dialogue dataset, perform model optimization on the pre-trained original large language model to obtain the initial first large language model.
[0049] In this step, after obtaining the original large language model M, it can be optimized. The optimization process is based on high-quality dialogue datasets local to each computing node (i.e., the first dialogue dataset on the master computing node and the second dialogue dataset on each slave computing node). Specifically, after completing the federated training of the original large language model M, each node prepares its own high-quality dialogue dataset related to its bond business; the master computing node prepares the first dialogue dataset, and each slave computing node prepares the second dialogue dataset. These high-quality dialogue datasets contain questions that may be asked in various bond business scenarios, as well as the answers to those questions. In this embodiment, the data in the high-quality dialogue datasets are all multimodal data. For example, the input question may contain scanned copies of documents related to a certain business. Such multimodal input data is converted into Chinese text using the aforementioned multimodal data conversion toolset, which will not be elaborated further here.
[0050] Step S204: Based on the initial first large language model obtained through training, determine several first target hidden layers and obtain the initial first sub-parameters corresponding to each first target hidden layer; In the specific implementation process, this step involves determining the global score of each hidden layer in the initial first-largest language model, and then determining several first-target hidden layers to be adjusted based on the global scores of each hidden layer. Specifically, for the initial first-largest language model, the correlation matrix between the parameters of any two hidden layers can be determined; based on the eigenvalues of the correlation matrices corresponding to the same hidden layer, the global score of each hidden layer is determined; several hidden layers with global scores greater than a predetermined score threshold are determined as first-target hidden layers, or alternatively, a predetermined number of hidden layers with the highest global scores can be used as first-target hidden layers.
[0051] In other words, during the feedback tuning process for each batch of data, the score of each layer of the large language model is dynamically calculated using the correlation matrix of the hidden states. Appropriate layers are selected for feedback tuning, avoiding the enormous computational and bandwidth consumption of full parameter updates in traditional methods. This reduces communication costs and local computational costs for each participant in the federated learning process, while simultaneously improving the accuracy of model training. Specifically, during each feedback tuning process, for any two hidden layers in the large model, their correlation can be represented by a matrix K composed of the cosine similarity between the parameters of the two hidden layers, and the eigenvalues λ of matrix K... KijThis represents the magnitude of the difference between two hidden layers i and j in this batch of data. Therefore, the summation of the eigenvalues of the correlation matrix K between hidden layer i and all other layers yields a score that represents the global score of hidden layer i; a higher score indicates a higher similarity between this layer and other layers. The global scores of all hidden layers are calculated, and a predetermined number of hidden layers with the highest scores are selected as the layers to be optimized in this iteration. The predetermined number is a hyperparameter that can be pre-determined based on the square root of the total number of hidden layers.
[0052] Step S205: Receive the initial second sub-parameters sent by each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; In this step, each slave computing node also optimizes its local original large language model based on the second dialogue dataset, obtaining an initial second large language model. Then, it determines the global score of each hidden layer in the initial second large language model and identifies several second target hidden layers to be adjusted based on the global scores of each hidden layer. The specific process for determining the second target hidden layers is similar to that of determining the first target hidden layer in the above embodiment and will not be repeated here. After determining the second target hidden layers, each slave computing node can send the initial second sub-parameters corresponding to the second target hidden layers to the master computing node.
[0053] Step S206: At least based on each initial first sub-parameter and each initial second sub-parameter, perform parameter aggregation to obtain initial aggregated parameters; When performing parameter aggregation, the initial parameters of each hidden layer can be combined to perform parameter aggregation. That is, based on the initial parameters of each hidden layer, each initial first sub-parameter, and each initial second sub-parameter, parameter aggregation is performed to obtain the initial aggregated parameters.
[0054] Step S207: Send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. Receive the current second sub-parameter sent by each slave computing node until the predetermined tuning conditions are met. Stop the model tuning, use the current aggregation parameters as the target aggregation parameters, and obtain the target large language model.
[0055] In this step, after obtaining the original large language model M, it can be optimized. The optimization process is based on high-quality dialogue datasets local to each computing node (i.e., the first dialogue dataset on the master computing node and the second dialogue dataset on each slave computing node). Specifically, after completing the federated training of the original large language model M, each node prepares its own high-quality dialogue data related to its bond business (the master computing node prepares the first dialogue dataset, and each slave computing node prepares the second dialogue dataset), containing questions that may be asked in various bond businesses, as well as the answers to those questions. Using the same distributed training process and intermediate parameter interaction process based on the federated learning framework as steps S201-S202, the model is optimized with hints, ultimately obtaining the target large language model M_SFT. In this embodiment, the input data is all multimodal data. For example, the input question may contain scanned copies of documents related to a certain business. Such multimodal input data is converted into Chinese text using the aforementioned multimodal data conversion toolset, which will not be elaborated further here.
[0056] In this step, after parameter aggregation is completed, the master compute node can send the aggregated parameters to each slave compute node, so that the slave compute nodes can perform the next round of model tuning based on the aggregated parameters until the predetermined model conditions are met. In this embodiment, the predetermined model tuning conditions can be that the number of tuning rounds is greater than a predetermined number of rounds or that the error corresponding to the current aggregated parameters is less than a predetermined error threshold.
[0057] Step S208: Based on the constructed first question-answer set, perform federated model training on the initial reward model to obtain the target reward model; In this step, after obtaining the target large language model, in order to make the final target large language model more accurate and reliable, a target reward model can be further trained so that subsequent federated reinforcement learning can be performed on the target large language model based on the target reward model.
[0058] In this step, the first question-and-answer set includes several partial order pairs of dialogues. Each partial order pair contains a sample question, several sample answers corresponding to the sample question, and the answer level of each sample answer. Each computing node (the master computing node and each slave computing node) prepares partial order pairs of dialogues within its respective business domain, which are sorted sequences of manually labeled question-and-answer samples, including the order of prompt words and corresponding answers.
[0059] Specifically, the master computing node prepares the security dataset, and each slave computing node prepares the availability dataset. The partially ordered questions in the datasets are designed by each node from its own business perspective. The answers are generated by the federated large language model, which has already undergone prompt optimization in step S207, after multiple calls. The generated answers are then manually sorted. The partially ordered dialogues in the availability dataset are mainly used to improve the accuracy of the large model's output. For example: [Question] What materials are needed to open an interbank bond market settlement account? [Optimal Answer] 1. China Central Depository & Clearing Co., Ltd. (CCDC) account business application form 2. CCDC market participant basic information form 3. Business license copy 4. Financial business license (if any) 5. Bond issuance, registration, and agency redemption service agreement signing page 6. ... [Inferior Answer] When preparing to open an interbank bond market settlement account, a series of related materials and documents are usually required to ensure the smooth opening of the account and subsequent trading activities. These required materials may vary depending on the requirements of different banks and regulatory agencies, and typically include legal person identification, legal representative certificate, financial reports, credit rating reports, trader information, account management agreements, other relevant documents, and other materials required by the bank.
[0060] In this step, the master computing node can train the initial reward model based on the first question-answer set to obtain the initial first reward parameters. Then, the master computing node can receive the initial second reward parameters sent from each computing node. Based on the initial first intermediate parameters and each initial second reward parameter, the master computing node aggregates the parameters to obtain the initial aggregated reward parameters and sends the initial aggregated reward parameters to each slave computing node so that the slave computing nodes can perform the next round of model training based on the initial aggregated reward parameters. The master computing node also receives the current second reward parameters sent from each slave computing node until the predetermined model reward stopping condition is met. Then, the current aggregated reward parameters are used as the target aggregated reward parameters to obtain the target reward model M_RM.
[0061] In this step, the loss function of the reward model is:
[0062] Where x is the input prompt (i.e., the question), and y is the answer. c Represents the sorting ratio y r The sentence that comes first, i.e., the sentence with the better response, is represented by m(r), which indicates the degree of difference in quality between the two sentences. There are four selectable values: 3, 2, 1, and 0. For example, for two responses y... I and y J y I Compared to y J In many cases, the reward model assigns scores to two responses that satisfy L(y). I )>L(y J), and at this time m(r)=3.
[0063] Step S209: Perform federated reinforcement learning on the target large language model based on the target reward model to obtain the reinforced target large language model; In this step, after training and obtaining the target reward model M_RM, the master compute node can score the answers output by the target large language model based on the target reward model, and then further strengthen the target large language model based on the score of the answer. Specifically, the master compute node uses the local training dataset to generate results / answers / responses based on the target large language model M_SFT, uses the target reward model M_RM to calculate the reward (i.e., score) of the answer / response, and adjusts the model parameters according to the reward to obtain the initial first strengthening parameters. At the same time, the master compute node receives the initial second strengthening parameters sent by each slave compute node, performs parameter aggregation based on the initial first strengthening parameters and each initial second strengthening parameter, obtains the initial aggregated strengthening parameters, and sends the initial aggregated strengthening parameters to each slave compute node so that the slave compute nodes can perform the next round of model strengthening based on the initial aggregated strengthening parameters. The master compute node also receives the current second strengthening parameters sent by each slave compute node until the predetermined model strengthening stopping condition is met, and then uses the current aggregated strengthening parameters as the target aggregated strengthening parameters to obtain the strengthened target large language model M_PPO.
[0064] In this embodiment, after training the target large language model M_PPO is completed, each computing node can save the trained M_PPO on each federated learning node, which can be directly used for inference. Therefore, business clients of each organization can directly input multimodal data through the front end developed by each node, convert it into Chinese text locally on the node using a multimodal data conversion toolset, and then call M_PPO to obtain the required output. At this time, there is no need for data interaction between nodes, thus protecting the data privacy of business clients from leakage.
[0065] In this embodiment, during each training round, the master computing node determines the first target hidden layer from the first large language model obtained from the training, and uses the parameters of each first target hidden layer as parameters to be tuned. Simultaneously, the master computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the master node can aggregate the parameters of each first target hidden layer and each second target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for subsequent precise model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then aggregating parameters based on the target hidden layer to tune it, the final tuning result is more accurate. At the same time, it avoids the huge computational power consumption and bandwidth occupation of full parameter updates in traditional methods, reducing the communication cost and local computational cost of each participant in the federated learning process.
[0066] Another embodiment of this application provides a model training method based on federated learning, which can be specifically applied to the various participants / slave computing nodes in federated learning, such as... Figure 4 As shown, the method in this embodiment includes the following steps: Step S301: Based on the initial second large language model obtained through training, determine several second target hidden layers and obtain the initial second sub-parameters corresponding to each second target hidden layer; In this step, the computing node can first perform model optimization on the pre-trained original large language model based on the second dialogue dataset to obtain an initial second large language model. The second dialogue dataset consists of high-quality dialogue data related to bond business, located locally on the computing node. Specifically, the second dialogue data contains several question-and-answer data points.
[0067] In this step, after the computing node obtains the initial second large language model through training, it can determine the global score of each hidden layer in the initial large language model. Then, based on the global score of each hidden layer, it can determine several second target hidden layers to be adjusted, which facilitates subsequent parameter aggregation processing based on the initial second sub-parameters of each second target hidden layer. That is, if the initial large language model contains, for example, 10 hidden layers, the method in this application can determine 3 hidden layers as the first target hidden layers from these 10 hidden layers, and then the parameters corresponding to each second target hidden layer in the initial second training parameters obtained during training are used as the initial second sub-parameters.
[0068] Step S302: Send each initial second sub-parameter to the master computing node so that the master computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain the initial aggregated parameters; In this step, after the slave computing nodes send the initial second sub-parameters corresponding to each second target hidden layer to the master computing node, the master computing node receives the initial second sub-parameters corresponding to each second target hidden layer sent by the slave computing nodes. Then, the master computing node performs parameter aggregation based on its local first initial sub-parameters and the initial second sub-parameters of each slave computing node to obtain the initial aggregated parameters. Specifically, when performing parameter aggregation, the master computing node can combine the initial parameters of each hidden layer with the initial first sub-parameters and initial second sub-parameters to obtain the initial aggregated parameters. In this step, after more than one training epoch, the tuning parameters of each hidden layer obtained from the previous model tuning epoch can also be used as initial parameters for parameter aggregation. That is, the current aggregated parameters are obtained based on the tuning parameters of each hidden layer obtained from the previous model tuning epoch, the current first sub-parameters, and the current second sub-parameters.
[0069] For example, a large language model has ten layers, labeled from 1 to 8, with initial parameters H1, H2, H3...H8 for each layer. Each computation node (master and slave nodes) selects several target hidden layers with the highest scores and labels them a, b... respectively.
[0070] For example, the main computing node a determines three first target hidden layers: a1, a2, and a3. The first target hidden layers a1, a2, and a3 correspond to the initial first sub-parameters A1, A2, and A3, respectively.
[0071] Three second target hidden layers are determined from the computing node b: b2, b3 and b4. The second target hidden layers b2, b3 and b4 correspond to the initial second sub-parameters B2, B3 and B4, respectively.
[0072] Three second target hidden layers are determined from the computing node c: c1, c3 and c5; the second target hidden layers c1, c3 and c5 correspond to the initial second sub-parameters C1, C3 and C5 respectively.
[0073] Therefore, after collecting the sub-parameters of all target hidden layers, the master compute node can obtain the optimized model parameters by performing an arithmetic average on the layers that need updating, labeled as M1 to M8. Thus, M1 = (H1 + A1 + C1) ÷ 3. M2 = (H2 + A2 + B2) ÷ 3. M3 = (H3 + A3 + B3 + C3) ÷ 4. M4 = (H4 + B4) ÷ 2. M5 = (H5 + C5) ÷ 2. M6 = H6. M7 = H7. M8 = H8.
[0074] Step S303: Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameter obtained by tuning to the master computing node until the predetermined tuning conditions are met. Stop the model tuning, and use the received current aggregation parameters as the target aggregation parameters to obtain the target large language model.
[0075] After receiving the initial aggregation parameters, the computing node can perform the next round of model tuning based on the initial aggregation parameters until the predetermined model conditions are met. In this embodiment, the predetermined model tuning conditions can be that the number of tuning rounds is greater than a predetermined number of rounds or that the error corresponding to the current aggregation parameters is less than a predetermined error threshold.
[0076] This embodiment presents a federated learning-based model training method. In each training round, the computing nodes determine the second target hidden layer from the second large language model obtained through training. The parameters of each second target hidden layer are sent to the master computing node as parameters to be tuned. The master computing node can then aggregate the parameters of each second target hidden layer with the locally determined parameters of the first target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for precise model training in the next round based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer and then aggregating parameters based on the target hidden layer for tuning, the final tuning result is more accurate. Simultaneously, it avoids the huge computational power consumption and bandwidth occupation of full parameter updates in traditional methods, reducing the communication cost and local computational cost of each participant in the federated learning process.
[0077] Based on the above embodiments, another embodiment of this application provides a model training method based on federated learning, which can be applied to the participants / from the computing nodes in federated learning. The method in this embodiment includes the following steps: Step S401: Train the predetermined base model based on the second original vocabulary to obtain the initial second intermediate parameters; In this step, when training the base model from the computing node, the specific steps include the following: Step S401-1, Installation and deployment of the multimodal data conversion toolset.
[0078] In this step, the lead / master computing node in the joint modeling process can pre-install and deploy a multimodal data transformation toolset, and simultaneously send the multimodal data transformation toolset to each slave computing node. Thus, each slave computing node can pre-install and deploy the multimodal data transformation toolset, and process data of various modalities based on the multimodal data transformation toolset to construct a second original vocabulary.
[0079] Step S401-2, Installation and initialization of the federated learning collaboration network.
[0080] In this step, the master compute node can send the federated learning components to each slave compute node, allowing each slave compute node to install the federated learning components. The framework of the federated learning collaborative network can be implemented using other software products with similar functionality, such as FATE, SecretFlow, and PaddleFL.
[0081] After the federated learning component is deployed, the slave compute nodes can receive the base large language model / base model initial parameters sent by the master compute node. In this embodiment, other large language models with the same functionality, such as Chinese-LLaMA and OpenChineseLLaMA, can be used as base models, and the vocabulary expansion is performed using the SentencePiece tool.
[0082] Step S401-3: Construct the second original vocabulary.
[0083] In this step, the compute node can process local multimodal bond data into Chinese text data / Chinese corpus training data using a deployed multimodal data conversion toolset. Then, it trains a second original vocabulary (chn_nodei.model) based on the Chinese text data / Chinese corpus and sends it to the master compute node. The leading party / master compute node for joint modeling then constructs the target vocabulary based on the first original vocabulary (chn_center.model) and each of the second original vocabulary (chn_nodei.model).
[0084] Step S401-4: Construct the second target topology.
[0085] In this step, the slave computing node can construct a second target topology corresponding to the slave computing node based on the second server corresponding to the slave computing node and the graphics processor deployed on each second server.
[0086] Step S401-5: Train the predetermined base model.
[0087] In this step, each slave computing node can use a ring-distributed learning approach and a second original vocabulary to train a predetermined base model for the second target topology.
[0088] Step S402: Send the initial second intermediate parameters to the master computing node so that the master computing node can perform parameter aggregation based on each initial second intermediate parameter and the initial first intermediate parameter of the master computing node to obtain the initial aggregated intermediate parameters. Step S403: Receive the initial aggregate intermediate parameters sent by the main computing node, perform the next round of model training based on the initial aggregate intermediate parameters, and send the current second intermediate parameters obtained from the training to the main computing node until the predetermined training conditions are met. Stop the model training, use the received current aggregate intermediate parameters as the target aggregate intermediate parameters, and obtain the original large language model.
[0089] In this step, after all slave computing nodes perform a batch of training locally, they send the intermediate parameters (i.e., matrices A and B) to the master computing node via GPRC. The master computing node reads the intermediate parameters from all slave computing nodes, aggregates them using the FedAvg algorithm, and then sends the aggregated parameters back to all slave computing nodes to start a new batch (next round) of training. This process is repeated until the model training stopping condition is met, resulting in the original large language model M. The stopping condition can be the completion of a predetermined batch / predetermined number of training rounds, or the acquisition of a satisfactory loss or the occurrence of certain errors. The predetermined batch / predetermined number of rounds and the loss condition are preset hyperparameters, while errors causing training to stop may include network communication problems such as timeouts between federated learning nodes or abnormal node states.
[0090] Step S404: Based on the second dialogue dataset, perform model optimization on the pre-trained original large language model to obtain the initial second large language model.
[0091] In this step, after obtaining the original large language model M, it can be optimized. The optimization process is based on high-quality dialogue datasets local to each computing node (i.e., the first dialogue dataset on the master computing node and the second dialogue dataset on each slave computing node). Specifically, after completing the federated training of the original large language model M, each node prepares its own high-quality dialogue dataset related to its bond business; the master computing node prepares the first dialogue dataset, and each slave computing node prepares the second dialogue dataset. These high-quality dialogue datasets contain questions that may be asked in various bond business scenarios, as well as the answers to those questions. In this embodiment, the data in the high-quality dialogue datasets are all multimodal data. For example, the input question may contain scanned copies of documents related to a certain business. Such multimodal input data is converted into Chinese text using the aforementioned multimodal data conversion toolset, which will not be elaborated further here.
[0092] Step S405: Based on the initial second large language model obtained through training, determine several second target hidden layers and obtain the initial second sub-parameters corresponding to each second target hidden layer; In the specific implementation process, this step involves determining the global score of each hidden layer in the initial second-largest language model, and then determining several second-target hidden layers to be adjusted based on the global scores of each hidden layer. Specifically, for the initial second-largest language model, the correlation matrix between the parameters of any two hidden layers can be determined; based on the eigenvalues of the correlation matrices corresponding to the same hidden layer, the global score of each hidden layer is determined; and several hidden layers with global scores greater than a predetermined score threshold are determined as second-target hidden layers, or alternatively, a predetermined number of hidden layers with the highest global scores can be used as second-target hidden layers.
[0093] In other words, during the feedback tuning process for each batch of data, the score of each layer of the large language model is dynamically calculated using the correlation matrix of the hidden states. Appropriate layers are selected for feedback tuning, avoiding the enormous computational and bandwidth consumption of full parameter updates in traditional methods. This reduces communication costs and local computational costs for each participant in the federated learning process, while simultaneously improving the accuracy of model training. Specifically, during each feedback tuning process, for any two hidden layers in the large model, their correlation can be represented by a matrix K composed of the cosine similarity between the parameters of the two hidden layers, and the eigenvalues λ of matrix K... Kij This represents the magnitude of the difference between two hidden layers i and j in this batch of data. Therefore, the summation of the eigenvalues of the correlation matrix K between hidden layer i and all other layers yields a score that represents the global score of hidden layer i; a higher score indicates a higher similarity between this layer and other layers. The global scores of all hidden layers are calculated, and a predetermined number of hidden layers with the highest scores are selected as the layers to be optimized in this iteration. The predetermined number is a hyperparameter that can be pre-determined based on the square root of the total number of hidden layers.
[0094] Step S406: Send each initial second sub-parameter to the master computing node so that the master computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain the initial aggregated parameters; In this step, after determining several second target hidden layers, each slave computing node can send the initial second sub-parameters corresponding to the second target hidden layers to the master computing node. Thus, the master computing node can perform parameter aggregation based on the initial parameters of each hidden layer, each initial first sub-parameter, and each initial second sub-parameter to obtain the initial aggregated parameters. (The previous round's parameters, initial first sub-parameters, and initial second sub-parameters are then used to perform parameter aggregation to obtain the initial aggregated parameters.)
[0095] Step S407: Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameter obtained by tuning to the master computing node until the predetermined tuning conditions are met. Stop the model tuning, and use the received current aggregation parameters as the target aggregation parameters to obtain the target large language model. In this embodiment, the predetermined tuning conditions can be that the number of tuning rounds is greater than the predetermined number of rounds or that the error corresponding to the current aggregation parameter is less than the predetermined error threshold.
[0096] Step S408: Based on the constructed second question-answer set, perform federated model training on the initial reward model to obtain the target reward model; In this step, after obtaining the target large language model, in order to make the final target large language model more accurate and reliable, the target reward model can be further trained from the computing node, so that the target large language model can be federated reinforcement learning based on the target reward model in the future.
[0097] In this step, the second question-and-answer set includes several partial order pairs of dialogues. Each partial order pair contains a sample question, several sample answers corresponding to the sample question, and the answer level of each sample answer. The process of constructing the second question-and-answer set from each computing node in this embodiment is similar to the process of constructing the first question-and-answer set from the main computing node in the above embodiment, and will not be described again here.
[0098] In this step, the slave computing nodes train the initial reward model based on the second question-answer set to obtain initial second reward parameters. Then, they send these initial second reward parameters to the master computing node. The master computing node then aggregates these parameters based on the initial second reward parameters sent from each computing node and the initial first reward parameters obtained through local training to obtain initial aggregated reward parameters, which are then sent to each slave computing node. Each slave computing node then performs the next round of model training based on these initial aggregated reward parameters, sending the current second reward parameters obtained during training to the master computing node. This process continues until predetermined training conditions are met, at which point model training stops, and the received current aggregated reward parameters are used as the target aggregated reward parameters to obtain the target reward model M_RM.
[0099] In this step, the loss function of the reward model is:
[0100] Where x is the input prompt (i.e., the question), and y is the answer. c Represents the sorting ratio y rThe sentence that comes first, i.e., the sentence with the better response, is represented by m(r), which indicates the degree of difference in quality between the two sentences. There are four selectable values: 3, 2, 1, and 0. For example, for two responses y... I and y J y I Compared to y J In many cases, the reward model assigns scores to two responses that satisfy L(y). I )>L(y J ), and at this time m(r)=3.
[0101] Step S409: Perform federated reinforcement learning on the target large language model based on the target reward model to obtain the reinforced target large language model; In this step, after training and obtaining the target reward model M_RM, the slave computing nodes can score the answers output by the target large language model based on the target reward model, and then further reinforce the target large language model based on the score of the answer. Specifically, the slave computing nodes use the local training dataset to generate results / answers / responses based on the target large language model M_SFT, use the target reward model M_RM to calculate the reward (i.e., score) of the answer / response, adjust the model parameters based on the reward, obtain the initial second reinforcement parameters, and send the initial second reinforcement parameters to the master computing node. The master computing node can then aggregate the parameters based on the initial second reinforcement parameters sent by each slave computing node and the initial first reinforcement parameters obtained through local training to obtain the initial aggregated reinforcement parameters, and send the initial aggregated reinforcement parameters to each slave computing node. Then, the computing node can receive the initial aggregation enhancement parameters sent by the master computing node, and perform the next round of model enhancement based on the initial aggregation enhancement parameters. The current second enhancement parameter obtained by model enhancement is sent to the master computing node until the predetermined enhancement conditions are met. Then, the model enhancement stops, and the received current aggregation enhancement parameters are used as the target aggregation enhancement parameters to obtain the enhanced target large language model M_PPO.
[0102] In this embodiment, after training the target large language model M_PPO is completed, each computing node can save the trained M_PPO on each federated learning node, which can be directly used for inference. Therefore, business clients of each organization can directly input multimodal data through the front end developed by each node, convert it into Chinese text locally on the node using a multimodal data conversion toolset, and then call M_PPO to obtain the required output. At this time, there is no need for data interaction between nodes, thus protecting the data privacy of business clients from leakage.
[0103] In this embodiment, during each training round, the computing nodes determine the second target hidden layer from the second large language model obtained from the training. The parameters of each second target hidden layer are then sent to the main computing node as parameters to be tuned. The main computing node can then aggregate the parameters of each second target hidden layer with the locally determined parameters of the first target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for precise model training in the next round based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer and then aggregating parameters based on the target hidden layer for tuning, the final tuning result is more accurate. Simultaneously, it avoids the huge computational cost and bandwidth consumption of full parameter updates in traditional methods, reducing communication costs and local computational costs for each participant in the federated learning process.
[0104] Another embodiment of this application provides a model training device based on federated learning, such as... Figure 5 As shown, it includes: The first determining module 11 is used to determine several first target hidden layers based on the initial first large language model obtained through training, and to obtain the initial first sub-parameters corresponding to each first target hidden layer. The first receiving module 12 is used to receive the initial second sub-parameters sent by each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein the initial second large language model is obtained by training the original large language model from the computing node; Aggregation module 13 is used to perform parameter aggregation based at least on each initial first sub-parameter and each initial second sub-parameter to obtain initial aggregated parameters; The first sending module 14 is used to send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters, and receive the current second sub-parameter sent by each slave computing node until the predetermined tuning conditions are met, then stop the model tuning, take the current aggregation parameters as the target aggregation parameters, and obtain the target large language model.
[0105] In this embodiment, the federated learning-based model training device further includes a first training module. The first training module is used to perform model optimization processing on the pre-trained original large language model based on the first dialogue dataset before determining a number of first target hidden layers, so as to obtain an initial first large language model.
[0106] In this embodiment, the first determining module is specifically used to: determine the correlation matrix between any two hidden layer parameters for the initial first large language model; determine the global score of each hidden layer based on the feature values of each correlation matrix corresponding to the same hidden layer; and determine several hidden layers with global scores greater than a predetermined score threshold as the first target hidden layers.
[0107] In this embodiment, the federated learning-based model training device further includes a first pre-training module, which is used to: pre-train the original large language model, specifically for: The base model is trained based on the target vocabulary to obtain the initial first intermediate parameters; wherein, the target vocabulary is obtained by merging the second original vocabulary from the computing node and the first original vocabulary local to the main computing node; Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, parameter aggregation is performed to obtain initial aggregated intermediate parameters. These initial aggregated intermediate parameters are then sent to the slave computing nodes so that they can perform the next round of model training based on them. The system also receives the current second intermediate parameters sent by each slave computing node until a predetermined model training stopping condition is met. Finally, the current aggregated intermediate parameters are used as the target aggregated intermediate parameters to obtain the original large language model.
[0108] In this embodiment, the federated learning-based model training device further includes a first construction module, which is used to: construct a first target topology structure corresponding to the main computing node based on the first server corresponding to the main computing node and the graphics processors deployed on each of the first servers; The first pre-training module is specifically used to: train a predetermined base model using a ring-distributed learning method and a target vocabulary for the first target topology.
[0109] In this embodiment, the federated learning-based model training device further includes a first reward model training module and a first reinforcement module; the first reward model training module is used to: perform federated model training on the initial reward model based on the constructed first question-answer set to obtain the target reward model; The first reinforcement module is used to: perform federated reinforcement learning on the target large language model based on the target reward model to obtain the reinforced target large language model.
[0110] In this embodiment, during each training round, the master computing node determines the first target hidden layer from the first large language model obtained through training, and uses the parameters of each first target hidden layer as parameters to be tuned. Simultaneously, the master computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the master node can aggregate the parameters of each first target hidden layer and each second target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for subsequent precise model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then aggregating parameters based on the target hidden layer to tune it, the final tuning result is more accurate. At the same time, it avoids the huge computational power consumption and bandwidth occupation of full parameter updates in traditional methods, reducing the communication cost and local computational cost of each participant in the federated learning process.
[0111] Another embodiment of this application provides a model training device based on federated learning, such as... Figure 6 As shown, it includes: The second determining module 21 is used to determine several second target hidden layers based on the initial second large language model obtained through training, and to obtain the initial second sub-parameters corresponding to each second target hidden layer. The second sending module 22 is used to send each initial second sub-parameter to the main computing node, so that the main computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain the initial aggregated parameters; The second receiving module 23 is used to receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the master computing node until the predetermined tuning conditions are met, then stop the model tuning, and use the received current aggregation parameters as the target aggregation parameters to obtain the target large language model.
[0112] In this embodiment, the second determining module is specifically used to: determine the correlation matrix between any two hidden layer parameters for the initial second large language model; determine the global score of each hidden layer based on the feature values of each correlation matrix corresponding to the same hidden layer; and determine several hidden layers with global scores greater than a predetermined score threshold as the second target hidden layers.
[0113] In this embodiment, the federated learning-based model training device further includes a second pre-training module. The second pre-training module is used to: pre-train the original large language model, specifically for: The model is trained on a predetermined base model based on the second original vocabulary to obtain initial second intermediate parameters. The initial second intermediate parameters are then sent to the main computing node, which aggregates the parameters based on the initial second intermediate parameters and the initial first intermediate parameters of the main computing node to obtain initial aggregated intermediate parameters. The model receives the initial aggregated intermediate parameters sent by the main computing node and performs the next round of model training based on the initial aggregated intermediate parameters. The current second intermediate parameters obtained from the training are then sent to the main computing node. When the predetermined training conditions are met, the model training is stopped, and the received current aggregated intermediate parameters are used as the target aggregated intermediate parameters to obtain the original large language model.
[0114] In this embodiment, the federated learning-based model training device further includes a second construction module; the second construction module is used to: construct a second target topology corresponding to the slave computing node based on the second server corresponding to the slave computing node and the graphics processor deployed on each second server; The second pre-training module is used to: train a predetermined base model using a ring-distributed learning approach and a second original vocabulary for the second target topology.
[0115] In this embodiment, the federated learning-based model training device further includes a second reward model training module and a second reinforcement module. The second reward model training module is used to: train the initial reward model using a federated model based on the constructed second question-and-answer set, and obtain the target reward model; The second reinforcement module is used to: perform federated reinforcement learning on the target large language model based on the target reward model to obtain the reinforced target large language model.
[0116] In this embodiment, during each training round, the computing nodes determine the second target hidden layer from the second large language model obtained through training, and send the parameters of each second target hidden layer as parameters to be tuned to the main computing node. The main computing node can then aggregate the parameters of each second target hidden layer with the parameters of the locally determined first target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for subsequent precise model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer and then aggregating parameters based on the target hidden layer for tuning, the final tuning result is more accurate. Simultaneously, it avoids the huge computational power consumption and bandwidth occupation of full parameter updates in traditional methods, reducing the communication cost and local computational cost of each participant in the federated learning process.
[0117] Another embodiment of this application provides an electronic device, such as... Figure 7 As shown, it includes at least a memory 1 and a processor 2. The memory 1 stores a computer program, and the processor 2 performs the following method steps when executing the computer program in the memory 1: Step 1: Based on the initial first large language model obtained through training, determine several first target hidden layers and obtain the initial first sub-parameters corresponding to each first target hidden layer; Step 2: Receive the initial second sub-parameters sent by each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; Step 3: Based on at least each initial first sub-parameter and each initial second sub-parameter, perform parameter aggregation to obtain initial aggregated parameters; Step 4: Send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. Receive the current second sub-parameter sent by each slave computing node until the predetermined tuning conditions are met. Stop the model tuning, use the current aggregation parameters as the target aggregation parameters, and obtain the target large language model.
[0118] Alternatively, implement the following steps: Step 1: Based on the initial second-largest language model obtained through training, determine several second-target hidden layers and obtain the initial second sub-parameters corresponding to each second-target hidden layer; Step 2: Send each initial second sub-parameter to the master compute node so that the master compute node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master compute node to obtain the initial aggregated parameters; Step 3: Receive the initial aggregation parameters sent by the master node, and perform the next round of model tuning based on the initial aggregation parameters. Send the current second sub-parameter obtained from the tuning to the master computing node until the predetermined tuning conditions are met. Stop the model tuning, and use the received current aggregation parameters as the target aggregation parameters to obtain the target large language model.
[0119] The specific implementation process of the above method steps can be found in any of the above embodiments of federated learning-based model training methods, and will not be repeated here.
[0120] In this embodiment, during each training round, the main computing node determines the first target hidden layer from the first large language model obtained from the training, and uses the parameters of each first target hidden layer as parameters to be tuned. Simultaneously, the main computing node also receives the parameters of each second target hidden layer sent by each slave computing node. Thus, the main node can aggregate the parameters of each first target hidden layer and each second target hidden layer to obtain aggregated parameters. This makes the aggregated parameters more reasonable and accurate, laying the foundation for subsequent precise model training based on the aggregated parameters. In this application, by determining the target hidden layer to be tuned from each hidden layer, and then aggregating parameters based on the target hidden layer to tune it, the final tuning result is more accurate. At the same time, it avoids the huge computational power consumption and bandwidth occupation of full parameter updates in traditional methods, reducing the communication cost and local computational cost of each participant in the federated learning process.
[0121] The above embodiments are merely exemplary embodiments of this application and are not intended to limit this application. The scope of protection of this application is defined by the claims. Those skilled in the art can make various modifications or equivalent substitutions to this application within its substance and scope of protection, and such modifications or equivalent substitutions should also be considered to fall within the scope of protection of this application.
Claims
1. A method for model training based on federated learning, applied to a master computing node, comprising: include: Based on the first dialogue dataset, the pre-trained original large language model is optimized to obtain the initial first large language model. The first dialogue dataset includes: several question and answer data related to bond business on the main computing node. Based on the initial first large language model obtained from training, determine several first target hidden layers, and obtain the initial first sub-parameters corresponding to each first target hidden layer; Receive the initial second sub-parameters sent from each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; Based at least on each initial first sub-parameter and each initial second sub-parameter, parameter aggregation is performed to obtain initial aggregated parameters; The initial aggregation parameters are sent to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. The current second sub-parameters sent by each slave computing node are received until the predetermined tuning conditions are met. Then the model tuning is stopped and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model. The master and slave compute nodes process Chinese text data. The initial first large language model obtained based on training determines several first target hidden layers, specifically including: For the initial first language model, determine the correlation matrix between any two hidden layer parameters; Based on the eigenvalues of the correlation matrices corresponding to the same hidden layer, the global score of each hidden layer is determined, and several first target hidden layers are determined based on the global scores of each hidden layer. Before performing model tuning on the pre-trained original large language model based on the first dialogue dataset, the method further includes: pre-training the original large language model, specifically including: The base model is trained based on the target vocabulary to obtain the initial first intermediate parameters; wherein, the target vocabulary is obtained by merging the second original vocabulary from the computing node and the first original vocabulary local to the main computing node; Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, parameter aggregation is performed to obtain initial aggregated intermediate parameters. These initial aggregated intermediate parameters are then sent to the slave computing nodes so that they can perform the next round of model training based on them. The system also receives the current second intermediate parameters sent by each slave computing node until a predetermined model training stopping condition is met. Finally, the current aggregated intermediate parameters are used as the target aggregated intermediate parameters to obtain the original large language model.
2. The method as described in claim 1, characterized in that, The determination of several first target hidden layers based on the global scores of each hidden layer specifically includes: Several hidden layers whose global scores are greater than a predetermined score threshold are identified as the first target hidden layers.
3. The method as described in claim 1, characterized in that, Before training the predetermined base model based on the target vocabulary, the method further includes: Based on the first server corresponding to the main computing node and the graphics processors deployed on each first server, construct the first target topology corresponding to the main computing node; The training of the predetermined base model based on the target vocabulary specifically includes: For the first target topology, a ring-based distributed learning approach is adopted, and the target vocabulary is used to train the predetermined base model.
4. The method as described in claim 1, characterized in that, After obtaining the target large language model, the method further includes: Based on the constructed first question-and-answer set, the initial reward model is trained using a federated model to obtain the target reward model; Federated reinforcement learning is applied to the target large language model based on the target reward model to obtain the reinforced target large language model.
5. A model training method based on federated learning, applied from computing nodes, characterized in that, include: Based on the second dialogue dataset, the pre-trained original large language model is optimized to obtain the initial second large language model. The second dialogue dataset includes: several question-and-answer data related to bond business from the local computing node. Based on the initial second-largest language model obtained through training, several second-target hidden layers are determined, and the initial second sub-parameters corresponding to each second-target hidden layer are obtained. Each initial second sub-parameter is sent to the master computing node so that the master computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the master computing node to obtain the initial aggregated parameters; Receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameter obtained by tuning to the master computing node until the predetermined tuning conditions are met. Then stop the model tuning, take the received current aggregation parameters as the target aggregation parameters, and obtain the target large language model. The master and slave compute nodes process Chinese text data. The initial second-largest language model obtained through training determines several second-target hidden layers, specifically including: For the initial second-largest language model, the correlation matrix between any two hidden layer parameters can be determined separately; Based on the eigenvalues of the correlation matrices corresponding to the same hidden layer, the global score of each hidden layer is determined, and several second target hidden layers are determined based on the global scores of each hidden layer. Before performing model tuning on the pre-trained original large language model based on the second dialogue dataset, the method further includes: pre-training the original large language model, specifically including: The pre-defined base model is trained based on the second original vocabulary to obtain the initial second intermediate parameters; The initial second intermediate parameters are sent to the master computing node so that the master computing node can perform parameter aggregation based on each initial second intermediate parameter and the initial first intermediate parameter of the master computing node to obtain the initial aggregated intermediate parameters. The system receives the initial aggregate intermediate parameters sent by the master computing node, performs the next round of model training based on the initial aggregate intermediate parameters, and sends the current second intermediate parameters obtained from the training to the master computing node. When the predetermined training conditions are met, the model training stops, and the received current aggregate intermediate parameters are used as the target aggregate intermediate parameters to obtain the original large language model.
6. A model training device based on federated learning, characterized in that, include: The first training module is used to perform model optimization on the pre-trained original large language model based on the first dialogue dataset to obtain the initial first large language model. The first dialogue dataset includes: several question and answer data related to bond business on the main computing node. The first determining module is used to determine several first target hidden layers based on the initial first large language model obtained through training, and to obtain the initial first sub-parameters corresponding to each first target hidden layer; The first receiving module is used to receive the initial second sub-parameters sent by each computing node, which correspond to the hidden layers of each second target in the initial second large language model; wherein, the initial second large language model is obtained by training the original large language model from the computing node; The aggregation module is used to aggregate parameters based on at least each initial first sub-parameter and each initial second sub-parameter to obtain initial aggregated parameters; The first sending module is used to send the initial aggregation parameters to each slave computing node so that each slave computing node can perform the next round of model tuning based on the initial aggregation parameters. It also receives the current second sub-parameter sent by each slave computing node until the predetermined tuning conditions are met, at which point the model tuning stops, and the current aggregation parameters are used as the target aggregation parameters to obtain the target large language model. The master and slave compute nodes process Chinese text data. The first determining module is specifically used to: for the initial first large language model, determine the correlation matrix between any two hidden layer parameters; Based on the eigenvalues of the correlation matrices corresponding to the same hidden layer, the global score of each hidden layer is determined, and several first target hidden layers are determined based on the global scores of each hidden layer. The federated learning-based model training device further includes: a first pre-training module, which is used for: The base model is trained based on the target vocabulary to obtain the initial first intermediate parameters; wherein, the target vocabulary is obtained by merging the second original vocabulary from the computing node and the first original vocabulary local to the main computing node; Based on the initial first intermediate parameters and the initial second intermediate parameters sent from the computing nodes, parameter aggregation is performed to obtain initial aggregated intermediate parameters. These initial aggregated intermediate parameters are then sent to the slave computing nodes so that they can perform the next round of model training based on them. The system also receives the current second intermediate parameters sent by each slave computing node until a predetermined model training stopping condition is met. Finally, the current aggregated intermediate parameters are used as the target aggregated intermediate parameters to obtain the original large language model.
7. A model training device based on federated learning, characterized in that, include: The second determination module is used to determine several second target hidden layers based on the initial second large language model obtained through training, and to obtain the initial second sub-parameters corresponding to each second target hidden layer. The initial second large language model is obtained by performing model optimization on the pre-trained original large language model based on the second dialogue dataset, wherein the second dialogue dataset includes: several question-and-answer data related to bond business from the local computing node. The second sending module is used to send each initial second sub-parameter to the main computing node, so that the main computing node can perform parameter aggregation based on each initial second sub-parameter and the initial first sub-parameter of the main computing node to obtain the initial aggregated parameters; The second receiving module is used to receive the initial aggregation parameters sent by the master node, perform the next round of model tuning based on the initial aggregation parameters, and send the current second sub-parameters obtained by tuning to the master computing node until the predetermined tuning conditions are met. Then, the model tuning is stopped, and the received current aggregation parameters are used as the target aggregation parameters to obtain the target large language model. The master and slave compute nodes process Chinese text data. The second determining module is specifically used for: For the initial second-largest language model, the correlation matrix between any two hidden layer parameters can be determined separately; Based on the eigenvalues of the correlation matrices corresponding to the same hidden layer, the global score of each hidden layer is determined, and several second target hidden layers are determined based on the global scores of each hidden layer. The federated learning-based model training device further includes: a second pre-training module, which is used for: The pre-defined base model is trained based on the second original vocabulary to obtain the initial second intermediate parameters; The initial second intermediate parameters are sent to the master computing node so that the master computing node can perform parameter aggregation based on each initial second intermediate parameter and the initial first intermediate parameter of the master computing node to obtain the initial aggregated intermediate parameters. The system receives the initial aggregate intermediate parameters sent by the master computing node, performs the next round of model training based on the initial aggregate intermediate parameters, and sends the current second intermediate parameters obtained from the training to the master computing node. When the predetermined training conditions are met, the model training stops, and the received current aggregate intermediate parameters are used as the target aggregate intermediate parameters to obtain the original large language model.
8. An electronic device, characterized in that, It includes at least a memory and a processor, wherein the memory stores a computer program, and the processor, when executing the computer program in the memory, implements the steps of the model training method based on federated learning as described in any one of claims 1-4 or 5.
Citation Information
Patent Citations
Model aggregation method, device and equipment, federal learning system and storage medium
CN117808125A
Internet of vehicles federal dynamic sparse training method based on cluster expansion
CN119906969A