Training method and device of hybrid expert model, electronic equipment and storage medium
By converting the feedforward neural network of the large language model into multiple parallel expert units and using multi-domain and single-domain training samples to perform multi-level training on the gating layer, the problem of low training efficiency of existing large language models is solved, and more efficient training and better data processing performance is achieved.
Patent Information
- Application Number
- CN202510004762.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-02
- Publication Date
- 2025-05-09
Smart Images

Figure CN119962604A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large language models, and in particular to a training method, device, electronic equipment and storage medium for a hybrid expert model. Background Art
[0002] Pre-trained Large Language Models (PLMs) have made significant progress in the field of natural language processing. Through the self-supervised learning process of massive text data, they can not only grasp complex language structures, but also perform well in various downstream tasks. However, as the model capabilities are enhanced, the parameter scale of PLMs models is also increasing dramatically, which leads to two major problems: first, the high cost of training; second, deploying these large models in practical application scenarios has become increasingly challenging.
[0003] The Mixture of Experts (MoE) model can achieve a balance between model performance and computational efficiency. Compared with the traditional single large language model, the MoE architecture can significantly reduce the demand for computing resources while maintaining or even improving model accuracy.
[0004] However, training large-scale PLMs or mixed expert models from scratch is an extremely complex and expensive process, especially when there is no large amount of high-quality labeled data. In summary, the training efficiency of existing large language models is low. Summary of the invention
[0005] The present invention provides a training method, device, electronic device and storage medium for a hybrid expert model, which are used to solve the defect of low training efficiency of large language models in the prior art and improve the training efficiency of large language models.
[0006] The present invention provides a training method for a hybrid expert model, comprising: converting a feedforward neural network of a preset large language model PLM into multiple parallel expert units, adding a gating layer before the feedforward neural network, and obtaining a first preset hybrid expert model; training multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, and obtaining a second preset hybrid expert model, the multi-domain training samples are training samples matching all expert units, and one output probability represents the probability that the multi-domain training samples are sent to one expert unit; training multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, and obtaining a third preset hybrid expert model; training the third preset hybrid expert model based on single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, and obtaining a hybrid expert model, the single-domain training samples are training samples matching some expert units.
[0007] According to the training method of the hybrid expert model provided by the present invention, the feedforward neural network includes an input layer, an input fully connected layer and an output fully connected layer, and the feedforward neural network of the preset large language model PLM is converted into multiple parallel expert units, including: dividing the multiple neurons of the input fully connected layer into multiple parallel neuron groups; constructing different initial expert units based on different neuron groups, the initial expert unit includes neurons in the neuron group, and the initial expert unit connects the input layer and the output fully connected layer; determining the initial parameters of the corresponding initial expert unit based on the parameters of each neuron group, and obtaining multiple parallel expert units.
[0008] According to the training method of the hybrid expert model provided by the present invention, multiple output probabilities of the gating layer are trained based on multi-domain training samples until the multiple output probabilities are the same, and a second preset hybrid expert model is obtained, including: constructing a target loss function based on the multiple output probabilities, adding the target loss function to the initial loss function of the first preset hybrid expert model, and obtaining the loss function of the first preset hybrid expert model; fixing other parameters in the first preset hybrid expert model, training multiple output probabilities based on multi-domain training samples, when the change value of the loss function of the first preset hybrid expert model is less than the set change value, determining that the multiple output probabilities are the same, terminating the training, and obtaining the second preset hybrid expert model, the other parameters are parameters other than the output probabilities.
[0009] According to the training method of the hybrid expert model provided by the present invention, multiple output probabilities of the gated layer are trained based on multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, and a third preset hybrid expert model is obtained, including: removing the target loss function from the initial loss function of the second preset hybrid expert model to obtain the loss function of the second preset hybrid expert model; fixing other parameters in the second preset hybrid expert model, training multiple output probabilities based on multi-domain training samples until the loss function of the second preset hybrid expert model reaches a minimum loss value, determining that the output error of the second preset hybrid expert model reaches a first minimum value, ending the training, and obtaining the third preset hybrid expert model.
[0010] According to the training method of the hybrid expert model provided by the present invention, the third preset hybrid expert model is trained based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value to obtain the hybrid expert model, including: selecting a fixed number of expert units as target expert units of the third preset hybrid expert model based on the sorting results of multiple output probabilities; training, verifying and testing the third preset hybrid expert model based on the single-domain training samples, determining the output error of the third preset hybrid expert model based on the sample prediction data of the third preset hybrid expert model and the answer label of the single-domain training samples, the sample prediction data being obtained by weighted summation of the sample prediction value of each target expert unit and the output probability corresponding to the target expert unit; when the output error of the third preset hybrid expert model reaches the second minimum value, the training is terminated to obtain the hybrid expert model.
[0011] According to the training method of the hybrid expert model provided by the present invention, the third preset hybrid expert model is trained based on the single-domain training samples until the output error of the third preset hybrid expert model reaches the second minimum value. After the hybrid expert model is obtained, it also includes: inputting the single-domain text question into the hybrid expert model to obtain the target answer to the single-domain text question output by the hybrid expert model; wherein the target answer is obtained by weighted summation of the prediction data output by the target expert unit and the output probability corresponding to the target expert unit.
[0012] According to the training method of the hybrid expert model provided by the present invention, the multi-domain training samples are determined based on the following steps: obtaining single-domain training samples with labels based on the question text and answer labels of the single domain of the sample; mixing the single-domain training samples into the open source fine-tuning dataset to obtain multi-domain training samples with labels, the open source fine-tuning dataset including question texts and answer labels of multiple other domains, the other domains being domains different from the single domain of the sample.
[0013] The present invention also provides a training device for a hybrid expert model, comprising: a preset model determination module, used to convert a feedforward neural network of a preset large language model PLM into multiple parallel expert units, and add a gating layer before the feedforward neural network to obtain a first preset hybrid expert model; a first layer training module, used to train multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, and obtain a second preset hybrid expert model, the multi-domain training samples are training samples that match all expert units, and one output probability represents the probability that the multi-domain training samples are sent to one expert unit; a second layer training module, used to train multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, and obtain a third preset hybrid expert model; a third layer training module, used to train the third preset hybrid expert model based on single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, and obtain a hybrid expert model, the single-domain training samples are training samples that match some expert units.
[0014] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, a training method for any of the hybrid expert models described above is implemented.
[0015] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the training method of any of the hybrid expert models described above is implemented.
[0016] The training method, device, electronic device and storage medium of the hybrid expert model provided by the present invention improve the reasoning efficiency and deployment flexibility of the first preset hybrid expert model by converting the feedforward neural network into multiple expert units. Multiple output probabilities are trained to be the same through multi-field training samples, so that the second preset hybrid expert model can converge quickly. The output probabilities are retrained through multi-field training samples to further optimize the performance of the gating layer. The third preset hybrid expert model is trained through single-field training samples, so that the hybrid expert model achieves optimal performance in a specific field. The present invention converts the feedforward neural network into multiple expert units and then trains the gating layer at three different levels, which not only improves the training efficiency of the hybrid expert model, but also improves the data processing accuracy and robustness of the hybrid expert model in a specific field. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0018] Figure 1 It is a flow chart of the training method of the hybrid expert model provided by the present invention.
[0019] Figure 2 It is a partial structural diagram of the preset large language model provided by the present invention.
[0020] Figure 3 It is a structural schematic diagram of the hybrid expert model provided by the present invention.
[0021] Figure 4 It is a structural schematic diagram of the feedforward neural network and the expert unit provided by the present invention.
[0022] Figure 5 It is a partial structural schematic diagram of the Qwen2-7B instruction model provided by the present invention.
[0023] Figure 6 It is a structural schematic diagram of the training device of the hybrid expert model provided by the present invention.
[0024] Figure 7 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0025] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0026] Combine the following Figure 1-Figure 7 The present invention describes the training method, device and electronic device of the hybrid expert model.
[0027] Figure 1 It is a flow chart of the training method of the hybrid expert model provided by the present invention, such as Figure 1 As shown, the training method of the hybrid expert model includes S100 to S400, and each step is specifically as follows.
[0028] S100: Convert the feedforward neural network of the preset large language model PLM into multiple parallel expert units, add a gating layer before the feedforward neural network, and obtain a first preset hybrid expert model.
[0029] like Figure 2 As shown in the figure, PLM consists of stacked decoding layers, each of which contains an attention layer, a normalization layer (LayerNorm) and a feedforward neural network (FFN). FFN occupies a considerable proportion of model parameters and computing resource requirements in PLM, and is crucial to the optimization of FFN.
[0030] The first preset hybrid expert model is different from PLM in that FFN. Figure 3 As shown, in the first preset hybrid expert model, FFN is replaced by multiple parallel expert units, each of which is an independent FFN structure.
[0031] After converting the FFN layer of PLM to the MoE structure, a gating layer is inserted into the first preset hybrid expert model to determine how the input samples (multi-domain training samples or single-domain training samples) are routed to each expert unit for processing. The gating layer consists of a fully connected layer and an activation function, and its output dimension (the number of output probabilities) is equal to the number of expert units n.
[0032] S200: Training multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, thereby obtaining a second preset hybrid expert model.
[0033] The multi-domain training samples are training samples that match all expert units, and an output probability represents the probability of the multi-domain training samples being sent to an expert unit.
[0034] The multi-domain training samples are determined based on the following steps: obtaining single-domain training samples with labels based on the question text and answer labels of the single domain of the sample; mixing the single-domain training samples into the open source fine-tuning dataset to obtain multi-domain training samples with labels. The open source fine-tuning dataset includes question texts and answer labels in multiple other domains, and the other domains are different from the single domain of the sample.
[0035] It should be noted that the training method of the hybrid expert model proposed in the present invention is not limited to specific application scenarios, and can be widely used in multiple fields such as question answering, code generation, and reading comprehension.
[0036] The training process of the present invention in the smart water conservancy scenario is now provided. Based on the private knowledge base provided by the company in the smart water conservancy field, 20,000 high-quality data are annotated to obtain single-field training samples (smart water conservancy field training samples), involving water conservancy projects, water resources management, water environment and ecological protection, smart water conservancy technology and other aspects.
[0037] Each expert unit is used to process data in a specific field, for example, an expert unit in the field of water conservancy engineering, an expert unit in the field of cargo transportation, an expert unit in the field of environmental protection, etc. The field of the multi-field training sample involves the processing field of all expert units of the first preset hybrid expert model. The single-field training sample is a sample that matches the processing field of some expert units.
[0038] For example, a single-domain training sample is a smart water conservancy field training sample, which matches the expert unit in the water conservancy engineering field and the expert unit in the environmental protection field. Each smart water conservancy field training sample includes a question text in the smart water conservancy field (the question text in the sample single domain) and an answer label. For example, the smart water conservancy field training sample is as follows. Question text: How to measure river flow? Answer label: current meter method, buoy method, ultrasonic method and electromagnetic method.
[0039] The training samples in the field of smart water conservancy are mixed into open source fine-tuning datasets, such as the Fine-tuned Language Net (FLAN), WikiDialog and other open source fine-tuning datasets, to obtain multi-field training samples.
[0040] Ideally, when the output probability of each expert unit output by the gating layer is always the same, the output data of the original PLM is equivalent to the sum of the output data of all expert units in the first preset hybrid expert model after conversion. However, since the parameters (output probability) of the newly inserted gating layer are not trained, it cannot correctly output the probability of the expert unit being selected. In addition, the expert unit after direct conversion cannot distinguish between tasks in different fields and choose to activate or close. Therefore, in order to make the converted first preset hybrid expert model adapt to different input data, the first preset hybrid expert model needs to be fine-tuned and trained.
[0041] In the first layer training, only the multiple output probabilities of the gated layer are trained based on the multi-domain training samples until the multiple output probabilities are the same, which can ensure the basic stability and effectiveness of the obtained second preset hybrid expert model. During the training process, the other parameters of the first preset hybrid expert model are fixed.
[0042] S300: Training multiple output probabilities of the gating layer based on multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, thereby obtaining a third preset hybrid expert model.
[0043] In the second layer training, the multiple output probabilities of the gated layer are freely trained according to the multi-domain training samples, allowing the gated layer to freely adjust each of its output probabilities according to the multi-domain training samples (cancel the constraint of the same output probability), so that the obtained third preset hybrid expert model can adapt to a wider data distribution. During the training process, the other parameters of the second preset hybrid expert model are fixed.
[0044] S400: Training the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, thereby obtaining a hybrid expert model, wherein the single-domain training samples are training samples matching some expert units.
[0045] In the third layer of training, the fixed state of other parameters of the third preset hybrid expert model is released, and all parameters of the third preset hybrid expert model are fine-tuned using only single-domain training samples, so that the final hybrid expert model can better adapt to the data in a single domain (such as data in the field of smart water conservancy), thereby improving the task processing capabilities of the hybrid expert model in the field of smart water conservancy.
[0046] The hybrid expert model has significant advantages over the traditional large language model, especially in the reasoning stage. The hybrid expert model only needs to activate some experts for calculation, and these experts can be executed in parallel, which significantly improves the reasoning efficiency.
[0047] The training method of the hybrid expert model provided in the embodiment of the present invention improves the reasoning efficiency and deployment flexibility of the first preset hybrid expert model by converting the feedforward neural network into multiple expert units. Multiple output probabilities are trained to be the same through multi-domain training samples, so that the second preset hybrid expert model can converge quickly. The output probabilities are retrained through multi-domain training samples to further optimize the performance of the gating layer. The third preset hybrid expert model is trained through single-domain training samples, so that the hybrid expert model achieves optimal performance in a specific field. The present invention converts the feedforward neural network into multiple expert units and then trains the gating layer at three different levels, which not only improves the training efficiency of the hybrid expert model, but also improves the data processing accuracy and robustness of the hybrid expert model in a specific field.
[0048] Based on the above embodiment, the feedforward neural network includes an input layer, an input fully connected layer and an output fully connected layer, and the feedforward neural network of the preset large language model PLM is converted into multiple parallel expert units, including S110 to S130, and each step is specifically as follows.
[0049] S110: Divide the multiple neurons of the input fully connected layer into multiple parallel neuron groups.
[0050] S120: construct different initial expert units based on different neuron groups, where the initial expert units include neurons in the neuron groups, and the initial expert units connect the input layer and the output fully connected layer.
[0051] S130: Determine initial parameters of the corresponding initial expert unit based on the parameters of each neuron group to obtain multiple parallel expert units.
[0052] like Figure 4 As shown in Figure 1, the feedforward neural network (FNN) consists of an input layer, an input fully connected layer (FC1), a nonlinear activation layer (not shown in Figure 4 ) and the output fully connected layer (FC2).
[0053] Determine the number of expert units n. Divide the multiple neurons of the input fully connected layer into multiple parallel neuron groups. For example, if the FC1 layer has d neurons, the d neurons of the FC1 layer are evenly divided into n neuron groups, each of which includes d / n neurons. For each neuron group, construct an initial expert unit, which contains the neurons in the neuron group and connects the input layer and the FC2 layer. According to the grouping situation, use the parameters of the original FFN layer to initialize the initial parameters of each initial expert unit, ensuring that the initial parameters of each expert unit are consistent with the parameters of the corresponding part of the original FFN layer.
[0054] For example, the PLM model includes the Qwen2-7B instruction model. The Qwen2-7B instruction model has about 7 billion parameters and consists of 28 multi-head attention layers, where the hidden layer dimension size is 3584 and the FFN layer intermediate dimension size is 18944. The FFN layer structure is as follows Figure 5 As shown in the figure, the input fully connected layer FC1 and the output fully connected layer FC2 are used to increase the intermediate dimension 3584 to 18944 to increase the model's expressiveness. The output of the FC1 layer is then multiplied element-wise with the output of the FC2 layer through the SiLU activation function, and finally reduced back to 3584 through the fully connected layer FC3.
[0055] First, the parameters of the FFN layer are grouped. The neurons of the FC1 layer are evenly divided into 16 parallel neuron groups. Each neuron after grouping is connected to the FC2 layer. Each neuron group represents an expert unit.
[0056] Then a gating layer is inserted before the expert unit. The gating layer includes a fully connected layer and an activation function. The input dimension size is 3584, and the output dimension size is the number of expert units 16. The output is converted into output probability through the softmax function. The output probability is used to determine the contribution of each expert unit, that is, dynamically select the appropriate expert unit (target expert unit) to handle the task according to different inputs.
[0057] The present invention divides the input fully connected layer into multiple parallel neuron groups to obtain multiple parallel expert units, thereby ensuring that the first preset hybrid expert model has the accuracy of PLM and utilizing the advantage of high reasoning efficiency of the hybrid expert model, thereby improving the reasoning efficiency and deployment flexibility of the first preset hybrid expert model.
[0058] Based on the above embodiment, multiple output probabilities of the gating layer are trained based on multi-domain training samples until the multiple output probabilities are the same, and a second preset hybrid expert model is obtained, including S210 to S220, and each step is specifically as follows.
[0059] S210: constructing a target loss function based on multiple output probabilities, and adding the target loss function to the initial loss function of the first preset hybrid expert model to obtain the loss function of the first preset hybrid expert model.
[0060] S220: Fix other parameters in the first preset hybrid expert model, train multiple output probabilities based on multi-field training samples, and when the change value of the loss function of the first preset hybrid expert model is less than the set change value, determine that the multiple output probabilities are the same, end the training, and obtain the second preset hybrid expert model, where the other parameters are parameters other than the output probabilities.
[0061] In the first layer training, a target loss function is constructed based on multiple output probabilities. The target loss function is defined as the negative of entropy.
[0062] ; in, is the target loss function, is the number of the expert unit, For the The output probability corresponding to each expert unit.
[0063] During the training process, the target loss function is added to the initial loss function of the first preset hybrid expert model to obtain the loss function of the first preset hybrid expert model. The other parameters of the first preset hybrid expert model except the output probability of the gating layer are fixed, and only the output probability of the gating layer is trained to promote the rapid convergence of the first preset hybrid expert model.
[0064] Determine the training parameters during the training process. For example, the training parameters are a batch size of 32, 4096 tokens per data, an initial learning rate of 0.001, a cosine learning rate with a weight decay of 0.1, and training for 2000 steps.
[0065] Divide the multi-domain training samples into training set, validation set and test set according to the preset ratio. For example, the ratios of training set, validation set and test set are 70%, 15% and 15% respectively. Train the multiple output probabilities of the gated layer according to the training set, validation set and test set. When the change value of the loss function of the first preset hybrid expert model is less than the set change value, it means that the change of the target loss function tends to be stable, and the multiple output probabilities are the same, the training is terminated, and the second preset hybrid expert model is obtained.
[0066] The present invention realizes monitoring of output probability changes through the target loss function. By fixing other parameters in the first preset hybrid expert model and training the output probability separately, the convergence efficiency of the first preset hybrid expert model is improved.
[0067] Based on the above embodiment, multiple output probabilities of the gating layer are trained based on multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, and a third preset hybrid expert model is obtained, including S310 to S320, and each step is specifically as follows.
[0068] S310: removing the target loss function from the initial loss function of the second preset hybrid expert model to obtain the loss function of the second preset hybrid expert model.
[0069] S320: Fix other parameters in the second preset hybrid expert model, train multiple output probabilities based on multi-field training samples, until the loss function of the second preset hybrid expert model reaches the minimum loss value, determine that the output error of the second preset hybrid expert model reaches the first minimum value, end the training, and obtain the third preset hybrid expert model.
[0070] In the second layer training, the target loss function is removed from the initial loss function of the second preset hybrid expert model to obtain the loss function of the second preset hybrid expert model, so that the gating layer can freely adjust each output probability according to the multi-domain training samples.
[0071] The other parameters in the second preset hybrid expert model are fixed, and the multiple output probabilities are trained based on the multi-domain training samples until the loss function of the second preset hybrid expert model reaches a minimum loss value. The loss function of the second preset hybrid expert model is determined based on the difference between the sample prediction data of the second preset hybrid expert model and the answer label of the multi-domain training samples.
[0072] Furthermore, the sample prediction data of the second preset hybrid expert model is obtained by weighted summing of the sample prediction values of all expert units and the corresponding output probabilities.
[0073] ; in, The gated layer The output probability, For the The sample prediction value of expert units, and are the different output values of the fully connected layer of the gating layer, The sample prediction data for the second preset hybrid expert model.
[0074] When the output error of the second preset hybrid expert model reaches the first minimum value, it is determined that the training of the second preset hybrid expert model is completed, and a third preset hybrid expert model is obtained.
[0075] The present invention freely trains the output probability of the gating layer through multi-field training samples, so that the obtained third preset hybrid expert model can better adapt to the input data of various fields, further improving the performance of the gating layer.
[0076] Based on the above embodiment, the third preset hybrid expert model is trained based on the single-domain training samples until the output error of the third preset hybrid expert model reaches the second minimum value, and a hybrid expert model is obtained, including S410 to S430, and each step is specifically as follows.
[0077] S410: Selecting a fixed number of expert units as target expert units of a third preset hybrid expert model based on the sorting results of the multiple output probabilities.
[0078] S420: Train, verify and test the third preset hybrid expert model based on the single-domain training samples, determine the output error of the third preset hybrid expert model based on the sample prediction data of the third preset hybrid expert model and the answer label of the single-domain training samples, and the sample prediction data is obtained by the weighted sum of the sample prediction value of each target expert unit and the output probability corresponding to the target expert unit.
[0079] S430: When the output error of the third preset hybrid expert model reaches the second minimum value, the training is terminated to obtain the hybrid expert model.
[0080] In the third-level training, the fixed state of other parameters is canceled, and all parameters of the third preset hybrid expert model are trained with full-parameter fine-tuning.
[0081] The third preset hybrid expert model is trained, verified and tested based on the single-domain training samples. During the training process, the multiple output probabilities output by the gated layer are sorted, for example, in descending order of output probability, and the expert units corresponding to the top fixed number (for example, 4) of output probabilities are selected as target expert units. The sample prediction value of each target expert unit and the output probability corresponding to the target expert unit are weighted and summed to obtain the sample prediction data of the third preset hybrid expert model.
[0082] The output error of the third preset hybrid expert model is determined according to the sample prediction data of the third preset hybrid expert model and the answer label of the single-domain training sample.
[0083] When the output error of the third preset hybrid expert model reaches the second minimum value, the training is terminated to obtain the hybrid expert model.
[0084] The present invention performs weighted summation on the sample prediction values and corresponding output probabilities of the target expert units to obtain sample prediction data, thereby improving the accuracy of the hybrid expert model in processing data in a specific field and improving the precision and robustness of the hybrid expert model in tasks in a specific field.
[0085] Based on the above embodiment, the third preset hybrid expert model is trained based on the single-domain training samples until the output error of the third preset hybrid expert model reaches the second minimum value, and after the hybrid expert model is obtained, S500 is also included.
[0086] S500: Input a single-domain text question into the hybrid expert model to obtain a target answer to the single-domain text question output by the hybrid expert model.
[0087] The target answer is obtained based on the weighted sum of the prediction data output by the target expert unit and the output probability corresponding to the target expert unit.
[0088] After training the hybrid expert model, the hybrid expert model is applied. When the hybrid expert model is applied, the gating layer selects a fixed number of expert units with larger output probabilities as target expert units. The target expert units are activated, and each target expert unit performs task processing on a single-domain text question to obtain prediction data. The prediction data output by the target expert unit and the output probability corresponding to the target expert unit are weighted and summed to obtain the target answer.
[0089] In order to compare the benefits of the present invention in actual deployment scenarios, the original PLM was fine-tuned in all parameters in the field of smart water conservancy. Through comparative tests, when the fixed number (k value) of the hybrid expert model obtained by the present invention is set to 4, the accuracy of the hybrid expert model is close to 98% of the original PLM model, and the speed is increased by 1.4 times.
[0090] Furthermore, different k values can be set according to the actual deployment environment constraints to obtain a suitable balance.
[0091] The present invention obtains target answers based on multiple parallel target expert units of the hybrid expert model, thereby improving the reasoning efficiency and reasoning accuracy for single-domain text questions.
[0092] The training device of the hybrid expert model provided by the present invention is described below. The training device of the hybrid expert model described below and the training method of the hybrid expert model described above can refer to each other.
[0093] like Figure 6 As shown, a training device for a hybrid expert model includes: a preset model determination module 601, which is used to convert a feedforward neural network of a preset large language model PLM into multiple parallel expert units, and add a gating layer before the feedforward neural network to obtain a first preset hybrid expert model.
[0094] The first layer training module 602 is used to train multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, and obtain a second preset hybrid expert model. The multi-domain training samples are training samples that match all expert units, and one output probability represents the probability of the multi-domain training samples being sent to one expert unit.
[0095] The second layer training module 603 is used to train multiple output probabilities of the gated layer based on multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, thereby obtaining a third preset hybrid expert model.
[0096] The third layer training module 604 is used to train the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches the second minimum value, thereby obtaining a hybrid expert model. The single-domain training samples are training samples that match some expert units.
[0097] The training device for the hybrid expert model provided in the embodiment of the present invention improves the reasoning efficiency and deployment flexibility of the first preset hybrid expert model by converting the feedforward neural network into multiple expert units. Multiple output probabilities are trained to be the same through multi-domain training samples, so that the second preset hybrid expert model can converge quickly. The output probabilities are retrained through multi-domain training samples to further optimize the performance of the gating layer. The third preset hybrid expert model is trained through single-domain training samples, so that the hybrid expert model achieves optimal performance in a specific field. The present invention converts the feedforward neural network into multiple expert units and then trains the gating layer at three different levels, which not only improves the training efficiency of the hybrid expert model, but also improves the data processing accuracy and robustness of the hybrid expert model in a specific field.
[0098] In one embodiment, a feedforward neural network includes an input layer, an input fully connected layer, and an output fully connected layer, and the preset model determination module 601 is used to: divide multiple neurons of the input fully connected layer into multiple parallel neuron groups; construct different initial expert units based on different neuron groups, the initial expert unit includes neurons in the neuron group, and the initial expert unit connects the input layer and the output fully connected layer; determine the initial parameters of the corresponding initial expert unit based on the parameters of each neuron group to obtain multiple parallel expert units.
[0099] In one embodiment, the first layer training module 602 is used to: construct a target loss function based on multiple output probabilities, add the target loss function to the initial loss function of the first preset hybrid expert model, and obtain the loss function of the first preset hybrid expert model; fix other parameters in the first preset hybrid expert model, and train multiple output probabilities based on multi-field training samples. When the change value of the loss function of the first preset hybrid expert model is less than the set change value, it is determined that the multiple output probabilities are the same, and the training is terminated to obtain a second preset hybrid expert model, and the other parameters are parameters other than the output probabilities.
[0100] In one embodiment, the second layer training module 603 is used to: remove the target loss function from the initial loss function of the second preset hybrid expert model to obtain the loss function of the second preset hybrid expert model; fix other parameters in the second preset hybrid expert model, train multiple output probabilities based on multi-field training samples until the loss function of the second preset hybrid expert model reaches the minimum loss value, determine that the output error of the second preset hybrid expert model reaches the first minimum value, end the training, and obtain the third preset hybrid expert model.
[0101] In one embodiment, the third-layer training module 604 is used to: select a fixed number of expert units as target expert units of the third preset hybrid expert model based on the sorting results of multiple output probabilities; train, verify and test the third preset hybrid expert model based on the single-domain training samples, and determine the output error of the third preset hybrid expert model based on the sample prediction data of the third preset hybrid expert model and the answer labels of the single-domain training samples, wherein the sample prediction data is obtained by weighted summation of the sample prediction value of each target expert unit and the output probability corresponding to the target expert unit; when the output error of the third preset hybrid expert model reaches the second minimum value, the training is terminated to obtain the hybrid expert model.
[0102] In one embodiment, the training device of the hybrid expert model also includes an application module, which is used to: input a single-domain text question into the hybrid expert model to obtain a target answer to the single-domain text question output by the hybrid expert model; wherein the target answer is obtained based on the weighted sum of the prediction data output by the target expert unit and the output probability corresponding to the target expert unit.
[0103] In one embodiment, the first-layer training module 602 is used to: obtain single-domain training samples with labels based on the question text and answer labels of the single-domain of the sample; mix the single-domain training samples into the open source fine-tuning dataset to obtain multi-domain training samples with labels, and the open source fine-tuning dataset includes question texts and answer labels in multiple other domains, and the other domains are domains different from the single domain of the sample.
[0104] Figure 7 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 7As shown, the electronic device may include: a processor (processor) 710 , a communication interface (Communications Interface) 720 , a memory (memory) 730 and a communication bus 740 , wherein the processor 710 , the communication interface 720 , and the memory 730 communicate with each other through the communication bus 740 . The processor 710 can call the logic instructions in the memory 730 to execute the training method of the hybrid expert model, which includes: converting the feedforward neural network of the preset large language model PLM into multiple parallel expert units, adding a gating layer before the feedforward neural network, and obtaining a first preset hybrid expert model; training multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, and obtaining a second preset hybrid expert model, the multi-domain training samples are training samples that match all expert units, and one output probability represents the probability that the multi-domain training samples are sent to one expert unit; training multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, and obtaining a third preset hybrid expert model; training the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, and obtaining a hybrid expert model, the single-domain training samples are training samples that match some expert units.
[0105] In addition, the logic instructions in the above-mentioned memory 730 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present invention can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.
[0106] On the other hand, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which is implemented when the processor executes the training method of the hybrid expert model provided by the above methods, the method comprising: converting the feedforward neural network of the preset large language model PLM into multiple parallel expert units, adding a gating layer before the feedforward neural network, and obtaining a first preset hybrid expert model; training multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, and obtaining a second preset hybrid expert model, the multi-domain training samples are training samples that match all expert units, and one output probability represents the probability that the multi-domain training samples are sent to one expert unit; training multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, and obtaining a third preset hybrid expert model; training the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, and obtaining a hybrid expert model, the single-domain training samples are training samples that match some expert units.
[0107] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.
[0108] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0109] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A training method for a hybrid expert model, characterized in that: include: Converting a feedforward neural network of a preset large language model PLM into multiple parallel expert units, adding a gating layer before the feedforward neural network to obtain a first preset hybrid expert model; Training multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, thereby obtaining a second preset hybrid expert model, wherein the multi-domain training samples are training samples that match all the expert units, and one output probability represents the probability that the multi-domain training sample is sent to one of the expert units; Training the multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, thereby obtaining a third preset hybrid expert model; The third preset hybrid expert model is trained based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, thereby obtaining a hybrid expert model, wherein the single-domain training samples are training samples matching some expert units.
2. The training method of the hybrid expert model according to claim 1, characterized in that: The feedforward neural network includes an input layer, an input fully connected layer and an output fully connected layer. The feedforward neural network that converts the preset large language model PLM into multiple parallel expert units includes: Dividing the multiple neurons of the input fully connected layer into multiple parallel neuron groups; Constructing different initial expert units based on different neuron groups, wherein the initial expert units include neurons in the neuron groups, and the initial expert units connect the input layer and the output fully connected layer; Based on the parameters of each neuron group, the initial parameters of the corresponding initial expert unit are determined to obtain a plurality of parallel expert units.
3. The training method of the hybrid expert model according to claim 1, characterized in that: The method of training the multiple output probabilities of the gating layer based on the multi-domain training samples until the multiple output probabilities are the same to obtain a second preset hybrid expert model includes: constructing a target loss function based on the multiple output probabilities, and adding the target loss function to the initial loss function of the first preset hybrid expert model to obtain the loss function of the first preset hybrid expert model; Fix other parameters in the first preset hybrid expert model, train the multiple output probabilities based on the multi-field training samples, and when the change value of the loss function of the first preset hybrid expert model is less than the set change value, determine that the multiple output probabilities are the same, end the training, and obtain the second preset hybrid expert model, where the other parameters are parameters other than the output probabilities.
4. The training method of the hybrid expert model according to claim 3, characterized in that: The method of training the multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value to obtain a third preset hybrid expert model comprises: Removing the target loss function from the initial loss function of the second preset hybrid expert model to obtain a loss function of the second preset hybrid expert model; Fix other parameters in the second preset hybrid expert model, train the multiple output probabilities based on the multi-field training samples until the loss function of the second preset hybrid expert model reaches the minimum loss value, determine that the output error of the second preset hybrid expert model reaches the first minimum value, end the training, and obtain the third preset hybrid expert model.
5. The training method of the hybrid expert model according to claim 3, characterized in that: The step of training the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value to obtain the hybrid expert model comprises: Selecting a fixed number of expert units as target expert units of the third preset hybrid expert model based on the sorting results of the plurality of output probabilities; The third preset hybrid expert model is trained, verified and tested based on the single-domain training samples, and the output error of the third preset hybrid expert model is determined based on the sample prediction data of the third preset hybrid expert model and the answer label of the single-domain training samples, wherein the sample prediction data is obtained based on the weighted sum of the sample prediction value of each target expert unit and the output probability corresponding to the target expert unit; When the output error of the third preset hybrid expert model reaches the second minimum value, the training is terminated to obtain the hybrid expert model.
6. The training method of the hybrid expert model according to claim 5, characterized in that: The method further includes: training the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches a second minimum value, and obtaining the hybrid expert model; Inputting a single-domain text question into the hybrid expert model, and obtaining a target answer to the single-domain text question output by the hybrid expert model; The target answer is obtained based on the weighted sum of the prediction data output by the target expert unit and the output probability corresponding to the target expert unit.
7. The training method of the hybrid expert model according to claim 1, characterized in that: The multi-domain training samples are determined based on the following steps: Acquire the single-domain training sample with labels based on the question text and answer labels of the single-domain sample; The single-domain training samples are mixed into an open-source fine-tuning dataset to obtain the multi-domain training samples with labels, wherein the open-source fine-tuning dataset includes question texts and answer labels in multiple other domains, and the other domains are domains different from the single domain of the samples.
8. A training device for a hybrid expert model, characterized in that: include: A preset model determination module, used to convert a feedforward neural network of a preset large language model PLM into multiple parallel expert units, and add a gating layer before the feedforward neural network to obtain a first preset hybrid expert model; A first layer training module is used to train multiple output probabilities of the gating layer based on multi-domain training samples until the multiple output probabilities are the same, thereby obtaining a second preset hybrid expert model, wherein the multi-domain training samples are training samples that match all the expert units, and one output probability represents the probability that the multi-domain training sample is sent to one of the expert units; A second layer training module is used to train the multiple output probabilities of the gating layer based on the multi-domain training samples until the output error of the second preset hybrid expert model reaches a first minimum value, thereby obtaining a third preset hybrid expert model; The third-layer training module is used to train the third preset hybrid expert model based on the single-domain training samples until the output error of the third preset hybrid expert model reaches the second minimum value to obtain the hybrid expert model, and the single-domain training samples are training samples that match some expert units.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the training method of the hybrid expert model according to any one of claims 1 to 7 is implemented.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the training method of the hybrid expert model as claimed in any one of claims 1 to 7 is implemented.
Citation Information
Cited By
Multi-expert mechanism-based chapter structure analysis system
CN120180242A