Language model training method and device, storage medium and product
By collaboratively training the large language model and the small language model between the first device and the second device, using public text data and private text data to select appropriate text composition units, the problem of too many tokens or low importance in the federated size model collaborative fine-tuning method is solved, and the natural language processing capability of the large language model is improved.
Patent Information
- Application Number
- CN202510284828.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-11
- Publication Date
- 2025-06-27
AI Technical Summary
The existing federal size model collaborative fine-tuning method may lead to excessive tokens or low-importance tokens being introduced during sentence prediction, affecting the fine-tuning effect and reducing the natural language processing ability of large language models.
By deploying the large language model and text composition unit selection model in the first device, deploying the small language model in the second device, using public text data and private text data for collaborative training, optimizing the large language model and small language model, and selecting the appropriate text composition unit as the target object for the model's prediction probability distribution.
With limited computing resources in enterprises, the large language model can absorb the field knowledge of each enterprise's field, improve the natural language processing capabilities of the large language model, avoid the disadvantages of limited computing resources, and make full use of the computing resources of each device.
Smart Images

Figure CN120218178A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and particularly to a method, device, storage medium, and product for training a language model. Background Art
[0002] With the rapid development of large language model (LLM) technology, their powerful performance has brought significant productivity improvements in various industries. For small and medium-sized enterprises, due to limited computing power resources, they cannot meet the fine-tuning computing power requirements of large language models with particularly large scales, and can only run large language models with relatively limited model parameter scales (for distinction, also referred to as small language models); for large enterprises, since domain data is distributed in other enterprises, the open-source data sets they own cannot meet the fine-tuning of specific domains. Based on this, the industry usually adopts the method of collaborative fine-tuning of large and small models in the federation, enabling large enterprises to learn the domain data of other enterprises, and small and medium-sized enterprises to absorb the capabilities of large language models through federated learning, thereby improving the effects of their respective small language models.
[0003] However, in the current method of collaborative fine-tuning of large and small models in the federation, when performing sentence prediction, there may be a situation where there are too many tokens (text generation units) to be probabilistically predicted, or a situation where tokens with low importance are introduced into the collaborative fine-tuning of large and small models in the federation, affecting the fine-tuning effect, thereby reducing the natural language processing ability of large language models.
[0004] Therefore, how to improve the natural language processing ability of large language models is an urgent problem to be solved currently.
[0005] The above content is only used to assist in understanding the technical solution of this application, and does not represent an admission that the above content is prior art. Summary of the Invention
[0006] The main objective of this application is to provide a method, device, equipment, storage medium, and computer program product for training a language model, aiming to solve the technical problem of how to improve the natural language processing ability of large language models.
[0007] To achieve the above objective, this application proposes a method for training a language model, and the method includes:
[0008] Obtain public text data, input the public text data into the text composition unit selection model, and determine the first text composition unit corresponding to each of the multiple text composition unit positions in the public text data;
[0009] Send each of the first text composition units to each of the second devices, so that the second devices perform vocabulary mapping based on each of the first text composition units to obtain their respective corresponding second text composition units;
[0010] Combine each of the second devices to optimize the large language model and the small language model according to each of the first text composition units and each of the second text composition units, and obtain a trained large language model.
[0011] Optionally, the step of combining each of the second devices to optimize the large language model and the small language model according to each of the first text composition units and each of the second text composition units to obtain a trained large language model includes:
[0012] Receive the second prediction results respectively sent by each of the second devices, where the second prediction result is the result obtained by the second device training the small language model with local private text data and then inputting the public text data into the small language model to predict the probability distribution of each of the second text composition units;
[0013] Input the public text data into the large language model to predict the probability distribution of each of the first text composition units, and obtain a first prediction result;
[0014] Calculate a first prediction loss according to the first prediction result and each of the second prediction results, and optimize the large language model according to the first prediction loss;
[0015] Input the public text data into the optimized large language model to predict the probability distribution of each of the first text composition units, obtain a third prediction result, and send the third prediction result to each of the second devices respectively, so that the second devices calculate a second prediction loss according to the second prediction result and the third prediction result, and optimize the small language model according to the second prediction loss;
[0016] Return to execute the step of receiving the second prediction results respectively sent by each of the second devices, and obtain a trained large language model until a preset training end condition is met.
[0017] Optionally, the step of inputting the public text data into the text composition unit selection model to determine the first text composition unit corresponding to each of the multiple text composition unit positions in the public text data includes:
[0018] Input the public text data into the text component unit selection model to predict the probability distribution of the text component units corresponding to each of the multiple text component unit positions in the public text data, and for each text component unit position among the multiple text component unit positions, select multiple text component units whose probability distribution meets the preset first threshold condition as the first text component units, where the text component unit selection model is an isomorphic model of the large language model; or,
[0019] Input the public text data into the text component unit selection model to classify the text component units corresponding to each of the multiple text component unit positions in the public text data to obtain a classification result, and obtain the first text component units corresponding to each of the multiple text component unit positions in the public text data according to the classification result, where the text component unit selection model is a pre-trained multi-classification model, and the multi-classification model is trained according to a dataset obtained in advance, and the dataset is composed of the labeled text component units in the public text data; or,
[0020] Input the public text data into the text component unit selection model to predict the probability distribution of the text component units corresponding to each of the multiple text component unit positions in the public text data, and for each text component unit position among the multiple text component unit positions, select multiple text component units whose probability distribution meets the preset second threshold condition as the first text component units, where the text component unit selection model is a pre-trained causal language model, and the loss function of the causal language model is iteratively updated using the Top-K truncation method.
[0021] Optionally, the first prediction result includes a first evaluation result, the second prediction result includes a second evaluation result, the first evaluation result is the result obtained by comparing and evaluating the first prediction result with the standard text in the public text data by the first device, and the second evaluation result is the result obtained by comparing and evaluating the second prediction result with the standard text in the public text data by the second device;
[0022] The step of calculating a first prediction loss according to the first prediction result and each of the second prediction results and optimizing the large language model according to the first prediction loss includes:
[0023] Compare the first evaluation result and each of the second evaluation results respectively to determine the target language model corresponding to the optimal evaluation result from the large language model and each of the small language models;
[0024] Taking the target language model as the teacher model and the large language model as the student model, calculate the first prediction loss corresponding to the teacher model and the student model according to the first prediction result and the second prediction result, so as to optimize the large language model according to the first prediction loss.
[0025] In addition, to achieve the above object, the present application proposes another language model training method, which is applied to any one of a plurality of second devices, each of the second devices is respectively communicatively connected to a first device, a large language model to be trained and a preset text component selection model are deployed in the first device, and a small language model to be trained is respectively deployed in each of the second devices. The method includes:
[0026] Receiving a first text component corresponding to each position of a plurality of text components in the public text data sent by the large language model, wherein the first text component is obtained by the first device inputting the pre-acquired public text data into the text component selection model;
[0027] Performing vocabulary mapping according to each of the first text components to obtain a corresponding second text component;
[0028] Jointly with the first device, optimizing the large language model and the small language model according to each of the first text components and each of the second text components to obtain a trained small language model.
[0029] Optionally, the step of performing vocabulary mapping according to each of the first text components to obtain a corresponding second text component includes:
[0030] In the case where the small language model is heterogeneous from the large language model, mapping the first text component to a text component corresponding to the model structure of the small language model according to a preset text component alignment method to obtain the second text component.
[0031] In addition, to achieve the above object, the present application also proposes a language model training device, the device includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, the computer program is configured to implement the steps of the language model training method as described above.
[0032] In addition, to achieve the above object, the present application also proposes a storage medium, the storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium, and when the computer program is executed by a processor, the steps of the language model training method as described above are implemented.
[0033] In addition, to achieve the above object, the present application further provides a computer program product, the computer program product includes a computer program, and when the computer program is executed by a processor, the steps of the language model training method described above are implemented.
[0034] One or more technical solutions proposed by the present application have at least the following technical effects:
[0035] By deploying a large language model to be trained on a first device and a small language model to be trained on a second device, the first device uses public text data to collaborate with each second device to use public text data and their respective private text data to co-train the large language model and the small language model, realizing a co-training architecture for the large language model and the small language model, enabling small and medium-sized enterprises with limited computing power resources to also participate in the training of the large language model, and using the local private text data in their respective specific fields to enhance the natural language processing ability of the large language model; wherein, by inputting the public text data into a preset text component selection model, the first text component corresponding to each position of multiple text components in the public text data is determined, thereby realizing the selection of the text components for which probability distribution prediction is required, wherein, through the text component selection model; by sending each of the first text components to each of the second devices for the second devices to perform vocabulary mapping according to each of the first text components to obtain the corresponding second text components respectively, each of the second devices can obtain the second text components corresponding to the first text components, realizing the unification of the text components for which probability distribution prediction is required between the first device and each second device, ensuring that the first device and each second device predict the probability distribution of the text generation unit for the same object; by jointly optimizing the large language model and the small language model by each of the second devices according to each of the first text components and each of the second text components, a trained large language model is obtained; the text information in the specific fields corresponding to each enterprise learned by the small language model (which can be understood as a kind of "knowledge" acquired by the model) is transmitted to the large language model, so that the large language model can absorb the domain knowledge in the fields where each enterprise is located. At the same time, since this process is carried out in the first device, it is possible to avoid the shortcoming of limited computing power resources of most enterprises, and at the same time make full use of the computing resources of each second device to disperse the computational pressure of optimizing the large language model concentrated on one device; therefore, in this embodiment, by making full use of the computing resources of each device and the text data in the specific fields private to the corresponding enterprises, and by selecting appropriate text components as the target objects for model probability distribution prediction, the training of the model's natural language processing ability is realized, and under the condition of limited enterprise computing power resources, the large language model is enabled to absorb the domain knowledge in the fields where each enterprise is located to improve the natural language processing ability of the large language model. Brief Description of the Drawings
[0036] The accompanying drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0037] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a schematic flowchart provided for the first embodiment of the language model training method of the present application;
[0039] Figure 2 It is a schematic flowchart provided for the second embodiment of the language model training method of the present application;
[0040] Figure 3 It is a schematic flowchart provided for the third embodiment of the language model training method of the present application;
[0041] Figure 4 It is a schematic diagram of the device structure of the hardware operating environment involved in the language model training method in the embodiments of the present application.
[0042] The implementation, functional features and advantages of the purpose of the present application will be further described with reference to the embodiments and the accompanying drawings. Detailed implementation manners
[0043] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.
[0044] The following presents the first embodiment of the language model training method of the present application. Refer to Figure 1 , Figure 1Schematic flowchart of the first embodiment of the language model training method of the present application. In this embodiment, the language model training method is applied to a first device, and the first device is respectively communicatively connected to each second device. Both the first device and the second device can be a computing service device with data processing, network communication, and program running functions, such as a server, a tablet computer, a personal computer, a mobile phone, etc. The first device and the second device can be devices deployed in different enterprises, that is, each enterprise uses its own device to jointly train a language model to perform natural language processing tasks through the trained language model. In the specific implementation, specific natural language processing tasks can be set according to the specific needs of each enterprise. For example, in a dialogue scenario, a task of generating a reply to the above text can be set, and in a summary generation scenario, a task of generating a text summary for the text can be set. The number of second devices is at least two, and the operations performed or functions implemented by the first device and the second device during the training of the language model are different, and the operations performed or functions implemented by each second device are the same. Therefore, in the following embodiments, the two types of devices are distinguished as "first" and "second". In some application scenarios, the first device can also be referred to as a server or a coordination end, etc., and the second device can be referred to as a client or a participating end, etc.
[0045] Deploy the large language model to be trained in the first device, and deploy the small language model to be trained in each second device. A language model is a model that can be used to perform natural language processing tasks, such as LLaMa2 13B, LLaMa2 7B, OPT 7B, etc. The large language model and the small language model in this embodiment both belong to the language model. The structure of the large language model is more complex and has more parameters, while the model structure of the small language model is relatively simple and has fewer parameters. For each enterprise, the language models that can be deployed or supported may be heterogeneous or homogeneous, and this is not restricted in this embodiment. Homogeneous language models have similarities in internal structure and use the same neural network architecture, data processing flow, or training algorithm. Due to the similarity of structure and function, knowledge and technology between homogeneous language models can often be transferred and shared. However, heterogeneous language models refer to language models that differ in structure, function, or characteristics. These models can adopt different principles, algorithms, or frameworks when processing language data. In this embodiment, for example, in a feasible implementation, LLaMa2 13B can be deployed in the first device, and OPT 7B can be deployed in the second device. Due to the differences in structure and the number of parameters, the large language model after training is generally better than the small language model after training in terms of the ability to process natural language tasks. At the same time, however, the training process and the running process after training of the large language model require more computing resources than the small language model. Therefore, the computing resources of the first device and the second device used can be different. The computing resources of the first device can be better than those of the second device, that is, the computing power is stronger than that of the second device. The superiority or inferiority of the computing resources is specifically reflected by some parameters. For example, the number of CPU cores, clock frequency, memory bandwidth, and capacity, etc., and this is not restricted in this embodiment. In the specific implementation process, an enterprise with powerful computing power resources can provide the device as the first device, and an enterprise with text data in each subdivision field but limited computing power resources can provide the device as the second device. Each enterprise jointly conducts the training of the language model, and uses the advantages of both parties to jointly train to obtain a language model with higher natural language processing ability.For example, enterprises such as banks and insurance companies need to train language models to perform natural language processing tasks in their business scenarios to improve the effectiveness of business services. Although enterprises such as banks and insurance companies have text data in sub-domains such as the banking business domain and the insurance business domain, their computing power resources are limited. It is difficult for these enterprises to separately train large language models that require powerful computing power using their respective devices, making it difficult for them to successfully deploy language models that can improve business service effectiveness in business scenarios. Enterprises or institutions with powerful computing power lack text data in these sub-domains and are difficult to train language models with stronger natural language processing capabilities. Therefore, the devices of these enterprises can be jointly used to collaboratively train language models, leveraging the advantages of both parties to jointly train and obtain language models with higher natural language processing capabilities. For example, a large language model for customer service Q&A can be deployed in a first device, and a small language model for customer service Q&A can be deployed in a second device. The large language model for customer service Q&A and the small language model for customer service Q&A are used to reply to the above text given by the customer to implement the function of a robot customer service. By collaboratively training the large language model for customer service Q&A in the first device and the small language model for customer service Q&A in the second device, it is possible to jointly use the devices of these enterprises and leverage the advantages of both parties to jointly train and obtain a customer service Q&A language model with higher Q&A accuracy.
[0046] The model to be trained refers to that the model parameters in the model need to be optimized through the training process. Its initial model parameters can be obtained after model pre-training (or it can be said that the pre-trained model can be used as the model to be trained), or its initial model parameters can also be initialized according to experience or randomly. Using the pre-trained model as the model to be trained, the process of training the model to be trained on specific natural language processing tasks is also called "fine-tuning". In the following embodiments, the process is represented by "training". In the specific implementation, using a pre-trained language model as the language model to be trained and further training the language model to be trained can obtain better training effects and also enable the trained language model to have stronger natural language processing capabilities. In this embodiment, it is not limited to using a pre-trained language model as the language model to be trained, nor is it limited to specific pre-training methods.
[0047] Text data for training a language model is respectively deployed in the first device and the second device. For example, publicly available text data and local private text data. Generally speaking, there can be multiple pieces of text data for training the language model, and each sentence in the text data is composed of tokens. In addition, for each piece of text data in the text data, it can be divided into input text data as the model input and label text data as the model training label. For example, in a dialogue scenario, if it is necessary to train a language model to generate corresponding answers based on specific questions, then the input text data can be the questions, and the label text data can be the answers. Publicly available text data for training the language model (hereinafter referred to as "public text data") that has been made public in the industry can be deployed in the first device and the second device for jointly training large language models and small language models. Private text data of each enterprise (hereinafter referred to as private text data) is respectively deployed in each second device for local training of the small language model in the second device. The private text data of the enterprise can be business data from the enterprise. Since the specific business directions of each enterprise are different, the private text data of each enterprise belongs to different sub - fields. Each second device participates in the training of the language model through the private text data in its respective sub - field, which can provide knowledge in different sub - fields for the language model, so that the finally trained language model has stronger natural language processing capabilities. For example, the first device can be a device whose computing power can support the training and operation of large language models, provided by a few enterprises or institutions that can support this device; the public text data can be publicly available text data in the industry for training language models; each second device can be deployed in enterprises such as banks and insurance companies, and the private text data can be business data from each enterprise, such as dialogue business data when customer service staff in banks and insurance companies serve customers. These private text data belong to text data in sub - fields such as the banking business field and the insurance business field.
[0048] In this embodiment, the language model training method includes steps S10 to S30:
[0049] Step S10, obtain the public text data, input the public text data into the text composition unit selection model, and determine the first text composition unit corresponding to each of the multiple text composition unit positions in the public text data;
[0050] With the accelerated implementation of large language models in major enterprises, the requirements for large language models are also getting higher and higher. Large language models require high computing power and corpus data. However, many small and medium-sized enterprises have relatively limited computing resources (such as GPUs), and the corpus data of small and medium-sized enterprises also cannot meet the needs of large language model training. At this time, it is a common practice in the industry to select a small amount of corpus data in its own field (also called domain data) to fine-tune and adapt to the downstream tasks of its own enterprise. However, when performing fine-tuning, in the case of limited computing power of small and medium-sized enterprises, it cannot meet the direct fine-tuning computing power requirements of too large language models, and large language models will result in lower task performance in the inference stage. Therefore, it is more inclined to use small-scale small language models. For enterprises with large computing power, although they have the ability to fine-tune larger models, due to the distribution of domain data in major companies, the open-source dataset cannot meet the fine-tuning of specific domain tasks.
[0051] In this embodiment, by adopting the method of collaborative fine-tuning of large and small models, the computing resources of each device and the text data of the private sub-domains of the corresponding enterprises are fully utilized to realize the training of the model language, and it is realized that in the case of limited computing resources of enterprises, the large language model can absorb the domain knowledge of each enterprise's field.
[0052] Exemplarily, in the case where the first device is a server and each second device is a client, a large language model is deployed on the server with relatively strong computing power; small language models are deployed on each client with relatively weak computing power. The large model on the server cannot access the domain data of each client, but can learn the domain capabilities of each client through collaborative learning, while the client hopes to absorb the capabilities of the large model on the server through federated learning, so as to improve the effect of its own small model.
[0053] However, the existing federated large and small model collaborative fine-tuning algorithm based on open-source datasets will use the prediction probability of each text component unit (or the text component units with the top k largest probability distributions) in the vocabulary when predicting each sentence. The problems that may be caused by this approach include: predicting the probability distribution of all text component units in the vocabulary, which will result in too many text component units to be learned when doing knowledge transfer, leading to a decrease in the model effect; at the same time, it will also lead to an increase in data transmission volume; if the text component units with the top k highest are selected for probability distribution prediction, some important text component units may be omitted; whether all text component units or the text component units with the top k highest are selected, if unimportant text component units are also introduced for learning, it may be negative for the effect. It should be noted that text component units are the basic units that make up the text, which are artificially divided, and each text component unit divided based on a set of division rules constitutes a vocabulary.
[0054] Based on the above considerations, a preset text composition unit selection model is also deployed in the first device. Among them, the text composition unit selection model can be applicable to different task requirements. The public text data is input into the text composition unit selection model to determine the corresponding first text composition units for each of the multiple text composition unit positions in the public text data. The public text data is the open-source dataset, and the text unit composition positions can be determined through position identifiers. In order to maintain a high model quality and considering computing power, the training of the text composition unit selection model can be completed by the server. Considering that the server has sufficient computing power, a large amount of public data can be selected for training.
[0055] Step S20: Send each of the first text composition units to each of the second devices for the second devices to perform vocabulary mapping based on each of the first text composition units to obtain the corresponding second text composition units.
[0056] It should be noted that each of the second devices needs to form a mapping relationship with the first text composition units of the first device so that when performing collaborative fine-tuning of the federated large and small models, the text composition unit objects for which probability distribution prediction needs to be performed can be unified. Among them, the first text composition units and the second text composition units can be sentences, phrases, word sequences, or other appropriate text segments. Specifically, the first device (which may be a central server or a coordination node) sends the text composition units to multiple second devices (which may be edge devices, clients, or other nodes participating in federated learning). These second devices use these text composition units to perform vocabulary mapping to generate the corresponding second text composition units. The result of the mapping is the second text composition units. These units are semantically corresponding to the first text composition units but may be different in representation form. In addition, it should be noted that in federated learning, security and privacy are very important considerations. Therefore, appropriate encryption and privacy protection measures need to be taken when sending and receiving text composition units.
[0057] Step S30: Jointly optimize the large language model and the small language model by each of the second devices based on each of the first text composition units and each of the second text composition units to obtain a trained large language model.
[0058] After the first device generates the first text composition unit, the first device distributes the first text composition unit to each second device. The second device receives the first text composition unit and performs vocabulary mapping through a local vocabulary or a word embedding model to generate a second text composition unit. Each second device locally trains a small language model using the data it owns (including the second text composition unit and possibly other relevant data). The training process may include forward propagation to calculate the loss function, backward propagation to update the model parameters, and related steps such as model evaluation. After the local training is completed, each second device calculates the updated parameters of its model (such as weight updates or gradients), and makes predictions on the public dataset based on the updated parameters to obtain the prediction results of each second device on the public dataset. The first device will receive the prediction results of each second device on the public dataset. The first device selects the model with the best prediction effect among each device for each text unit position in the public dataset according to its own prediction results on the public dataset and the prediction results of each second device on the public dataset, and learns the probability distribution prediction of the model, so as to optimize the large language model, and repeats the above steps until the preset training end condition is met, and obtains the trained large language model.
[0059] In a feasible implementation manner, the step S10 of inputting the public text data into the text composition unit selection model to determine the first text composition unit corresponding to each of the multiple text composition unit positions in the public text data may include:
[0060] Input the public text data into the text composition unit selection model, predict the probability distribution of the text composition units corresponding to each of the multiple text composition unit positions in the public text data, and select multiple text composition units whose probability distribution meets the preset first threshold condition as the first text composition unit for each text composition unit position among the multiple text composition unit positions, where the text composition unit selection model is an isomorphic model of the large language model.
[0061] In this embodiment, first, a model isomorphic to the large language model needs to be selected. It should be noted that the large language model in this embodiment can be GPT-3, BERT, T5, or more advanced models such as GPT-J, GPT-NeoX, etc. That is, this embodiment does not limit the model for selecting text composition units to the large language model already deployed in the first device, and it can also be another isomorphic large language model. Then, leveraging the advantages of the large language model, set a prompt like the following: "Given the sentence...[mask]..., please fill in [mask] and give the top K most likely filling words." Here, [mask] is the mask for the position of the text composition unit to be predicted, and this prompt is used to clearly output the task to the large language model. The sentence here is for the sentence to be predicted. For each position (i.e., the position of the text composition unit) in this sentence, use [mask] to cover the word at this position, and then let the large model give the K words with the highest probability distribution values.
[0062] For example, for the sentence "I like to eat apples", convert it to "I [mask] eat [mask]". For each [mask] position, provide the top 3 most likely words.
[0063] In this embodiment, the model for selecting text composition units does not need to be trained. Just call any open-source large language model isomorphic to the large language model deployed in the first device. Since there is no need for training, the already open-source and widely verified model can be directly used, so it can be quickly deployed and integrated into the existing system. This significantly shortens the development cycle and improves the response speed of the project. In summary, by adopting a model for selecting text composition units that does not need to be trained and calling an open-source large language model isomorphic to the large language model deployed in the first device, technical effects such as rapid deployment, reduced resource consumption, guaranteed model quality, increased flexibility, lowered technical threshold, and rich community support can be achieved.
[0064] In another feasible embodiment, the step of inputting the publicly available text data into the model for selecting text composition units and determining the first text composition unit corresponding to each of the multiple text composition unit positions in the publicly available text data, step S10 may include:
[0065] Input the public text data into the text component unit selection model, classify the text components corresponding to each of the multiple text component positions in the public text data to obtain a classification result, and obtain the first text component corresponding to each of the multiple text component positions in the public text data according to the classification result. Among them, the text component unit selection model is a pre-trained multi-classification model, and the multi-classification model is trained according to a pre-obtained data set, and the data set is composed of the labeled text component units in the public text data.
[0066] In this embodiment, a model with a relatively small scale is selected for training, which is converted into a multi-classification problem, and the size of the classification category is the size of the model vocabulary that needs to use this text component unit selection model.
[0067] Specifically, some data can be manually labeled first. Each line of the data is a sentence. Randomly replace a certain position in the sentence with [mask], and then find the top K most likely words in the model vocabulary. The fitting target is a 0 / 1 vector (the size of the vocabulary), and the positions with 1 indicate that the word is one of the top K words.
[0068] In this embodiment, a relatively small model can be used for fine-tuning, or a non-linguistic model can be used for fine-tuning to obtain the text component unit selection model. Since the text component unit selection model is relatively small, it has the characteristics of high model performance ratio. At the same time, the relatively small model scale is also easy to be fine-tuned and optimized to adapt to different application scenarios and requirements. In addition, by training with a manually labeled data set, the model can learn the intentions and judgment criteria of the labelers, thereby improving the accuracy of classification.
[0069] In another feasible embodiment, the step of inputting the public text data into the text component unit selection model to determine the first text component corresponding to each of the multiple text component positions in the public text data, step S10 may include:
[0070] Input the public text data into the text component unit selection model, predict the probability distribution of the text components corresponding to each of the multiple text component positions in the public text data, and for each text component position among the multiple text component positions, select multiple text components whose probability distribution satisfies a preset second threshold condition as the first text component. Among them, the text component unit selection model is a pre-trained causal language model, and the loss function of the causal language model is iteratively updated using the Top-K truncation method.
[0071] In this embodiment, the CLM (Causal Language Modeling) can be selected as the text component unit selection model. Among them, the CLM model can generate a text sequence according to the given context. Specifically, the model attempts to predict the next word in the context where word prediction in a sentence needs to be performed. The context usually includes all words before the current word. When calculating the loss function for each position, only the K words with the highest probability distribution values participate in the calculation of the loss function, and other words do not participate.
[0072] The loss function can be used to calculate and represent; among them, Top K[i] is the indicator function, logits[i] is the model output array, and labels is the label for each text component unit position. That is, if the logits (model confidence) of the array token[i] corresponding to the text component unit belong to the largest K, it is 1, otherwise it is 0. By modifying the loss function, the model can be enabled to have the ability to select text component units.
[0073] In this embodiment, by deploying a large language model to be trained on a first device and a small language model to be trained on a second device, the first device uses public text data to collaborate with each second device to use public text data and their respective private text data to co-train the large language model and the small language model, realizing a co-training architecture for the large language model and the small language model, enabling small and medium-sized enterprises with limited computing power resources to also participate in the training of the large language model, and using their private local private text data in specific fields to enhance the natural language processing ability of the large language model; wherein, by inputting the public text data into a preset text component selection model, the first text component corresponding to each of the multiple text component positions in the public text data is determined, thereby realizing the selection of the text components for which probability distribution prediction is required, wherein, through the text component selection model; by sending each of the first text components to each of the second devices for the second devices to perform vocabulary mapping according to each of the first text components to obtain their respective corresponding second text components, each of the second devices can obtain the second text components corresponding to the first text components, realizing the unification of the text components for which probability distribution prediction is required between the first device and each second device, ensuring that the first device and each second device predict the probability distribution of the text generation unit for the same object; by jointly optimizing the large language model and the small language model by each of the second devices according to each of the first text components and each of the second text components, a trained large language model is obtained; the text information in the specific fields corresponding to each enterprise learned by the small language model (which can be understood as a kind of "knowledge" learned by the model) is transmitted to the large language model, so that the large language model can absorb the domain knowledge of each enterprise's field. At the same time, since this process is carried out on the first device, it is possible to avoid the shortcoming of limited computing power resources of most enterprises, and at the same time make full use of the computing resources of each second device to disperse the computational pressure of optimizing the large language model on one device; therefore, in this embodiment, by making full use of the computing resources of each device and the text data in the specific fields private to the corresponding enterprises, and by selecting appropriate text components as the target objects for model prediction probability distribution, the training of the model's natural language processing ability is realized, and under the condition of limited enterprise computing power resources, the large language model is enabled to absorb the domain knowledge of each enterprise's field to improve the natural language processing ability of the large language model.
[0074] Based on the first embodiment of the present application, a second embodiment of the language model training method of the present application is proposed. In the second embodiment of the present application, the same or similar content as in the above-mentioned first embodiment can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 2 , Figure 2This is a flowchart of the second embodiment of the language model training method of the present application. In this embodiment, the step of jointly optimizing the large language model and the small language model by each of the second devices according to each of the first text composition units and each of the second text composition units to obtain a trained large language model, step S30 includes:
[0075] Step S301, receiving second prediction results respectively sent by each of the second devices, where the second prediction results are results obtained by the second devices predicting the probability distributions of each of the second text composition units by inputting the public text data into the small language model after training the small language model with local private text data;
[0076] The first device and each of the second devices jointly perform multiple rounds of iterative training on the large language model to be trained and the small language model to be trained based on the computing power resources they each possess. Each round of training is carried out on the basis of the previous round of training. During a round of training, each of the second devices trains its own small language model using its own local private text data based on the computing power resources it possesses, optimizes the small language model, and then inputs the input text data in the public text data into the optimized small language model. After being processed by the small language model, the predicted results are output (hereinafter referred to as "second prediction results" for distinction). In this embodiment, the data form of the second prediction results is not limited. For example, the second prediction results can be the probability distribution corresponding to the text, that is, the probability distribution predicted by the small language model for predicting the public text data.
[0077] Since the second prediction results are only the prediction results of the small language model trained and optimized with private text data in the second device for the public text data, the second device does not need to send its private text data during the collaborative training with the first device, so it will not directly expose the private text data to the first device. Thus, while using the private text data in the second device to help optimize the large language model in the first device and enabling the large language model to learn the text information in each sub - field, the privacy and security of the local private text data in the second device are ensured. And because the model structure of the small language model is relatively simple and the number of parameters is relatively small, the computing power requirements for the second device are also relatively low. As a result, most enterprises with limited computing power resources can also participate in the training of the language model based on the text data in their respective sub - fields they own to obtain a language model with stronger natural language processing capabilities.
[0078] For example, each second device deployed in enterprises such as banks and insurance companies locally trains its respective small language model by using private text data from the dialogue service data of enterprises such as banks and insurance companies. After that, each small language model learns the text information in sub-domains such as the banking business domain and the insurance business domain, and sends the obtained second training result to the first device for training the large language model in the first device. This also requires a relatively low computing power for the second devices of each enterprise, so that the text data in sub-domains such as the banking business domain and the insurance business domain can participate in the training of the large language model, and a language model with stronger natural language processing ability can be obtained.
[0079] Step S302: Input the public text data into the large language model to predict the probability distribution of each of the first text components, and obtain a first prediction result.
[0080] Based on the computing power resources it has, the first device inputs the input text data in the public text data into the large language model, and through the processing of the large language model, outputs the predicted result (hereinafter referred to as the "first prediction result" for distinction). In this embodiment, the data form of the first prediction result is not limited. For example, the first prediction result may be the probability distribution of each word that may appear at a certain position to be filled in any sentence in the public text data.
[0081] It should be noted that the public text data deployed in the first device and the second device is the same, that is, the first device and the second device respectively use their respective language models to process the same text data to obtain prediction results. At the same time, since the large language model and the small language model are trained with different training data, obviously, for the same public text data, the prediction results of each language model may be different.
[0082] Step S303: Calculate a first prediction loss according to the first prediction result and each of the second prediction results, and optimize the large language model according to the first prediction loss.
[0083] After obtaining the first prediction result and the second prediction result, the computing power resources available can be utilized to perform knowledge distillation from the small language model to the large language model based on the first prediction result and the second prediction result. That is, the small language model is used as the teacher model and the large language model is used as the student model. It is possible to further optimize the large language model by using the small language model that has been locally optimized on the second device, and transfer the knowledge learned by the small language model in the sub-field to the large language model, so that the large language model can also learn the text information in the sub-field corresponding to each enterprise. The first prediction loss can be the prediction loss of the knowledge distillation from the small language model to the large language model. In addition, it should be noted that the method for calculating the first prediction loss is not limited in this embodiment. For example, the first prediction loss can be obtained by calculating the cross-entropy or KL (Kullback-Leibler) distance between the first prediction result and the second prediction result.
[0084] The first device optimizes the large language model according to the first prediction loss. Specifically, the model parameters in the large language model can be updated with the goal of reducing the first prediction loss. For example, the gradient value corresponding to the model parameters in the large language model can be calculated using the gradient descent algorithm according to the first prediction loss, and the model parameters in the large language model can be updated according to the gradient value, so as to achieve the purpose of optimizing the large language model.
[0085] The first prediction result and the second prediction result are the results obtained by the large language model and the small language model respectively for processing the same publicly available text data. The first prediction loss represents the difference between the first prediction result and the second prediction result. The greater the difference, the greater the first prediction loss. By reducing the first prediction loss, the difference between the prediction result output by the large language model and the prediction result output by the small language model can be reduced, so that the natural language processing ability of the large language model can be optimized.
[0086] Step S304: Input the publicly available text data into the optimized large language model to predict the probability distribution of each of the first text components, obtain a third prediction result, and send the third prediction result to each of the second devices respectively, so that the second device can calculate a second prediction loss according to the second prediction result and the third prediction result, and optimize the small language model according to the second prediction loss;
[0087] If the preset training end condition is satisfied, then step S305 is executed to obtain the trained large language model;
[0088] If the preset training end condition is not satisfied, then return to execute step S301.
[0089] After optimizing the large language model, the first device inputs the public text data into the optimized large language model. After being processed by the large language model, the predicted result is output (hereinafter referred to as the "third prediction result" for distinction). In this embodiment, the data form of the third prediction result is not limited. For example, the third prediction result can be the probability distribution corresponding to the text, that is, the probability distribution predicted by the optimized large language model for the public text data in this round of iteration.
[0090] The second device uses the second prediction result and the third prediction result to perform knowledge distillation from the large language model to the small language model. That is, the large language model is used as the teacher model and the small language model is used as the student model, which can enhance the natural language processing ability of the small language model through the large language model, enabling enterprises with limited computing resources to also obtain the small language model enhanced by the large language model. Finally, after multiple rounds of training, a small language model with stronger natural language processing ability is obtained.
[0091] After the second device calculates the second knowledge distillation loss, it can optimize the small language model according to the second knowledge distillation loss. Specifically, it can update the model parameters in the small language model with the goal of reducing the second knowledge distillation loss. For example, it can calculate the gradient value corresponding to the model parameters in the small language model using the gradient descent algorithm according to the second knowledge distillation loss, and update the model parameters in the small language model according to the gradient value, so as to achieve the purpose of optimizing the small language model.
[0092] After the first device sends the first training result to the second device, it can detect whether the preset training end condition is met. The preset training end condition is the condition that needs to be met for ending the training pre-trained, which can be set as needed and is not limited in this embodiment. For example, in the specific implementation, it can be set that the loss function of the large language model converges or the loss function of the small language model converges, or the number of training rounds reaches the set number of rounds, or the training duration reaches the set duration, etc.
[0093] If the first device detects that the preset training end condition is met, it can end the iterative training process and use the finally updated large language model as the large language model after training. Correspondingly, each second device also ends the iterative training process and uses the finally updated small language model as the small language model after training. The first device can use the trained large language model to complete the corresponding natural language processing tasks, and each second device can use the trained small language model to complete the corresponding natural language processing tasks.
[0094] If the first device detects that the preset training end condition is not met, it can enter the next round of the training process. That is, the second device performs the next round of local training based on the small language model after knowledge distillation in this round, obtains the second training result again based on the public text data, and returns it to the first device. The first device receives the second training result again and executes the subsequent training process. This loop iterates until it detects that the preset training end condition is met.
[0095] In a feasible implementation manner, the first prediction result includes a first evaluation result, the second prediction result includes a second evaluation result. The first evaluation result is the result obtained by the first device comparing and evaluating the first prediction result and the standard text in the public text data, and the second evaluation result is the result obtained by the second device comparing and evaluating the second prediction result and the standard text in the public text data.
[0096] The step of calculating the first prediction loss according to the first prediction result and each second prediction result and optimizing the large language model in step S303 may include steps S3031 to S3032:
[0097] Step S3031: Compare the first evaluation result and each second evaluation result respectively to determine the target language model corresponding to the optimal evaluation result from the large language model and each small language model.
[0098] After obtaining the first evaluation result and each second evaluation result, it is necessary to compare the first evaluation result with each second evaluation result. And in each round of iteration, the first device compares the prediction results of the large language model on the public text data and the prediction results of each small language model on the public text data, and selects the language model with the best model effect as the teacher model for model optimization. Among them, the first evaluation result refers to the result obtained by the first device comparing and evaluating the first prediction result and the standard text in the public text data, and each second evaluation result is the result obtained by the second device comparing and evaluating the second prediction result and the standard text in the public text data. The purpose of the comparison is to find out which model performs best among these models, so as to determine the target language model.
[0099] It should be noted that this embodiment does not limit the evaluation index corresponding to the comparison evaluation result. For example, the evaluation index can be accuracy, recall rate, etc. These indexes can quantify the performance of the model on the task.
[0100] Step S3032: Take the target language model as the teacher model and the large language model as the student model. Calculate the first prediction loss corresponding to the teacher model and the student model according to the first prediction result and the second prediction result, so as to optimize the large language model according to the first prediction loss.
[0101] After obtaining the target pre-language model, take the target language model as the teacher model and the large language model as the student model. The teacher model refers to the model with better performance, while the student model is the model to be optimized. According to the first prediction result and the second prediction result, the first prediction loss corresponding to the teacher model and the student model can be calculated, and the large language model can be optimized based on the first prediction loss.
[0102] It should be noted that this embodiment does not limit the specific implementation manner of optimizing the large language model based on the first prediction loss. For example, based on the prediction results of the large language model and each small language model on the public text data, the cross-entropy loss is calculated as the first prediction loss. Then, the Adam optimization algorithm is used to minimize the first prediction loss, thereby updating the parameters of the student model (i.e., the large language model). After multiple iterations, the performance of the large language model has been significantly improved.
[0103] In this embodiment, by deploying a large language model to be trained on a first device and a small language model to be trained on a second device, the first device uses public text data to cooperate with each second device to use public text data and its respective private text data to co-train the large language model and the small language model, realizing a co-training architecture for the large language model and the small language model, enabling small and medium-sized enterprises with limited computing power resources to also participate in the training of the large language model, and using the local private text data in their respective specific fields to enhance the natural language processing ability of the large language model; wherein, by inputting the public text data into a preset text component selection model, the first text component corresponding to each position of multiple text components in the public text data is determined, thereby realizing the selection of the text components for which probability distribution prediction is required, wherein, through the text component selection model; by sending each of the first text components to each of the second devices for the second devices to perform vocabulary mapping according to each of the first text components to obtain the corresponding second text components respectively, each of the second devices can obtain the second text components corresponding to the first text components, realizing the unification of the text components for which probability distribution prediction is required between the first device and each second device, ensuring that the first device and each second device predict the probability distribution of the text generation unit for the same object; by jointly optimizing the large language model and the small language model by each of the second devices according to each of the first text components and each of the second text components, a trained large language model is obtained; the text information in the specific fields corresponding to each enterprise learned by the small language model (which can be understood as a kind of "knowledge" learned by the model) is transmitted to the large language model, so that the large language model can absorb the domain knowledge of the fields where each enterprise is located. At the same time, since this process is carried out in the first device, it is possible to avoid the disadvantage of limited computing power resources of most enterprises, and at the same time make full use of the computing resources of each second device to disperse the computational pressure of optimizing the large language model concentrated on one device; therefore, in this embodiment, by making full use of the computing resources of each device and the text data in the specific fields private to the corresponding enterprises, and by selecting appropriate text components as the target objects for model prediction probability distribution, the training of the model's natural language processing ability is realized, and under the condition of limited enterprise computing power resources, the large language model is enabled to absorb the domain knowledge of the fields where each enterprise is located to improve the natural language processing ability of the large language model.
[0104] Based on the first and second embodiments of the present application, the third embodiment of the language model training method of the present application is proposed. In the third embodiment of the present application, the content that is the same as or similar to the above-mentioned first and second embodiments can be referred to the above introduction and will not be repeated hereinafter. On this basis, please refer to Figure 3 , Figure 3This is a schematic flowchart of the third embodiment of the language model training method of the present application. In this embodiment, the language model training method is applied to any one of multiple second devices, that is, the execution entity is any second device, and each of the second devices is communicatively connected to the first device. The first device deploys a large language model to be trained and a preset text composition unit selection model, and each of the second devices deploys a small language model to be trained. The method includes steps A10 to A30:
[0105] Step A10: Receive the first text composition units corresponding to the positions of multiple text composition units in the public text data sent by the large language model, where the first text composition units are obtained by the first device inputting the pre-acquired public text data into the text composition unit selection model;
[0106] Step A20: Perform vocabulary mapping according to each of the first text composition units to obtain the corresponding second text composition units;
[0107] Step A30: Jointly with the first device, optimize the large language model and the small language model according to each of the first text composition units and each of the second text composition units to obtain a trained small language model.
[0108] For the specific implementation manners of the above steps A10 to A30, reference may be made to the specific implementation manners of steps S10 to S30 in the above first embodiment, and the achieved effects may also be referred to the above first embodiment, which will not be elaborated here.
[0109] In a feasible implementation manner, the step of jointly with the first device, optimizing the large language model and the small language model according to each of the first text composition units and each of the second text composition units to obtain a trained small language model includes:
[0110] Step A301: Train the small language model using local private text data;
[0111] Step A302: Input the public text data into the small language model trained locally to predict the probability distribution of each of the second text composition units to obtain a second prediction result;
[0112] Step A303: Send the second prediction result to the first device. After the first device receives the second prediction results respectively sent by each of the second devices, input the public text data into the large language model to predict the probability distribution of each of the first text components, obtain a first prediction result, calculate a first prediction loss according to the first prediction result and each of the second prediction results, optimize the large language model according to the first prediction loss, and send a third prediction result to each of the second devices respectively.
[0113] Step A304: Receive the third prediction result sent by the first device, and calculate a second prediction loss according to the second prediction result and the third prediction result, where the third prediction result is obtained by the first device inputting the public text data into the optimized large language model to predict the probability distribution of each of the first text components.
[0114] Step A305: Optimize the small language model according to the second prediction loss, and return to execute the step of training the small language model using the local private text data until the preset training end condition is met, and then obtain a trained small language model.
[0115] For the specific implementation manners of the above steps A301 - A305, reference may be made to the specific implementation manners of steps S301 - S305 in the above first embodiment, and the achieved effects may also be referred to the above second embodiment, which will not be elaborated here.
[0116] In another feasible implementation manner, for the step of performing vocabulary mapping according to each of the first text components to obtain their respective corresponding second text components, step A20 may include:
[0117] In the case where the small language model and the large language model are heterogeneous, map the first text components to the text components corresponding to the model structure of the small language model according to a preset text component alignment method to obtain the second text components. For the sake of distinction, in this embodiment, the vocabulary of the large language model is referred to as the first vocabulary, and the vocabulary of the small language model is referred to as the second vocabulary.
[0118] Specifically, this embodiment provides two solutions for the text component alignment method, as follows:
[0119] Solution 1: The second device searches for the text component unit with the minimum edit distance from the first text component unit in the vocabulary corresponding to the small language model deployed by it, and determines it as the second text component unit. Among them, the minimum edit distance is an algorithm for measuring the similarity between two strings, and the minimum edit distance is obtained by calculating the minimum number of operations required to convert one string into another string. These operations include insertion, deletion, and substitution. Then, in this embodiment, the minimum edit distance between the text component unit in the first vocabulary and the text component unit in the second vocabulary is obtained by calculating the minimum number of operations required to convert the text component unit in the first vocabulary into the text component unit in the second vocabulary.
[0120] Solution 2: The second device uses the Embedding (embedding) technology to map the first text component unit into the vector space to obtain the vector form of the first text component unit. Then, the cosine similarity or Euclidean distance is calculated according to the mapped vector of the first text component unit. Thus, the vector with the highest similarity degree to the mapped vector of the first text component unit is found in the vector space according to the pre-similarity or Euclidean distance, and the corresponding text component unit is determined as the second text component unit.
[0121] This application provides a language model training device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the language model training method in the first embodiment above.
[0122] Next, refer to Figure 4 , which shows a schematic structural diagram of a language model training device suitable for implementing the embodiments of the present application. The language model training device in the embodiments of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 4 The language model training device shown is only an example and should not impose any limitations on the functions and usage scope of the embodiments of the present application.
[0123] As Figure 4As shown in the figure, the language model training device may include a processing device 1001 (such as a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM: Read Only Memory) 1002 or the program loaded from the storage device 1003 into the random access memory (RAM: Random Access Memory) 1004. In the RAM 1004, various programs and data required for the operation of the language model training device are also stored. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other through a bus 1005. The input / output (I / O) interface 1006 is also connected to the bus. Generally, the following systems can be connected to the I / O interface 1006: an input device 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD: Liquid Crystal Display), a speaker, a vibrator, etc.; a storage device 1003 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 1009. The communication device 1009 can allow the language model training device to communicate with other devices wirelessly or wiredly to exchange data. Although the figure shows a language model training device having various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be alternatively implemented or had.
[0124] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for executing the methods shown in the flowcharts. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device, or installed from the storage device 1003, or installed from the ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the methods of the embodiments disclosed in the present application are executed.
[0125] The language model training device provided by the present application adopts the language model training method in the above-mentioned embodiment. Compared with the prior art, the beneficial effects of the language model training device provided by the present application are the same as those of the language model training method provided by the above-mentioned embodiment, and other technical features in the language model training device are the same as the features disclosed in the method of the previous embodiment, and will not be elaborated here.
[0126] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in a suitable manner in any one or more embodiments or examples.
[0127] As described above, the above is only the specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed in this application, and all of them should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
[0128] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., computer programs) stored thereon, and the computer-readable program instructions are used to execute the language model training method in the above embodiments.
[0129] The computer-readable storage medium provided by this application can be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM) or flash memory, optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the above. In this embodiment, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or combined with an instruction execution system, device, or device. The program code contained on the computer-readable storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.
[0130] The above computer-readable storage medium can be included in the language model training device; it can also exist separately and not be assembled into the language model training device.
[0131] The above computer-readable storage medium carries one or more programs. When the above one or more programs are executed by the language model training device, the language model training device is caused to execute the above functions defined in the method of the disclosed embodiments of this application.
[0132] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The above-mentioned programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network (LAN) or a wide area network (WAN), or it can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).
[0133] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of the code, and this module, program segment, or part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks can occur in a different order than that marked in the accompanying drawings. For example, two consecutively represented blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, as well as the combination of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.
[0134] The modules described in the embodiments of this application can be implemented in software or in hardware. Among them, the name of the module does not constitute a limitation on the unit itself in some cases.
[0135] The readable storage medium provided by this application is a computer-readable storage medium, and the computer-readable storage medium stores computer-readable program instructions (i.e., computer programs) for performing the above-mentioned language model training method. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by this application are the same as those of the language model training method provided by the above embodiments, and will not be elaborated here.
[0136] The present application also provides a computer program product, including a computer program which, when executed by a processor, implements the steps of the language model training method as described above.
[0137] The computer program product provided by the present application can solve the technical problems of language model training. Compared with the prior art, the beneficial effects of the computer program product provided by the present application are the same as those of the language model training method provided in the above embodiments, and will not be elaborated herein.
[0138] The foregoing are only partial embodiments of the present application, and thus do not limit the patent scope of the present application. Any equivalent structural transformation made under the technical concept of the present application by using the content of the specification and drawings of the present application, or any direct / indirect application in other related technical fields, is included in the patent protection scope of the present application.
Claims
1. A language model training method, characterized in that: Applied to a first device, the first device is respectively connected to each second device for communication, a large language model to be trained and a preset text component unit selection model are deployed in the first device, and a small language model to be trained is respectively deployed in each second device, the method comprising: Acquire public text data, input the public text data into the text component unit selection model, and determine the first text component units corresponding to the positions of multiple text component units in the public text data; Sending each of the first text component units to each of the second devices, so that the second devices can perform vocabulary mapping according to each of the first text component units to obtain the second text component units corresponding to each of them; The second devices are combined to optimize the large language model and the small language model according to the first text component units and the second text component units to obtain a trained large language model.
2. The method according to claim 1, characterized in that The step of optimizing the large language model and the small language model according to the first text component units and the second text component units by combining the second devices to obtain the trained large language model comprises: Receiving second prediction results respectively sent by each of the second devices, wherein the second prediction results are obtained by the second device training the small language model with local private text data, inputting the public text data into the small language model to predict the probability distribution of each of the second text component units; Inputting the public text data into the large language model to predict the probability distribution of each of the first text component units to obtain a first prediction result; Calculating a first prediction loss according to the first prediction result and each of the second prediction results, and optimizing the large language model according to the first prediction loss; Input the public text data into the optimized large language model to predict the probability distribution of each of the first text component units to obtain a third prediction result, and send the third prediction result to each of the second devices respectively, so that the second device calculates a second prediction loss according to the second prediction result and the third prediction result, and optimizes the small language model according to the second prediction loss; Return to the step of receiving the second prediction results sent by each of the second devices respectively, until a preset training end condition is met, to obtain a trained large language model.
3. The method according to claim 1, characterized in that The step of inputting the public text data into the text component unit selection model to determine the first text component units corresponding to each of the plurality of text component unit positions in the public text data comprises: Inputting the public text data into the text component unit selection model, predicting the probability distribution of the text component units corresponding to the multiple text component unit positions in the public text data, so as to select multiple text component units whose probability distribution meets the preset first threshold condition as the first text component unit for each text component unit position in the multiple text component unit positions, wherein the text component unit selection model is an isomorphic model of the large language model; or Inputting the public text data into the text component unit selection model, classifying the text component units corresponding to the multiple text component unit positions in the public text data to obtain classification results, and obtaining the first text component units corresponding to the multiple text component unit positions in the public text data according to the classification results, wherein the text component unit selection model is a pre-trained multi-classification model, which is trained based on a pre-acquired data set, and the data set is composed of manually annotated label text component units in the public text data; or, The public text data is input into the text component unit selection model, and the probability distribution of the text component units corresponding to each of the multiple text component unit positions in the public text data is predicted, so as to select, for each of the multiple text component unit positions, multiple text component units whose probability distribution meets a preset second threshold condition as the first text component unit, wherein the text component unit selection model is a pre-trained causal language model, and the loss function of the causal language model is iteratively updated using the Top-K truncation method.
4. The method according to claim 2, characterized in that The first prediction result includes a first evaluation result, and the second prediction result includes a second evaluation result, wherein the first evaluation result is a result obtained by the first device comparing and evaluating the first prediction result with a standard text in the public text data, and the second evaluation result is a result obtained by the second device comparing and evaluating the second prediction result with the standard text in the public text data; The step of calculating a first prediction loss according to the first prediction result and each of the second prediction results, and optimizing the large language model according to the first prediction loss comprises: Respectively comparing the first evaluation result and each of the second evaluation results to determine a target language model corresponding to an optimal evaluation result from the large language model and each of the small language models; The target language model is used as a teacher model, and the large language model is used as a student model. According to the first prediction result and the second prediction result, the first prediction loss corresponding to the teacher model and the student model is calculated to optimize the large language model according to the first prediction loss.
5. A language model training method, characterized in that: Applied to any one of a plurality of second devices, each of the second devices is respectively connected to a first device for communication, a large language model to be trained and a preset text component unit selection model are deployed in the first device, and a small language model to be trained is respectively deployed in each of the second devices, the method comprising: Receiving first text component units corresponding to respective positions of a plurality of text component units in the public text data sent by the large language model, wherein the first text component units are obtained by the first device inputting pre-acquired public text data into the text component unit selection model; Perform vocabulary mapping according to each of the first text component units to obtain the second text component units corresponding to each of them; The first device is combined to optimize the large language model and the small language model according to each of the first text component units and each of the second text component units to obtain a trained small language model.
6. The method according to claim 5, characterized in that The step of optimizing the large language model and the small language model according to each of the first text component units and each of the second text component units by combining with the first device to obtain a trained small language model comprises: Using local private text data to train the small language model; Inputting the public text data into the locally trained small language model, predicting the probability distribution of each of the second text component units, and obtaining a second prediction result; sending the second prediction result to the first device, so that after receiving the second prediction results respectively sent by each of the second devices, the first device inputs the public text data into the large language model to predict the probability distribution of each of the first text component units to obtain a first prediction result, and calculates a first prediction loss according to the first prediction result and each of the second prediction results, optimizes the large language model according to the first prediction loss, and sends a third prediction result to each of the second devices respectively; receiving the third prediction result sent by the first device, and calculating a second prediction loss according to the second prediction result and the third prediction result, wherein the third prediction result is obtained by the first device inputting the public text data into the optimized large language model to predict the probability distribution of each of the first text component units; The small language model is optimized according to the second prediction loss, and the step of training the small language model using local private text data is returned to be executed until a preset training end condition is met, thereby obtaining a trained small language model.
7. The method according to claim 5, characterized in that The step of performing vocabulary mapping according to each of the first text component units to obtain the second text component units corresponding to each of the first text component units comprises: In the case where the small language model and the large language model are heterogeneous, the first text component unit is mapped to a text component unit corresponding to the model structure of the small language model according to a preset text component unit alignment method to obtain the second text component unit.
8. A language model training device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the language model training method according to any one of claims 1 to 7.
9. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the language model training method according to any one of claims 1 to 7 are implemented.
10. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the steps of the language model training method according to any one of claims 1 to 7 are implemented.