Language model training method and device, text processing method and device, equipment and medium

By sparsely compressing the probability matrix and dynamic vocabulary mapping of large language models, the problems of high storage cost and low efficiency of distillation training in the existing technology are solved, and efficient language model training and deployment are achieved.

CN120146200AActive Publication Date: 2025-06-13IFLYTEK CO LTD

Patent Information

Application Number
CN202510617411.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-14
Publication Date
2025-06-13
Estimated Expiration
2045-05-14

AI Technical Summary

Technical Problem

The distillation training methods of existing large language models are difficult to deploy on resource-constrained devices due to their high storage costs and low training efficiency.

Method used

By predicting the first probability matrix of each data unit in the sample text based on the teacher model and sparsely compressing according to the probability value size, the second probability matrix is ​​obtained. Then, the word list of the student model is aligned using the word elements corresponding to each probability value in the second probability matrix to obtain the third word list. Finally, the student model is distilled and trained based on the third vocabulary list and the second probability matrix to obtain the target language model.

Benefits of technology

It significantly reduces the problem of excessive storage costs caused by the growth of the vocabulary scale, improves the efficiency of distillation training, and enables the target language model to better adapt to different model architectures and text processing scenarios while maintaining high performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120146200A_ABST
    Figure CN120146200A_ABST
Patent Text Reader

Abstract

The invention provides a language model training method and device, a text processing method and device, equipment and a medium, and relates to the technical field of natural language processing, and the method comprises the steps: predicting a first probability matrix corresponding to each data unit in a sample text based on a teacher model; the first probability matrix comprises a probability value of each data unit belonging to each lexical element in the first word list; according to the numerical value of each probability value in the first probability matrix, compressing the first probability matrix to obtain a second probability matrix corresponding to each data unit; performing alignment operation on the second word list according to lexical elements corresponding to probability values in the second probability matrix to obtain a third word list; and according to the third word list and the second probability matrix, carrying out distillation training on the student model to obtain a target language model, thereby reducing the storage cost, improving the distillation training efficiency, and enabling the target language model trained according to the method to better adapt to different model architectures and text processing scenes while keeping high performance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular, to a method for training a language model, a method for text processing, an apparatus, a device, and a medium. Background Art

[0002] Large language models have demonstrated excellent performance and broad application prospects in text processing. However, these large language models usually have a huge parameter scale and complex structures, resulting in problems such as high computational costs, making it difficult to directly deploy them on resource-constrained devices. Therefore, related technologies use distillation training to transfer the capabilities of large language models to lightweight student models, so that the student models trained accordingly can not only learn how to judge the categories of correct samples from labeled data at a low computational cost, but also learn the relationships between classes from the teacher model.

[0003] In the existing distillation training method for large language models, a teacher model constructed using a large language model is usually used to generate full soft labels for all tokens (data units) in all training data in advance, and the full soft labels of each token in each training data under the vocabulary of the teacher model and the hard labels are saved locally together for direct loading during the training of the student model; however, with the explosive growth of the vocabulary size, the storage quantity of the soft labels of each token in a single training data also increases accordingly, resulting in too high storage costs and reduced model training efficiency. Summary of the Invention

[0004] The present invention provides a method for training a language model, a method for text processing, an apparatus, a device, and a medium to solve the defects in the prior art.

[0005] The present invention provides a method for training a language model, including: Based on a teacher model, predicting a first probability matrix corresponding to each data unit in a sample text; the first probability matrix includes the probability values of each data unit belonging to each token in a first vocabulary, and the first vocabulary is the vocabulary of the teacher model; Compressing the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain a second probability matrix corresponding to each data unit; Aligning a second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of the student model; Performing distillation training on the student model according to the third vocabulary and the second probability matrix to obtain a target language model.

[0006] A language model training method provided by the present invention, the method of compressing the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain a second probability matrix corresponding to each data unit includes: Receiving user input information, and determining a target quantity according to the compression mode in the user input information; Sorting the probability values in the first probability matrix in descending order according to the numerical magnitudes; In the first probability matrix, selecting the probability values of the target quantity with the front sorting positions to construct the second probability matrix.

[0007] A language model training method provided by the present invention, the method of determining a target quantity according to the compression mode in the user input information includes: When the compression mode is an adaptive compression mode, successively performing cumulative calculation on the probability values in the first probability matrix according to the descending order sorting result until the cumulative calculation value is greater than or equal to a preset threshold, and determining the target quantity according to the number of probability values participating in the cumulative calculation; When the compression mode is a fixed compression mode, determining the preset quantity as the target quantity.

[0008] A language model training method provided by the present invention, the method of performing distillation training on the student model according to the third vocabulary and the second probability matrix to obtain a target language model includes: Constructing an empty matrix according to the number of tokens in the third vocabulary; Filling the probability values in the second probability matrix into the empty matrix according to the first indexes of the tokens corresponding to the probability values in the second probability matrix to obtain a first reconstructed probability matrix corresponding to each data unit; the first index is the index of the token corresponding to each probability value in the second probability matrix in the third vocabulary; Performing normalization processing on the probability values in the first reconstructed probability matrix to obtain a second reconstructed probability matrix corresponding to each data unit; Based on the student model, predicting a third probability matrix corresponding to each data unit; the third probability matrix includes the probability values of each data unit belonging to each token in the third vocabulary; Performing distillation training on the student model according to the second reconstructed probability matrix, the third probability matrix, and the token labels corresponding to each data unit to obtain the target language model.

[0009] A language model training method provided by the present invention, the method of performing alignment operation on the second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary includes: In the second probability matrix, obtain each first probability value and each second probability value; there is no mapping relationship between the token corresponding to the first probability value and all tokens in the second token list; there is a mapping relationship between the token corresponding to the second probability value and one token in the second token list; Based on the tokenizer of the student model, map and generate the mapping index of the token corresponding to each first probability value; Determine the second index of the target token that has a mapping relationship with the token corresponding to each second probability value as the mapping index of the token corresponding to each second probability value; the second index is the index of the target token in the second token list; Fill the tokens corresponding to each first probability value into the empty token list according to the mapping index of the tokens corresponding to each first probability value, and fill the tokens corresponding to each second probability value into the empty token list according to the mapping index of the tokens corresponding to each second probability value; Obtain the third token list according to the filling result.

[0010] According to a language model training method provided by the present invention, the method further includes: When the second probability matrix is obtained and the training instruction of the student model is not received, encode and store each data unit, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix to the disk; the third index is the index of each probability value in the second probability matrix in the first token list; When the training instruction of the student model is received, parse and obtain each data unit, each probability value in the second probability matrix, and the third index in the disk.

[0011] According to a language model training method provided by the present invention, the encoding and storing each data unit, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix to the disk includes: Perform quantization encoding on each probability value in the second probability matrix to obtain a first encoding result; Perform quantization encoding on the third index to obtain a second encoding result; Perform hash encoding on each data unit to obtain a third encoding result; Store the first encoding result, the second encoding result, and the third encoding result in the disk in the form of a compressed file.

[0012] A language model training method provided by the present invention, in the disk, parsing and obtaining each of the data units, each probability value in the second probability matrix, and the third index, includes: Loading the compressed file in the disk; Performing inverse quantization decoding on the first encoding result in the compressed file to obtain each probability value in the second probability matrix; Performing inverse quantization decoding on the second encoding result in the compressed file to obtain the third index; Performing a hash reverse lookup operation on the third encoding result in the compressed file to obtain each of the data units.

[0013] The present invention also provides a text processing method, including: Obtaining the text to be processed; Based on the target language model, performing token prediction on each data unit in the text to be processed to obtain a token prediction result corresponding to the text to be processed; Performing text processing on the text to be processed according to the token prediction result; Wherein, the text processing includes text generation, code generation, machine translation or text classification; the target language model is trained based on the language model training method described in any one of the above.

[0014] The present invention also provides a language model training device, including: A first prediction unit, configured to predict a first probability matrix corresponding to each data unit in the sample text based on a teacher model; the first probability matrix includes probability values of each data unit belonging to each token in a first vocabulary, and the first vocabulary is the vocabulary of the teacher model; A compression unit, configured to compress the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain a second probability matrix corresponding to each data unit; A mapping unit, configured to perform an alignment operation on a second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of the student model; A training unit, configured to perform distillation training on the student model according to the third vocabulary and the second probability matrix to obtain a target language model.

[0015] The present invention also provides a text processing device, including: An acquisition unit, configured to acquire the text to be processed; A second prediction unit, configured to perform token prediction on each data unit in the text to be processed based on the target language model to obtain a token prediction result corresponding to the text to be processed; A processing unit, configured to perform text processing on the text to be processed according to the token prediction result; Wherein, the text processing includes text generation, code generation, machine translation or text classification; the target language model is trained based on the language model training method described in any one of the above.

[0016] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the language model training method described in any one of the above is implemented.

[0017] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the language model training method described in any one of the above is implemented.

[0018] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, the language model training method described in any one of the above is implemented.

[0019] The language model training method, text processing method, device, equipment and medium provided by the present invention are based on the teacher model to predict the first probability matrix corresponding to each data unit in the sample text, so as to obtain the full probability distribution of each data unit, and perform sparse adaptive compression on the first probability matrix according to the magnitudes of the probability values in the first probability matrix to generate a second probability matrix with a lower storage capacity, effectively reducing the data volume and storage requirements; then, using the tokens corresponding to the probability values in the second probability matrix to perform an alignment operation on the vocabulary of the student model to obtain a third vocabulary, so as to solve the problem of inconsistent vocabularies between the teacher model and the student model through dynamic vocabulary mapping, avoiding the problem of knowledge transfer failure caused by vocabulary differences. Finally, the student model is distilled and trained through the third vocabulary and the second probability matrix to obtain the target language model. Since the distillation training is performed through sparse compression of the probability matrix and vocabulary alignment, it not only significantly reduces the problem of excessive storage costs caused by the growth of the vocabulary size, effectively improves the distillation training efficiency, but also enables the target language model trained accordingly to better adapt to different model architectures and text processing scenarios while maintaining high performance. Description of the Drawings

[0020] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0021] Figure 1 It is one of the flow schematic diagrams of the language model training method provided by the present invention.

[0022] Figure 2 It is the second of the flow schematic diagrams of the language model training method provided by the present invention.

[0023] Figure 3 It is the third of the flow schematic diagrams of the language model training method provided by the present invention.

[0024] Figure 4 It is the fourth of the flow schematic diagrams of the language model training method provided by the present invention.

[0025] Figure 5 It is the fifth of the flow schematic diagrams of the language model training method provided by the present invention.

[0026] Figure 6 It is the flow schematic diagram of the text processing method provided by the present invention.

[0027] Figure 7 It is the structural schematic diagram of the language model training device provided by the present invention.

[0028] Figure 8 It is the structural schematic diagram of the text processing device provided by the present invention.

[0029] Figure 9 It is the structural schematic diagram of the electronic device provided by the present invention. Detailed implementation manners

[0030] To make the objectives, technical solutions and advantages of the present invention clearer, the technical solutions in the present invention will be clearly and completely described below with reference to the accompanying drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention belong to the scope of protection of the present invention.

[0031] Large language models have demonstrated excellent performance and broad application prospects in text processing. However, these large language models usually have a huge parameter scale and complex structure, resulting in problems such as high computational costs, making it difficult to directly deploy them on resource-constrained devices. Therefore, in related technologies, distillation training is used to transfer the capabilities of large language models to lightweight student models, that is, the soft labels (probability matrices) generated by the teacher model are used to guide the training of the student model, so that the student model trained accordingly can not only learn how to judge the categories of correct samples from labeled data at a lower computational cost, but also learn the inter-class relationships from the teacher model.

[0032] In the existing distillation training method for large language models, usually, a teacher model built by a large language model is first used to generate full soft labels for all tokens in all training data, and the full soft labels of each token in each training data under the vocabulary of the teacher model are saved together with the hard labels to the local for direct loading during the training of the student model. Since the dimension of the full soft label is directly related to the vocabulary size, that is, the dimension of the full soft label is 1×vocab_size, where vocab_size is the number of tokens in the vocabulary of the teacher model. Therefore, with the explosive growth of the vocabulary size, the storage amount of the full soft labels of tokens in a single training data also increases. For example, when the vocabulary of the teacher model contains 150,000 tokens, a single data contains 3,000 tokens, and is stored in float32 type, it takes about 180MB of space to store the full soft labels of only one data, resulting in problems of too high storage cost and low training efficiency during the model training process.

[0033] If traditional compression methods, such as compression package (ZIP) compression or quantization coding (such as directly converting the data from float32 type to float16), are used to compress the full soft labels of all tokens in all training data, this method not only has limited compression rate, but also needs to be fully loaded into memory after decompression, and the data distribution format after decompression is inconsistent with the data distribution format required for distillation training, requiring additional conversion processing, resulting in problems of too high storage cost and low training efficiency during the model training. If the hard label method is directly used to replace the full soft label, although the problem of too high storage cost can be avoided, since the hard label cannot effectively represent the relationship between classes, the relative relationship information between classes in the training data of the student model is lost, easily leading to low performance of the student model trained accordingly.

[0034] In view of this, the present application proposes a language model training method. This method can be widely applied to the lightweight model deployment in scenarios such as text generation, code generation, machine translation, text classification or intelligent customer service, etc., to provide key technical support for the implementation of large-scale artificial intelligence models.

[0035] Figure 1 is one of the schematic flowcharts of the language model training method provided by the present invention. As Figure 1 shown, this method includes step 110, step 120, step 130 and step 140.

[0036] Step 110, based on the teacher model, predict the first probability matrix corresponding to each data unit in the sample text; the first probability matrix includes the probability values of each data unit belonging to each token in the first vocabulary, and the first vocabulary is the vocabulary of the teacher model.

[0037] The teacher model here is a large-scale model constructed based on a large language model. The large language model (LLM) here, also known as a large model or a general pre-trained large model, etc., refers to a natural language processing (NLP) model with a huge number of parameters, and the number of model parameters and / or the complexity of the model structure exceed a set threshold. This model is pre-trained with a large amount of text data and has a high semantic understanding ability and the ability to generate natural language. This large language model can be a Bidirectional Encoder Representations from Transformers (BERT), a Generative Pre-trained Transformer, a large language model Meta AI, etc. This embodiment does not make specific limitations on this.

[0038] The sample text here refers to the text data used to train the student model, which can be text corpora in various text processing tasks, such as news articles in text classification tasks.

[0039] Figure 2 It is the second flow diagram of the language model training method provided by the present invention. Figure 3 It is the third flow diagram of the language model training method provided by the present invention.

[0040] As Figure 2 and Figure 3 shown, during the model training process, the sample text can be first input into the teacher model, and the teacher model can predict the probability values of each data unit (token) in the sample text belonging to all word elements in the first vocabulary (hereinafter also referred to as the teacher vocabulary or the vocabulary of the teacher model) through forward propagation inference, so as to obtain the first probability matrix corresponding to each data unit (hereinafter also referred to as the full soft label), that is, obtain the full soft label available for training the student model. In this process, specifically, the sample text can be divided into multiple tokens by using the tokenizer of the teacher model, and forward propagation calculations are performed on each token after tokenization processing to obtain the scores of each token unit belonging to each word element in the first vocabulary, and then a normalization function (such as the softmax function, etc.) is used to normalize the scores of each token unit belonging to each word element in the first vocabulary, and the probability values of each token belonging to each word element in the first vocabulary can be obtained, thereby forming the first probability matrix corresponding to each data unit.

[0041] Among them, the dimension of the first probability matrix is . Among them, is the number of word elements included in the first vocabulary.

[0042] Exemplarily, the inference steps of the teacher model can be specifically implemented through the following code: {def generate_soft_labels(teacher_model, input_text): tokenized_input = teacher_tokenizer(input_text) # Tokenization logits = teacher_model( tokenized_input).logits # Forward propagation probs = torch.softmax(logits, dim=-1) # Normalization return probs}。

[0043] Step 120: Compress the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain a second probability matrix corresponding to each of the data units.

[0044] Optionally, after obtaining the first probability matrix, the first probability matrix can be compressed according to the numerical magnitudes of the probability values in the first probability matrix to retain some probability values in the first probability sequence and their corresponding token information (such as the index of the token in the first vocabulary), thereby reducing its dimension and further reducing its storage space.

[0045] It should be noted that during the compression process, probability values in the first probability matrix that are greater than a target value and their corresponding token information can be selected to construct a second probability sequence; or the probability values in the first probability sequence can be sorted in descending order of magnitude, and the top target number of probability values and their corresponding token information can be selected from them to form a second probability sequence, etc. This embodiment does not make specific limitations in this regard. The target value and the target number here can be fixedly set according to actual needs, or adaptively determined according to the distribution characteristics of the probability values in the first probability sequence, etc. This embodiment does not make specific limitations in this regard.

[0046] Step 130: Align the second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of the student model.

[0047] The student model here is a lightweight network model, which can be a network model with the same model architecture as the teacher model but with a much smaller scale of model parameters than the teacher model, or a network model with a different model architecture from the teacher model and with a much smaller scale of model parameters than the teacher model, etc. This embodiment does not make specific limitations on this.

[0048] It should be noted that during the knowledge distillation process, when the vocabulary of the teacher model is inconsistent with that of the student model, the indices of the word tokens corresponding to the probability values in the full soft labels generated based on the teacher model will not be directly alignable with the vocabulary of the student model. In the existing solutions, the vocabulary is forced to be unified for vocabulary forced alignment, such as unifying the tokenizer, sharing the embedding layer between the student model and the teacher model, or discarding the probability values of the tokens that do not appear in the student vocabulary in the full soft labels generated by the teacher model and only retaining the probability values of the common word tokens, etc. However, this method of forcibly unifying the vocabulary cannot effectively solve the index misalignment problem caused by subword combination differences, which in turn leads to the student model being unable to effectively utilize the soft labels of the teacher model for effective learning, that is, resulting in the failure of knowledge transfer. Moreover, experiments show that this method of forcibly unifying the vocabulary will increase the perplexity of the student model in the generation task by 18% - 25%, greatly limiting the distillation application in cross-model architecture or cross-language scenarios.

[0049] To overcome the above problems, in this embodiment, an alignment operation of dynamically remapping the vocabulary of the student model is performed by using the probability distribution characteristics of the full soft labels to solve the probability distribution misalignment caused by tokenization differences and support lossless distillation between heterogeneous models. The specific implementation steps are as follows: According to the word tokens corresponding to the probability values in the second probability matrix and the mapping relationship of the word tokens between the first vocabulary and the second vocabulary, an alignment operation is performed on the second vocabulary to obtain a third vocabulary. For example, for any probability value in the second probability matrix, if the word token corresponding to this probability value has a mapping with a certain word token in the second vocabulary, then determine the mapping index of the word token corresponding to this probability value as the index of the word token in the second vocabulary that has a mapping relationship with the word token corresponding to this probability value, and add it to the initially constructed empty vocabulary according to the mapping index of the word token corresponding to this probability value; if the word token corresponding to this probability value has no mapping with all the word tokens in the second vocabulary, then generate an index of an unknown word token for the word token corresponding to this probability value through the tokenizer of the student model to determine the mapping index of the word token corresponding to this probability value, and add the word token corresponding to this probability value as an unknown word token to the initially constructed empty vocabulary according to the index of the unknown word token. Thus, by traversing each probability value in the second probability matrix in this step, the third vocabulary can be obtained.

[0050] Step 140, perform distillation training on the student model according to the third vocabulary and the second probability matrix to obtain a target language model.

[0051] The target language model here is the student model trained by distillation. It can learn the knowledge of the teacher model as much as possible with a smaller model scale, so as to effectively predict the tokens of the input text, and then obtain high-performance token prediction results to complete corresponding text processing tasks, such as text generation, code generation, machine translation or text classification, etc. This embodiment does not make specific limitations on this.

[0052] Optionally, after obtaining the third vocabulary, the sample text can be input into the student model, and the student model predicts the probability values of each data unit in the sample text belonging to each token in the third vocabulary through forward propagation inference, so as to form a third probability matrix corresponding to each data unit.

[0053] And reconstruct the second probability matrix according to the structure and quantity of the third vocabulary to obtain a reconstructed probability matrix associated with the third probability matrix; for example, adjust the positions of the probability values in the second probability matrix according to the structure and quantity of the third vocabulary, and fill in zeros at necessary positions to align it with the third vocabulary, so as to obtain a reconstructed probability matrix associated with the third probability sequence matrix, to ensure that the soft label knowledge of the teacher model can match the output of the student model.

[0054] Immediately, use the third probability matrix and the reconstructed probability matrix to calculate the difference between the output of the student model and the output of the teacher model. For example, KL divergence loss can be used to measure the difference between the output of the student model and the output of the teacher model to obtain the first loss; use the difference between the third probability matrix and the token label (i.e., hard label) corresponding to each data unit. For example, cross-entropy loss can be used to measure the difference between the output of the student model and the hard label to obtain the second loss.

[0055] Immediately, update the student model using the combined first loss and second loss through the backpropagation algorithm. By minimizing the first loss, prompt the learning output of the student model to be as close as possible to the output of the teacher model, so as to effectively transfer the knowledge of the teacher model to the student model, and by minimizing the second loss, prompt the student model to learn the fitting ability for the true label. Thus, prompt the student model to be able to learn the soft label knowledge of the teacher model and retain a good fit for the hard label during the distillation process, so as to obtain a target language model that can perform text processing quickly and with high performance.

[0056] The method provided in this embodiment predicts the first probability matrix corresponding to each data unit in the sample text based on the teacher model to obtain the full probability distribution of each data unit, and performs sparse adaptive compression on the first probability matrix according to the size of each probability value in the first probability matrix to generate a second probability matrix with lower storage volume, which effectively reduces the amount of data and storage requirements; then, the word element corresponding to each probability value in the second probability matrix is ​​used to align the vocabulary of the student model to obtain a third vocabulary, so as to solve the problem of inconsistency between the vocabulary of the teacher model and the student model through dynamic vocabulary mapping, and avoid the problem of knowledge transfer failure caused by vocabulary differences. Finally, the student model is distilled and trained through the third vocabulary and the second probability matrix to obtain the target language model. Since distillation training is performed by sparse compression of probability matrix and vocabulary alignment, it not only significantly reduces the problem of excessive storage cost caused by the growth of vocabulary scale, effectively improves the efficiency of distillation training, but also enables the target language model trained thereby to better adapt to different model architectures and text processing scenarios while maintaining high performance.

[0057] Figure 4 This is a fourth flow chart of the language model training method provided by the present invention; Figure 4 As shown, the method includes step 410 , step 420 , step 430 and step 440 .

[0058] Step 410, based on the teacher model, predict a first probability matrix corresponding to each data unit in the sample text; the first probability matrix includes probability values ​​of each data unit belonging to each word in a first vocabulary, and the first vocabulary is the vocabulary of the teacher model.

[0059] Optionally, during the model training process, the sample text can be first input into the teacher model, and the teacher model predicts the probability value of each token in the sample text belonging to all the words in the first vocabulary through forward propagation reasoning, so as to obtain the first probability matrix corresponding to each data unit, that is, to obtain the full amount of soft labels that can be used for student model training. The specific implementation steps can be found in step 110, which will not be repeated here.

[0060] Step 420, receiving user input information, determining the target quantity according to the compression mode in the user input information; sorting the probability values ​​in the first probability matrix in descending order according to the numerical value; in the first probability matrix, selecting the probability value of the target quantity with a forward sorting position to construct the second probability matrix.

[0061] The user input information here is the information for selecting the compression mode for the full soft labels output by the teacher model input by the user on the front-end interface. This information can be input through input forms such as command-line interface input, graphical interface input, touch input, drop-down selection input, voice input, and gesture input. This embodiment does not specifically limit this.

[0062] Optionally, after receiving the user input information, the compression mode can be parsed from the user input information, and according to the compression mode, the corresponding number of probability values to be retained can be determined, thereby obtaining the target number. The compression modes here include fixed compression mode or adaptive compression mode, etc., and the number of probability values retained corresponding to different compression modes is different. For example, the number of probability values retained corresponding to the fixed compression mode is a fixed selected number, while the number of probability values retained corresponding to the adaptive compression mode is a number dynamically determined based on the probability distribution characteristics.

[0063] In a possible implementation manner, determining the target number according to the compression mode in the user input information includes: When the compression mode is the adaptive compression mode, the probability values in the first probability matrix are successively accumulated and calculated in descending order until the accumulated calculation value is greater than or equal to a preset threshold, and the target number is determined according to the number of probability values participating in the accumulated calculation; When the compression mode is the fixed compression mode, the preset number is determined as the target number.

[0064] As Figure 2 and Figure 3 shown, when the compression mode is the adaptive compression mode, the probability values in the first probability matrix can be sorted in descending order according to the numerical size, and according to the descending order result, the probability values in the first probability matrix are successively accumulated and calculated until the accumulated value is greater than or equal to the preset threshold, then the cumulative calculation step is stopped; then, the target number K is dynamically determined according to the number of probability values participating in the accumulated calculation to balance the information retention rate and the compression efficiency.

[0065] The preset threshold here can be defined according to actual needs, such as defined as 0.95 or 0.9, etc.; it can also be dynamically determined according to the performance requirements of the model. For example, if it is required that the student model has a high accuracy after distillation training, the preset threshold can be set to a higher value, such as 0.9 or 0.95, to retain more high-probability token information. If there are high requirements for the inference speed of the model, the preset threshold can be appropriately reduced to reduce the retained information volume, thereby reducing the computational complexity. This embodiment does not specifically limit the determination method of this preset threshold.

[0066] When the compression mode is the fixed compression mode, the number of reserved probability values directly specified according to task requirements (such as high-precision distillation requirements or low storage overhead requirements) (i.e., the preset number) is determined as the target number K. For example, the target number is set to 50, etc., and this embodiment does not make specific limitations on this.

[0067] In summary, during the compression process, by flexibly determining the target number according to the compression mode in the user input information to perform soft label compression, it can not only improve the user experience and enhance the flexibility of soft label compression, but also achieve flexible balancing of model accuracy and training efficiency while reducing storage and computing costs.

[0068] After obtaining the target number K through the above steps, the probability values in the first probability matrix can be sorted in descending order according to their numerical magnitudes, and according to the descending order sorting result, in the first probability matrix, the top K probability values and their corresponding third indices are selected to construct the second probability matrix. Thus, the structured compression method of Top-K probability value truncation is used to retain the high probability values and their indices, so as to reduce the storage amount of the soft label from O(vocab_size) to O(K), reduce the storage space of a single piece of data by more than 99%, and thus significantly reduce the storage cost, improve the data transmission efficiency, accelerate the model training and inference processes, and enhance the deployment efficiency and performance of the model.

[0069] Exemplarily, the compression step here can be specifically implemented through the following code: {def compress_probs(probs: Tensor, mode: str, param: float) ->Tuple[Tensor, Tensor]: if mode == "fixed_k": k = int(param) top_values, top_indices = torch.topk(probs, k) elif mode == "adaptive_threshold": sorted_probs, _ = torch.sort(probs, descending=True) cum_sum = torch.cumsum(sorted_probs, dim=0) k = torch.argmax(cum_sum>= param) + 1 top_values = sorted_probs[:k] top_indices = torch.argsort(probs, descending=True)[:k] return top_values, top_indices}。

[0070] Step 430: Align the second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of the student model.

[0071] Optionally, after obtaining the second probability matrix, the second vocabulary can be aligned according to the tokens corresponding to the probability values in the second probability matrix and the mapping relationship of the tokens between the first vocabulary and the second vocabulary to obtain a third vocabulary. For the specific implementation steps, refer to Step 130, which will not be elaborated here.

[0072] Step 440: Distill and train the student model according to the third vocabulary and the second probability matrix to obtain a target language model.

[0073] Optionally, after obtaining the third vocabulary, the sample text can be input into the student model. The student model predicts the probability values of each data unit in the sample text belonging to each token in the third vocabulary through forward propagation inference to form a third probability matrix corresponding to each data unit; and reconstruct the second probability matrix according to the structure and quantity of the third vocabulary to obtain a reconstructed probability matrix associated with the third probability matrix; then, calculate the difference between the output of the student model and the output of the teacher model using the third probability matrix and the reconstructed probability matrix to obtain a first loss; calculate the difference between the third probability matrix and the token labels corresponding to each data unit to obtain a second loss; then, jointly update the student model using the first loss and the second loss through the backpropagation algorithm to obtain a target language model that can process text quickly and with high performance. For the specific implementation steps, refer to Step 140, which will not be elaborated here.

[0074] The method provided in this embodiment performs efficient, cross-frame, and cross-vocabulary model distillation training through sparse compression of the probability matrix and dynamic vocabulary mapping, not only significantly reducing the problem of excessive storage costs caused by the growth of the vocabulary size, effectively improving the model training efficiency, but also enabling the target language model trained accordingly to better adapt to different model architectures and text processing scenarios while maintaining high performance.

[0075] In some embodiments, the distilling and training the student model according to the third vocabulary and the second probability matrix to obtain a target language model includes: Constructing an empty matrix according to the number of word-units in the third vocabulary; According to the first index of the word element corresponding to each probability value in the second probability matrix, each probability value in the second probability matrix is ​​filled into the empty matrix to obtain a first reconstructed probability matrix corresponding to each data unit; the first index is the index of the word element corresponding to each probability value in the second probability matrix in the third vocabulary; Normalizing each probability value in the first reconstruction probability matrix to obtain a second reconstruction probability matrix corresponding to each of the data units; Based on the student model, predicting a third probability matrix corresponding to each of the data units; the third probability matrix includes probability values ​​of each of the data units belonging to each word in the third vocabulary; According to the second reconstruction probability matrix, the third probability matrix, and the word-unit label corresponding to each of the data units, the student model is distilled and trained to obtain the target language model.

[0076] like Figure 2 and Figure 3 As shown in Figure 1, when training the student model, in order to ensure the integrity of the distillation effect, the second probability matrix is ​​first reconstructed. Specifically, according to the number of words in the third vocabulary, an empty matrix is ​​constructed. The matrix can be ,in, is the number of word units in the third vocabulary. Then, according to the index of the word unit corresponding to each probability value in the second probability matrix in the third vocabulary (that is, the first index), each probability value in the second probability matrix is ​​filled into the empty matrix to obtain the first reconstruction probability matrix corresponding to each data unit, so that the first reconstruction probability matrix reconstructed thereby is adapted to the output of the student model; at the same time, the first reconstruction probability matrix is ​​normalized (such as rescaling of the softmax function, etc.) to ensure that the sum of all probability values ​​in the second reconstruction probability matrix constructed thereby is 1, so as to avoid the deviation in loss calculation caused by truncation of probability values, and experimentally verify that this method can reduce the distillation loss by 32%.

[0077] Exemplarily, the reconstruction step of the second probability matrix can be implemented based on the following code: {def reconstruct_probs(top_values, student_indices, vocab_S): probs = torch.zeros(vocab_S) # empty matrix construction probs.scatter_(0, student_indices, top_values) # probability value filling if config.normalize: probs = probs / probs.sum() # Normalization processing return probs}。

[0078] In addition, the sample text is input into the student model, and the student model predicts the probability values of each data unit in the sample text belonging to all tokens in the third vocabulary through forward propagation inference, so as to form a third probability matrix corresponding to each data unit.

[0079] Subsequently, using the third probability matrix and the second reconstruction probability matrix, calculate the difference between the output of the student model and the output of the teacher model. For example, KL divergence loss can be used to measure the difference between the output of the student model and the output of the teacher model, and the first loss (also called soft label loss) is obtained, so as to minimize the first loss in the follow-up to make the learning output of the student model as close as possible to the output of the teacher model, thereby effectively transferring the knowledge of the teacher model to the student model; use the difference between the third probability matrix and the token labels corresponding to each data unit. For example, cross-entropy loss can be used to measure the difference between the output of the student model and the hard labels, and the second loss (also called hard label loss) is obtained, so as to minimize the second loss in the follow-up to enable the student model to learn the fitting ability to the true labels.

[0080] Subsequently, fuse the first loss and the second loss (such as direct addition or weighted addition, etc.) to obtain a mixed loss, so as to update the student model through the mixed loss using the backpropagation algorithm, aiming to enable the student model to learn the soft label knowledge of the teacher model and retain a good fit to the hard labels during the distillation process, and finally obtain a target language model that can perform text processing quickly and with high performance.

[0081] Exemplarily, the calculation steps of the mixed loss can be implemented based on the following code: {def distillation_loss(student_logits, hard_labels, soft_probs, alpha=0.7): loss_ce = F.cross_entropy(student_logits, hard_labels) # Hard label loss log_probs = F.log_softmax(student_logits, dim=-1) # Soft label loss loss_kl = F.kl_div(log_probs, soft_probs, reduction="batchmean") total_loss = alpha loss_kl + (1 - alpha) loss_ce # Weighted sum return total_loss}。

[0082] The method provided in this embodiment makes the reconstructed second reconstruction probability matrix adapt to the output of the student model by reconstructing and normalizing the second probability matrix formed by compression, avoiding the loss calculation deviation caused by probability truncation, ensuring that the student model can effectively learn the soft label knowledge of the teacher model, and through the fusion of multiple losses, prompting the student model to retain the fitting ability to hard labels while learning the knowledge of the teacher model, effectively achieving the training of a high-performance target language model.

[0083] In some embodiments, the operation of aligning the second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary includes: In the second probability matrix, obtain each first probability value and each second probability value; there is no mapping relationship between the token corresponding to the first probability value and all tokens in the second vocabulary; there is a mapping relationship between the token corresponding to the second probability value and one token in the second vocabulary; Based on the tokenizer of the student model, map and generate the mapping indexes of the tokens corresponding to each of the first probability values; Determine the second index of the target token that has a mapping relationship with the token corresponding to each of the second probability values as the mapping index of the token corresponding to each of the second probability values; the second index is the index of the target token in the second vocabulary; Fill the tokens corresponding to each of the first probability values into the empty vocabulary according to the mapping indexes of the tokens corresponding to each of the first probability values, and fill the tokens corresponding to each of the second probability values into the empty vocabulary according to the mapping indexes of the tokens corresponding to each of the second probability values; Obtain the third vocabulary according to the filling result.

[0084] Such as Figure 3As shown, when performing vocabulary alignment, it can be determined whether there is a mapping relationship between the word elements corresponding to the probability values in the second probability matrix and the word elements in the second vocabulary. For any probability value in the second probability matrix, if the word element corresponding to the probability value has a mapping with a certain word element in the second vocabulary, it is used as the second probability value, and the mapping index of the word element corresponding to the probability value is obtained by mapping according to the second index of the target word element in the second vocabulary that has a mapping relationship with each second probability value, and it is added to the initially constructed empty vocabulary according to the mapping index of the word element corresponding to the probability value; if the word element corresponding to the probability value has no mapping with all the word elements in the second vocabulary, it is used as the first probability value, and the index of the unknown word element (UNK) is generated for it through the tokenizer of the student model, and the index of the unknown word element is used as the mapping index of the word element corresponding to the probability value, and the word element corresponding to the probability value is added to the initially constructed empty vocabulary in the form of the unknown word element. By traversing each probability value in the second probability matrix in this way, the third vocabulary can be obtained.

[0085] Exemplarily, the vocabulary alignment operation here can be specifically implemented through the following code: {def map_to_student_vocab(raw_tokens, student_tokenizer): student_indices = [] # Initialize an empty vocabulary for tok in raw_tokens: if tok in student_tokenizer.get_vocab(): idx = student_tokenizer.convert_tokens_to_ids(tok) student_indices.append(idx) else: if config.oov_handle == "discard": continue else: student_indices.append(student_tokenizer.unk_token_id) return torch.tensor(student_indices)}。

[0086] In summary, the method provided in this embodiment performs a dynamic mapping alignment operation on the second vocabulary according to the mapping relationship between each token in each probability value in the second probability matrix obtained by compression and each token in the second vocabulary, effectively alleviating the problem of knowledge transfer failure caused by vocabulary differences and improving the success rate of cross-model distillation. Moreover, through simulation verification, it can be seen that this method can increase the success rate of cross-model distillation to more than 95%.

[0087] Figure 5 is the fifth flow diagram of the language model training method provided by the present invention; as Figure 5 shown, the method includes steps 510, 520, 530, 540, and 550.

[0088] Step 510: Based on the teacher model, predict the first probability matrix corresponding to each data unit in the sample text; the first probability matrix includes the probability values of each of the data units belonging to each token in the first vocabulary, and the first vocabulary is the vocabulary of the teacher model.

[0089] Optionally, during the model training process, the sample text can be first input into the teacher model, and the teacher model can predict the probability values of each token in the sample text belonging to all tokens in the first vocabulary through forward propagation inference to obtain the first probability matrix corresponding to each data unit, that is, obtain the full soft labels available for training the student model. The specific implementation steps can refer to step 110 and will not be elaborated here.

[0090] Step 520: Compress the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain the second probability matrix corresponding to each data unit.

[0091] Optionally, after obtaining the first probability matrix, the first probability matrix can be compressed according to the numerical magnitudes of the probability values in the first probability matrix to retain some probability values and their corresponding token information in the first probability sequence, thereby reducing its dimension and further reducing its storage space. The specific implementation steps can refer to step 120 and will not be elaborated here.

[0092] Step 530: When the second probability matrix is obtained and the training instruction of the student model has not been received, encode and store each data unit, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix to the disk; the third index is the index of each probability value in the second probability matrix in the first vocabulary; when the training instruction of the student model is received, parse and obtain each data unit, each probability value in the second probability matrix, and the third index in the disk.

[0093] It should be noted that in the prior art, real-time inference is usually adopted for distillation training. That is, each time during forward propagation, the teacher model is called to perform immediate inference on the sample data to generate soft labels (probability distributions), which are directly used for training the student model. However, this training method relies on the deep coupling of the teacher model and the student model training framework, requires synchronous processing of model inference and gradient backpropagation, and requires the teacher model and the student model to share computing resources. However, usually the teacher model and the student model are often based on different frameworks, resulting in frequent switching of the training device between inference and training tasks, leading to resource scheduling conflicts and a significant decrease in resource utilization, and the data transmission overhead between heterogeneous frameworks significantly reduces the training efficiency.

[0094] In response to this, in this embodiment, by designing a soft label pre-compression and decompression training mechanism, the inference of the teacher model and the training of the student model are completely decoupled to effectively support cross-frame hybrid training, eliminate resource competition, and improve training efficiency. The specific implementation steps are as follows: As Figure 2 and Figure 3 shown, when the second probability matrix is obtained and the training instruction of the student model is not received, it indicates that the training step of the learning model has not been triggered. At this time, each data unit, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix can be encoded and then stored offline on the disk, thereby further reducing the storage cost.

[0095] The encoding here can be encoding at least one of the data in each data unit, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix. The encoding here can be hash encoding, quantization encoding, etc., and this embodiment does not make specific limitations on this.

[0096] It should be noted that when encoding the data unit, probability value, and third index, a unified encoding method can be used for encoding, or different encoding methods can be used for different data. For example, hash encoding can be used for the data unit, and quantization encoding can be used for the probability value and the third index.

[0097] When a training instruction for the student model is received, the representation learning model training step is triggered. At this time, in the disk, the corresponding decoding method of the encoding method can be used to read the encoded data in the disk and perform corresponding decoding to obtain each data unit, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix. Furthermore, based on the parsed data, the student model can be assisted in distillation training. For example, if hash encoding is used for the data unit during the encoding process, the data unit can be decoded through hash reverse look-up operation or the inverse process of the hash algorithm; if the probability value and the third index use quantization encoding, the probability value and the third index can be decoded through inverse quantization operation.

[0098] In a possible implementation manner, the encoding and storing of each of the data units, each probability value in the second probability matrix, and the third index of the token corresponding to each probability value in the second probability matrix to the disk includes: Performing quantization encoding on each probability value in the second probability matrix to obtain a first encoding result; Performing quantization encoding on the third index to obtain a second encoding result; Performing hash encoding on each of the data units to obtain a third encoding result; Storing the first encoding result, the second encoding result, and the third encoding result in the disk in the form of a compressed file.

[0099] As Figure 3 shown, during the encoding and storage process, specifically, quantization encoding can be performed on each probability value in the second probability matrix. For example, each probability value in the second probability matrix is quantized from float32 type to float16 type, etc., to reduce the storage space requirement of each probability value in the second probability matrix, thereby significantly reducing the storage cost.

[0100] In addition, quantization encoding can also be performed on the third index of the token corresponding to each probability value in the second probability matrix. For example, differential quantization encoding is performed on the third index of the token corresponding to each probability value of float16 type to map the third index to a smaller quantization interval, so as to significantly reduce the storage space requirement of the third index while maintaining data integrity. In addition, hash encoding can also be performed on each data unit to encode each data unit into a binary vector, thereby reducing the storage space requirement of each data unit.

[0101] Exemplarily, the encoding step here can be specifically implemented by the following code: { def encode_for_storage(top_values, top_indices, raw_tokens): values_quantized = top_values.to(torch.float16) # Quantization encoding indices_quantized = top_indices.to(torch.int32) # Quantization encoding token_hashes = [hash(tok) % 2 32 for tok in raw_tokens] # Hash encoding return (values_quantized, indices_quantized, token_hashes)}。

[0102] After completing the multi-level data encoding of the compressed data, the encoded data (i.e., the first encoding result, the second encoding result, and the third encoding result) can be further written to the disk in the form of a compressed file according to formats such as HDF5 / Parquet. Thus, by optimizing the storage of the compressed data through multi-level data encoding, the storage space requirement and cost are further reduced, and the model training efficiency is improved.

[0103] Correspondingly, in a possible implementation manner, parsing and obtaining each of the data units, each probability value in the second probability matrix, and the third index in the disk includes: Loading the compressed file in the disk; Performing inverse quantization decoding on the first encoding result in the compressed file to obtain each probability value in the second probability matrix; Performing inverse quantization decoding on the second encoding result in the compressed file to obtain the third index; Performing a hash reverse look-up operation on the third encoding result in the compressed file to obtain each of the data units.

[0104] Such as Figure 3As shown, when a training instruction for the student model is received, the compressed file can be read from the disk, and an inverse quantization decoding is performed on the first encoded result in the compressed file using a decoding method corresponding to the encoding method of the probability value. For example, the first encoded result is inverse quantization decoded from the float16 type to the float32 type to obtain each probability value in the second probability matrix. In addition, an inverse quantization decoding can also be performed on the second encoded result in the compressed file using a decoding method corresponding to the encoding method of the third index. For example, the second encoded result is inverse quantization decoded from the differential encoding to the absolute index of the float16 type to obtain the third index corresponding to each probability value in the second probability matrix. In addition, a hash reverse lookup operation can be performed on the third encoded result in the compressed file to obtain each data unit. Thus, by performing inverse decoding on the compressed file stored on the disk in this step, the data can be efficiently and completely recovered from the compressed file, obtaining each probability value in the complete second probability sequence matrix, its corresponding third index, and each data unit, so as to effectively transfer the knowledge of the teacher model to the student model, improve the performance of the student model, and at the same time decouple the inference step of the teacher model from the training step of the learning model, eliminate resource competition, thereby reducing the memory requirement, improving the training efficiency, and enhancing the training effect of the student model.

[0105] Exemplarily, the decoding step here can be specifically implemented by the following code: { values = loaded_values.to(torch.float32)# Inverse quantization indices = loaded_indices.to(torch.int64) # Inverse quantization raw_tokens = decode_hashes(loaded_hashes) # Hash reverse lookup}

[0106] Step 540, perform an alignment operation on the second vocabulary according to the tokens corresponding to each probability value in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of the student model.

[0107] Optionally, perform an alignment operation on the second vocabulary according to the tokens corresponding to each probability value in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of the student model.

[0108] Optionally, after obtaining the second probability matrix, an alignment operation can be performed on the second vocabulary according to the tokens corresponding to each probability value in the second probability matrix and the mapping relationship between the tokens in the first vocabulary and the second vocabulary to obtain a third vocabulary. The specific implementation steps can refer to step 130 and will not be elaborated here.

[0109] Step 550: Distill and train the student model according to the third vocabulary and the second probability matrix to obtain the target language model.

[0110] Optionally, after obtaining the third vocabulary, the sample text can be input into the student model, and the student model predicts the probability values of each data unit in the sample text belonging to each token in the third vocabulary through forward propagation inference to form a third probability matrix corresponding to each data unit; and reconstruct the second probability matrix according to the structure and quantity of the third vocabulary to obtain a reconstructed probability matrix associated with the third probability matrix; then, use the third probability matrix and the reconstructed probability matrix to calculate the difference between the output of the student model and the output of the teacher model to obtain the first loss; use the difference between the third probability matrix and the token labels corresponding to each data unit to obtain the second loss; then, jointly update the student model using the first loss and the second loss through the backpropagation algorithm to obtain a target language model that can perform text processing quickly and with high performance. The specific implementation steps can refer to Step 140 and will not be elaborated here.

[0111] The method provided in this embodiment first predicts the first probability matrix of each data unit in the sample text based on the teacher model, and compresses it according to the probability value size to obtain the second probability matrix; when the training instruction of the student model is not received, the data unit, the probability value in the second probability matrix, and the third index encoding are stored on the disk to reduce the storage space requirement; when the training instruction of the student model is received, the encoded data is read from the disk and the data is restored using the corresponding decoding method, then the third vocabulary is obtained by aligning the second vocabulary according to the third index, and finally the student model is distilled and trained in combination with the third vocabulary and the second probability matrix to obtain the target language model. Throughout the training process, it not only decouples the inference of the teacher model and the training of the student model, but also reduces the memory requirement and improves the training efficiency, and enables the target language model trained accordingly to better adapt to different model architectures and language processing scenarios while maintaining high performance.

[0112] The following specifically describes the effectiveness of the method provided in this embodiment with experimental comparison results.

[0113] During the experimental comparison process, to ensure the fairness of the comparison, the student model (i.e., the target language model) trained by the method provided in this embodiment was comprehensively compared with the student model trained by traditional hard labels in the text generation task on the basis of maintaining the same compression rate. The experimental results show that at the same compression rate, the bilingual evaluation substitution index of the student model trained by the method provided in this embodiment in the text generation task is improved by 12.7% compared with the student model trained by traditional hard labels, indicating that the student model trained by the method provided in this embodiment can generate higher-quality text content when generating text.

[0114] Meanwhile, in terms of storage cost, the storage cost of the method provided in this embodiment can be reduced to 1 / 100 of the traditional hard label training scheme.

[0115] As can be seen from the above, compared with the prior art, the method provided in this embodiment significantly reduces the space required for model storage while improving the model performance, and also improves the efficiency and applicable scope of distillation training.

[0116] Figure 6 is a schematic flowchart of the text processing method provided by the present invention. As Figure 6 shown, the method includes step 610, step 620, and step 630.

[0117] Step 610, obtain the text to be processed; Step 620, based on the target language model, perform token prediction on each data unit in the text to be processed, and obtain the token prediction result corresponding to the text to be processed; Step 630, perform text processing on the text to be processed according to the token prediction result; wherein, the text processing includes text generation, code generation, machine translation, or text classification; the target language model is trained based on a language model training method.

[0118] The text to be processed here can be the text that needs to be processed, which can be the text corpus required for various text processing tasks, such as news articles in text classification tasks. The text to be processed here can be obtained by file reading or user input, etc.

[0119] Optionally, when performing text processing, the target language model can be first trained using the training method as Figure 1 shown, so as to use the target language model to analyze each data unit in the text, predict the token corresponding to each data unit, and obtain the token prediction result; then, use the token prediction result to perform text processing tasks such as text generation, code generation, machine translation, or text classification on the text to be processed, so as to realize the intelligent processing of the text to be processed.

[0120] For example, in the text generation task, new text content can be gradually generated specifically according to the token prediction result. In the code generation task, code segments that conform to grammar and logic can be generated specifically according to the token prediction result; in the machine translation task, the token prediction result of the source language can be converted into the token of the target language, thereby generating the translated text. In the text classification task, the category to which the text belongs, such as news category, sentiment analysis category, etc., can be determined according to the token prediction result.

[0121] The method provided in this embodiment uses a sparsified compression probability matrix and word table alignment to distill and train a lightweight target language model for token prediction, and realizes intelligent processing of tasks such as text generation, code generation, machine translation, or text classification based on the prediction results, significantly improving the accuracy, efficiency, and adaptability of text processing.

[0122] The language model training device provided by the present invention will be described below. The language model training device described below can be mutually referred to the language model training method described above.

[0123] Figure 7 is a schematic structural diagram of the language model training device provided by the present invention; as Figure 7 shown, the device includes: The first prediction unit 710 is used to predict the first probability matrix corresponding to each data unit in the sample text based on the teacher model; the first probability matrix includes the probability values of each data unit belonging to each token in the first word table, and the first word table is the word table of the teacher model; The compression unit 720 is used to compress the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain the second probability matrix corresponding to each data unit; The mapping unit 730 is used to perform an alignment operation on the second word table according to the tokens corresponding to the probability values in the second probability matrix to obtain the third word table; the second word table is the word table of the student model; The training unit 740 is used to distill and train the student model according to the third word table and the second probability matrix to obtain the target language model.

[0124] The device provided in this embodiment predicts the first probability matrix corresponding to each data unit in the sample text based on the teacher model to obtain the full probability distribution of each data unit, and performs sparse adaptive compression on the first probability matrix according to the magnitudes of the probability values in the first probability matrix to generate a second probability matrix with a lower storage capacity, effectively reducing the data volume and storage requirements; then, using the tokens corresponding to the probability values in the second probability matrix, an alignment operation is performed on the word table of the student model to obtain the third word table, so as to solve the problem of inconsistent word tables between the teacher model and the student model through dynamic word table mapping, and avoid the problem of knowledge transfer failure caused by word table differences. Finally, the student model is distilled and trained according to the third word table and the second probability matrix to obtain the target language model. Since distillation training is performed through sparsified compression of the probability matrix and word table alignment, not only the problem of excessively high storage costs caused by the growth of the word table scale is significantly reduced, the distillation training efficiency is effectively improved, but also the target language model trained accordingly can better adapt to different model architectures and text processing scenarios while maintaining high performance.

[0125] In some embodiments, the compression unit is specifically configured to: receive user input information, and determine a target quantity according to the compression mode in the user input information; sort the probability values in the first probability matrix in descending order according to the numerical size; in the first probability matrix, select the probability values of the target quantity with the front sorting positions to construct the second probability matrix.

[0126] In some embodiments, the compression unit is further configured to: when the compression mode is an adaptive compression mode, sequentially perform cumulative calculation on the probability values in the first probability matrix according to the descending order sorting result until the cumulative calculation value is greater than or equal to a preset threshold, and determine the target quantity according to the number of probability values participating in the cumulative calculation; when the compression mode is a fixed compression mode, determine the preset quantity as the target quantity.

[0127] In some embodiments, the training unit is specifically configured to: construct an empty matrix according to the number of tokens in the third vocabulary; fill the probability values in the second probability matrix into the empty matrix according to the first indexes of the tokens corresponding to the probability values in the second probability matrix to obtain the first reconstruction probability matrix corresponding to each data unit; the first index is the index of the token corresponding to each probability value in the second probability matrix in the third vocabulary; perform normalization processing on the probability values in the first reconstruction probability matrix to obtain the second reconstruction probability matrix corresponding to each data unit; predict the third probability matrix corresponding to each data unit based on the student model; the third probability matrix includes the probability values of each data unit belonging to each token in the third vocabulary; perform distillation training on the student model according to the second reconstruction probability matrix, the third probability matrix, and the token labels corresponding to each data unit to obtain the target language model.

[0128] In some embodiments, the mapping unit is specifically configured to: obtain each first probability value and each second probability value in the second probability matrix; there is no mapping relationship between the token corresponding to the first probability value and all tokens in the second vocabulary; there is a mapping relationship between the token corresponding to the second probability value and one token in the second vocabulary; map and generate the mapping indexes of the tokens corresponding to each first probability value based on the tokenizer of the student model; determine the second index of the target token having a mapping relationship with the token corresponding to each second probability value as the mapping index of the token corresponding to each second probability value; the second index is the index of the target token in the second vocabulary; fill the tokens corresponding to each first probability value into an empty vocabulary according to the mapping indexes of the tokens corresponding to each first probability value, and fill the tokens corresponding to each second probability value into the empty vocabulary according to the mapping indexes of the tokens corresponding to each second probability value; obtain the third vocabulary according to the filling result.

[0129] In some embodiments, the device further includes an encoding unit and a decoding unit; the encoding unit is specifically configured to, when the second probability matrix is obtained and the training instruction of the student model is not received, encode and store each of the data units, each probability value in the second probability matrix, and a third index of a token corresponding to each probability value in the second probability matrix to a disk; the third index is an index of each probability value in the second probability matrix in the first vocabulary; the decoding unit is specifically configured to, when the training instruction of the student model is received, parse and obtain each of the data units, each probability value in the second probability matrix, and the third index in the disk.

[0130] In some embodiments, the encoding unit is further configured to: perform quantization encoding on each probability value in the second probability matrix to obtain a first encoding result; perform quantization encoding on the third index to obtain a second encoding result; perform hash encoding on each of the data units to obtain a third encoding result; and store the first encoding result, the second encoding result, and the third encoding result in the disk in the form of a compressed file.

[0131] In some embodiments, the decoding unit is further configured to: load the compressed file in the disk; perform inverse quantization decoding on the first encoding result in the compressed file to obtain each probability value in the second probability matrix; perform inverse quantization decoding on the second encoding result in the compressed file to obtain the third index; and perform a hash reverse look-up operation on the third encoding result in the compressed file to obtain each of the data units.

[0132] Figure 8 is a schematic structural diagram of a text processing device provided by the present invention; as Figure 8 shown, the device includes: An obtaining unit 810 is configured to obtain a text to be processed; A second prediction unit 820 is configured to perform token prediction on each data unit in the text to be processed based on a target language model to obtain a token prediction result corresponding to the text to be processed; A processing unit 830 is configured to perform text processing on the text to be processed according to the token prediction result; wherein the text processing includes text generation, code generation, machine translation, or text classification; the target language model is trained based on a language model training method.

[0133] The device provided in this embodiment predicts tokens by using a lightweight target language model that performs distillation training through sparse compression of the probability matrix and word table alignment, and implements intelligent processing of tasks such as text generation, code generation, machine translation, or text classification based on the prediction results, significantly improving the accuracy, efficiency, and adaptability of text processing.

[0134] The device provided by the present invention is used to execute the above-mentioned method embodiments. For the specific process and detailed content, please refer to the above embodiments and will not be elaborated here.

[0135] Figure 9 An example of the physical structure diagram of an electronic device is shown as Figure 9 As shown, the electronic device may include: a processor 910, a communication interface 920, a memory 930, and a communication bus 940. Among them, the processor 910, the communication interface 920, and the memory 930 complete mutual communication through the communication bus 940. The processor 910 can call the logical instructions in the memory 930 to execute the language model training method, which includes: predicting a first probability matrix corresponding to each data unit in the sample text based on a teacher model; the first probability matrix includes the probability values of each data unit belonging to each token in a first word table, and the first word table is the word table of the teacher model; compressing the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain a second probability matrix corresponding to each data unit; performing an alignment operation on a second word table according to the tokens corresponding to the probability values in the second probability matrix to obtain a third word table; the second word table is the word table of a student model; performing distillation training on the student model according to the third word table and the second probability matrix to obtain a target language model, or a text processing method, which includes: obtaining a text to be processed; predicting tokens for each data unit in the text to be processed based on the target language model to obtain a token prediction result corresponding to the text to be processed; performing text processing on the text to be processed according to the token prediction result; where the text processing includes text generation, code generation, machine translation, or text classification; the target language model is obtained by training based on the language model training method.

[0136] In addition, when the logical instructions in the above-mentioned memory 930 are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), magnetic disks, or optical discs.

[0137] On the other hand, the present invention also provides a computer program product. The computer program product includes a computer program that can be stored on a non-transitory computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the language model training method provided by the above-mentioned various methods. The method includes: based on a teacher model, predicting a first probability matrix corresponding to each data unit in a sample text; the first probability matrix includes probability values of each data unit belonging to each token in a first vocabulary, and the first vocabulary is the vocabulary of the teacher model; according to the numerical magnitudes of the probability values in the first probability matrix, compressing the first probability matrix to obtain a second probability matrix corresponding to each data unit; according to the tokens corresponding to the probability values in the second probability matrix, performing an alignment operation on a second vocabulary to obtain a third vocabulary; the second vocabulary is the vocabulary of a student model; according to the third vocabulary and the second probability matrix, performing distillation training on the student model to obtain a target language model, or a text processing method, which includes: obtaining a text to be processed; based on the target language model, performing token prediction on each data unit in the text to be processed to obtain a token prediction result corresponding to the text to be processed; according to the token prediction result, performing text processing on the text to be processed; wherein the text processing includes text generation, code generation, machine translation, or text classification; the target language model is trained based on the language model training method.

[0138] In another aspect, the present invention further provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the language model training method provided by the above-mentioned various methods. The method includes: predicting, based on a teacher model, a first probability matrix corresponding to each data unit in a sample text; the first probability matrix includes probability values of each data unit belonging to each token in a first vocabulary, and the first vocabulary is the vocabulary of the teacher model; compressing the first probability matrix according to the numerical magnitudes of the probability values in the first probability matrix to obtain a second probability matrix corresponding to each data unit; performing an alignment operation on a second vocabulary according to the tokens corresponding to the probability values in the second probability matrix to obtain a third vocabulary; the second vocabulary is the vocabulary of a student model; training the student model by distillation according to the third vocabulary and the second probability matrix to obtain a target language model, or a text processing method, which includes: obtaining a text to be processed; predicting tokens for each data unit in the text to be processed based on the target language model to obtain a token prediction result corresponding to the text to be processed; performing text processing on the text to be processed according to the token prediction result; wherein the text processing includes text generation, code generation, machine translation, or text classification; the target language model is trained by the language model training method.

[0139] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0140] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence, or the part that contributes to the prior art can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0141] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A language model training method, characterized in that: include: Based on the teacher model, predict the first probability matrix corresponding to each data unit in the sample text; The first probability matrix includes probability values ​​of each data unit belonging to each word in a first vocabulary, and the first vocabulary is a vocabulary of the teacher model; Compressing the first probability matrix according to the numerical values ​​of each probability value in the first probability matrix to obtain a second probability matrix corresponding to each of the data units; According to the word elements corresponding to the probability values ​​in the second probability matrix, the second word list is aligned to obtain a third word list; the second word list is the word list of the student model; The student model is subjected to distillation training according to the third vocabulary and the second probability matrix to obtain a target language model.

2. The language model training method according to claim 1, characterized in that: The compressing the first probability matrix according to the numerical values ​​of each probability value in the first probability matrix to obtain a second probability matrix corresponding to each data unit includes: receiving user input information, and determining a target quantity according to a compression mode in the user input information; According to the numerical values, the probability values ​​in the first probability matrix are sorted in descending order; In the first probability matrix, the probability values ​​of the target quantities with higher sorting positions are selected to construct the second probability matrix.

3. The language model training method according to claim 2, characterized in that: The step of determining the target quantity according to the compression mode in the user input information includes: When the compression mode is an adaptive compression mode, the probability values ​​in the first probability matrix are cumulatively calculated in descending order until the cumulative calculation value is greater than or equal to a preset threshold value, and the target number is determined according to the number of probability values ​​involved in the cumulative calculation; When the compression mode is a fixed compression mode, a preset number is determined as the target number.

4. The language model training method according to any one of claims 1 to 3, characterized in that: The step of performing distillation training on the student model according to the third vocabulary and the second probability matrix to obtain a target language model includes: Constructing an empty matrix according to the number of word-units in the third vocabulary; According to the first index of the word element corresponding to each probability value in the second probability matrix, each probability value in the second probability matrix is ​​filled into the empty matrix to obtain a first reconstructed probability matrix corresponding to each data unit; the first index is the index of the word element corresponding to each probability value in the second probability matrix in the third vocabulary; Normalizing each probability value in the first reconstruction probability matrix to obtain a second reconstruction probability matrix corresponding to each of the data units; Based on the student model, predicting a third probability matrix corresponding to each of the data units; the third probability matrix includes probability values ​​of each of the data units belonging to each word in the third vocabulary; According to the second reconstruction probability matrix, the third probability matrix, and the word-unit label corresponding to each of the data units, the student model is distilled and trained to obtain the target language model.

5. The language model training method according to any one of claims 1 to 3, characterized in that: The step of performing an alignment operation on the second vocabulary according to the word elements corresponding to the probability values ​​in the second probability matrix to obtain a third vocabulary includes: In the second probability matrix, each first probability value and each second probability value are obtained; there is no mapping relationship between the word element corresponding to the first probability value and all the word elements in the second vocabulary; there is a mapping relationship between the word element corresponding to the second probability value and a word element in the second vocabulary; A word segmenter based on the student model maps and generates a mapping index of the word element corresponding to each of the first probability values; Determine the second index of the target word element that has a mapping relationship with the word element corresponding to each of the second probability values ​​as the mapping index of the word element corresponding to each of the second probability values; the second index is the index of the target word element in the second vocabulary; According to the mapping index of the word-gram corresponding to each first probability value, fill the word-gram corresponding to each first probability value into the empty word list, and according to the mapping index of the word-gram corresponding to each second probability value, fill the word-gram corresponding to each second probability value into the empty word list; According to the filling result, the third vocabulary is obtained.

6. The language model training method according to any one of claims 1 to 3, characterized in that: The method further comprises: When the second probability matrix is ​​obtained and the training instruction of the student model is not received, each of the data units, each probability value in the second probability matrix, and a third index of the word element corresponding to each probability value in the second probability matrix are encoded and stored on disk; the third index is the index of each probability value in the second probability matrix in the first vocabulary; When the training instruction of the student model is received, each of the data units, each probability value in the second probability matrix, and the third index are parsed and obtained in the disk.

7. The language model training method according to claim 6, characterized in that: The encoding and storing each of the data units, each probability value in the second probability matrix, and a third index of a word element corresponding to each probability value in the second probability matrix to a disk includes: quantize and encode each probability value in the second probability matrix to obtain a first encoding result; Performing quantization encoding on the third index to obtain a second encoding result; Performing hash coding on each of the data units to obtain a third coding result; The first encoding result, the second encoding result and the third encoding result are stored in the disk in the form of a compressed file.

8. The language model training method according to claim 7, characterized in that: The step of parsing and obtaining each of the data units, each probability value in the second probability matrix, and the third index in the disk includes: Loading the compressed file in the disk; Performing inverse quantization decoding on the first encoding result in the compressed file to obtain each probability value in the second probability matrix; Performing inverse quantization decoding on the second encoding result in the compressed file to obtain the third index; A hash reverse table lookup operation is performed on the third encoding result in the compressed file to obtain each of the data units.

9. A text processing method, characterized in that: include: Get the text to be processed; Based on the target language model, word unit prediction is performed on each data unit in the text to be processed to obtain a word unit prediction result corresponding to the text to be processed; According to the word unit prediction result, performing text processing on the text to be processed; The text processing includes text generation, code generation, machine translation or text classification; the target language model is trained based on the language model training method according to any one of claims 1 to 8.

10. A language model training device, characterized in that: include: A first prediction unit, used to predict a first probability matrix corresponding to each data unit in the sample text based on the teacher model; The first probability matrix includes probability values ​​of each data unit belonging to each word in a first vocabulary, and the first vocabulary is a vocabulary of the teacher model; a compression unit, configured to compress the first probability matrix according to the numerical values ​​of each probability value in the first probability matrix to obtain a second probability matrix corresponding to each of the data units; A mapping unit, configured to perform an alignment operation on the second vocabulary according to the word elements corresponding to the probability values ​​in the second probability matrix to obtain a third vocabulary; the second vocabulary is a vocabulary of the student model; A training unit is used to perform distillation training on the student model according to the third vocabulary and the second probability matrix to obtain a target language model.

11. A text processing device, characterized in that: include: An acquisition unit, used for acquiring text to be processed; A second prediction unit is used to perform word unit prediction on each data unit in the text to be processed based on the target language model to obtain a word unit prediction result corresponding to the text to be processed; A processing unit, configured to perform text processing on the text to be processed according to the word unit prediction result; The text processing includes text generation, code generation, machine translation or text classification; the target language model is trained based on the language model training method according to any one of claims 1 to 8.

12. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the computer program, the language model training method according to any one of claims 1 to 8 is implemented, or the text processing method according to claim 9 is implemented.

13. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the language model training method according to any one of claims 1 to 8 is implemented, or the text processing method according to claim 9 is implemented.

Citation Information

Patent Citations

  • Model training method and device based on knowledge distillation, and electronic equipment

    CN113837308A

  • Model training method and device, natural language processing method and device and storage medium

    CN117216544A

  • Word list construction method and device, storage medium and electronic equipment

    CN118133818A

  • Power load accurate prediction system based on artificial intelligence

    CN119250259A

  • Model distillation method, apparatus, medium, apparatus and computer program product

    CN119378646A

Cited By

  • Neural network calculation method and device, electronic equipment and storage medium

    CN121070445A

  • Computing method and device of neural network, electronic equipment and storage medium

    CN121070445B

  • Inference model training method and device, equipment, medium and product

    CN122154840A

  • Psychological pollution name classification method based on multi-agent self-correction and thinking chain distillation

    CN122196187A