Language model optimization method based on autoregression training and related equipment

By dynamically adjusting the proportion and quantity limits of real data and synthetic data in autoregressive training, combining distribution difference measurements, and optimizing language model training, the model distribution drift and crash problems are solved, and the stability and robustness of the model are improved.

CN120258047AActive Publication Date: 2025-07-04SHANG HAI JIE YUE XING CHEN ZHI NENG KE JI YOU XIAN GONG SI
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510741088.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-05
Publication Date
2025-07-04
Estimated Expiration
2045-06-05

AI Technical Summary

Technical Problem

Language models in existing autoregressive training are prone to distribution drift and collapse, resulting in a decrease in model output diversity and accuracy, making it difficult to maintain stability and robustness in multiple rounds of training.

Method used

By building an autoregressive training framework, the ratio of real data to synthetic data is dynamically adjusted, and the number of synthetic data is limited, and the training process of language models is optimized by combining distribution difference measurement and training algebraic control.

Benefits of technology

It effectively suppresses the distribution drift of the language model during multiple rounds of autoregressive training, prevents abnormal offset or degradation of the output distribution, ensures the controllability and performance reliability of model training, and improves the stability and robustness of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258047A_ABST
    Figure CN120258047A_ABST
Patent Text Reader

Abstract

The invention provides a language model optimization method based on autoregression training and related equipment, and the method comprises the steps: constructing an autoregression training framework of a language model, predicting the probability distribution of subsequent words based on an input historical context, and enabling the probability distribution in each generation of autoregression training to be used for generating the next generation of synthetic data; in each generation of training, real data and synthetic data generated by a previous generation of model are mixed according to a set proportion to serve as training data, and the proportion of the real data is dynamically reduced along with increase of training algebra; training a language model by using the mixed data, updating model parameters, and generating synthetic data required by next-generation training; after each generation of training, measuring the difference between the data distribution generated by the current model and the real data distribution, and dynamically adjusting the proportion of the real data according to the difference; the autoregression training process of the language model is continuously optimized through the steps. And the problem of distribution drift of the language model in the multi-round autoregression training process is effectively inhibited.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a method for optimizing a language model based on autoregressive training and related devices. Background Art

[0002] In recent years, with the wide application of Generative Pre-trained Models (GPT) and other large language models (LLMs), the proportion of synthetic data generated based on such models in the training corpus has gradually increased. Especially in natural language processing tasks such as text generation, machine translation, and dialogue systems, the language model has formed a loop of training - generation mechanism (Train - Generate Loop) by continuously generating new data and incorporating it into the training corpus.

[0003] However, this loop process brings potential performance degradation problems, namely the so - called "model collapse". Model collapse refers to the situation in autoregressive training where the model output gradually deviates from the true data distribution, and ultimately loses the diversity or accuracy of generating language. For example, after multiple iterations, the model may tend to generate limited and repetitive outputs, which reflects the loss of information in the tail of the original distribution and seriously affects the actual application effect of the model.

[0004] To address this problem, existing technologies have attempted to reduce the negative impact of autoregressive training to a certain extent by introducing a mixed training strategy of real data and synthetic data. For example, it quantitatively describes the impact of the increase in the proportion of synthetic data on the model performance. A certain proportion of real data is introduced in each iteration of autoregressive training to slow down the distribution shift. However, the research results show that when the proportion of synthetic data exceeds a certain threshold, the model still inevitably collapses.

[0005] Existing technologies provide a certain theoretical basis and experimental analysis for understanding the model collapse phenomenon, but there are still many limitations in how to effectively avoid the model collapse problem caused by autoregressive training, making it difficult to fundamentally solve the model collapse problem in autoregressive training. Summary of the Invention

[0006] Aiming at the deficiencies of the existing technology, this application provides a method for optimizing a language model based on autoregressive training and related devices, at least to solve the problems that the model output distribution drifts severely and is prone to distribution collapse in existing autoregressive training, thereby leading to model degradation.

[0007] To achieve the above objectives and other advantages, some embodiments of the present application provide the following aspects: In a first aspect, some embodiments of the present application provide a method for optimizing a language model based on autoregressive training, including: Construct an autoregressive training framework for the language model, and predict the probability distribution of subsequent words based on the input historical context. Among them, the probability distribution in each generation of autoregressive training is used to generate synthetic data for the next generation; In each generation of training, mix real data and synthetic data generated by the previous generation model as training data according to a set ratio. Among them, the proportion of real data in the training data decreases dynamically as the number of training generations increases; Use the mixed data to train the language model, update the model parameters, and generate synthetic data required for the next generation of training. Among them, limit the number of the synthetic data not to exceed a set synthetic data ratio control condition relative to the number of the real data; After each generation of training, measure the difference between the data distribution generated by the current model and the real data distribution, and dynamically adjust the proportion of the real data according to the difference; Based on the synthetic data generated by the probability distribution, dynamic data ratio control, and synthetic data quantity constraint, continuously optimize the autoregressive training process of the language model.

[0008] In a second aspect, some embodiments of the present application further provide an electronic device, which includes: One or more processors; and a memory storing computer program instructions, and when the computer program instructions are executed, the processors execute the method for optimizing a language model based on autoregressive training as described in any one of the above.

[0009] In a third aspect, some embodiments of the present application further provide a computer-readable storage medium, on which computer programs and / or instructions are stored, and when the computer programs and / or instructions are executed by a processor, the method for optimizing a language model based on autoregressive training as described in any one of the above is implemented.

[0010] In a fourth aspect, some embodiments of the present application further provide a computer program product, including computer programs and / or instructions, and when the computer programs / instructions are executed by a processor, the method for optimizing a language model based on autoregressive training as described in any one of the above is implemented.

[0011] Compared with related technologies, in the solution provided by the embodiments of the present application, by introducing a control strategy of dynamically adjusting the ratio between real data and synthetic data based on the number of training generations and setting an upper limit on the number of synthetic data in each generation of training, the adaptive data balance of the model in different training stages is achieved, effectively reducing the risk that synthetic data overly affects model training. This can effectively suppress the distribution drift problem generated during the multi-round autoregressive training process of the language model, prevent the output distribution of the language model from abnormally deviating or degenerating into a single-category output, ensure the controllability of the model training process and the performance reliability of the final model, and significantly improve the stability and robustness of model training. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other embodiments can be obtained based on these drawings.

[0013] Figure 1 is a flowchart showing an optimization method for a language model based on autoregressive training provided by an embodiment of the present application; Figure 2 is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments of the present application belong to the scope of protection of the present application.

[0015] First Embodiment The first embodiment of the present application relates to an optimization method for a language model based on autoregressive training. Referring to Figure 1 as shown, the method may include the following steps: Step S1: Construct an autoregressive training framework for the language model, and predict the probability distribution of subsequent words based on the input historical context. Among them, the probability distribution in each generation of autoregressive training is used to generate synthetic data for the next generation.

[0016] Regarding step S1, specifically, the language model is the core basic model in natural language processing. It is a model used to model the probability relationship between contexts in a sequence of words or terms. Its goal is to predict the probability distribution of the next word under the current context conditions. Common language models include RNN language models such as LSTM and GRU, and language models based on self-attention mechanisms such as Transformer, etc.

[0017] Autoregressive training means decomposing a sequence prediction task into multiple prediction subtasks that progress sequentially in time. The model uses the historical known context information at each step to predict the next word and takes the current output as part of the input for the next step. In this framework, the model predicts the probability of the next word based on the historical context and uses the prediction result as the input for the next step, continuously iterating for training.

[0018] In each generation of training, after the model is trained based on real data and synthetic data, it will form an understanding of the language distribution. This understanding is reflected in the probability distribution generated by the model, that is, each context input corresponds to a probability distribution of candidate words. These probability distributions can be used to sample and generate the next generation of synthetic data, compare with the real distribution, conduct drift metric analysis, and then control the risk of model collapse.

[0019] By constructing a language model with an autoregressive structure, it is possible to achieve the generation-by-generation iteration and optimization of the generation probability in language model training, providing support for subsequent synthetic data generation.

[0020] Step S2: In each generation of training, mix the real data and the synthetic data generated by the previous generation model according to a set ratio as the training data. Among them, the proportion of real data in the training data decreases dynamically as the number of training generations increases.

[0021] Regarding step S2, specifically, in each generation of training process of the model, mix the statically collected real corpus data and the synthetic text data generated by the previous generation language model according to a preset ratio to form the training data input of this generation. Among them, the proportion of real data is denoted as α m , and this proportion value decreases dynamically as the number of training generations m increases, gradually guiding the language model from the training stage mainly based on real data to the self-supervised generation training stage mainly based on synthetic data.

[0022] By dynamically regulating the proportion of real data in the mixed data, it is possible to effectively alleviate the distribution drift problem in the training process, improve the stability and robustness of model training, as well as the diversity and generalization ability of the final generation results.

[0023] Step S3: Train a language model using the mixed data, update the model parameters, and generate synthetic data required for the next generation of training, where the number of synthetic data is restricted not to exceed a set synthetic data ratio control condition relative to the number of real data.

[0024] Regarding step S3, specifically, in each generation of training, the training data mixed and constructed in step S2 is input into the language model, and a model training process based on an autoregressive structure is executed, and the model parameters are updated based on an optimization objective function (such as cross-entropy loss). After the current generation of training is completed, based on the updated model, by inputting the historical context sequence, several target word sequences are generated using an autoregressive prediction mechanism as the synthetic data generated by the current generation of the model.

[0025] To prevent the synthetic data from overly occupying the proportion during the training process, resulting in the model's generated results deviating from the real corpus distribution or experiencing distribution collapse, it is necessary to set an upper limit constraint on the number of synthetic data to form a synthetic data ratio control condition. Specifically, based on the total amount of real data samples and the size of the vocabulary used in the current training, the upper limit value of the allowed synthetic data generation is dynamically calculated, thereby controlling the influence degree of the synthetic data on the training process, and avoiding problems such as model distribution drift or output degradation caused by too high a proportion of synthetic data, and ensuring the stability of the training process and the effectiveness of the generated data.

[0026] Step S4: After each generation of training, measure the difference between the data distribution generated by the current model and the real data distribution, and dynamically adjust the proportion of the real data according to the difference.

[0027] Regarding step S3, specifically, the distribution difference refers to the statistical deviation between the data distribution generated by the model in the current generation and the real data distribution, and the L1 norm can be used as a metric for the distribution difference. By introducing a distribution difference measurement and feedback adjustment mechanism based on the L1 norm, the degree of fit between the model's generation ability and the data distribution can be continuously monitored during the model training process. If the generated distribution deviates too much from the real distribution, the system automatically triggers a real data proportion adjustment mechanism to increase the proportion of real data for correction, guiding the training back to a reasonable distribution range, thereby preventing the model from falling into a state of output degradation or collapse.

[0028] Step S5: Continuously optimize the autoregressive training process of the language model based on the synthetic data generated by the probability distribution, dynamic data ratio control, and synthetic data quantity constraint.

[0029] Specifically for step S5, after completing the sub - processes of each generation of training, the system continuously generates new synthetic data based on the conditional probability distribution generated by the current - generation model; meanwhile, an adaptive adjustment of the mixing ratio between real data and synthetic data is carried out by adopting a real - data proportion control strategy based on the number of training generations; and in combination with the synthetic - data proportion control condition, an upper limit is imposed on the total amount of synthetic data that can be used in each generation of training to prevent the risk of distribution collapse during model training. The above - mentioned mechanism forms a recursively optimized training loop, enabling the model to achieve a stable and controllable transition from real - data - driven to synthetic - data - dominated in each generation of training, improving the model training efficiency while ensuring the stability and generalization ability of the model output distribution.

[0030] It is not difficult to find that, compared with the related technologies, in the solution provided by the embodiments of the present application, by introducing a strategy of dynamically adjusting the ratio between real data and synthetic data based on the number of training generations and adopting a control strategy of setting an upper limit on the number of synthetic data in each generation of training, an adaptive data balance of the model in different training stages is achieved, effectively reducing the risk that synthetic data overly affects model training. This can effectively suppress the distribution drift problem generated during the multi - round autoregressive training process of the language model, prevent the output distribution of the language model from abnormally deviating or degenerating into a single - category output, ensure the controllability of the model training process and the performance reliability of the final model, and significantly improve the stability and robustness of model training.

[0031] Second Embodiment The second embodiment of the present application relates to an optimization method for a language model based on autoregressive training. The second embodiment is an improvement based on the first embodiment. The specific improvement lies in that: in the second embodiment of the present application, a specific implementation manner of dynamically adjusting the proportion of real data based on the number of training generations is provided. That is, the process in which the proportion of real data in the training data dynamically decreases as the number of training generations increases is as follows: As the number of training generations increases, the proportion of real data decreases according to a preset decreasing strategy; the decreasing strategy is specifically: the proportion of real data in the initial training stage is close to 1, and as the number of training generations increases, the proportion of real data gradually decreases in a manner inversely proportional to the logarithmic function of the number of training generations, and always does not exceed 1, and the decreasing rate of the proportion of real data is determined by a control parameter set according to the scale of the training data and the quality of the generated data.

[0032] Specifically, in the process of each generation of training, the training data is composed of a mixture of real data and synthetic data. For example, in the m - th generation of training, the training data is composed of a mixture of real data and synthetic data generated by the previous - generation model, and their proportions are α m and 1 - α m , forming the following combination method:

[0033] Among them, D real represents the real data; D synthetic represents the synthetic data generated by the previous-generation model.

[0034] The proportion of real data in the training data decreases dynamically as the number of training epochs increases, thus constituting an adaptable evolution decreasing strategy. Exemplarily, the decreasing strategy for the adaptable evolution of real data satisfies the following calculation formula:

[0035] Among them, α m represents the proportion of real data in the m-th generation of training; β is a control parameter, which is related to the sample size of real data and the quality evaluation result of synthetic data; m is the number of epochs in the current recursive training process.

[0036] The specific process of dynamically adjusting the proportion of real data is as follows: in the initial training stage, the proportion of real data is close to 1 to ensure that the model establishes stable language representation ability in the early stage; as the number of training epochs gradually increases, the proportion of real data decreases in the way of the above logarithmic function, realizing a natural transition from real-data dominance to synthetic-data dominance in the training process; at the same time, by setting the control parameter β, the decreasing rate can be adaptively adjusted according to different data scales and task scenarios to achieve fine-grained control of the autoregressive training process of the model.

[0037] It should be noted that the control parameter is a key factor for adjusting the decreasing speed of the proportion of real data, and its value is associated with the sample size of real data and the quality evaluation result of synthetic data in the current training task. When the real data sample is sufficient or the quality of synthetic data is high, β can be set to a smaller value to accelerate the transition of the model to self-supervised training; otherwise, the value of β needs to be increased to extend the training stage dominated by real data to ensure the stability and reliability of the initial stage of model training.

[0038] It is not difficult to find that in the solution provided by the embodiment of the present application, the above improvement enables the model to balance stability and efficiency throughout the training process, significantly reducing risks such as distribution drift and model collapse caused by synthetic data dominating too early, and thus improving the generalization ability and output quality of the language model.

[0039] The Third Embodiment The third embodiment of the present application relates to an optimization method for a language model based on autoregressive training. The third embodiment is an improvement on the basis of the first embodiment. The specific improvement lies in: in the third embodiment of the present application, a specific implementation manner of dynamically adjusting the proportion of real data based on distribution drift feedback control is provided. That is, step S4 can further include the following steps: Step S401: Calculate the L1 norm between the data distribution generated by the current-generation model and the initial true data distribution as a measure of distribution drift. The L1 norm is the sum of the absolute values of the probability differences between the two across all categories. Step S402: Compare the measure with a preset distribution deviation threshold. Step S403: If the measure exceeds the distribution deviation threshold, increase the proportion of true data in the next-generation training data and correspondingly reduce the proportion of synthetic data to reduce the risk of model distribution drift. Step S404: If the measure does not exceed the distribution deviation threshold, adjust the proportion of true data according to a preset decreasing strategy and continue the next-generation training.

[0040] Specifically, in autoregressive training, assume that the initial true data distribution is After m times of autoregressive training, the generations of the model can be expressed as:

[0041] where α m is the proportion of the current-generation true data in the mixed training data; is the estimated distribution of the (m - 1)-th generation model; is the synthetic data distribution finally generated by the current-generation model.

[0042] Based on the data distribution generated by the current-generation model and the initial true data distribution

[0043] calculate the L1 norm between the two as a measure of distribution drift. This norm is defined as the sum of the absolute values of the probability differences between the two distributions for each category over the entire vocabulary space s, i.e.: where ‖•‖1 is the L1 norm, representing the total difference between the two distributions. The larger the value, the more severely the two distributions deviate; is the generation probability of the k-th word in the m-th generation;

[0044] Compare the above L1 norm measure with the distribution deviation threshold set by the system to determine whether there is a large deviation in the current-generation model. The distribution deviation threshold is the maximum tolerable range preset by the system for determining whether the current-generation model deviates from the true semantic distribution. If the deviation value exceeds this threshold, trigger the true data proportion adjustment mechanism to guide the training back to a reasonable distribution range.

[0045] If the metric value exceeds the distribution deviation threshold, then in the next generation of training, increase the proportion α of real data m , and correspondingly reduce the proportion of synthetic data to suppress the trend of the expansion of distribution drift, so as to maintain the proximity between the model output distribution and the real corpus distribution. If the metric value does not exceed the distribution deviation threshold, it is considered that the current training state is within an acceptable range, and continue to adjust the proportion of real data according to the original decreasing strategy, and execute the next generation of training process.

[0046] Furthermore, the strategy for controlling the degree of distribution drift between the data distribution generated by the current generation model and the initial real data distribution includes: by setting a constant related to the scale of real data, and recording the number of training generations of the current autoregressive training, ensure that after each generation of training, the metric value of the distribution drift between the data distribution generated by the current generation model and the initial real data distribution shows a decreasing trend, so as to form a dynamic balance relationship between the number of training generations and distribution drift control during the model training process, and ensure that the distribution drift during the training process is always controllable.

[0047] Exemplarily, by setting a constant term C related to the scale of real data, and combining with the current training generation m, the following upper bound constraint of distribution drift is formed:

[0048] where represents the L1 norm deviation between the generated distribution of the current generation and the initial real distribution.

[0049] Through this control strategy, as the number of training generations increases, the theoretical upper limit of distribution drift converges step by step in the form of to form a dynamic drift suppression mechanism driven by the number of training generations. This mechanism can ensure that during the recursive iteration process of the model, the deviation between the data distribution generated by the current generation and the initial real data distribution is always within the theoretically controllable range. This strategy can be used in conjunction with the decreasing proportion of α m data to form a closed-loop control system during the recursive training process of the model, so as to effectively prevent the generated distribution from deviating from the real semantic distribution prematurely, and ensure that the model maintains stable convergence and semantic consistency during multiple rounds of training.

[0050] It is not difficult to find that in the solution provided by the embodiments of the present application, by introducing the L1 norm as a quantization index of distribution drift, the system can accurately judge whether the model deviates from the real semantic distribution; when the deviation is too large, increase the proportion of real data in time for correction to avoid problems such as distribution collapse or semantic degradation caused by synthetic data dominating the training. On the contrary, when the deviation is controllable, continue to promote the decrease of the proportion of real data, which helps to improve the training efficiency and enhance the self-learning ability of the model.

[0051] It should be noted that the third embodiment of this application can also be an improvement based on any one or more of the first embodiment to the second embodiment.

[0052] Fourth Embodiment The fourth embodiment of this application relates to an optimization method for a language model based on autoregressive training. The fourth embodiment is an improvement based on the first embodiment. The specific improvement lies in that: in the fourth embodiment of this application, a specific implementation manner for controlling the limited number of synthetic data is provided. That is, the condition for restricting the number of synthetic data relative to the number of real data not to exceed a set synthetic data ratio control condition is specifically: The number of synthetic data needs to satisfy not exceeding a dynamic upper limit jointly determined by the number of samples of real data and the vocabulary size of the language model, that is, as the number of training data samples increases, the number of allowed synthetic data can increase accordingly, and at the same time, as the vocabulary size expands, the allowed number of synthetic data is reduced to prevent the risk of probability distribution collapse during the model training process due to the excessive proportion of synthetic data.

[0053] Specifically, in each generation of training process, this embodiment restricts the number of synthetic data generated by the current generation model so that it does not exceed a dynamic upper limit jointly determined by the total amount of training data samples and the vocabulary size of the language model. Exemplarily, the constraint condition for the number of synthetic data can be expressed as:

[0054] where n represents the number of samples of synthetic data generated in the current generation; N is the total amount of training data samples; s is the vocabulary size of the language model.

[0055] According to this constraint condition, it can be seen that the upper limit n of synthetic data is proportional to the total amount of training data N. As the total amount of training data samples N increases, the system can appropriately increase the upper limit of generating synthetic data, thereby improving the training efficiency and self-supervised ability. At the same time, the upper limit n of synthetic data is inversely proportional to the logarithm of the vocabulary size s, indicating that as the vocabulary size of the language model expands, its generated semantic space becomes sparser and more complex. To prevent the rapid deviation of the semantic distribution, it is necessary to impose more strict restrictions on the number of synthetic data generated to avoid problems such as mode collapse and semantic degradation of the model in the high-dimensional space.

[0056] It is not difficult to find that in the solution provided by the embodiment of this application, by restricting this proportional relationship, it is possible to effectively prevent the proportion of synthetic data in the training data from being too high, avoid the deviation of the model learning process from the semantic distribution of real data, and thus prevent the risk of model collapse or degradation of generation quality.

[0057] It should be noted that the fourth embodiment of this application can also be an improvement based on any one or more of the first to third embodiments.

[0058] Fifth Embodiment The fifth embodiment of this application relates to a method for optimizing a language model based on autoregressive training. The fifth embodiment is an improvement based on the first embodiment. The specific improvement lies in: in the fifth embodiment of this application, a specific implementation manner is provided for extending the effective training cycle of the model by setting training stability conditions to avoid model collapse during model training. That is, it includes the following steps: In each generation of training, according to the diversity of the synthetic data generated by the current model and the real data, combined with the dispersion metric value of the initial real data, dynamically adjust the proportion of real data to reduce the degradation rate of the synthetic data generated by the model and delay the speed at which the data distribution output by the model converges to a single distribution state.

[0059] Specifically, in each generation of training of this embodiment, by introducing a theoretical lower bound constraint mechanism for the model collapse time, to guide the dynamic adjustment of the proportion α of real data m The model collapse time T represents the number of generations when the model generation distribution converges to a single Dirac distribution (i.e., semantic degradation to a unique output). The lower bound of its model collapse time can be expressed as:

[0060] where T is the time of model collapse, that is, the number of training rounds when the model generation distribution degrades to a single-point distribution; S0 is the dispersion metric of the initial real data distribution; γ is the degradation rate of the generated data, and its magnitude is related to the real data ratio α m and the diversity of real data.

[0061] On the basis of introducing the theoretical lower bound constraint mechanism for the model collapse time, in each generation of training, evaluate the diversity difference between the synthetic data generated by the current language model and the real data; combined with the dispersion metric value S0 of the initial real data, dynamically adjust the proportion α of real data in the current generation of training data m ; if it is detected that the quality of the synthetic data degrades or the semantic concentration increases, then increase α m , enhance the guiding role of real data in training; by controlling the increasing trend of γ, achieve the goal of delaying the model collapse time T, thereby effectively extending the effective training cycle of the model in the diversity-maintaining state.

[0062] It is not difficult to find that in the solution provided by the embodiments of the present application, a real data proportion regulation mechanism oriented to training stability is constructed. While ensuring the model generation efficiency, it improves the model's ability to maintain semantic diversity and generation generalization ability, and effectively prevents the language model from falling into convergence collapse or content repetition problems in deep autoregressive training.

[0063] It should be noted that the fifth embodiment of the present application can also be an improvement based on any one or more of the first to fourth embodiments.

[0064] Sixth Embodiment The sixth embodiment of the present application relates to a method for optimizing a language model based on autoregressive training. The sixth embodiment is an improvement based on the first embodiment. The specific improvement lies in: in the sixth embodiment of the present application, a specific implementation manner of autoregressive training modeling is provided. That is, step S1 can further include the following steps: Step S101: The language model receives the context sequence of the training sample as input, and sets the vocabulary size, context length, and number of samples. Step S102: Based on the currently input context sequence, calculate the conditional probability distribution of all candidate words by taking the inner product of the model weight matrix and the context vector, and the conditional probability distribution is normalized by the softmax function. Step S103: Use the cross-entropy loss function as the objective function for language model training, and calculate the error between the conditional probability distribution and the actual target word label. Step S104: By optimizing the objective function, iteratively update the model parameters until the objective function reaches the convergence condition, and generate the synthetic data required for the next generation of training based on the conditional probability distribution result after the current generation of model training is completed.

[0065] Specifically, using the modeling method of the statistical language model, a conditional probability distribution framework for predicting the next word item based on the context is constructed. First, context modeling and input structure setting are performed. Set the vocabulary size of the language model to s, which represents the number of recognizable word items in the language model; at the same time, set the context length to ℓ, which represents the context window length for predicting the next word in each training sample. The possible combination number of the context is c. Each training sample consists of the context sequence X and the target word Y, where represents the discrete random variable of the historical context, represents the discrete random variable of the next generated word.

[0066] The conditional probability distribution between the context and the target word of the synthetic data is defined as:

[0067] During the autoregressive training process, the estimated values of these conditional probability distributions are denoted as , which can be obtained by optimizing the following objective function through a Softmax classifier, and its expression is as follows:

[0068] where M represents the number of training samples; is the weight matrix of the model; represents the context input vector of the -th sample; represents its corresponding target word; is the Softmax function, which represents converting the linear score into a conditional probability distribution.

[0069] During the autoregressive training process, the conditional probability distribution of the m-th generation model can be expressed as:

[0070] where represents the conditional probability of generating word k given context j in the m-th generation model; represents the weight vector of the k-th word in the m-th generation model; represents the embedding vector representation of context j; s represents the size of the vocabulary; the Softmax function ensures that the sum of the probabilities of all is 1.

[0071] To optimize the learning effect of the model, in this embodiment, the cross-entropy loss function is used as the training objective to minimize the difference between the probability distribution output by the model and the true target:

[0072] where is the weight matrix of the model; is the Softmax function, and are the sample context vector and the target word vector, respectively.

[0073] By minimizing this loss function, the fitting ability of the conditional probability distribution can be effectively improved. After the training objective function converges, the current generation of the language model has learned the conditional probability distribution for word item prediction. Based on this probability distribution, the model can generate output word items in two ways: one is to determine the most likely target word by selecting the maximum probability, and the other is to perform random sampling according to the probability values to enhance the generation diversity. Generate target word items from the vocabulary and concatenate them word by word to form a new sentence. The data thus generated can be used as synthetic data in the next generation of the training process and mixed with real data to form a new training set. Therefore, after each generation of training, the model can generate synthetic data for the next generation of training, constructing a closed-loop mechanism of prediction, generation, and retraining. Combining various adjustment strategies, the recursive evolution and self-supervised enhancement of the language model are achieved.

[0074] It is not difficult to find that in the solution provided by the embodiments of the present application, by performing an inner product operation based on the context vector and the word item weight matrix, constructing a conditional probability distribution over the entire vocabulary range, and normalizing with the Softmax function, the probability of generating words in the current context can be expressed more accurately, which helps to improve the model's context perception ability and prediction accuracy.

[0075] It should be noted that the sixth embodiment of the present application can also be an improvement based on any one or more of the first to fifth embodiments.

[0076] The step division of the above various methods is only for clear description. When implemented, they can be combined into one step or some steps can be split into multiple steps. As long as the same logical relationship is included, they are all within the protection scope of the present application; adding insignificant modifications to the algorithm or process or introducing insignificant designs, but not changing the core design of its algorithm and process, are all within the protection scope of this application.

[0077] In addition, some embodiments of the present application also provide an electronic device. The electronic device can be various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and so on. The electronic device can also be various forms of mobile devices, such as personal digital processors, cellular phones, smart phones, wearable devices, and other similar computing devices.

[0078] The electronic device includes: one or more processors; and a memory storing computer program instructions, which when executed cause the processors to execute a method for optimizing a language model based on autoregressive training provided by any one or more of the above embodiments. Figure 2An exemplary structural diagram of the electronic device is disclosed. The electronic device includes: one or more processors 1101, a memory 1102, and interfaces for connecting the components, including a high-speed interface and a low-speed interface. Each component is interconnected using different buses and can be mounted on a common motherboard or otherwise installed as required. The processor can process instructions executed within the electronic device, including instructions stored in the memory or on the memory to display graphical information of the GUI on an external input / output device (such as a display device coupled to the interface). In some other embodiments, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories if needed. Similarly, multiple electronic devices can be connected, and each device provides some necessary operations. Among them, the components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described herein and / or required herein.

[0079] The electronic device may further include: an input device 1103 and an output device 1104. The processor 1101, the memory 1102, the input device 1103, and the output device 1104 can be connected via a bus or other means. Figure 2 Taking connection via a bus as an example.

[0080] The input device 1103 can receive input digital or character information and generate key signal inputs related to the user settings and function controls of the electronic device, such as input devices like a touch screen, a keypad, a mouse, a trackpad, a touchpad, a pointing stick, one or more mouse buttons, a trackball, a joystick, etc. The output device 1104 can include a display device, an auxiliary lighting device (such as an LED), and a haptic feedback device (such as a vibration motor), etc. The display device can include, but is not limited to, a liquid crystal display, a light-emitting diode display, and a plasma display. In some embodiments, the display device can be a touch screen.

[0081] To provide interaction with the user, the electronic device can be a computer. The computer has: a display device for displaying information to the user (such as a cathode ray tube or an LCD monitor); and a keyboard and a pointing device (such as a mouse), through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (such as visual feedback, auditory feedback); and the input from the user can be received in any form (such as voice input or tactile input).

[0082] In the embodiments of the present application, a computer program / instruction is stored on a computer-readable medium. When the computer program / instruction is executed by a processor, it implements an optimization method for a language model based on autoregressive training provided by any one or more of the above embodiments. The computer-readable medium may be included in the electronic device described in the above embodiments; or it may exist separately without being assembled into the device. The above computer-readable medium carries one or more computer-readable instructions.

[0083] The memory 1102 can be used as a non-transitory computer-readable storage medium, and can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules. By running the non-transitory software programs, instructions, and modules stored in the memory 1102, the processor 1101 executes various functional applications and data processing of the server, so as to implement the program instructions / modules corresponding to the methods provided by any one or more of the above embodiments in the embodiments of the present application.

[0084] The memory 1102 may include a program storage area and a data storage area. Among them, the program storage area may store an operating system and application programs required for at least one function; the data storage area may store data created according to the use of the electronic device, etc. In addition, the memory 1102 may include a high-speed random access memory, and may also include a non-transitory memory, such as at least one magnetic disk storage device, a flash memory device, or other non-transitory solid-state storage devices. In some embodiments, the memory 1102 may optionally include a memory remotely set relative to the processor 1101, and these remote memories may be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0085] It should be noted that the computer-readable medium described in the present application may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the above two. The computer-readable medium may be, for example, but not limited to: an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory, an optical fiber, a portable compact disk read-only memory, an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, the computer-readable medium may be any tangible medium that contains or stores a program, and this program can be used by or in combination with an instruction execution system, apparatus, or device.

[0086] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media, and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of the computer's storage media include, but are not limited to, phase change memory, static random access memory, dynamic random access memory, other types of random access memory, read-only memory, electrically erasable programmable read-only memory, flash memory or other memory technologies, compact disc read-only memory, digital versatile disc or other optical storage, magnetic cassette tapes, magnetic disk storage or other magnetic storage devices, or any other non-transmission medium that can be used to store information accessible by a computing device.

[0087] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as C language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any kind of network, including a local area network or a wide area network, or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0088] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. For example, a dedicated integrated circuit, a general-purpose computer, or any other similar hardware device can be used. In some embodiments, the software program of this application can be executed by a processor to implement the above steps or functions. Similarly, the software program of this application (including related data structures) can be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, or a floppy disk and similar devices. Additionally, some steps or functions of this application can be implemented by hardware, for example, as a circuit that cooperates with a processor to execute each step or function.

[0089] The computer program product provided by the embodiments of the present application includes one or more computer programs / instructions. When the computer programs / instructions are executed by a processor, they entirely or partially generate the processes or functions described in the embodiments of the present application. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium may be any available medium that can be accessed by a computer, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid-state drive), etc.

[0090] The flowcharts or block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of devices, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and the combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.

[0091] The scope of the present application is defined by the appended claims rather than the above description. Therefore, all changes that fall within the meaning and scope of the equivalent elements of the claims are intended to be included in the present application. Any reference signs in the claims should not be construed as limiting the claimed subject matter. In addition, it is obvious that the word "comprising" does not exclude other elements or steps, and the singular does not exclude the plural. The multiple units or devices recited in the apparatus claims may also be implemented by one unit or device through software or hardware. The terms "first", "second", etc. are only used for descriptive distinction and do not represent any specific order, nor can they be construed as indicating or implying relative importance.

[0092] As described above, these are only specific embodiments of the present application. However, the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed in the present application can easily make changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims, and the above embodiments should be regarded as exemplary and non-limiting.

Claims

1. A method for optimizing a language model based on autoregressive training, characterized in that, Including: Construct an autoregressive training framework for a language model to predict the probability distribution of subsequent words based on the input historical context. Among them, the probability distribution in each generation of autoregressive training is used to generate synthetic data for the next generation; In each generation of training, mix real data and synthetic data generated by the previous generation model as training data according to a set ratio. Among them, the proportion of real data in the training data decreases dynamically as the number of training generations increases; Use the mixed data to train the language model, update the model parameters, and generate synthetic data required for the next generation of training. Among them, limit the number of the synthetic data not to exceed the set synthetic data ratio control condition relative to the number of the real data; After each generation of training, measure the difference between the data distribution generated by the current model and the real data distribution, and dynamically adjust the proportion of the real data according to the difference; Based on the synthetic data generated based on the probability distribution, dynamic data ratio control, and synthetic data quantity constraint, continuously optimize the autoregressive training process of the language model.

2. The method for optimizing a language model based on autoregressive training according to claim 1, wherein The process that the proportion of real data in the training data decreases dynamically as the number of training generations increases is as follows: as the number of training generations increases, the proportion of the real data decreases according to a preset decreasing strategy; the specific decreasing strategy is: the proportion of real data in the initial training stage is close to 1. As the number of training generations increases, the proportion of real data decreases gradually in a way inversely proportional to the logarithmic function of the number of training generations, and it never exceeds 1. Moreover, the decreasing rate of the proportion of real data is determined by a control parameter set according to the scale of the training data and the quality of the generated data.

3. The method for optimizing a language model based on autoregressive training according to claim 1 or 2, characterized in that The steps of measuring the difference between the data distribution generated by the current model and the real data distribution, and dynamically adjusting the proportion of the real data according to the difference include: Based on the data distribution generated by the current generation model and the initial real data distribution, calculate the L1 norm between the two as a measure of distribution drift. The L1 norm is the sum of the absolute values of the probability differences between the two in all categories; Compare the measured value with a preset distribution deviation threshold; If the measured value exceeds the distribution deviation threshold, increase the proportion of real data in the next generation of training data and correspondingly reduce the proportion of synthetic data to reduce the risk of model distribution drift; If the measured value does not exceed the distribution deviation threshold, adjust the proportion of real data according to the preset decreasing strategy and continue the next generation of training.

4. The method for optimizing a language model based on autoregressive training according to claim 3, wherein The strategy for controlling the degree of distribution drift between the data distribution generated by the current generation model and the initial real data distribution includes: by setting a constant related to the scale of the real data and recording the number of training generations of the current autoregressive training, ensure that after each generation of training, the measured value of the distribution drift between the data distribution generated by the current generation model and the initial real data distribution shows a decreasing trend, so as to form a dynamic balance relationship between the number of training generations and distribution drift control during the model training process, and ensure that the distribution drift during the training process is always controllable.

5. The method for optimizing a language model based on autoregressive training according to claim 1, wherein Restricting the quantity of the synthetic data not to exceed a set synthetic data ratio condition relative to the quantity of the real data. Specifically, the quantity of the synthetic data needs to satisfy not exceeding a dynamic upper limit jointly determined by the sample quantity of the real data and the vocabulary size of the language model, that is, as the quantity of the training data samples increases, the quantity of the synthetic data is allowed to increase correspondingly, and at the same time, as the vocabulary scale expands, the allowed quantity of the synthetic data is reduced, so as to prevent the risk of probability distribution collapse during the model training process due to an excessive proportion of synthetic data.

6. The method for optimizing a language model based on autoregressive training according to claim 1, characterized in that The method further includes: during the model training process, extending the effective training cycle of the model by setting training stability conditions to avoid model collapse, specifically including: In each generation of training, according to the diversity situation between the synthetic data generated by the current model and the real data, combined with the dispersion metric value of the initial real data, dynamically adjust the proportion of the real data, so as to reduce the degradation rate of the synthetic data generated by the model and delay the speed at which the data distribution output by the model converges to a single distribution state.

7. The method for optimizing a language model based on autoregressive training according to claim 1, wherein The step of constructing the autoregressive training framework of the language model and predicting the probability distribution of subsequent words based on the input historical context includes: The language model receives the context sequence of the training sample as input, and sets the vocabulary size, context length, and sample quantity; Based on the currently input context sequence, calculate the conditional probability distribution of all candidate words by taking the inner product of the model weight matrix and the context vector, and the conditional probability distribution realizes probability normalization through the softmax function; Use the cross-entropy loss function as the objective function for language model training, and calculate the error between the conditional probability distribution and the actual target word label; By optimizing the objective function, iteratively update the model parameters until the objective function reaches the convergence condition, and generate the synthetic data required for the next generation of training based on the conditional probability distribution result after the current generation of model training is completed.

8. An electronic device, characterized in that, The electronic device includes: One or more processors; and a memory storing computer program instructions, and when the computer program instructions are executed, the processors execute the autoregressive training-based language model optimization method according to any one of claims 1-7.

9. A computer-readable storage medium having a computer program and / or instructions stored thereon, characterized in that, When the computer program and / or instructions are executed by the processor, the autoregressive training-based language model optimization method according to any one of claims 1-7 is implemented.

10. A computer program product, comprising a computer program and / or instructions, characterized in that, When the computer program and / or instructions are executed by the processor, the autoregressive training-based language model optimization method according to any one of claims 1-7 is implemented.

Citation Information

Patent Citations

  • Method and device for training language model based on multiple word vectors, equipment and medium

    CN111737995A

  • Long text multi-label classification model optimization method and device

    CN118133076A

  • Continuous learning training method and device of large language model, medium and equipment

    CN118313482A

  • Large language model construction method and system supporting streaming input

    CN119398097A

  • Large model-based scheme generation and optimization model training method and device

    CN119721940A