A training method and device of a target large language model, equipment and medium
Patent Information
- Application Number
- CN202610853931.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-12
- Publication Date
- 2026-08-28
AI Technical Summary
[0004]现有的掩码扩散式大语言模型采用固定加权策略,对所有时间步t对应的损失赋予相同权重,忽略了不同掩码比例带来的任务难度差距
[0013]本说明书实施例提供的目标大语言模型的训练方法,通过获取掩码序列集合;将各个掩码序列输入目标大语言模型得到被掩码位置的预测分布;基于各个掩码序列的时间步确定随时间步增大而单调递减的重新加权因子;基于预测分布计算基础掩码扩散损失;再基于重新加权因子与基础掩码扩散损失计算目标损失值;最后基于目标损失值训练模型。本说明书实施例中,利用与时间步单调递减的重新加权因子,使得时间步越小、掩码比例越低、重建难度越小的样本的权重越大;反之,时间步越大、重建难度越大的样本的权重越小,进而使得,在反向传播时自动放大低噪声样本的梯度贡献、压缩高噪声样本的梯度贡献,避免了训练初期高噪声样本的随机梯度对优化方向的干扰,进而提升了目标大语言模型继续预训练的梯度稳定性和收敛平滑性。
Smart Images

Figure CN122655892A_ABST
Abstract
Description
Technical Field
[0001] The embodiments in this specification relate to the field of model training technology, and in particular to a training method, apparatus, device and medium for a target large language model. Background Technology
[0002] Diffusion-based large language model (DLLM) is an emerging large language model (LLM). Unlike traditional autoregressive (AR) large language models, DLLM generates text by iterative denoising. It can complete context modeling by bidirectional attention and decode multiple mask positions in parallel, resulting in better inference efficiency and flexibility.
[0003] The current mainstream approach to building a diffusion-based large language model is to use a pre-trained autoregressive large language model as a foundation, and then continue training it using the training objective of a masked diffusion-based large language model (MDM), gradually adapting the original autoregressive model to the diffusion generation paradigm. During the continued training process, a time step t is randomly sampled from the interval [0,1]. According to the numerical proportion of time step t, the words in the input sequence are masked and replaced, allowing the model to reconstruct the original masked words based on the unmasked context.
[0004] Existing masked diffusion-based large language models employ a fixed weighting strategy, assigning the same weight to the loss at all time steps t, ignoring the differences in task difficulty caused by different mask ratios. In the early stages of training, the original autoregressive pre-trained model has not yet mastered the diffusion denoising logic. Unifying the weights forces the model to learn difficult samples with high mask ratios and high uncertainty using equal training weights, potentially leading to problems such as training gradient oscillations and excessively large fluctuations in model convergence, ultimately negatively impacting the overall performance of the diffusion-based large language model. Summary of the Invention
[0005] In view of this, embodiments of this specification provide a training method for a target large language model. One or more embodiments of this specification also relate to a text generation method based on a target large language model, a training apparatus for a target large language model, a text generation apparatus based on a target large language model, a computing device, a computer-readable storage medium, and a computer program product, to address the technical deficiencies existing in the prior art.
[0006] According to a first aspect of the embodiments of this specification, a method for training a target large language model is provided, the target large language model being used to generate word sequences in parallel, the training method comprising: Obtain a set of mask sequences; the proportion of masked words in each mask sequence of the set of mask sequences is positively correlated with the time step; Each of the mask sequences is input into the target large language model to obtain the predicted distribution of the masked words in each of the mask sequences; Based on the time steps of each of the mask sequences, a reweighting factor for each of the mask sequences is determined; within a preset time step interval, the reweighting factor monotonically decreases as the time step increases; Based on the predicted distribution, the basic mask diffusion loss of each of the mask sequences is calculated; the basic mask diffusion loss is used to reflect the difference between the predicted distribution and the true label; The target loss value is calculated based on the reweighting factor of each of the mask sequences and the base mask diffusion loss of each of the mask sequences; The target large language model is trained based on the target loss value.
[0007] According to a second aspect of the embodiments of this specification, a text generation method based on a target large language model is provided, wherein the target large language model is used to generate word sequences in parallel, and the text generation method includes: Obtain the initial noise sequence; The initial noise sequence is input into the target large language model to generate the target sequence corresponding to the initial noise sequence, wherein the target large language model is a model trained based on the method described above.
[0008] According to a third aspect of the embodiments of this specification, a training apparatus for a target large language model is provided, the target large language model being used to generate word sequences in parallel, the training apparatus comprising: The acquisition module is configured to acquire a set of mask sequences; the proportion of masked words in each mask sequence of the mask sequence set is positively correlated with the time step; The prediction module is configured to input each of the mask sequences into the target large language model to obtain the predicted distribution of the masked words in each of the mask sequences; The determination module is configured to determine the reweighting factor of each of the mask sequences based on the time steps of each of the mask sequences; within a preset time step interval, the reweighting factor decreases monotonically as the time step increases; The basic loss calculation module is configured to calculate the basic mask diffusion loss for each of the mask sequences based on the predicted distribution; the basic mask diffusion loss is used to reflect the difference between the predicted distribution and the true label; The target loss calculation module is configured to calculate the target loss value based on the reweighting factor of each of the mask sequences and the basic mask diffusion loss of each of the mask sequences; The training module is configured to train the target large language model based on the target loss value.
[0009] According to a fourth aspect of the embodiments of this specification, a text generation apparatus based on a target large language model is provided, the target large language model being used to generate word sequences in parallel, the text generation apparatus comprising: The noise sequence acquisition module is configured to acquire an initial noise sequence; The text generation module is configured to input the initial noise sequence into the target large language model to generate the target sequence corresponding to the initial noise sequence, wherein the target large language model is a model trained based on the method described above.
[0010] According to a fifth aspect of the embodiments of this specification, a computing device is provided, comprising: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, they implement the steps of the above-mentioned training method for the target large language model or the text generation method based on the target large language model.
[0011] According to a sixth aspect of the embodiments of this specification, a computer-readable storage medium is provided that stores computer-executable instructions, which, when executed by a processor, implement the steps of the above-described training method for a target large language model or a text generation method based on a target large language model.
[0012] According to a seventh aspect of the embodiments of this specification, a computer program product is provided, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described training method for a target large language model, or a text generation method based on a target large language model.
[0013] The training method for a target large language model provided in this specification involves: acquiring a set of mask sequences; inputting each mask sequence into the target large language model to obtain the predicted distribution of the masked positions; determining a reweighting factor that monotonically decreases with increasing time step based on the time step of each mask sequence; calculating the basic mask diffusion loss based on the predicted distribution; calculating the target loss value based on the reweighting factor and the basic mask diffusion loss; and finally training the model based on the target loss value. In this specification, the reweighting factor, which monotonically decreases with time step, ensures that samples with smaller time steps, lower mask ratios, and easier reconstruction have larger weights; conversely, samples with larger time steps and greater reconstruction difficulty have smaller weights. This automatically amplifies the gradient contribution of low-noise samples and compresses the gradient contribution of high-noise samples during backpropagation, avoiding interference from the stochastic gradients of high-noise samples on the optimization direction in the early stages of training. This improves the gradient stability and convergence smoothness of the target large language model during further pre-training.
[0014] Furthermore, the reweighting factor is directly defined as a monotonically decreasing function of time step t. Since time step t is already a variable that must be sampled during mask diffusion training, no additional computational or storage overhead is required, achieving zero-cost improvement. Without changing the model structure or increasing the learnable parameters, the stability and final performance of the target large language model training are significantly improved. Attached Figure Description
[0015] Figure 1 This is a schematic diagram illustrating an application scenario of a training method for a target large language model provided in one embodiment of this specification; Figure 2 This is a flowchart illustrating a training method for a target large language model according to one embodiment of this specification; Figure 3 This is a flowchart illustrating a text generation method based on a target large language model, as provided in one embodiment of this specification. Figure 4 This is a flowchart of a text generation method based on a target large language model provided in one embodiment of this specification; Figure 5 This is an overall flowchart of a training method for a target large language model provided in one embodiment of this specification; Figure 6 This is a schematic diagram of the structure of a training device for a target large language model provided in one embodiment of this specification; Figure 7 This is a schematic diagram of the structure of a text generation device based on a target large language model provided in one embodiment of this specification; Figure 8 This is a structural block diagram of a computing device provided in one embodiment of this specification. Detailed Implementation
[0016] Many specific details are set forth in the following description to provide a full understanding of this specification. However, this specification can be implemented in many other ways than those described herein, and those skilled in the art can make similar extensions without departing from the spirit of this specification. Therefore, this specification is not limited to the specific implementations disclosed below.
[0017] The terminology used in one or more embodiments of this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the one or more embodiments of this specification. The singular forms “a,” “described,” and “the” as used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more associated listed items.
[0018] It should be understood that although the terms first, second, etc., may be used to describe various information in one or more embodiments of this specification, such information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, first may also be referred to as second without departing from the scope of one or more embodiments of this specification, and similarly, second may also be referred to as first. Depending on the context, the word "if" as used herein may be interpreted as "when," "when," or "in response to a determination."
[0019] Furthermore, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in one or more embodiments of this specification are obtained through open-source datasets or public datasets that comply with their license agreements, or are obtained with full authorization from the relevant parties. Moreover, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0020] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, typically containing hundreds of millions, tens of billions, hundreds of billions, trillions, or even tens of trillions of model parameters. A large model can also be called a foundation model. It is pre-trained using large-scale unlabeled corpora to produce a pre-trained model with hundreds of millions of parameters. Such models can adapt to a wide range of downstream tasks and have good generalization ability. Examples include Large Language Models (LLMs) and multi-modal pre-training models.
[0021] In practical applications, large models only require a small number of samples to fine-tune the pre-trained model before they can be applied to different tasks. Large models can be widely used in fields such as Natural Language Processing (NLP) and Computer Vision. Specifically, they can be applied to computer vision tasks such as Visual Question Answering (VQA), Image Captioning (IC), and Image Generation, as well as natural language processing tasks such as text-based sentiment classification, text summarization, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0022] First, the terms and concepts used in one or more embodiments of this specification will be explained.
[0023] Diffusion Large Language Model (DLLM) is a large language model based on a diffusion process. Its core generation mechanism employs iterative denoising, starting from an initial state composed entirely of masked markers and progressively predicting and replacing mask positions to ultimately form a complete text sequence. Unlike traditional autoregressive large language models that generate text word by word, DLLM utilizes a bidirectional attention mechanism to comprehensively model the context and performs parallel predictions of multiple masked positions during decoding, exhibiting unique advantages in inference efficiency and generation flexibility. DLLM is suitable for text generation scenarios requiring high throughput and low latency, such as real-time translation, intelligent customer service, and emotional chat.
[0024] Iterative denoising (ID) refers to the process of starting with a fully masked noisy text sequence, and then iteratively predicting and replacing the content at the masked positions through multiple rounds to continuously remove sequence noise, repair semantic information, and finally output complete text. Iterative denoising can support parallel denoising at multiple positions and, combined with bidirectional context modeling capabilities, complete the overall sequence optimization.
[0025] Autoregressive Large Language Models (AR-LLMs) are large language models based on an autoregressive generative paradigm. The core mechanism of AR-LLMs is to predict the next word in a sequence word by word from left to right. The generation of each word depends only on previously generated historical words and cannot access future information. AR-LLMs typically employ causal attention mechanisms to achieve this unidirectional dependency. AR-LLMs excel at natural language generation tasks such as text continuation, dialogue generation, and code writing, and are widely used in scenarios such as intelligent customer service, emotional companionship, and machine translation.
[0026] A token is a fundamental semantic unit in natural language processing and large language models, and it is also the smallest operational unit in model text processing. Before processing natural language text, the system breaks down continuous strings, Chinese characters, words, punctuation marks, special symbols, and other raw content into independent tokens using rules such as word segmentation and sub-word splitting. Token segmentation is not limited to complete Chinese characters or words; it can be segmented into single characters, phrases, word roots, character fragments, and also includes model-specific special symbols such as mask markers and sequence end markers. All text input to the large language model is first converted into corresponding token codes, and the model performs semantic understanding, feature calculation, mask prediction, and text generation operations entirely based on tokens.
[0027] The Masked Diffusion Model (MDM), also known as the Absorbed State Discrete Diffusion Model, is a non-autoregressive generation paradigm for generating discrete data (such as text and code). Its forward process corrupts the data by gradually replacing tokens with special [MASK] markers (absorbed states), while the backward process learns to iteratively predict the original content from full [MASK] or partial [MASK] states.
[0028] Training samples (TS) are data units used to train machine learning models. They typically contain input data and corresponding expected outputs (labels) or feedback information. By learning the data features and inherent patterns in a large number of training samples, the model gradually optimizes its own parameters to acquire corresponding prediction, classification, or generation capabilities. TS is the fundamental data support for building effective machine learning models.
[0029] Bidirectional attention is a form of attention mechanism that allows each position in a sequence to simultaneously pay attention to all positions to its left (historical) and right (future), thereby obtaining complete contextual information. Bidirectional attention enables models to fully utilize the semantic connections between the preceding and following text when predicting the content at a given position, improving their overall understanding of the text.
[0030] Random masking is a data augmentation and training technique. Its core operation involves independently replacing each position in a sequence with a mask marker with a certain probability. Through random masking, models are trained to recover the complete original sequence from partially masked input, thereby learning the text's inherent structure and contextual relationships.
[0031] Masking probability is the probability value that each position in a sequence will be independently replaced with a mask label during random masking. Masking probabilities are typically sampled from a predefined numerical range (e.g., between 0 and 1), and the magnitude of the masking probability directly determines the expected number of masked positions in the sequence. Generally, a higher masking probability means more positions in the sequence are expected to be masked, requiring the model to recover the original content from less information, thus increasing the training difficulty. In the training of diffusion-based large language models, masking probabilities are usually randomly sampled according to a uniform distribution, enabling the model to adapt to inputs with varying levels of noise and improving generalization ability.
[0032] A loss function is a function used in machine learning models to measure the difference between the model's predicted output and the target output. The loss function takes the model's predicted values and the true labels as input and outputs a scalar value called the loss value. During training, the optimization algorithm iteratively updates the model parameters to minimize the loss value, gradually bringing the model's predictions closer to the true target.
[0033] Loss value is a scalar numerical value calculated by the loss function based on the model's predicted output and the true labels. It quantitatively reflects the magnitude of the model's prediction error on the current training samples. A smaller loss value indicates that the model's prediction is closer to the true content; a larger loss value indicates a more severe prediction bias. In batch training, the loss values of all samples within a batch are typically averaged or summed to obtain an overall loss value, which is then used for backpropagation and parameter updates. The decreasing trend of the loss value is an important indicator of whether the model has converged.
[0034] Cross-entropy loss is an application of the concept of cross-entropy from information theory to machine learning, widely used for loss calculation in classification tasks. In language model training, the model's predicted output for each position is a probability distribution with respect to the vocabulary size, while the true label is a one-hot vector. Cross-entropy loss calculates the difference between these two distributions, specifically using the negative log-likelihood formula. The closer the model's predicted probability distribution is to the true label, the smaller the cross-entropy loss value.
[0035] Diffusion-based large language model (DLLM) is an emerging large language model (LLM). Unlike traditional autoregressive (AR) large language models, DLLM generates text by iterative denoising. It can complete context modeling by bidirectional attention and can decode multiple masked positions in parallel, resulting in better inference efficiency and flexibility.
[0036] The current mainstream approach to building a diffusion-based large language model is to use a pre-trained autoregressive large language model as a foundation, and then continue training it using the training objective of a masked diffusion-based large language model (MDM), gradually adapting the original autoregressive model to the diffusion generation paradigm. During the continued training process, a time step t is randomly sampled from the interval [0,1]. According to the numerical proportion of time step t, the words in the input sequence are masked and replaced, allowing the model to reconstruct the original masked words based on the unmasked context.
[0037] Existing masked diffusion-based large language models employ a fixed weighting strategy, assigning the same weight to the loss at all time steps t, ignoring the differences in task difficulty caused by different mask ratios. In the early stages of training, the original autoregressive pre-trained model has not yet mastered the diffusion denoising logic. Unifying the weights forces the model to learn difficult samples with high mask ratios and high uncertainty using equal training weights, potentially leading to problems such as training gradient oscillations and excessively large fluctuations in model convergence, ultimately negatively impacting the overall performance of the diffusion-based large language model.
[0038] To address the aforementioned problems, this specification provides a training method for a target large language model. One or more embodiments of this specification also relate to a text generation method based on a target large language model, a training apparatus for a target large language model, a text generation apparatus based on a target large language model, a computing device, a computer-readable storage medium, and a computer program product, which will be described in detail in the following embodiments.
[0039] The training method for a target large language model provided in one embodiment of this specification can be executed in a training device. Figure 1 This diagram illustrates an application scenario of a training method for a target large language model provided in one embodiment of this specification. For example... Figure 1 As shown: The training device 101 can perform iterative training on the target large language model 102 for a preset number of rounds. Specifically, within one round, the training device 101 can pre-acquire a set of original sequences. For each original sequence in the set, it randomly samples the time steps corresponding to each original sequence within a preset time step interval. It then masks each original sequence according to its corresponding time step, obtaining a set of masked sequences. Next, it inputs each masked sequence into the target large language model 102 to obtain the predicted distribution of masked words in each masked sequence output by the target large language model 102. Based on the time steps of each masked sequence, it determines the reweighting factor for each masked sequence. Within the preset time step interval, the reweighting factor monotonically decreases as the time step increases. Based on the predicted distribution, it calculates the basic mask diffusion loss for each masked sequence. Finally, based on the reweighting factors and the basic mask diffusion loss of each masked sequence, it calculates the target loss value. The target large language model is then trained based on the target loss value. Among them, the target large language model can be a large language model that generates word sequences in parallel through iterative denoising.
[0040] The training device 101 can be a terminal device or a server. The terminal can include, but is not limited to, smartphones, tablets, laptops, PDAs, personal computers, smart home devices, and in-vehicle devices. The server can include, but is not limited to, any device, equipment, platform, or device cluster with computing and processing capabilities. The server can connect to one or more terminal devices via a local area network (LAN), a wide area network (WAN), the Internet, or other types of data networks.
[0041] See Figure 2 , Figure 2 A flowchart of a training method for a target large language model provided in one embodiment of this specification is shown, which specifically includes the following steps.
[0042] Step S202: Obtain a set of mask sequences; the proportion of masked words in each mask sequence of the set of mask sequences is positively correlated with the time step.
[0043] In the embodiments of this specification, the target large language model is used to generate word sequences in parallel. That is, the target large language model can generate word sequences by performing parallel predictions on multiple masked positions in the sequence. Specifically, the target large language model can be a diffusion-type large language model, or it can be any other large language model that can generate word sequences in parallel; no specific limitation is made in this regard.
[0044] In the embodiments of this specification, a sequence can refer to an ordered set composed of discrete symbols or tokens arranged in sequence. A masked sequence refers to a noisy sequence obtained by randomly replacing some tokens with mask symbols [MASK] at a certain time step t from an original sequence, i.e., a clean, unmasked discrete token sequence. The time step t can be a value randomly sampled within a preset time step interval. The time step t determines the masking strength; the larger the time step t, the higher the proportion of masked tokens in the masked sequence and the higher the noise level; the smaller the time step t, the more original information is retained, and the closer the masked sequence is to the original sequence.
[0045] The mask sequence set consists of training samples for the current batch. The mask sequence set can include multiple mask sequences, and the proportion of masked words in each mask sequence is positively correlated with the time step corresponding to each mask sequence, meaning that the proportion of masked words in each mask sequence can be different.
[0046] In the embodiments described in this specification, by introducing a masking ratio positively correlated with the time step, the target large language model can learn the denoising and reconstruction task at different noise levels. If the time step is small, the model faces a nearly complete context, making learning easier; if the time step is large, the model needs to infer the masked words under sparse cues, making learning more difficult. This allows the target large language model to gradually master the ability to recover a complete sequence from partial observations.
[0047] Step S204: Input each of the mask sequences into the target large language model to obtain the predicted distribution of the masked lexical units in each of the mask sequences.
[0048] In the embodiments of this specification, the target large language model is a large language model used for parallel generation of word sequence. Specifically, the target large language model can be a diffusion-type large language model. Here, a diffusion-type large language model refers to a neural network model that uses an iterative denoising method for text generation.
[0049] All masked sequences of the current batch can be input into the target large language model at one time in parallel. The target large language model performs forward computation on each masked position in each masked sequence, and finally outputs the prediction distribution of each masked position in each masked sequence.
[0050] Wherein, the prediction distribution may refer to the probability distribution of all tokens in the vocabulary given by the target large language model for a masked position, which is usually output by the Softmax function, and the sum of all probabilities is 1. For example, for the masked sequence [I] [love] [MASK] [fruit], the prediction distribution that the model may output at the third position (originally "app") is: {"app": 0.85, "ban": 0.10, "big": 0.03, others: 0.02}, where 0.85 indicates that the target large language model considers that the probability of this position being "app" has the highest confidence.
[0051] Step S206: determining the reweighting factor of each said masked sequence based on the time step of each said masked sequence; within a preset time step interval, said reweighting factor monotonically decreases as said time step increases.
[0052] In the embodiments of this specification, the preset time step interval refers to the value range of the sampling time step t. The time step t is a value randomly sampled within the preset time step interval. The reweighting factor may be a function dependent on the time step t, which may be denoted as w(t). The reweighting factor can be used to reflect the training weight of masked sequences under different noise levels.
[0053] In practical applications, the reweighting factor corresponding to each masked sequence can be calculated through a preset functional relationship according to the time step t corresponding to the masked sequence. The preset functional relationship enables the reweighting factor to monotonically decrease as the time step increases within the preset time step interval. That is, the smaller the time step, the lower the masking ratio, the closer the masked sequence is to the original sequence, and the larger the reweighting factor; the larger the time step, the higher the masking ratio, the stronger the noise, and the smaller the reweighting factor.
[0054] In an optional implementation of this embodiment, the reweighting factor may be one of the following; wherein, is the time step, , , are constants greater than 0. For example, if the form w(t)=1 / t is adopted, when t=0.1, w(0.1)=10; when t=0.9, w(0.9)≈1.11.
[0055] It can be understood that the reweighting factor can also be in other forms, such as This can be achieved simply by ensuring that the reweighting factor decreases monotonically with increasing time step, where... It is a constant greater than 0.
[0056] By introducing a reweighting factor, the target large language model can amplify the gradient contribution of low-noise samples and compress the gradient contribution of high-noise samples during backpropagation. This avoids the target large language model being dominated by the messy gradients of difficult samples in the early stage of training, making the optimization process of the target large language model more stable.
[0057] In the embodiments of this specification, the time step t directly determines the proportion of masked words in the masked sequence, i.e., the amount of information destroyed in the original sequence. Therefore, the time step t is an intrinsic physical quantity in the training task of the target large language model, strictly positively correlated with the denoising difficulty. Thus, a weighting factor based on t can accurately match the actual reconstruction difficulty of the model under different noise levels: when t is small, the masking proportion is low, the context is almost complete, and the target large language model has low difficulty in recovering the masked words; when t is large, most words are masked, and the target large language model needs to perform highly uncertain inferences under sparse cues, significantly increasing the difficulty. Therefore, determining a reweighting factor based on the time step t allows this weighting factor to automatically match the actual reconstruction difficulty of the model under different noise levels.
[0058] In the embodiments described in this specification, the reweighting factor is directly defined as a monotonically decreasing function of time step t. Since time step t is already a variable that must be sampled in mask diffusion training, no additional computational or storage overhead is required, achieving zero-cost improvement. This method significantly improves the stability and final performance of the target large language model training without changing the model structure or increasing the learnable parameters, and has high engineering practical value.
[0059] Step S208: Based on the predicted distribution, calculate the basic mask diffusion loss for each of the mask sequences; the basic mask diffusion loss is used to reflect the difference between the predicted distribution and the true label.
[0060] In the embodiments of this specification, the real tag refers to the original, unmasked word element at the masked position in the masked sequence, that is, the word element at the corresponding position in the original sequence.
[0061] In practical applications, the prediction distribution of each masked position can be compared with the real label of the masked position to quantify the difference between the two, so as to obtain the basic mask diffusion loss. In practical applications, Negative Log-Likelihood (NLL for short) loss can be used to calculate the basic mask diffusion loss. Specifically, for each masked position, the negative logarithm of the probability value corresponding to the real label in the prediction distribution is obtained, and then the negative logarithms at all masked positions are summed or averaged to calculate the mask diffusion loss. For example, for the mask sequence [I] [love] [MASK] [fruit], the real label is "apple", and the prediction distribution of the target large language model at this position is {"apple": 0.85, "pear": 0.10, "big": 0.03, ...}, then the negative log-likelihood is -log(0.85)≈0.162. If there are multiple masked positions in the sequence, the negative log-likelihoods of each masked position can be added and then divided by the total number of masked positions, or multiplied by a normalization factor related to the time step, to obtain the basic mask diffusion loss of the mask sequence.
[0062] Step S210: Calculating a target loss value based on the reweighting factor of each said mask sequence and the basic mask diffusion loss of each said mask sequence.
[0063] In the embodiments of this specification, the target loss value is a loss value obtained after reweighting adjustment based on the reweighting factor, and used for updating the model parameters of the target large language model. In practical applications, the target loss value can be obtained by performing operations on the basic mask diffusion loss of each mask sequence and the corresponding reweighting factor of each mask sequence, and then aggregating the operation results.
[0064] In the embodiments of this specification, the reweighting factor is used to further adjust the relative contribution of each sample (i.e., mask sequence) to the total loss. For small time steps, that is, easy samples, the basic mask diffusion loss is amplified; for large time steps, that is, hard samples, the basic mask diffusion loss is reduced. During backpropagation, the gradient signal of easy samples is made to dominate, guiding the target large language model to prioritize consolidating the nearly mastered low-noise denoising ability; the gradient of hard samples is relatively suppressed, avoiding gradient oscillation caused by highly uncertain reconstruction tasks in the early training stage. Thus, the target loss value becomes a unified training objective that takes into account both difficulty adaptation and optimization stability, and provides a gradient source with clear direction and stable scale for the subsequent update of model parameters.
[0065] Step S212: Training the target large language model based on the target loss value.
[0066] In practical applications, the calculated target loss value can be used to calculate the gradient of each trainable parameter in the target large language model through the backpropagation algorithm, and the model parameters can be updated according to the gradient. This allows the target large language model to gradually reduce the loss value in subsequent iterations, thereby improving the accuracy of reconstructing the masked position.
[0067] Compared to traditional training methods that apply equal weights to all time steps, this approach exhibits more stable convergence and significantly reduced oscillations in the early stages of pre-training. With the same number of training steps, the final model achieves stable improvements in metrics such as text generation quality and performance on downstream tasks. For example, in a continued pre-training experiment with approximately 220 billion words, the target large language model trained using this approach outperformed the unweighted control model on multiple evaluation tasks.
[0068] The training method for a target large language model provided in this specification involves: acquiring a set of mask sequences; inputting each mask sequence into the target large language model to obtain the predicted distribution of the masked positions; determining a reweighting factor that monotonically decreases with increasing time step based on the time step of each mask sequence; calculating the basic mask diffusion loss based on the predicted distribution; calculating the target loss value based on the reweighting factor and the basic mask diffusion loss; and finally training the model based on the target loss value. In this specification, the reweighting factor, which monotonically decreases with time step, ensures that samples with smaller time steps, lower mask ratios, and easier reconstruction have larger weights; conversely, samples with larger time steps and greater reconstruction difficulty have smaller weights. This automatically amplifies the gradient contribution of low-noise samples and compresses the gradient contribution of high-noise samples during backpropagation, avoiding interference from the stochastic gradients of high-noise samples on the optimization direction in the early stages of training. This improves the gradient stability and convergence smoothness of the target large language model during further pre-training.
[0069] Furthermore, the reweighting factor is directly defined as a monotonically decreasing function of time step t. Since time step t is already a variable that must be sampled during mask diffusion training, no additional computational or storage overhead is required, achieving zero-cost improvement. Without changing the model structure or increasing the learnable parameters, the stability and final performance of the target large language model training are significantly improved.
[0070] For ease of understanding, one or more embodiments of this specification also provide specific methods for obtaining the mask sequence set.
[0071] In one optional implementation of this embodiment, obtaining the mask sequence set may specifically include: Get the original set of sequences.
[0072] For each original sequence in said original sequence set, determining the respective corresponding time step of each said original sequence in the original sequence set from said preset time step interval.
[0073] Performing masking processing on each said original sequence according to the respective corresponding time step of each said original sequence, to obtain a masked sequence set.
[0074] In the embodiments of this specification, an original sequence is a clean text sequence that has not been masked (i.e., a real token sequence). An original sequence may be a sequence formed by sequentially arranging discrete tokens, where each token corresponds to a real semantic unit (such as a Chinese character, a word, an English word, etc.). The original sequence can be used as the source of the masked sequence, and can also be used as the correct answer or target output during the training process of the target large language model. The original sequence does not contain any mask symbols ([MASK]). For example, an original sequence can be [I][love] [ar][tifi][cial][intelligence]. The original sequence set may be a set formed by all original sequences in the current training batch, and each original sequence corresponds to one masked sequence to be generated.
[0075] In practical applications, the preset time step interval refers to the value range of the sampled time step t. A value can be independently sampled from the preset time step interval for each original sequence as the time step t corresponding to that original sequence. In the embodiments of this specification, the sampling process of time step t can be random, so as to ensure that the target large language model can be exposed to samples with various noise levels.
[0076] In practical applications, according to the time step corresponding to each original sequence, tokens in the original sequence can be randomly replaced according to the proportion indicated by the time step: each position is replaced with the mask symbol [MASK] with probability t, with probability 1 t the original token is retained, thereby obtaining a noisy masked sequence.
[0077] In this specification, the time step t determines the noise intensity of each sample. In the embodiments described, by randomly sampling time steps t for each original sequence, the coverage of the training data in the noise space is ensured. This allows the target large language model to simultaneously encounter samples of varying difficulty levels, from low to high noise, within the same training batch, thereby enabling the target large language model to generalize to masked sequences with various masking ratios. In other words, diverse time step sampling strategies help the target large language model learn continuous denoising capabilities, rather than overfitting only at a fixed noise level. On the other hand, since each original sequence is sampled independently at a time step, training instability or a preference for a particular noise level that might occur when using the same t for the entire batch is avoided. For example, in a batch with three original sequences: sequence A is assigned t=0.1, with only a few positions masked; sequence B is assigned t=0.5, with approximately half of its positions masked; and sequence C is assigned t=0.9, with the vast majority of its positions masked. The target large language model handles these three difficulty levels simultaneously in the same forward propagation, and assigns different weights to different difficulties through subsequent reweighting factors during loss calculation, thereby achieving efficient and stable continued pre-training.
[0078] In practical applications, to further reduce training costs, a pre-trained autoregressive large language model can be used as the starting point for further training of the target model.
[0079] In one optional implementation of this embodiment, the initial weights of the target large language model are the weights of a pre-trained autoregressive large language model.
[0080] In the embodiments of this specification, the initial weights can refer to the initial values of the network parameters of each layer before the target large language model begins further training. A pre-trained autoregressive large language model is a model that has been pre-trained on a large-scale corpus using a traditional autoregressive objective, such as the GPT model or the LLaMA model.
[0081] One implementation approach is to copy all trainable parameters of the pre-trained autoregressive large language model to the target large language model, using this as the starting point for training the target large language model. Another implementation approach is to copy only the model parameters other than the output layer if the output layer dimensions of the autoregressive large language model and the target large language model do not match (e.g., different vocabulary), and then randomly initialize the output layer.
[0082] In the embodiments described in this specification, the autoregressive large language model has learned rich grammatical, semantic, and factual knowledge from a large-scale corpus. This knowledge is also crucial for the target large language model to understand context and generate text that conforms to language rules. Using the weights of the pre-trained autoregressive large language model as the initial weights of the target large language model is equivalent to allowing the target large language model to stand on the shoulders of giants, thereby avoiding the target large language model learning the language basics from scratch.
[0083] Compared to randomly initializing weights, using the weights of a pre-trained autoregressive large language model as the initial weights of the target large language model can significantly reduce training costs, accelerate training convergence, and enable the prior knowledge provided by the autoregressive model to effectively guide the reconstruction of the diffusion model on low-noise samples. This, combined with the reweighting factors, makes the transition of the model from autoregressive to the target model smoother, ultimately reducing training costs while significantly improving model performance.
[0084] In practical applications, if the time step t is too small (e.g., t=0.01), the reweighting factor may be too large, causing the basic mask diffusion loss of the corresponding mask sequence to dominate the target loss value, which in turn leads to gradient explosion, numerical instability and other problems, affecting training convergence.
[0085] In one optional implementation of this embodiment, the preset time step interval can be [0.1, 1].
[0086] The preset time step interval refers to the range of values for the sampling time step t, within which the time step t can be randomly sampled.
[0087] In the embodiments of this specification, the lower limit t of the preset time step interval is... min Setting the reweighting factor to 0.1 avoids the instability of the target loss value caused by excessively small values of t. Specifically, since the reweighting factor monotonically decreases with increasing time step, if t is very close to 0 (e.g., t=0.01), the reweighting factor will become very large. For example, assuming the reweighting factor is in the form of 1 / t, the reweighting factor would be 1 / 0.01=100, which could lead to an abnormally high proportion of the basic mask diffusion loss of the corresponding mask sequence in the target loss value, causing drastic fluctuations in the gradient scale during backpropagation, potentially triggering gradient explosion or making the optimization process overly sensitive. By truncating the lower bound to 0.1, the maximum value of the reweighting factor 1 / t is limited to within 10 (i.e., 1 / 0.1=10), which can control the basic mask diffusion loss of each mask sequence within a reasonable range, while still retaining the monotonically decreasing characteristic of larger weights at smaller time steps t and smaller weights at larger time steps t.
[0088] In practical applications, an extremely small step size t also means a particularly low masking ratio (e.g., only 1% of tokens are masked). Since the target large language model, after autoregressive pre-training, is already capable of processing relatively clean text, masked sequences with a very low masking ratio contribute little to the target large language model's learning of denoising ability, and forcibly assigning a large weight to such masked sequences may instead interfere with the target large language model's effective learning on medium- and high-noise samples. Therefore, setting the preset time step interval as [0.1, 1] can further improve the overall training efficiency of the target large language model and the learning effect of the target large language model on medium- and high-noise samples.
[0089] In practical applications, in order to avoid excessively high masking ratio caused by an excessively large time step t (e.g., t close to 1), where almost all elements in the masked sequence are mask tokens, which leads to overly sparse learning signals and unstable training, the upper limit t of the preset time step interval max can also be set to a value less than 1, for example, t max = 0.95.
[0090] For ease of understanding, one or more embodiments of this specification further provide specific methods for masking each original sequence, In an optional implementation of this embodiment, for any of said original sequences, said masking each original sequence according to the time step respectively corresponding to each said original sequence may specifically comprise: For any position i in said original sequence, generate a random number u i .
[0091] If said random number u i is less than said time step, replace the token at said position i with a mask token.
[0092] If said random number u i is greater than or equal to said time step, keep the original token at said position i unchanged.
[0093] In the embodiments of this specification, the random number u i may be a randomly generated floating point number following a uniform distribution U(0,1). In an original sequence, each position i corresponds to one random number u i .
[0094] In practical applications, for any original sequence, each position of the original sequence may be traversed sequentially, and a random number may be generated for each position, or alternatively, random numbers for all positions may be generated in parallel for each position of the original sequence. For each position, if the random number u generated for that position i < t, the token at that position is replaced with a mask token [MASK]; if the random number u generated for that position iIf ≥t, then the original word element remains unchanged.
[0095] For example, for an original sequence [A, B, C, D, E] of length L=5, with a time step t=0.4, assume the five generated random numbers are u1=0.5, u2=0.6, u3=0.3, u4=0.7, and u5=0.1. Since u1=0.5≥0.4, position 1 is retained; u2=0.6≥0.4, position 2 is retained; u3=0.3<0.4, position 3 is masked; u4=0.7≥0.4, position 4 is retained; and u5=0.1<0.4, position 5 is masked. The final masked positions are 3 and 5, a total of 2 positions, resulting in an actual masking ratio of 2 / 5=0.4=40%, which is 40%. In practical applications, due to randomness, the actual ratio may fluctuate around the expected value of 40%. For example, if u1 = 0.2, since 0.2 < 0.4, position 1 is masked. In this case, the final masked positions are 1, 3, and 5, a total of 3 positions, and the actual mask ratio is 3 / 5 = 0.6 = 60%. In practical applications, if an example closer to the expected value is needed, random numbers can be regenerated until the mask ratio equals t.
[0096] In the embodiments described in this specification, whether each position is masked is determined independently and randomly, and the probability of each position being masked is t. Therefore, the expected proportion of masked positions in the entire sequence is exactly equal to t or fluctuates around t. Furthermore, because each position is determined independently and randomly, the masked positions are not concentrated in a certain segment, but are distributed relatively evenly throughout the sequence, avoiding the problems that may arise from fixed patterns such as continuous masking.
[0097] In practical applications, the basic mask diffusion loss for each of the mask sequences can be calculated using the following formula:
[0098] in, Based on the diffusion loss of the base mask, Let L be the mathematical expectation, and L be the length of the mask sequence or the original sequence corresponding to the mask sequence. This is an indicator function; if the i-th position is masked (i.e., [MASK]), the value is 1. If the i-th position is not masked (the original word is preserved), the value is 0. This ensures that only the masked positions are considered for loss. For unmasked positions, multiplying by 0 eliminates the loss from the sum. θ represents the model parameters of the target large language model. The original sequence, It is the masked sequence after masking. As the inner normalization factor, This represents a target large language model with parameters θ, represented by a mask sequence. Given the input, predict whether the i-th position is a real word. The conditional probability distribution.
[0099] In the embodiments of this specification, here This is an inner normalization factor used to offset the dimensional differences caused by the varying number of masked positions over time step t. When time step t is small (e.g., 0.1), the number of masked words is small (e.g., 10%), and the summation result would be small without dividing by time step t. When time step t is large (e.g., 0.9), the number of masked words is large (e.g., 90%), and the summation result would be large without dividing by time step t. The inner normalization factor ensures that regardless of the time step t, the basic masking diffusion loss reflects the average prediction error per masked position.
[0100] In practical applications, after obtaining the basic mask diffusion loss of each mask sequence, the target loss value can be calculated based on the reweighting factor of each mask sequence and the basic mask diffusion loss of each mask sequence.
[0101] In one optional implementation of this embodiment, calculating the target loss value based on the reweighting factor of each of the mask sequences and the basic mask diffusion loss of each of the mask sequences may specifically include: For each of the mask sequences, the weighted loss of the mask sequence is obtained by multiplying the basic mask diffusion loss of the mask sequence by the reweighting factor of the mask sequence.
[0102] Based on the weighted loss of each of the mask sequences, the average value of the weighted loss is calculated to obtain the target loss value.
[0103] In the embodiments described in this specification, the weighted loss is the loss value obtained by multiplying the basic mask diffusion loss of a single mask sequence by the reweighting factor corresponding to that mask sequence. The weighted loss is used to reflect the degree of contribution that mask sequence is assigned to the total loss, i.e., the target loss value.
[0104] Specifically, the weighted loss i of a single mask sequence can be calculated using the following formula:
[0105] in, The weighted loss of the mask sequence, This is the reweighting factor for the mask sequence. It can be One of them; the physical meanings of the other parameters have been detailed above and will not be repeated here.
[0106] In the embodiments described in this specification, for In this case, As the role of outer layer factors and inner layer factors The functions are completely different, inner factor This is to normalize the dimensions of the loss term, outer factor. This is to dynamically adjust the relative contribution of the basic diffusion loss term of each mask sequence to the overall gradient based on the noise level.
[0107] In practical applications, the weighted losses of all mask sequences can be summed and then divided by the total number of mask sequences (i.e., the number of samples in the current batch) to obtain the target loss value. In other words, the target loss value is the average of all weighted losses.
[0108] Specifically, assuming the current batch contains N mask sequences, the basic loss of the j-th mask sequence is L. base(j) The reweighting factor is w(t) j If ), then the weighted loss of the j-th mask sequence is w(t). j )*L base(j) Then, the arithmetic mean of these N weighted losses can be calculated. For example, suppose the current batch has 3 mask sequences: Mask sequence A (t=0.2, base loss=0.5, weight=5) weighted loss=5*0.5=2.5; Sequence B (t=0.5, base loss=1.0, weight=2) weighted loss=1.0*2=2.0; Mask sequence C (t=0.8, base loss=2.0, weight=1.25) weighted loss=2.0*1.25=2.5; then the target loss value=(2.5+2.0+2.5) / 3≈2.33.
[0109] One or more embodiments in this specification calculate the target loss value using an averaging method, making the magnitude of the total gradient independent of the batch size. Regardless of whether the current batch contains 16 or 64 mask sequences, the contribution coefficient of a single sample decreases proportionally with the increase in the number of samples, keeping the expected magnitude of the total gradient stable. This achieves decoupling of the target loss value and gradient from the batch size, allowing the model to train stably without adjusting the learning rate and simplifying the hyperparameter tuning process.
[0110] As one implementation method, the sum of the weighted losses of each mask sequence can be directly used as the target loss value. In this case, as long as the number of mask sequences in each training batch remains consistent (e.g., fixed at 32, 64, or other constants), the dimensions of the target loss value and the gradient magnitude can remain relatively stable during training. Specifically, the weighted loss of each mask sequence can be calculated first as described above, and then these weighted losses can be summed to obtain a scalar sum value, which can be used as the target loss value for that batch.
[0111] In the initial stage of pre-training from an autoregressive large language model to a target large language model, the model lacks experience in handling high-noise samples, and its gradient direction may exhibit significant randomness. Although this application suppresses the gradient contribution of high-noise samples through a monotonically decreasing weighting factor, to avoid potential spikes caused by the superposition of gradients from multiple such samples, and thus to prevent drastic parameter updates or even gradient explosion, specific training methods for the target large language model are provided in the embodiments of this specification.
[0112] In one optional implementation of this embodiment, training the target large language model based on the target loss value may specifically include: Backpropagation is performed on the target loss value to calculate the original gradients of each training parameter in the target large language model.
[0113] Calculate the global gradient norm of the original gradient.
[0114] If the global gradient norm is greater than a preset clipping threshold, a scaling factor is calculated based on the ratio of the preset clipping threshold to the global gradient norm.
[0115] The original gradient is scaled using the scaling factor to obtain the clipped gradient.
[0116] If the global gradient norm is less than or equal to the preset clipping threshold, then the original gradient is used as the clipped gradient.
[0117] The model parameters of the target large language model are updated using the pruned gradient.
[0118] In the embodiments described in this specification, backpropagation refers to using the chain rule to calculate the partial derivatives of the loss with respect to each trainable parameter of the target large language model, starting from the target loss value and proceeding layer by layer. The original gradient is composed of the partial derivatives obtained above. The global gradient norm refers to the L2 norm of a vector after flattening the original gradient values of all trainable parameters into a one-dimensional vector. The global gradient norm is a scalar used to measure the overall magnitude of the gradients of all parameters in the model. The larger the global gradient norm, the larger the overall gradient magnitude of the target large language model (i.e., the larger the sum of the step sizes required to update all parameters); the smaller the global gradient norm, the smaller the gradient magnitude.
[0119] The preset clipping threshold is a pre-set positive number that serves as a boundary for judging whether the gradient is too large.
[0120] If the global gradient norm is less than or equal to the preset clipping threshold, it means that the current gradient size is within a safe range. The original gradient can be used directly to update the parameters without disrupting the original optimization direction.
[0121] If the global gradient norm exceeds the preset pruning threshold, it indicates that due to the uncertainty of the random sampling time step, the mask sequence set may contain a large number of high-noise samples. Although the gradients of these high-noise samples are suppressed by the weighting factor, the superposition of multiple high-noise samples still produces spikes, or the base mask diffusion loss of a certain sample is abnormally large. This leads to the overall magnitude of the original gradients of all parameters in the current batch being too large, exceeding the preset safety range. In practical applications, if no gradient processing is performed, the update step size will exceed a reasonable range, causing the parameters to exceed the optimum or even diverge.
[0122] If the global gradient norm exceeds the preset clipping threshold, the scaling factor can be obtained by dividing the preset clipping threshold by the global gradient norm. Then, the original gradient value of each parameter is multiplied by the scaling factor to scale the original gradient value of each parameter, resulting in the clipped gradient. The clipped gradient is then passed to the optimizer (such as AdamW) for parameter updates.
[0123] In the embodiments described in this specification, when the gradient exceeds the safe range, the gradient is pruned to ensure that the parameter update step size is limited to a controllable range, thereby preventing training divergence caused by gradient explosion. Furthermore, a scaling factor is used to scale the original gradients of all parameters proportionally, ensuring that the direction of the pruned gradient is completely consistent with the original gradient direction. This does not compromise the correctness of the optimization path and also preserves the relative proportions of gradients between layers.
[0124] As an alternative implementation, to avoid gradient spikes caused by too many high-noise samples (i.e., mask sequences with large time steps) in a training batch due to random sampling, the sampling distribution of time steps can be constrained when constructing the mask sequence set to control the proportion of high-noise samples. Specifically, an upper limit R can be preset. max (e.g. R) max =0.3), indicating that the time step in the current batch is greater than a certain high noise threshold t. high (e.g. t) high The number of mask sequences (=0.7) must not exceed R of the total batch size. max When sampling time steps, a batch of candidate time steps can be randomly sampled from the preset time step interval ([0,1]), and then the time steps greater than t can be counted. high The number of; if it exceeds R max If the sample size is *N, then the batch of samples is rejected. Here, N is the batch size. Time steps t can be resampled until the proportion constraint is met, or the time steps of the excess samples can be forcibly replaced with a smaller value, such as randomly selecting from [0, t]. highResampling is used in training batches to ensure that the number of noisy samples in each batch is limited to a controllable range. Even if the weighting factor has suppressed the gradient contribution of noisy samples, the total gradient after the superposition of multiple noisy samples will still not exceed the safety boundary, thus effectively preventing gradient explosion without using gradient clipping. It should be noted that this method is not mutually exclusive with gradient clipping, and the two can be used together to further improve robustness.
[0125] In the embodiments described in this specification, by controlling the upper limit of the proportion of high-noise samples in each batch, the concentrated occurrence of large gradient samples is limited from the source, thereby reducing the risk of gradient explosion.
[0126] In practical applications, the preset clipping threshold is usually set to a fixed constant during training. However, a fixed threshold often fails to account for the gradient characteristics of samples with different noise levels.
[0127] Based on this, in an optional implementation of this embodiment, the preset pruning threshold is determined according to the statistical distribution of the time steps corresponding to each mask sequence in the mask sequence set, and the method may further include: If the proportion of mask sequences in the mask sequence set whose time steps are less than a preset time step threshold is greater than a preset proportion threshold, then the preset pruning threshold is increased.
[0128] In the embodiments of this specification, the preset pruning threshold is a variable dynamically adjusted based on the time step information of all samples in the current training batch (i.e., the mask sequence set). The preset time step threshold can be a pre-set time step value (e.g., 0.3), used to distinguish between easy and hard samples: samples with a time step less than this threshold are considered low-noise easy samples, and samples with a time step greater than this threshold are considered high-noise hard samples. The preset ratio threshold is a pre-set ratio value (e.g., 0.7 or 0.8), used to determine whether easy samples dominate in the current batch.
[0129] In practical applications, we can first iterate through all mask sequences in the current batch (i.e., the set of mask sequences), count the number of samples with a time step less than a preset time step threshold (e.g., 0.3), and calculate their proportion. If the proportion is greater than a preset proportion threshold (e.g., 0.8), it means that simple samples account for the majority in the current batch, so we perform the operation of increasing the pruning threshold, such as multiplying the current pruning threshold by 1.2 or adding a fixed increment; otherwise, we keep the threshold unchanged or use the default value.
[0130] In practical applications, the gradient magnitude of simple samples is usually small and the direction is stable, making them less prone to gradient explosion. When simple samples dominate, the risk of the entire batch is low. In this case, the pruning restrictions can be appropriately relaxed (i.e., the threshold can be increased) so that the gradients of simple samples can participate in parameter updates with a more complete magnitude, avoiding the loss of effective learning signals due to overly strict pruning. Conversely, if there are many difficult samples, the threshold should be maintained or decreased to strictly prevent gradient spikes. For example, suppose a batch has 64 mask sequences, with a preset time step threshold of 0.3 and a preset proportion threshold of 0.6. Statistical analysis shows that 50 samples have a time step less than 0.3, accounting for about 78%, and more than 60%. The current pruning threshold can be increased from the default 1.0 to 1.2. At this time, the gradients of simple samples that were originally slightly greater than 1.0 will no longer be pruned, fully preserving their optimization contribution, while the gradients of difficult samples will still be controlled within 1.0, and training remains stable. This implementation complements the reweighting factor: the reweighting factor assigns higher weights to simple samples at the loss level, while the adaptive pruning threshold protects the gradients of these high-weight samples from being mispruned at the gradient level. The two work together to improve training performance.
[0131] In the embodiments described in this specification, the pruning threshold is adaptively adjusted to match the gradient pruning operation with the batch difficulty distribution, taking into account both training stability and learning efficiency; in addition, when simple samples dominate, unnecessary gradient suppression is avoided, thus accelerating model convergence.
[0132] The training method for the target large language model provided in the embodiments of this specification can enable the model to learn a "from easy to difficult, gradually denoising" mindset during the model training process. That is, it prioritizes mastering the reconstruction ability under low-noise samples and then smoothly transitions to high-noise samples. This can significantly improve the training stability and convergence efficiency of the model, avoid gradient oscillation or divergence caused by the superposition of gradients from high-noise samples, and thus obtain better model parameters.
[0133] The training method for a target large language model provided in this specification involves: acquiring a set of mask sequences; inputting each mask sequence into the target large language model to obtain the predicted distribution of the masked positions; determining a reweighting factor that monotonically decreases with increasing time step based on the time step of each mask sequence; calculating the basic mask diffusion loss based on the predicted distribution; calculating the target loss value based on the reweighting factor and the basic mask diffusion loss; and finally training the model based on the target loss value. In this specification, the reweighting factor, which monotonically decreases with time step, ensures that samples with smaller time steps, lower mask ratios, and easier reconstruction have larger weights; conversely, samples with larger time steps and greater reconstruction difficulty have smaller weights. This automatically amplifies the gradient contribution of low-noise samples and compresses the gradient contribution of high-noise samples during backpropagation, avoiding interference from the stochastic gradients of high-noise samples on the optimization direction in the early stages of training. This improves the gradient stability and convergence smoothness of the target large language model during further pre-training.
[0134] Furthermore, the reweighting factor is directly defined as a monotonically decreasing function of time step t. Since time step t is already a variable that must be sampled during mask diffusion training, no additional computational or storage overhead is required, achieving zero-cost improvement. Without changing the model structure or increasing the learnable parameters, the stability and final performance of the target large language model training are significantly improved.
[0135] Once the target large language model is trained, it can be used to generate text.
[0136] Corresponding to the above method embodiments, this specification also provides an embodiment of a text generation method based on a target large language model. Figure 3 A flowchart of a text generation method based on a target large language model, provided in one embodiment of this specification, is shown.
[0137] Step S302: Obtain the initial noise sequence.
[0138] In the embodiments of this specification, the initial noise sequence can be a discrete sequence composed of mask symbols [MASK] and optional known terms. In practical applications, for unconditional generation, such as randomly writing a paragraph, a fixed-length sequence consisting entirely of mask symbols [MASK] can be directly constructed; for conditional generation, such as text filling and code completion, the known context provided by the user is filled into the corresponding positions, and the remaining unknown positions are filled with [MASK], thereby forming the initial noise sequence.
[0139] Step S304: Input the initial noise sequence into the target large language model to generate the target sequence corresponding to the initial noise sequence, wherein the target large language model is trained by a training method based on the target large language model.
[0140] The target large language model is a model that has obtained the ability to recover complete sequences from noisy sequences through the foregoing training method. The target large language model is a large language model capable of generating token sequences in parallel. Specifically, the target large language model may be a diffusion large language model, or the target large language model may also be another large language model that can parallelly generate token sequences through iterative denoising, which is not specifically limited herein.
[0141] In practical applications, the generation of the target sequence by the target large language model is not completed in one step, but is an iterative denoising process. Starting from an initial noise sequence, the model repeatedly predicts the token probability distribution at all masked positions, and gradually replaces high-confidence [MASK] with specific tokens until the sequence no longer contains any mask symbols.
[0142] The target sequence is a complete token sequence obtained after complete denoising and does not contain any [MASK], that is, the finally generated text or code. For example, for an unconditional generation task, the initial noise sequence can be set as a fully masked sequence of length 10: [MASK][MASK] [MASK][MASK] [MASK][MASK] [MASK][MASK] [MASK][MASK]; after iterative denoising by the model, the target sequence [To][day] [wea][ther] [sun][ny] [,][sun] [ny] may be obtained, which corresponds to the natural language "Today is sunny and pleasant". For a text infilling task, if the user inputs "I like to eat [MASK][MASK]", the initial noise sequence is [I] [like] [to eat] [MASK][MASK], and the model can complete it as [I] [like] [to eat] [ap][ple].
[0143] By utilizing the denoising capability of the trained target large language model, high-quality text can be generated efficiently and flexibly, which is suitable for various downstream scenarios such as automatic writing, code completion, and dialogue generation.
[0144] The following description is combined with Figure 4 , and further illustrates the text generation method by taking the application of the text generation method provided in this specification on an intelligent writing assistance platform as an example. Figure 4 shows a sequential flowchart of a text generation method based on a target large language model provided in one embodiment of this specification.
[0145] The text generation method can be applied to an intelligent writing assistance platform, which comprises a client and a server. The client is a terminal device for a user to input a text fragment and receive the complete text returned by the server, such as a web editor and a mobile phone APP; the server is a backend system that provides text generation services.
[0146] A client, configured to obtain an original text segment input by a user and send the original text segment to a server.
[0147] In the embodiments of this specification, the original text segment may be several sentences input by a user through a keyboard, an incomplete sentence, or a segment of text to be continued selected from a document. For example, a user inputs "Once upon a time there was a brave kitten that decided to go...".
[0148] A server, configured to construct an initial noise sequence by using the original text segment as known context. Specifically, original tokens may be retained for the known part, and the unknown part may be filled with the [MASK] symbol.
[0149] Then, input the initial noise sequence into a target large language model to generate a complete text sequence, and return the generated target text to the client.
[0150] Wherein, the target large language model is obtained through pre-training using the training method of the target large language model described above. The specific training process is as follows: first, obtain a masked sequence set; wherein, in each masked sequence in the masked sequence set, the proportion of masked tokens is positively correlated with a time step; input each masked sequence into the target large language model to obtain a prediction distribution of the masked tokens in each masked sequence; determine a reweighting factor for each masked sequence based on the time step of each masked sequence; wherein, the reweighting factor decreases monotonically as the time step increases within a preset time step interval; calculate a basic mask diffusion loss for each masked sequence based on the prediction distribution; calculate a target loss value based on the reweighting factor of each masked sequence and the basic mask diffusion loss of each masked sequence; and train the target large language model based on the target loss value.
[0151] For example, the original text segment input by the user is "The weather is really nice today, let's go", the server constructs it as an initial noise sequence [Today] [weather] [is] [really] [nice] [,][we] [go][MASK][MASK]... (the rear part is filled with [MASK]), then calls the trained target large language model to perform multi-step iterative denoising. In each iteration, the model uses bidirectional attention to predict all masked positions at the same time, gradually replaces [MASK] with specific tokens, and finally generates the target text "The weather is really nice today, let's go for a walk in the park.", and returns it to the client to display to the user.
[0152] Through this application example, the training method of the present invention can significantly improve the fluency and accuracy of intelligent writing assistance, and has higher generation efficiency and flexibility especially when handling tasks such as long text completion, text filling, and code generation.
[0153] See Figure 5 , Figure 5 This document illustrates an overall flowchart of a training method for a target large language model according to one embodiment of this specification. Figure 5 As shown, the training method for a target large language model can include the following steps: Step 502: Obtain the original sequence set.
[0154] The original sequence set can be one or more clean text sequences from the current training batch, randomly sampled from a large-scale training corpus, and each original sequence consists of discrete tokens.
[0155] Step 504: For each original sequence in the original sequence set, randomly sample a time step t from the preset time step interval.
[0156] The time step t determines the masking ratio in subsequent masking processing. The larger the time step t, the higher the masking ratio, indicating a higher noise level; conversely, the smaller the time step t, the lower the masking ratio, indicating a lower noise level.
[0157] Step 506: Perform masking processing on each of the original sequences according to the time steps corresponding to each of the original sequences to obtain a set of masked sequences.
[0158] In practical applications, the original sequence is masked according to the time step tt corresponding to each original sequence to obtain a masked sequence, and then a set of masked sequences is obtained.
[0159] Specifically, for each position in the original sequence, the word at that position can be replaced with the mask symbol [MASK] with probability t, and with probability 1... t preserves the original words. The proportion of masked words in the masked sequence is positively correlated with time step t.
[0160] Step 508: Input each mask sequence into the target large language model to obtain the predicted distribution of the masked words in each mask sequence.
[0161] In practical applications, a target large language model can use forward propagation to calculate and output the predicted distribution for each masked position in each masked sequence. This predicted distribution represents the probability that the model infers that position represents a word in the vocabulary.
[0162] The target large language model is a large language model capable of generating word sequences in parallel. Specifically, the target large language model can be a diffusion-based large language model, or it can be any other large language model that can generate word sequences in parallel through iterative denoising; there are no specific limitations on this.
[0163] Step 510: Based on the predicted distribution of each mask sequence and the corresponding real label, calculate the basic mask diffusion loss for each mask sequence.
[0164] The true label refers to the original word element at the masked position in the original sequence.
[0165] The basic masking diffusion loss can be expressed as a negative log-likelihood, including an inner normalization factor of 1 / t. This inner normalization factor is used to eliminate the dimensional effects caused by differences in the number of masks.
[0166] Step 512: Determine the reweighting factor for each mask sequence based on the time step of each mask sequence.
[0167] Specifically, the reweighting factor w(t) of each mask sequence can be determined based on the time step t corresponding to each mask sequence.
[0168] Within the preset time step interval, w(t) monotonically decreases as time step t increases. The reweighting factor can be... One of them.
[0169] Step 514: Calculate the target loss value based on the reweighting factor of each of the mask sequences and the base mask diffusion loss of each of the mask sequences.
[0170] For each mask sequence, the base mask diffusion loss of that mask sequence can be multiplied by the reweighting factor of that mask sequence to obtain the weighted loss of that mask sequence. Then, the average of the weighted losses of all mask sequences in the current batch is calculated as the target loss value.
[0171] Step 516: Perform backpropagation on the target loss value to calculate the original gradients of each trainable parameter in the target large language model.
[0172] Step 518: Calculate the global gradient norm of the original gradient. This global gradient norm reflects the magnitude of the overall gradient of the model.
[0173] Step 520: Determine if the global gradient norm is greater than the preset clipping threshold. If yes, proceed to step 522; otherwise, proceed to step 524.
[0174] Step 522: Calculate the scaling factor based on the ratio of the preset clipping threshold to the global gradient norm, and use the scaling factor to scale the original gradient proportionally to obtain the clipped gradient.
[0175] Where the scaling factor is less than 1, the global gradient norm after scaling is equal to the preset clipping threshold.
[0176] Step 524: Use the original gradient or the clipped gradient as the final gradient, and use the optimizer to update the model parameters of the target large language model.
[0177] Step 526: Determine whether the preset total number of training steps has been reached or whether the model's performance on the validation set no longer improves. If yes, end the training; otherwise, return to step 502.
[0178] Corresponding to the above-described embodiments of the training method for the target large language model, this specification also provides embodiments of the training apparatus for the target large language model. Figure 6 This specification illustrates a schematic diagram of a training apparatus for a target large language model according to one embodiment. Figure 6 As shown, the device may include: The acquisition module 602 is configured to acquire a set of mask sequences; the proportion of masked words in each mask sequence of the set of mask sequences is positively correlated with the time step.
[0179] The prediction module 604 is configured to input each of the mask sequences into the target large language model to obtain the predicted distribution of the masked lexical units in each of the mask sequences.
[0180] The determining module 606 is configured to determine the reweighting factor of each of the mask sequences based on the time steps of each of the mask sequences; within a preset time step interval, the reweighting factor monotonically decreases as the time step increases.
[0181] The basic loss calculation module 608 is configured to calculate the basic mask diffusion loss of each of the mask sequences based on the predicted distribution; the basic mask diffusion loss is configured to reflect the difference between the predicted distribution and the true label.
[0182] The target loss calculation module 610 is configured to calculate the target loss value based on the reweighting factor of each of the mask sequences and the basic mask diffusion loss of each of the mask sequences.
[0183] Training module 612 is configured to train the target large language model based on the target loss value.
[0184] In an optional embodiment, the reweighting factor is One of them; among them, For the time step, , , It is a constant greater than 0.
[0185] In an optional embodiment, the acquisition module 602 may further include: The original sequence acquisition unit is configured to acquire a set of original sequences.
[0186] The time step sampling unit is configured to determine the time step corresponding to each original sequence in the original sequence set from the preset time step interval for each original sequence in the original sequence set.
[0187] The masking unit is configured to perform masking processing on each of the original sequences according to the time steps corresponding to each of the original sequences, so as to obtain a set of masked sequences.
[0188] In one optional embodiment, the preset time step interval is [0.1, 1].
[0189] In an optional embodiment, for any of the original sequences, the masking unit is further configured to: For any position i in the original sequence, generate a random number u. i .
[0190] If the random number u i If the time step is less than the specified time step, then the term at position i is replaced with a mask symbol.
[0191] If the random number u i If the time step is greater than or equal to the stated time step, then the original lexical unit at position i is retained unchanged.
[0192] In an optional embodiment, the target loss calculation module 610 is further configured to: For each of the mask sequences, the weighted loss of the mask sequence is obtained by multiplying the basic mask diffusion loss of the mask sequence by the reweighting factor of the mask sequence.
[0193] Based on the weighted loss of each of the mask sequences, the average value of the weighted loss is calculated to obtain the target loss value.
[0194] In an optional embodiment, the training module 612 is further configured to: Backpropagation is performed on the target loss value to calculate the original gradients of each training parameter in the target large language model.
[0195] Calculate the global gradient norm of the original gradient.
[0196] If the global gradient norm is greater than a preset clipping threshold, a scaling factor is calculated based on the ratio of the preset clipping threshold to the global gradient norm.
[0197] The original gradient is scaled using the scaling factor to obtain the clipped gradient.
[0198] If the global gradient norm is less than or equal to the preset clipping threshold, then the original gradient is used as the clipped gradient.
[0199] The model parameters of the target large language model are updated using the pruned gradient.
[0200] In one optional embodiment, the preset pruning threshold is determined based on the statistical distribution of the time steps corresponding to each mask sequence in the mask sequence set, such as... Figure 6 The apparatus shown may also include: The preset cropping threshold adjustment module is configured to increase the preset cropping threshold if the proportion of mask sequences in the mask sequence set whose time steps are less than a preset time step threshold is greater than a preset proportion threshold.
[0201] In one optional embodiment, the initial weights of the target large language model are the weights of a pre-trained autoregressive large language model.
[0202] The training apparatus for the target large language model provided in this specification's embodiments acquires a set of mask sequences; inputs each mask sequence into the target large language model to obtain the predicted distribution of the masked positions; determines a reweighting factor that monotonically decreases with increasing time step based on the time step of each mask sequence; calculates the basic mask diffusion loss based on the predicted distribution; calculates the target loss value based on the reweighting factor and the basic mask diffusion loss; and finally trains the model based on the target loss value. In this specification's embodiments, the reweighting factor, which monotonically decreases with time step, ensures that samples with smaller time steps, lower mask ratios, and lower reconstruction difficulty have larger weights; conversely, samples with larger time steps and higher reconstruction difficulty have smaller weights. This automatically amplifies the gradient contribution of low-noise samples and compresses the gradient contribution of high-noise samples during backpropagation, avoiding interference from the stochastic gradients of high-noise samples on the optimization direction in the early stages of training, thereby improving the gradient stability and convergence smoothness of the target large language model during continued pre-training.
[0203] Furthermore, the reweighting factor is directly defined as a monotonically decreasing function of time step t. Since time step t is already a variable that must be sampled during mask diffusion training, no additional computational or storage overhead is required, achieving zero-cost improvement. Without changing the model structure or increasing the learnable parameters, the stability and final performance of the target large language model training are significantly improved.
[0204] The above is an illustrative scheme of a training device for a target large language model according to this embodiment. It should be noted that the technical solution of this training device for a target large language model and the technical solution of the training method for a target large language model described above belong to the same concept. For details not described in detail in the technical solution of the training device for a target large language model, please refer to the description of the technical solution of the training method for a target large language model described above.
[0205] Corresponding to the above embodiments of the text generation method based on the target large language model, this specification also provides embodiments of the text generation apparatus based on the target large language model. Figure 7 This diagram illustrates a structural schematic of a text generation apparatus based on a target large language model, according to one embodiment of this specification. Figure 7 As shown, the device may include: The noise sequence acquisition module 702 is configured to acquire an initial noise sequence; The text generation module 704 is configured to input the initial noise sequence into the target large language model to generate the target sequence corresponding to the initial noise sequence, wherein the target large language model is the model trained by the method described above.
[0206] The above is an illustrative scheme of a text generation device based on a target large language model according to this embodiment. It should be noted that the technical solution of this text generation device based on a target large language model and the technical solution of the text generation method based on a target large language model described above belong to the same concept. For details not described in detail in the technical solution of the text generation device based on a target large language model, please refer to the description of the technical solution of the text generation method based on a target large language model described above.
[0207] Figure 8 A structural block diagram of a computing device 800 according to one embodiment of this specification is shown. The components of the computing device 800 include, but are not limited to, a memory 810 and a processor 820. The processor 820 is connected to the memory 810 via a bus 830, and a database 850 is used to store data.
[0208] The computing device 800 also includes an access device 840, which enables the computing device 800 to communicate via one or more networks 860. Examples of these networks include Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or combinations of communication networks such as the Internet. The access device 840 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Wi-MAX (Worldwide Interoperability for Microwave Access) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0209] In one embodiment of this specification, the above-described components of the computing device 800 and Figure 8 Other components, not shown, can also be connected to each other, for example, via a bus. It should be understood that... Figure 8 The block diagram of the computing device shown is for illustrative purposes only and is not intended to limit the scope of this specification. Those skilled in the art can add or replace other components as needed.
[0210] The computing device 800 can be any type of stationary or mobile computing device, including mobile computers or mobile computing devices (e.g., tablet computers, personal digital assistants, laptop computers, notebook computers, netbooks, etc.), mobile phones (e.g., smartphones), wearable computing devices (e.g., smartwatches, smart glasses, etc.) or other types of mobile devices, or stationary computing devices such as desktop computers or personal computers (PCs). The computing device 800 can also be a mobile or stationary server.
[0211] The processor 820 is configured to execute the following computer-executable instructions, which, when executed by the processor, implement the steps of the above-mentioned training method for the target large language model or the text generation method based on the target large language model.
[0212] The above is an illustrative scheme of a computing device according to this embodiment. It should be noted that the technical solution of this computing device belongs to the same concept as the technical solution of the target large language model training method or the text generation method based on the target large language model described above. For details not described in detail in the technical solution of the computing device, please refer to the description of the technical solution of the target large language model training method or the text generation method based on the target large language model described above.
[0213] An embodiment of this specification also provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the above-described training method for a target large language model or a text generation method based on a target large language model.
[0214] The above is an illustrative scheme of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium belongs to the same concept as the technical solution of the target large language model training method or the text generation method based on the target large language model described above. For details not described in detail in the technical solution of the storage medium, please refer to the description of the technical solution of the target large language model training method or the text generation method based on the target large language model described above.
[0215] An embodiment of this specification also provides a computer program product, including a computer program or instructions that, when executed by a processor, implement the steps of the above-described training method for a target large language model or a text generation method based on a target large language model.
[0216] The above is an illustrative scheme of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product belongs to the same concept as the above-described training method for a target large language model or the text generation method based on a target large language model. For details not described in detail in the technical solution of the computer program product, please refer to the description of the above-described training method for a target large language model or the text generation method based on a target large language model.
[0217] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0218] The computer instructions include computer program code, which may be in the form of source code, object code, executable file, or certain intermediate forms. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drive, portable hard drive, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium may be appropriately added or removed according to the requirements of patent practice. For example, in some regions, according to patent practice, computer-readable media may not include electrical carrier signals and telecommunication signals.
[0219] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments in this specification are not limited to the described order of actions, because according to the embodiments in this specification, some steps can be performed in other orders or simultaneously. Furthermore, those skilled in the art should also understand that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments in this specification.
[0220] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0221] The preferred embodiments disclosed above are merely illustrative of this specification. Optional embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the embodiments described in this specification. These embodiments are selected and specifically described in this specification to better explain the principles and practical applications of the embodiments, thereby enabling those skilled in the art to better understand and utilize this specification.
Claims
1. A training method for a target large language model, wherein the target large language model is used to generate word sequence in parallel, characterized in that, The training method includes: Obtain a set of mask sequences; the proportion of masked words in each mask sequence of the set of mask sequences is positively correlated with the time step; Each of the mask sequences is input into the target large language model to obtain the predicted distribution of the masked words in each of the mask sequences; Based on the time steps of each of the mask sequences, a reweighting factor for each of the mask sequences is determined; within a preset time step interval, the reweighting factor monotonically decreases as the time step increases; Based on the predicted distribution, the basic mask diffusion loss of each of the mask sequences is calculated; the basic mask diffusion loss is used to reflect the difference between the predicted distribution and the true label; The target loss value is calculated based on the reweighting factor of each of the mask sequences and the base mask diffusion loss of each of the mask sequences; The target large language model is trained based on the target loss value.
2. The training method for the target large language model as described in claim 1, characterized in that, The reweighting factor is One of them; among them, For the time step, , , It is a constant greater than 0.
3. The training method for the target large language model as described in claim 1, characterized in that, The acquisition of the mask sequence set includes: Obtain the original set of sequences; For each original sequence in the original sequence set, determine the time step corresponding to each original sequence in the original sequence set from the preset time step interval; Each original sequence is masked according to its corresponding time step to obtain a set of masked sequences.
4. The training method for the target large language model as described in claim 3, characterized in that, The preset time step interval is [0.1, 1].
5. The training method for the target large language model as described in claim 3, characterized in that, For any of the original sequences, the masking process performed on each of the original sequences according to the time step corresponding to each of the original sequences includes: For any position i in the original sequence, generate a random number u. i ; If the random number u i If the time step is less than the specified time step, then the word at position i is replaced with a mask symbol; If the random number u i If the time step is greater than or equal to the stated time step, then the original lexical unit at position i is retained unchanged.
6. The training method for the target large language model as described in claim 1, characterized in that, The calculation of the target loss value based on the reweighting factor of each of the mask sequences and the basic mask diffusion loss of each of the mask sequences includes: For each of the mask sequences, the weighted loss of the mask sequence is obtained by multiplying the basic mask diffusion loss of the mask sequence by the reweighting factor of the mask sequence; Based on the weighted loss of each of the mask sequences, the average value of the weighted loss is calculated to obtain the target loss value.
7. The training method for the target large language model as described in claim 6, characterized in that, The training of the target large language model based on the target loss value includes: Backpropagation is performed on the target loss value to calculate the original gradient of each training parameter in the target large language model; Calculate the global gradient norm of the original gradient; If the global gradient norm is greater than the preset clipping threshold, then the scaling factor is calculated based on the ratio of the preset clipping threshold to the global gradient norm. The original gradient is scaled using the scaling factor to obtain the clipped gradient; If the global gradient norm is less than or equal to the preset clipping threshold, then the original gradient is used as the clipped gradient. The model parameters of the target large language model are updated using the pruned gradient.
8. The training method for the target large language model as described in claim 7, characterized in that, The preset pruning threshold is determined based on the statistical distribution of the time steps corresponding to each mask sequence in the mask sequence set, and the method further includes: If the proportion of mask sequences in the mask sequence set whose time steps are less than a preset time step threshold is greater than a preset proportion threshold, then the preset pruning threshold is increased.
9. The training method for the target large language model as described in claim 1, characterized in that, The initial weights of the target large language model are the weights of the pre-trained autoregressive large language model.
10. The training method for the target large language model as described in claim 1, characterized in that, The target large language model is a diffusion-type large language model.
11. A text generation method based on a target large language model, wherein the target large language model is used to generate word sequences in parallel, characterized in that, The text generation method includes: Obtain the initial noise sequence; The initial noise sequence is input into the target large language model to generate the target sequence corresponding to the initial noise sequence, wherein the target large language model is a model trained based on the method of any one of claims 1-10.
12. A training apparatus for a target large language model, wherein the target large language model is used to generate word sequence in parallel, characterized in that, The training device includes: The acquisition module is configured to acquire a set of mask sequences; the proportion of masked words in each mask sequence of the mask sequence set is positively correlated with the time step; The prediction module is configured to input each of the mask sequences into the target large language model to obtain the predicted distribution of the masked words in each of the mask sequences; The determination module is configured to determine the reweighting factor of each of the mask sequences based on the time steps of each of the mask sequences; within a preset time step interval, the reweighting factor decreases monotonically as the time step increases; The basic loss calculation module is configured to calculate the basic mask diffusion loss for each of the mask sequences based on the predicted distribution; the basic mask diffusion loss is configured to reflect the difference between the predicted distribution and the true label; The target loss calculation module is configured to calculate the target loss value based on the reweighting factor of each of the mask sequences and the basic mask diffusion loss of each of the mask sequences; The training module is configured to train the target large language model based on the target loss value.
13. A text generation device based on a target large language model, wherein the target large language model is used to generate word sequences in parallel, characterized in that, The text generation device includes: The noise sequence acquisition module is configured to acquire an initial noise sequence; The text generation module is configured to input the initial noise sequence into the target large language model to generate a target sequence corresponding to the initial noise sequence, wherein the target large language model is a model trained based on the method of any one of claims 1-10.
14. A computing device, characterized in that, include: Memory and processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the method according to any one of claims 1 to 11.
15. A computer-readable storage medium, characterized in that, It stores computer-executable instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.
16. A computer program product, characterized in that, It includes a computer program or instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 11.