Model training method, text generation method, device, equipment and medium

CN118468976BActive Publication Date: 2026-10-09THUNDERSOFT (NANJING) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410674264.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-28
Publication Date
2026-10-09
Estimated Expiration
2044-05-28

AI Technical Summary

Technical Problem

[0005]在实际应用中,较多的迭代次数和训练时间不仅会增加大语言模型在训练过程中的数据传输量

Benefits of technology

[0059] In the technical solution of this application embodiment, during the model training process, the column dimensions of the query weight matrix and the keyword weight matrix are first normalized. Then, the column vectors in the normalized query weight matrix or the normalized keyword weight matrix are adjusted using the modulus weight vector.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118468976B_ABST
    Figure CN118468976B_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a model training method, a text generation method, an apparatus, a device and a medium. The model training method comprises: performing first module length normalization on a query weight matrix; performing module length adjustment on the normalized query weight matrix by using a first module length weight vector; performing second module length normalization on a keyword weight matrix; performing module length adjustment on the normalized keyword weight matrix by using a second module length weight vector; obtaining a first attention input matrix corresponding to an attention processing module for text data; and training a large language model according to the adjusted query weight matrix and the adjusted keyword weight matrix. The embodiments of the present application can improve the training sufficiency and overall training effect of the attention module of the large language model, can reduce the data transmission amount of the large language model in the training process, and thus can improve the hardware processing speed and response speed corresponding to the large language model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of artificial intelligence technology, and in particular to a model training method, a text generation method, an apparatus, a device, and a medium. Background Technology

[0002] As is well known, large language models are one of the feasible paths to general artificial intelligence. A large language model is a natural language processing model based on deep learning, with wide applications in summarization, machine translation, literary creation, and intelligent dialogue. A large language model typically includes an embedding layer and a processing layer. After tokens are input into the large language model, the embedding layer performs embedding processing on the tokens to obtain their corresponding embedding representations. The processing layer specifically includes a transformer layer, etc., used to determine the text information corresponding to the embedding representation.

[0003] The attention processing module in the transformer layer utilizes the attention mechanism. The attention mechanism is an important technique for improving model performance. It helps large language models better understand and analyze the relationships between different positions when processing long sequence data, greatly improving the processing speed and accuracy of large language models.

[0004] Since the attention module plays a crucial role in large language models, it is particularly important to train it effectively. In existing techniques, insufficient training of the attention module can negatively impact the overall training performance of the large language model. Furthermore, these techniques often require a significant number of iterations and training time to bring the large language model to convergence.

[0005] In practical applications, a large number of iterations and training time not only increase the amount of data transmitted during the training process of large language models, but also, because the operation of large language models relies on the support of computer hardware computing power, their low efficiency can easily lead to insufficient computing power, resulting in technical problems such as poor response speed and stuttering. Summary of the Invention

[0006] This application provides a model training method that can improve the training sufficiency and overall training effect of the attention module of a large language model, reduce the amount of data transmission during the training process of the large language model, and thus improve the hardware processing speed and response speed of the large language model.

[0007] Accordingly, embodiments of this application also provide a text generation method, a model training device, a text generation apparatus, an electronic device, and a machine-readable medium to ensure the implementation and application of the above methods.

[0008] To address the aforementioned problems, this application discloses a model training method for training a large language model. The parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The method includes:

[0009] The query weight matrix is ​​subjected to a first modulus normalization to obtain a normalized query weight matrix; the first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value;

[0010] The normalized query weight matrix is ​​adjusted using the first modulus weight vector to obtain the adjusted query weight matrix; an element of the first modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix.

[0011] The keyword weight matrix is ​​subjected to a second modulus normalization to obtain a normalized keyword weight matrix; the second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value;

[0012] The normalized keyword weight matrix is ​​adjusted using the second modulus weight vector to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector in the normalized keyword weight matrix.

[0013] For text data, obtain the first attention input matrix corresponding to the attention processing module;

[0014] Based on the adjusted query weight matrix and the adjusted keyword weight matrix, attention processing is performed on the first attention input matrix to obtain the first attention processing result;

[0015] Based on the first attention processing result, the text information of the large language model is determined;

[0016] Based on the text information, determine the loss information;

[0017] The parameters of the large language model are updated based on the loss information.

[0018] This application also discloses a text generation method applied to a large language model. The parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The method includes:

[0019] After training the large language model, the query weight matrix is ​​normalized by the first modulus to obtain the normalized query weight matrix. The first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value.

[0020] The normalized query weight matrix is ​​adjusted using the first modulus weight vector to obtain the adjusted query weight matrix; an element of the first modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix.

[0021] After training the large language model, the keyword weight matrix is ​​subjected to a second modulus normalization to obtain a normalized keyword weight matrix. The second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value.

[0022] The normalized keyword weight matrix is ​​adjusted using the second modulus weight vector to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector in the normalized keyword weight matrix.

[0023] For the text to be processed, obtain the second attention input matrix corresponding to the attention processing module;

[0024] Based on the adjusted query weight matrix and the adjusted keyword weight matrix, attention processing is performed on the second attention input matrix to obtain the second attention processing result;

[0025] Based on the result of the second attention processing, the target generated text output by the large language model is determined.

[0026] This application also discloses a model training apparatus for training a large language model. The parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The apparatus includes:

[0027] The first module normalization module is used to perform first module normalization on the query weight matrix to obtain a normalized query weight matrix; the first module normalization is used to normalize the module lengths of different column vectors in the query weight matrix to the same preset value.

[0028] The first module length adjustment module is used to adjust the normalized query weight matrix using the first module length weight vector to obtain the adjusted query weight matrix; an element of the first module length weight vector is used to adjust the module length of a column vector in the normalized query weight matrix.

[0029] The second module normalization module is used to perform second module normalization on the keyword weight matrix to obtain a normalized keyword weight matrix; the second module normalization is used to normalize the module lengths of different column vectors in the keyword weight matrix to the same preset value.

[0030] The second module length adjustment module is used to adjust the normalized keyword weight matrix using the second module length weight vector to obtain the adjusted keyword weight matrix; an element of the second module length weight vector is used to adjust the module length of a column vector in the normalized keyword weight matrix.

[0031] The attention input module is used to obtain the first attention input matrix corresponding to the attention processing module for text data;

[0032] An attention processing module is used to perform attention processing on the first attention input matrix according to the adjusted query weight matrix and the adjusted keyword weight matrix to obtain a first attention processing result;

[0033] The output determination module is used to determine the text information of the large language model based on the first attention processing result;

[0034] The loss determination module is used to determine loss information based on the text information;

[0035] The parameter update module is used to update the parameters of the large language model based on the loss information.

[0036] Optionally, the first mold length adjustment module includes:

[0037] The first extension module is used to expand the first modulus weight vector into a first modulus weight matrix using an automatic broadcasting mechanism; the size of the first modulus weight matrix is ​​the same as the size of the normalized query weight matrix.

[0038] The first dot product module is used to perform a dot product operation on the normalized query weight matrix and the first modulus weight matrix to obtain the adjusted query weight matrix.

[0039] Optionally, the second mold length adjustment module includes:

[0040] The second extension module is used to extend the second modulus weight vector into a second modulus weight matrix using an automatic broadcasting mechanism; the size of the second modulus weight matrix is ​​the same as the size of the normalized keyword weight matrix.

[0041] The second dot product module is used to perform a dot product operation on the normalized keyword weight matrix and the second modulus weight matrix to obtain the adjusted keyword weight matrix.

[0042] Optionally, the training includes: pre-training, wherein the text data includes: N word units; the text information represents: the prediction information of the i-th word unit; N and i are both positive integers; i is not greater than N;

[0043] The loss determination module includes:

[0044] The first loss determination module is used to determine the first cross-entropy loss information based on the conditional probability corresponding to the predicted information of the i-th word and the actual information of the i-th word in the text data; the conditional probability is the probability of the predicted information of the i-th word appearing under the condition of the first i-1 words in the text data.

[0045] Optionally, the training includes: fine-tuning training, wherein the text data includes: the input sequence and the label values ​​of the output sequence; and the text information includes: the predicted values ​​of the output sequence.

[0046] The loss determination module includes:

[0047] The second loss determination module is used to determine the second cross-entropy loss information or mean squared error loss information based on the predicted value and the label value.

[0048] This application also discloses a text generation apparatus applied to a large language model. The parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The apparatus includes:

[0049] The first module normalization module is used to perform first module normalization on the query weight matrix after the training of the large language model is completed, so as to obtain a normalized query weight matrix; the first module normalization is used to normalize the module lengths of different column vectors in the query weight matrix to the same preset value.

[0050] The first module length adjustment module is used to adjust the normalized query weight matrix using the first module length weight vector to obtain the adjusted query weight matrix; an element of the first module length weight vector is used to adjust the module length of a column vector in the normalized query weight matrix.

[0051] The second module normalization module is used to perform second module normalization on the keyword weight matrix after the training of the large language model is completed, so as to obtain a normalized keyword weight matrix. The second module normalization is used to normalize the module length of different column vectors in the keyword weight matrix to the same preset value.

[0052] The second module length adjustment module is used to adjust the normalized keyword weight matrix using the second module length weight vector to obtain the adjusted keyword weight matrix; an element of the second module length weight vector is used to adjust the module length of a column vector in the normalized keyword weight matrix.

[0053] The attention input module is used to obtain the second attention input matrix corresponding to the attention processing module for the text to be processed;

[0054] The attention processing module is used to perform attention processing on the second attention input matrix according to the adjusted query weight matrix and the adjusted keyword weight matrix to obtain the second attention processing result;

[0055] The text generation module is used to determine the target generated text output by the large language model based on the result of the second attention processing.

[0056] This application also discloses an electronic device, including: a processor; and a memory storing executable code thereon, which, when executed, causes the processor to perform the method described in this application.

[0057] This application also discloses a machine-readable medium storing executable code thereon, which, when executed, causes a processor to perform the method described in this application.

[0058] The embodiments of this application have the following advantages:

[0059] In the technical solution of this application embodiment, during the model training process, the column dimensions of the query weight matrix and the keyword weight matrix are first normalized. Then, the column vectors in the normalized query weight matrix or the normalized keyword weight matrix are adjusted using the modulus weight vector.

[0060] Since one element in the modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix or the normalized keyword weight matrix, during model training, a single weight parameter is used to adjust the modulus information of a column vector, which increases the difficulty of modulus adjustment. Given this increased difficulty, the embodiments of this application are more conducive to adjusting the angles between different column vectors in the query weight matrix or the keyword weight matrix, allowing for more thorough training of these angles. Therefore, the embodiments of this application can improve the training sufficiency and effectiveness of the attention module in the large language model, and further, improve the overall training effect of the large language model.

[0061] This application embodiment, by improving the training sufficiency of the attention module and the overall training effect of the large language model, can achieve convergence of the large language model while reducing the number of iterations and training time. Because it reduces the number of iterations and training time, this application embodiment can reduce the amount of data transmitted during the training process of the large language model, thereby improving the hardware processing speed and response speed of the corresponding large language model. In other words, this application embodiment can overcome technical problems such as insufficient computing power, poor model response speed, and stuttering to a certain extent, thereby improving the response speed of the large language model. Attached Figure Description

[0062] Figure 1 This is a schematic diagram of the structure of a large language model according to an embodiment of this application;

[0063] Figure 2 This is a schematic diagram of the converter layer in a large language model according to an embodiment of this application;

[0064] Figure 3 This is a schematic diagram of the structure of an attention processing module according to an embodiment of this application;

[0065] Figure 4 This is a schematic flowchart of the weight calculation of the attention processing module according to one embodiment of this application;

[0066] Figure 5 This is a schematic diagram of the reparameterization process of the query weight matrix according to one embodiment of this application;

[0067] Figure 6 This is a flowchart illustrating the reparameterization process of the keyword weight matrix according to an embodiment of this application;

[0068] Figure 7 This is a schematic diagram illustrating the principle of attention processing according to an embodiment of this application;

[0069] Figure 8 This is a schematic flowchart of the model training method according to an embodiment of this application;

[0070] Figure 9 This is a flowchart illustrating the steps of a text generation method according to an embodiment of this application;

[0071] Figure 10 This is a schematic diagram of the structure of a model training device according to an embodiment of this application;

[0072] Figure 11 This is a schematic diagram of the structure of a text generation apparatus according to an embodiment of this application;

[0073] Figure 12 This is a schematic diagram of the structure of an apparatus provided in one embodiment of this application. Detailed Implementation

[0074] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, this application will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0075] The large language model described in this application is an artificial intelligence model designed to understand and generate human language. Trained on large amounts of text data, the large language model can perform a wide range of tasks, including the text generation task described in this application. Large language models are typically based on deep learning architectures, such as transformer structures. This application does not limit the specific type of large language model; for example, the type of large language model may include GPT (Generative Pre-Trained Transformer) models and GLM (Generative Language Model), etc.

[0076] For example, the GPT model is a large-scale natural language generation model based on the Transformer architecture. It is pre-trained using massive amounts of text data and can generate high-quality natural language text, including articles, dialogues, and summaries. The core of the GPT model is its transformer structure, which employs an attention mechanism, enabling the GPT model to consider all positions in the input sequence simultaneously, thus better capturing long-distance dependencies. After pre-training, the GPT model can be fine-tuned to adapt to various downstream text generation tasks. During the pre-training phase, the GPT model uses an autoregressive language model training method, generating only one output at each time step, with subsequent time steps generating the next output based on the previous input and output.

[0077] The principle behind training autoregressive language models is as follows: given the preceding i words, predict the probability distribution of the (i+1)th word, where i can be a positive integer. In this way, large language models can generate semantically coherent text. During training, autoregressive language models use maximum likelihood estimation to learn the dependencies between words. Specifically, large language models update their parameters by maximizing the likelihood probability of the real text, thus making the generated new text more similar to the real text.

[0078] Text generation refers to the use of computer programs or algorithms to capture the syntax, semantics, structure, and style of text based on given input or contextual information, and then using this knowledge to generate new text content. This new text content possesses a reasonable grammatical structure, semantic accuracy, and contextual coherence, meeting the requirements of a specific task. The generated text can be a piece of natural language text, an article, a sentence, a paragraph, or other forms of text. Text generation can be used in various application scenarios, such as summarization, machine translation, literary creation, and intelligent dialogue.

[0079] The large language model in this application embodiment may include an embedding layer and a processing layer. The embedding layer is used to embed the input matrix corresponding to the lexical units of the input sequence to obtain an embedded representation. During the training phase, the input sequence may be text data. During the inference phase (text generation phase), the input sequence may be the text to be processed.

[0080] Processing layers can include transformer layers. Large language models typically include multiple transformer layers. A transformer layer can employ a transformer architecture, specifically including an attention mechanism and a feedforward neural network. Multiple transformer layers are used to perform multi-level representation learning on the embedding representations, thereby better capturing complex semantics and text structure.

[0081] Other processing layers can be connected to the transformer layer. Examples of other processing layers include: normalization layers, linear transformation layers, and normalization layers.

[0082] Reference Figure 1 The diagram shows a structural schematic of a large language model according to an embodiment of this application, specifically including: an embedding layer 101, N converter layers 102 (from converter layer 1 to converter layer N), a normalization layer 103, a linear conversion layer 104, and a normalization layer 105.

[0083] The embedding layer 101 is used to embed the input matrix corresponding to the lexical units of the input sequence to obtain the embedding representation. The embedding representation is a semantic vector that can represent the semantic information of lexical units and can serve as the basis for large language models to understand the context, nuances, and subtle meanings of words and phrases.

[0084] N transformer layers 102 are used for multi-level representation learning of the embedded representations, thereby better capturing complex semantics and text structures. The output of transformer layer N can be referred to as the first output.

[0085] The normalization layer 103 is used to normalize the first output to obtain the second output. Normalization can, to some extent, overcome the gradient explosion problem caused by excessively large or small parameter values ​​in a large language model. The parameter values ​​of a large language model can refer to the numerical values ​​of the parameters corresponding to the network layers of the large language model. Embedding layer 101, N transformer layers 102 (from transformer layer 1 to transformer layer N), normalization layer 103, linear transformation layer 104, and normalization layer 105 are all network layers of the large language model.

[0086] The linear transformation layer 104 is used to transform the second output into a probability distribution for predicting the next word or new text. Specifically, the linear transformation layer 104 performs a linear transformation on the second output, mapping it to a vector space of the same size as the vocabulary. The purpose of this is to predict the probability distribution of the next word. The vocabulary is used to store words that a large language model can recognize.

[0087] The normalization layer 105 normalizes the probability distribution output by the linear transformation layer 104 to obtain the target probability value between [0,1]. The target probability value can be the probability that each word in the vocabulary becomes the target word.

[0088] A token can be the smallest indivisible semantic unit. For example, "waterfall" can be broken down into two tokens: water and fall. Additionally, punctuation marks are also broken down into tokens because they affect the semantic understanding of the entire text. For instance, "I don't know." can be broken down into five tokens: "I", "don", "'t", "know", and ".".

[0089] Current large language models generally employ residual connections and pre-normalization. (See reference...) Figure 2 The diagram illustrates the structure of a converter layer in a large language model according to an embodiment of this application. Specifically, the converter layer includes an attention processing module 201 and a multilayer perceptron module 202.

[0090] The attention processing module 201 performs attention processing on the input. In the case of converter layer 1, the input can be an embedded representation. Attention processing can include self-attention processing and cross-attention processing. Taking self-attention processing as an example, the attention processing module performs attention calculations between each position in the input sequence and other positions. Taking cross-attention processing as an example, the attention processing module performs attention calculations between each position in the input sequence and a position in the output sequence.

[0091] The multilayer perceptron module 202 is used to perform nonlinear transformations and mappings on the input to capture higher-level features.

[0092] Figure 2 In This indicates a residual connection. When the converter layer is converter layer 1, the residual connection directly adds the embedded representation of the embedding layer output to the output of converter layer 1. Pre-normalization can refer to normalization performed before the residual connection. Specifically, the attention processing module 201 normalizes its own input before performing attention processing. Figure 2 (Not shown in the image), the multilayer perceptron module 202 also normalizes its own input before performing the nonlinear transformation and mapping of the embedded representation. Figure 2 (Not shown in the image). The normalization in the attention processing module 201 and the normalization in the multilayer perceptron module 202 are examples of pre-normalization.

[0093] Reference Figure 3 The diagram illustrates the structure of an attention processing module according to an embodiment of this application. This attention processing module is used for self-attention processing and specifically includes: an input normalization layer, a query weight matrix, a keyword weight matrix and a numerical weight matrix, RoPE (Rotary Position Embedding), an AttentionMask (used to prevent the attention mechanism from focusing on padding characters), a softmax layer, a multiplier F, and an O matrix. The multiplier F is used to perform matrix multiplication operations. Figure 3 In A represents addition.

[0094] The query weight matrix is ​​used to linearly transform the attention input matrix to obtain the query matrix Q. The keyword weight matrix is ​​used to linearly transform the input matrix to obtain the keyword matrix K. The numerical weight matrix is ​​used to linearly transform the input matrix to obtain the numerical matrix V. Furthermore, attention information can be calculated using formula (1). The attention input matrix can refer to the matrix input to the attention processing module.

[0095]

[0096] Where Attention(Q, K, V) represents attention information. Used to represent the scaling factor.

[0097] Reference Figure 4 The diagram illustrates a flowchart of the weight calculation process of the attention processing module according to an embodiment of this application, wherein the attention input matrix is ​​X and the query weight matrix is ​​M. Q The keyword weight matrix is ​​M K .

[0098] Assuming the attention input matrix has a size of s×c, and the query weight matrix M... Q If the size is c×c, then the attention input matrix X can be combined with the query weight matrix M. Q Perform matrix multiplication to obtain the query matrix Q, which can have a size of s×c. Similarly, the keyword weight matrix M... K The size is c×c, and the attention input matrix X can be combined with the keyword weight matrix M. K Perform matrix multiplication to obtain the keyword matrix K, which can have a size of s×c.

[0099] Assume the input matrix of the large language model has dimensions [b, s], where the text data is divided into multiple batches, b represents the number of text data items in a batch, and s represents the sequence length of the input matrix, i.e., the number of tokens in the input matrix. c represents the dimension of the embedding vector, that is, how many dimensions each token is represented as a vector. This value can be set according to the specific requirements of the task. Examples of c include 50, 100, 200, 300, etc. Typically, c < 0. <n。

[0100] Furthermore, the transpose of the keyword matrix K can be taken to obtain the transpose matrix K. T Next, for the query matrices Q and K T Matrix multiplication is performed on matrices. See formula (1) for details.

[0101] Figure 4 The weight calculation process of the attention processing module shown is used for self-attention processing. The attention input matrix for cross-attention processing specifically includes: attention input matrix X (s×c) and attention input matrix Y (l×c). Then, the attention input matrix X (s×c) and the query weight matrix M can be compared... Q Perform matrix multiplication (c×c) to obtain the query matrix Q(s×c). You can also perform matrix multiplication on Y(l×c) and the keyword matrix M. K Perform matrix multiplication on (c×c) to obtain the keyword matrix K(1×c). Then, perform matrix multiplication on the query matrix Q(s×c) and the transpose of the keyword matrix K(1×c). T (c×1) Perform matrix multiplication. Since the principles of cross-attention and self-attention are similar, they will not be elaborated here; please refer to each other for details.

[0102] Since the attention module plays a crucial role in large language models, it is particularly important to train the attention module more effectively.

[0103] This application provides a model training method for training a large language model. The parameters of the attention processing module in the large language model specifically include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The method specifically includes:

[0104] The query weight matrix is ​​subjected to a first modulus normalization to obtain a normalized query weight matrix; the first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value.

[0105] Using the first modulus weight vector, the modulus of the normalized query weight matrix is ​​adjusted to obtain the adjusted query weight matrix; one element of the first modulus weight vector is used to adjust the modulus of one column vector of the normalized query weight matrix.

[0106] The keyword weight matrix is ​​subjected to a second modulus normalization to obtain a normalized keyword weight matrix; the second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value.

[0107] The normalized keyword weight matrix is ​​adjusted using the second modulus weight vector to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector in the normalized keyword weight matrix.

[0108] For text data, obtain the first attention input matrix corresponding to the attention processing module;

[0109] Based on the above-mentioned adjusted query weight matrix and adjusted keyword weight matrix, attention processing is performed on the above-mentioned first attention input matrix to obtain the first attention processing result;

[0110] Based on the results of the first attention processing, the text information of the large language model is determined.

[0111] Based on the above text information, the loss information is determined;

[0112] Based on the aforementioned loss information, the parameters of the large language model are updated.

[0113] This application's embodiments utilize reparameterization technology to... Figure 3 The structures of the query weight matrix and keyword weight matrix in the attention processing module shown have been improved. The principle of reparameterization is as follows: first, a first series of structures (generally used for training) is constructed, and the parameters of the first series of structures are equivalently transformed into another set of parameters (generally used for inference), thereby transforming the first series of structures into the second series of structures.

[0114] Specifically, in the embodiments of this application, Figure 3 The structure of the attention processing module shown is an example of the second series of structures, which can be used in the inference stage of a large language model, such as in the text generation method applied to embodiments of this application.

[0115] Reference Figure 5 This diagram illustrates a flowchart of the reparameterization process for a query weight matrix according to an embodiment of this application, wherein the query weight matrix M can be reparameterized first. Q Perform first-order modulus normalization on the column dimensions; then, use the first-order weight vector I. Q For the normalized query weight matrix N Q Perform modulus adjustment to obtain the adjusted query weight matrix M'. Q .

[0116] Reference Figure 6 This diagram illustrates a flowchart of the reparameterization process for a keyword weight matrix according to an embodiment of this application. The keyword weight matrix M can be reparameterized first. K Perform second-order normalization of the column dimensions; then, use the second-order weight vector I. K For the normalized keyword weight matrix N k Adjust the modulus length to obtain the adjusted keyword weight matrix M'. K .

[0117] Figure 5 The reparameterized structure of the query weight matrix shown can be used as a first series of structures during the model training phase. After training the large language model, the query weight matrix undergoes a first modulus normalization to obtain a normalized query weight matrix. This first modulus normalization normalizes the modulus of different column vectors in the query weight matrix to the same preset value. Using the first modulus weight vector, the normalized query weight matrix is ​​adjusted to obtain an adjusted query weight matrix. An element of the first modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix. During the model inference phase, the adjusted query weight matrix described above can be used to replace... Figure 3 The query weight matrix in the database.

[0118] Similarly, Figure 6The reparameterized structure of the keyword weight matrix shown can be used as the first series of structures during the model training phase. After training the large language model, a second modulus normalization is performed on the keyword weight matrix to obtain a normalized keyword weight matrix. This second modulus normalization normalizes the modulus of different column vectors in the keyword weight matrix to the same preset value. Using the second modulus weight vector, the modulus of the normalized keyword weight matrix is ​​adjusted to obtain an adjusted keyword weight matrix. One element of the second modulus weight vector is used to adjust the modulus of one column vector in the normalized keyword weight matrix. During the model inference phase, the adjusted keyword weight matrix described above can be used to replace... Figure 3 Keyword weight matrix in the text.

[0119] Reference Figure 7 This diagram illustrates the principle of attention processing according to an embodiment of this application, wherein the transpose matrix K of the query matrix Q(s×c) and the keyword matrix K(s×c) is... T (c×s) performs matrix multiplication to obtain a result matrix with s rows and s columns.

[0120] For convenience, any row vector in the query matrix Q(s×c) is denoted as Transpose matrix K T Any column vector in the (c×s) matrix is ​​denoted as Attention processing involves processing c-dimensional row vectors. With s c-dimensional column vectors Perform an inner product operation to obtain s values, add these s values ​​to the attention mask, and then perform a normalized exponent operation to obtain the attention information.

[0121] Assuming the attention information is the attention score, then the magnitude of the attention score at each of the s positions directly depends on the row vector. With column vectors The size of the inner product operation result. And row vectors. and column vectors The values ​​depend on the query weight matrix M. Q and keyword weight matrix M K .according to Figure 4 It can be seen that the query weight matrix M Q and keyword weight matrix M K All vectors involved in the operation are column vectors. The inner product of two vectors can be understood not only as multiplying the values ​​of their corresponding dimensions and then adding all the results, but also as multiplying the product of the magnitudes of the two vectors and then multiplying by the cosine of the angle between the vectors.

[0122] In this embodiment of the application, the magnitudes of the query weight matrix and the keyword weight matrix are controlled during the model training process.

[0123] Compared to traditional techniques that control the modulus of all elements in the query weight matrix and keyword weight matrix, this embodiment first normalizes the column dimensions of the query weight matrix and keyword weight matrix during model training. Then, it uses modulus weight vectors to adjust the modulus of the column vectors in the normalized query weight matrix or the normalized keyword weight matrix.

[0124] Since one element in the modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix or the normalized keyword weight matrix, during model training, a single weight parameter is used to adjust the modulus information of a column vector, which increases the difficulty of modulus adjustment. Given this increased difficulty, the embodiments of this application are more conducive to adjusting the angles between different column vectors in the query weight matrix or the keyword weight matrix, allowing for more thorough training of these angles. Therefore, the embodiments of this application can improve the training sufficiency and effectiveness of the attention module in the large language model, and further, improve the overall training effect of the large language model.

[0125] This application embodiment, by improving the training sufficiency of the attention module and the overall training effect of the large language model, can achieve convergence of the large language model while reducing the number of iterations and training time. Because it reduces the number of iterations and training time, this application embodiment can reduce the amount of data transmitted during the training process of the large language model, thereby improving the hardware processing speed and response speed of the corresponding large language model. In other words, this application embodiment can overcome technical problems such as insufficient computing power, poor model response speed, and stuttering to a certain extent, thereby improving the response speed of the large language model.

[0126] Method Example 1

[0127] refer to Figure 8 This diagram illustrates a step-by-step flowchart of a model training method according to an embodiment of this application. The method is used to train a large language model. The parameters of the attention processing module in the large language model specifically include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The method specifically includes:

[0128] Step 801: Perform a first modulus normalization on the query weight matrix to obtain a normalized query weight matrix; the first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value.

[0129] Step 802: Using the first modulus weight vector, adjust the modulus of the normalized query weight matrix to obtain the adjusted query weight matrix; one element of the first modulus weight vector is used to adjust the modulus of one column vector of the normalized query weight matrix.

[0130] Step 803: Perform a second modulus normalization on the keyword weight matrix to obtain a normalized keyword weight matrix; the second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value.

[0131] Step 804: Using the second modulus weight vector, adjust the modulus of the normalized keyword weight matrix to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector of the normalized keyword weight matrix.

[0132] Step 805: For the text data, obtain the first attention input matrix corresponding to the attention processing module;

[0133] Step 806: Based on the above-mentioned adjusted query weight matrix and adjusted keyword weight matrix, perform attention processing on the above-mentioned first attention input matrix to obtain the first attention processing result;

[0134] Step 807: Based on the results of the first attention processing described above, determine the text information of the large language model described above;

[0135] Step 808: Determine the loss information based on the above text information;

[0136] Step 809: Update the parameters of the large language model based on the above loss information.

[0137] Figure 5 The first embodiment of the method illustrates updating the parameters of the large language model during its training. These parameters can include those corresponding to network layers such as the embedding layer and the transformer layer. Specifically, the embedding layer parameters include the embedding weight matrix. The transformer layer's attention processing module parameters specifically include the query weight matrix, the first modulus-length weight vector, the keyword weight matrix, and the second modulus-length weight vector.

[0138] The training process of a large language model can include forward propagation and backward propagation.

[0139] The forward propagation process calculates the final text information sequentially from the embedding layer to the processing layer, based on the parameters of the large language model. This text information is then used to determine the loss information.

[0140] Backpropagation, based on loss information, sequentially calculates and updates the parameters of a large language model, proceeding from the embedding layer to the processing layer. Large language models typically employ a neural network structure, and their parameters can include neural network weights and other parameters. During backpropagation, the gradient information of the large language model's parameters can be determined and used to update these parameters. For example, backpropagation can follow the chain rule in calculus, sequentially calculating and storing the gradient information of the large language model's parameters from the processing layer to the embedding layer.

[0141] This application embodiment can utilize a normal distribution initialization method and an equal-weight initialization method to initialize the weight values ​​of parameters such as the query weight matrix, thereby obtaining the initial weight values ​​of the query weight matrix and other parameters. Subsequently, a backpropagation process can be used to update the initial weight values ​​of the query weight matrix and other parameters to obtain the current weight values ​​of the query weight matrix and other parameters.

[0142] In step 801, based on the first current weight value of the query weight matrix, a first modulus normalization can be performed on the query weight matrix to obtain a normalized query weight matrix. This first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value; an example of the preset value can be 1. Of course, those skilled in the art can use other values ​​besides 1 as preset values ​​according to actual application requirements.

[0143] In practical applications, normalization techniques can be used to normalize the magnitude of the column vectors in the query weight matrix. Specific normalization techniques include L1 norm normalization and L2 norm normalization.

[0144] Taking L2 norm normalization as an example, in Euclidean space, the L2 norm of a vector can point to the square root of the sum of the squares of its elements. L2 norm normalization divides each element of a vector by its L2 norm, which can scale the vector's magnitude to a unit length of 1.

[0145] The first modulus normalization changes the query weight matrix M. Q The magnitude of each column vector in (c×c) is not changed, but the size of the query weight matrix is ​​not altered; therefore, the normalized query weight matrix N is... Q The dimensions remain c×c.

[0146] In step 802, the normalized query weight matrix can be adjusted using the first modulus weight vector to obtain the adjusted query weight matrix.

[0147] The process of adjusting the normalized query weight matrix using the first modulus weight vector specifically includes: expanding the first modulus weight vector into a first modulus weight matrix using an automatic broadcasting mechanism; the size of the first modulus weight matrix is ​​the same as the size of the normalized query weight matrix; and performing a dot product operation on the normalized query weight matrix and the first modulus weight matrix to obtain the adjusted query weight matrix.

[0148] Automatic broadcasting refers to the process of automatically expanding data to meet computational requirements during array operations.

[0149] The array operation corresponding to the modulus adjustment in this embodiment is specifically: the array operation corresponding to the normalized query weight matrix and the first modulus weight vector.

[0150] Assume the normalized query weight matrix is ​​a c x c matrix, and the first modulus weight vector is a 1 x c vector. To perform operations on the c x c normalized query weight matrix and the 1 x c first modulus weight vector, the training framework will automatically broadcast and copy the elements of each row of the first modulus weight vector c times to obtain a c x c first modulus weight matrix.

[0151] The aforementioned first modulus weight vector I Q One element is used to adjust the modulus of a column vector in the normalized query weight matrix described above. Assume the first modulus weight vector I... Q The elements in the matrix are [Q1, Q2, ..., Qj, ..., Qc], and Qj is used to adjust the modulus of the j-th column vector in the normalized query weight matrix.

[0152] The dot product operation is used for matrix multiplication. Assume matrices A and B are matrices of the same dimension, meaning the number of rows in A equals the number of rows in B, and the number of columns in A equals the number of columns in B. During the operation, corresponding elements of matrices A and B are multiplied. If both A and B are c x c matrices, their dot product will also result in a c x c matrix. This is achieved using the first modulus weight vector I. Q For the normalized query weight matrix N Q The modulus is adjusted to obtain the adjusted query weight matrix M'. Q It can be a c-row, c-column matrix.

[0153] In step 803, based on the second current weight value of the keyword weight matrix, a second modulus normalization can be performed on the keyword weight matrix to obtain a normalized keyword weight matrix. The aforementioned second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value; an example of the preset value can be 1.

[0154] In practical applications, normalization techniques can be used to normalize the magnitude of column vectors in the keyword weight matrix.

[0155] The second module length normalization changed the keyword weight matrix M. K The magnitude of each column vector in (c×c) is adjusted, but the size of the keyword weight matrix remains unchanged. Therefore, the normalized keyword weight matrix N... k The dimensions remain c×c.

[0156] In step 804, the normalized keyword weight matrix can be adjusted using the second modulus weight vector to obtain the adjusted keyword weight matrix.

[0157] The process of adjusting the normalized keyword weight matrix using the second modulus weight vector specifically includes: expanding the second modulus weight vector into a second modulus weight matrix using an automatic broadcasting mechanism; the size of the second modulus weight matrix is ​​the same as the size of the normalized keyword weight matrix; and performing a dot product operation on the normalized keyword weight matrix and the second modulus weight matrix to obtain the adjusted keyword weight matrix.

[0158] The dot product operation is used for matrix multiplication. It utilizes the second-order weight vector I. K For the normalized keyword weight matrix N k The modulus is adjusted to obtain the adjusted query weight matrix M'. Q It can be a c-row, c-column matrix.

[0159] The aforementioned second modulus weight vector I K One element is used to adjust the magnitude of a column vector in the normalized keyword weight matrix mentioned above. Assume the second magnitude weight vector I... K The elements in the matrix are [K1, K2, ..., Kj, ..., Kc], and Kj is used to adjust the magnitude of the j-th column vector in the normalized keyword weight matrix.

[0160] In step 805, the training of the large language model may include pre-training or fine-tuning. During the pre-training phase, a single piece of text data may include N lexical units. These N lexical units can correspond to one or more sentences. The pre-training text data may originate from one or more of the following pre-defined corpora: books, online articles, text data from social media platforms, scientific literature, dialogue data, dialogue texts from films and television dramas, and public datasets, etc.

[0161] During the fine-tuning training phase, a piece of text data includes: an input sequence and label values ​​for the output sequence. Taking an intelligent dialogue task as an example, the input sequence can be a question in the dialogue, and the label values ​​for the output sequence can be the answer to the dialogue.

[0162] In a practical implementation, a piece of text data can be segmented into tokens, and the input matrix corresponding to the tokens of the text data can be obtained.

[0163] Assuming the size of the input matrix is ​​[b, s], the embedding layer can perform mapping processing on the input matrix, and the size of the resulting embedding representation can be [b, s, c]. The embedding layer can then pass the embedding representation to the processing layer.

[0164] The embedding layer corresponds to an embedding weight matrix. The process of using the embedding layer to perform the first mapping process on the input matrix specifically includes: matching the token identifiers in the input matrix with the row indices in the embedding weight matrix to obtain the target row corresponding to the token identifiers in the input matrix; and determining the embedding representation corresponding to the input matrix based on the target row.

[0165] The embedding weight matrix has dimensions of n rows and c columns. Each row in the n rows corresponds to an index, which can be a word identifier. Therefore, by matching the word identifiers in the input matrix with the row indices in the embedding weight matrix, the target rows corresponding to the word identifiers in the input matrix can be obtained. s word identifiers in the input matrix can correspond to s target rows. Therefore, the size of the embedding representation obtained by the embedding layer can be [b, s, c].

[0166] The processing layer may include: Figure 1 The diagram shows N converter layers 102, 103, 104, and 105, including converter layer 1 to converter layer N. The processing layer can process the embedded representation to obtain text information.

[0167] like Figure 2 and Figure 3 As shown, the input normalization layer included in the attention processing module of converter layer 1 can first normalize the embedded representation to obtain a normalized embedded representation. Therefore, for converter layer 1, the first attention input matrix corresponding to the attention processing module can be: the normalized embedded representation. For other converter layers, those skilled in the art can determine the first attention input matrix corresponding to the attention processing module based on existing common sense. This application embodiment does not limit the specific determination process of the first attention input matrix corresponding to the attention processing module.

[0168] In step 806, the first query matrix Q(s×c) and the first keyword matrix K(s×c) corresponding to the first attention input matrix can be determined first based on the above-mentioned adjusted query weight matrix and adjusted keyword weight matrix. Then, the first attention processing result can be determined using formula (1) based on the first query matrix Q(s×c) and the first keyword matrix K(s×c).

[0169] Assuming the first attention input matrix is ​​X1(s×c), then the first attention input matrix X1(s×c) and the adjusted query weight matrix M' can be used to... Q Perform matrix multiplication to obtain the first query matrix Q(s×c). The first attention input matrix X1(s×c) can also be multiplied by adjusting the keyword weight matrix M'. K Perform matrix multiplication to obtain the first keyword matrix K(s×c). Then, the transpose of the first query matrix Q(s×c) and the first keyword matrix K(s×c) can be calculated. T (c×s) performs matrix multiplication to obtain the first attention processing result.

[0170] In step 807, the text information of the large language model can be determined based on the results of the first attention processing.

[0171] During the pre-training phase, the text information can be the predicted information of the i-th word, where i can be a positive integer not greater than N. During the fine-tuning training phase, the text information can be the predicted values ​​of the output sequence. Taking intelligent dialogue tasks as an example, the text information can be the predicted values ​​of the answers.

[0172] In step 808, loss information can be determined based on the training phase.

[0173] In the pre-training phase, the text data includes: N word units; the text information represents: the prediction information of the i-th word unit; N and i are both positive integers; i is not greater than N;

[0174] The process of determining the loss information based on the above text information specifically includes: determining the first cross-entropy loss information based on the conditional probability corresponding to the predicted information of the i-th word and the actual information of the i-th word in the text data; the conditional probability is specifically the probability of the predicted information of the i-th word appearing under the condition of the first i-1 words in the text data.

[0175] Formula (2) shows an example of the calculation process for the first cross-entropy loss information.

[0176]

[0177] Where L1 represents the first cross-entropy loss information, wi Let P(w) represent the i-th word in the text data. i |w1, w2..., w i-1 ) represents the probability of the predicted information of the i-th word given the first i-1 words in the text data. The conditional probability is 1 if the predicted information of the i-th word is the same as the actual information of the i-th word in the text data; or, the conditional probability is 0 if the predicted information of the i-th word is different from the actual information of the i-th word in the text data.

[0178] During the fine-tuning training phase, the text data includes: the input sequence and the label values ​​of the output sequence; the text information includes: the predicted value of the output sequence; then the process of determining the loss information based on the above text information specifically includes: determining the second cross-entropy loss information or the mean squared error loss information based on the predicted value and the label value.

[0179] Specifically, the second cross-entropy loss information can be determined based on the second conditional probability corresponding to the predicted value of the output sequence and the label value of the output sequence; the second conditional probability is specifically the probability of the predicted value occurring under the condition of the input sequence.

[0180] Formula (3) shows an example of the calculation process for the second cross-entropy loss information.

[0181]

[0182] Where L2 represents the second cross-entropy loss information, x i Let y represent the i-th word in the text data. j Let P(y) represent the predicted value corresponding to the j-th text data. j |x1, x2, ..., x s Let represent the probability of the predicted value corresponding to the j-th text data given the input sequence of the j-th text data. Specifically, the second conditional probability is 1 if the predicted value corresponding to the j-th text data is the same as the label value corresponding to the j-th text data; or, the second conditional probability is 0 if the predicted value corresponding to the j-th text data is different from the label value corresponding to the j-th text data.

[0183] Formula (4) shows an example of the calculation process for mean square error loss information.

[0184]

[0185] Where L3 represents the mean squared error loss information, R j Let y represent the i-th word in the text data. jLet P(y) represent the predicted value corresponding to the j-th text data. j |x1, x2, ..., x s ) represents the tag value corresponding to the j-th text data.

[0186] In step 809, based on the above loss information, the parameters of the network layers such as the query weight matrix, the first modulus weight vector, the keyword weight matrix, and the second modulus weight vector are updated to obtain the target parameter values ​​of the modulus weight vector and other parameters.

[0187] The embodiments of this application can characterize the mapping relationship between loss information and parameters of the large language model through any of the loss functions shown in formulas (2) to (4). In practical applications, partial derivatives of the parameters of the large language model can be calculated, and the obtained partial derivatives of the parameters can be written out in vector form. The vector corresponding to the partial derivatives can be called the gradient information of the parameters. The update amount of the parameters can be obtained based on the gradient information and the step size information.

[0188] When using gradient descent, batch gradient descent, stochastic gradient descent, or mini-batch gradient descent can be employed. In specific implementations, iterative training can be performed based on multiple batches. The convergence condition for the above iteration can be: the loss information meets the preset conditions. The preset conditions can be: the absolute value of the difference between the loss information and the target value is less than the difference threshold, or the number of iterations exceeds the number threshold, etc. For example, the target value of the mean squared error loss information shown in formula (4) is 0. In other words, the iteration can end when the loss information meets the preset conditions; in this case, the target parameter values ​​of the large language model can be obtained. In particular, the target parameter values ​​of parameters such as the modulus weight vector can be obtained to improve the training sufficiency of word units.

[0189] In summary, the model training method of this application first normalizes the column dimensions of the query weight matrix and the keyword weight matrix during the model training process, and then adjusts the column lengths of the normalized query weight matrix or the normalized keyword weight matrix using the modulus-length weight vector.

[0190] Since one element in the modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix or the normalized keyword weight matrix, during model training, a single weight parameter is used to adjust the modulus information of a column vector, which increases the difficulty of modulus adjustment. Given this increased difficulty, the embodiments of this application are more conducive to adjusting the angles between different column vectors in the query weight matrix or the keyword weight matrix, allowing for more thorough training of these angles. Therefore, the embodiments of this application can improve the training sufficiency and effectiveness of the attention module in the large language model, and further, improve the overall training effect of the large language model.

[0191] This application embodiment, by improving the training sufficiency of the attention module and the overall training effect of the large language model, can achieve convergence of the large language model while reducing the number of iterations and training time. Because it reduces the number of iterations and training time, this application embodiment can reduce the amount of data transmitted during the training process of the large language model, thereby improving the hardware processing speed and response speed of the corresponding large language model. In other words, this application embodiment can overcome technical problems such as insufficient computing power, poor model response speed, and stuttering to a certain extent, thereby improving the response speed of the large language model.

[0192] Method Example 2

[0193] Reference Figure 9 The diagram illustrates a step-by-step flowchart of a text generation method according to an embodiment of this application. The method is applied to a large language model, where the parameters of the attention processing module include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. The method specifically includes the following steps:

[0194] Step 901: After completing the training of the large language model, the query weight matrix is ​​normalized by the first modulus to obtain the normalized query weight matrix; the first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value.

[0195] Step 902: Using the first modulus weight vector, adjust the modulus of the normalized query weight matrix to obtain the adjusted query weight matrix; one element of the first modulus weight vector is used to adjust the modulus of one column vector of the normalized query weight matrix.

[0196] Step 903: After completing the training of the large language model, perform a second modulus normalization on the keyword weight matrix to obtain a normalized keyword weight matrix; the second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value.

[0197] Step 904: Using the second modulus weight vector, adjust the modulus of the normalized keyword weight matrix to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector of the normalized keyword weight matrix.

[0198] Step 905: For the text to be processed, obtain the second attention input matrix corresponding to the attention processing module;

[0199] Step 906: Based on the above-mentioned adjusted query weight matrix and adjusted keyword weight matrix, perform attention processing on the above-mentioned second attention input matrix to obtain the second attention processing result;

[0200] Step 907: Based on the results of the second attention processing described above, determine the target generated text output by the large language model.

[0201] Figure 9 The second embodiment of the method shown is used to generate target generated text corresponding to the text to be processed by utilizing a large language model that has been trained. Figure 9 The illustrated method embodiment two can be applied to scenarios such as summary generation, machine translation, literary creation, and intelligent dialogue. For example, in a summary generation scenario, the target generated text can be the summary text corresponding to the text to be processed. Similarly, in a machine translation scenario, the target generated text can be the translated text corresponding to the text to be processed. In a literary creation scenario, such as essay writing, the target generated text can be the essay text corresponding to the text to be processed. Or, in an intelligent dialogue scenario, the target generated text can be the answer text or reply text corresponding to the text to be processed.

[0202] One difference between step 901 and step 801 is that step 801 performs a first-order modulus normalization on the query weight matrix corresponding to the first current weight value during the training of the large language model; while step 901 performs a first-order modulus normalization on the query weight matrix corresponding to the first target weight value after the training of the large language model is completed. The first target weight value can be the weight value corresponding to the query weight matrix when training is complete.

[0203] One difference between step 902 and step 802 is that in step 802, during the training of the large language model, the normalized query weight matrix obtained based on the first current weight value is adjusted using the first modulus-length weight vector corresponding to the third current weight value; while in step 902, after the large language model is trained, the normalized query weight matrix obtained based on the first target weight value is adjusted using the first modulus-length weight vector corresponding to the third target weight value. The third target weight value can be the weight value corresponding to the first modulus-length weight vector after training is complete.

[0204] Similarly, one difference between step 903 and step 803 is that step 803 performs second modulus normalization on the keyword weight matrix corresponding to the second current weight value during the training of the large language model; while step 903 performs second modulus normalization on the keyword weight matrix corresponding to the second target weight value after the training of the large language model is completed. The second target weight value can be the weight value corresponding to the keyword weight matrix when training is complete.

[0205] One difference between step 904 and step 804 is that in step 804, during the training of the large language model, the normalized keyword weight matrix obtained based on the second current weight value is adjusted using the second modulus-length weight vector corresponding to the fourth current weight value; while in step 904, after the large language model is trained, the normalized keyword weight matrix obtained based on the second target weight value is adjusted using the second modulus-length weight vector corresponding to the fourth target weight value. The fourth target weight value can be the weight value corresponding to the second modulus-length weight vector after training is complete.

[0206] The adjusted query weight matrix obtained in step 902 can be used as... Figure 3 The query weight matrix is ​​used for inference in the large language model. The adjusted keyword weight matrix obtained in step 904 can be used as... Figure 3 The keyword weight matrix is ​​used for inference in large language models.

[0207] In step 905, the text to be processed can be segmented into tokens, and the input matrix corresponding to the tokens of the text to be processed can be obtained.

[0208] Assuming the size of the input matrix is ​​[b, s], the embedding layer can perform mapping processing on the input matrix, and the size of the resulting embedding representation can be [b, s, c]. The embedding layer can then pass the embedding representation to the processing layer.

[0209] The embedding layer corresponds to an embedding weight matrix. The process of using the embedding layer to perform the first mapping process on the input matrix specifically includes: matching the token identifiers in the input matrix with the row indices in the embedding weight matrix to obtain the target row corresponding to the token identifiers in the input matrix; and determining the embedding representation corresponding to the input matrix based on the target row.

[0210] The embedding weight matrix has dimensions of n rows and c columns. Each row in the n rows corresponds to an index, which can be a word identifier. Therefore, by matching the word identifiers in the input matrix with the row indices in the embedding weight matrix, the target rows corresponding to the word identifiers in the input matrix can be obtained. s word identifiers in the input matrix can correspond to s target rows. Therefore, the size of the embedding representation obtained by the embedding layer can be [b, s, c].

[0211] The processing layer may include: Figure 1 The diagram shows N converter layers 102, 103, 104, and 105, including converter layer 1 to converter layer N. The processing layer can process the embedded representation to obtain text information.

[0212] like Figure 2 and Figure 3 As shown, the input normalization layer included in the attention processing module of converter layer 1 can first normalize the embedded representation to obtain a normalized embedded representation. Therefore, for converter layer 1, the second attention input matrix corresponding to the attention processing module can be: the normalized embedded representation. For other converter layers, those skilled in the art can determine the second attention input matrix corresponding to the attention processing module based on existing common sense. This application embodiment does not limit the specific determination process of the second attention input matrix corresponding to the attention processing module. One difference between the second attention input matrix and the first attention input matrix is ​​that the second attention input matrix corresponds to the text to be processed, while the first attention input matrix corresponds to the text data.

[0213] In step 906, the second query matrix Q(s×c) and the second keyword matrix K(s×c) corresponding to the second attention input matrix can be determined first based on the above-mentioned adjusted query weight matrix and the above-mentioned adjusted keyword weight matrix. Then, the second attention processing result can be determined by using formula (1) based on the second query matrix Q(s×c) and the second keyword matrix K(s×c).

[0214] Assuming the second attention input matrix is ​​X2(s×c), then the second attention input matrix X2(s×c) and the adjusted query weight matrix M' can be used to... Q Perform matrix multiplication to obtain the second query matrix Q(s×c). The second attention input matrix X2(s×c) can also be compared with the adjusted keyword weight matrix M'. K Perform matrix multiplication to obtain the second keyword matrix K(s×c). Then, the second query matrix Q(s×c) and the transpose of the second keyword matrix K(s×c) can be multiplied. T (c×s) performs matrix multiplication to obtain the result of the second attention processing.

[0215] The N converter layers 102, 103, 104, and 105, from converter layer 1 to converter layer N, can determine the target text to be generated based on the results at the second attention point.

[0216] As mentioned earlier, the normalization layer 105 normalizes the probability distribution output by the linear transformation layer 104 to obtain target probability values ​​between [0,1]. The target probability value can be the probability that each word in the vocabulary becomes a target word. A single inference by the large language model yields a target word, which can be the word with the highest target probability value during the inference process, or one of the X largest candidate words with the highest target probability values, where X can be a positive integer such as 3. During text generation, the large language model can perform multiple inferences until it encounters a word that serves as a termination element. The target words obtained from these multiple inferences are then combined to obtain the target generated text.

[0217] The target generated text can be one or more. For example, X candidate lexical units obtained from multiple inferences can be combined, and one or more target generated texts can be selected from the combination results.

[0218] In practical applications, text can be generated and output from one or more targets. For example, in an intelligent dialogue scenario, the client can receive text to be processed submitted by the user, the server can receive the text to be processed sent by the client, and use the text generation method of this application embodiment to determine one or more target generated texts, and return one or more target generated texts to the client, so that the client can provide the user with the response text or answer text of the intelligent dialogue.

[0219] In summary, the text generation method of this application embodiment improves the accuracy and other performance of the large language model by training the large language model, thus improving the training effect of the large language model. Therefore, the text generation method of this application embodiment can improve the accuracy of the target generated text.

[0220] It should be noted that, for the sake of simplicity, the method embodiments are all described as a series of actions. However, those skilled in the art should understand that the embodiments of this application are not limited to the described order of actions, because according to the embodiments of this application, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also understand that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily required by the embodiments of this application.

[0221] Based on the above embodiments, this embodiment also provides a model training device for training a large language model. The parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector. (Refer to...) Figure 10 The aforementioned device includes:

[0222] The first module length normalization module 1001 is used to perform first module length normalization on the query weight matrix to obtain a normalized query weight matrix; the first module length normalization is used to normalize the module lengths of different column vectors in the query weight matrix to the same preset value.

[0223] The first module length adjustment module 1002 is used to adjust the module length of the normalized query weight matrix using the first module length weight vector to obtain the adjusted query weight matrix; one element of the first module length weight vector is used to adjust the module length of one column vector of the normalized query weight matrix.

[0224] The second module length normalization module 1003 is used to perform second module length normalization on the keyword weight matrix to obtain a normalized keyword weight matrix; the second module length normalization is used to normalize the module lengths of different column vectors in the keyword weight matrix to the same preset value.

[0225] The second module length adjustment module 1004 is used to adjust the module length of the normalized keyword weight matrix using the second module length weight vector to obtain the adjusted keyword weight matrix; one element of the second module length weight vector is used to adjust the module length of one column vector of the normalized keyword weight matrix.

[0226] The attention input module 1005 is used to obtain the first attention input matrix corresponding to the attention processing module for text data;

[0227] Attention processing module 1006 is used to perform attention processing on the first attention input matrix according to the above-mentioned adjusted query weight matrix and adjusted keyword weight matrix to obtain the first attention processing result;

[0228] The output determination module 1007 is used to determine the text information of the large language model based on the first attention processing result.

[0229] The loss determination module 1008 is used to determine loss information based on the above text information;

[0230] The parameter update module 1009 is used to update the parameters of the large language model based on the loss information mentioned above.

[0231] Optionally, the aforementioned first module length adjustment module 1002 specifically includes:

[0232] The first extension module is used to extend the first modulus weight vector into a first modulus weight matrix using an automatic broadcasting mechanism; the size of the first modulus weight matrix is ​​the same as the size of the normalized query weight matrix.

[0233] The first dot product module is used to perform a dot product operation on the normalized query weight matrix and the first modulus weight matrix to obtain the adjusted query weight matrix.

[0234] Optionally, the aforementioned second module length adjustment module 1004 specifically includes:

[0235] The second extension module is used to extend the second modulus weight vector into a second modulus weight matrix using an automatic broadcasting mechanism; the size of the second modulus weight matrix is ​​the same as the size of the normalized keyword weight matrix.

[0236] The second dot product module is used to perform a dot product operation on the normalized keyword weight matrix and the second modulus weight matrix to obtain the adjusted keyword weight matrix.

[0237] Optionally, the above training specifically includes: pre-training, the above text data includes: N word units; the above text information represents: the prediction information of the i-th word unit; N and i are both positive integers; i is not greater than N;

[0238] The aforementioned loss determination module 1008 specifically includes:

[0239] The first loss determination module is used to determine the first cross-entropy loss information based on the conditional probability corresponding to the predicted information of the i-th word and the actual information of the i-th word in the above text data; the conditional probability is the probability of the predicted information of the i-th word appearing under the condition of the first i-1 words in the above text data.

[0240] Optionally, the above training specifically includes: fine-tuning training, the above text data includes: the input sequence and the label values ​​of the output sequence; the above text information includes: the predicted values ​​of the output sequence;

[0241] The aforementioned loss determination module 1008 includes:

[0242] The second loss determination module is used to determine the second cross-entropy loss information or mean squared error loss information based on the above predicted value and label value.

[0243] In summary, the model training apparatus of this application first normalizes the column dimensions of the query weight matrix and the keyword weight matrix during the model training process, and then adjusts the column length of the normalized query weight matrix or the normalized keyword weight matrix using the modulus weight vector.

[0244] Since one element in the modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix or the normalized keyword weight matrix, during model training, the use of a single weight parameter to adjust the modulus information of a column vector increases the difficulty of modulus adjustment. Given this increased difficulty, the embodiments of this application are more conducive to adjusting the angles between different column vectors in the query weight matrix or the keyword weight matrix, allowing for more thorough training of these angles. Therefore, the embodiments of this application can improve the training effect of large language models.

[0245] Based on the above embodiments, this embodiment also provides a text generation device. This device is applied to a large language model, and the parameters of the attention processing module in the large language model specifically include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector; (Refer to...) Figure 11 The aforementioned device specifically includes the following modules:

[0246] The first module length normalization module 1101 is used to perform first module length normalization on the query weight matrix after the training of the large language model is completed, so as to obtain a normalized query weight matrix; the first module length normalization is used to normalize the module lengths of different column vectors in the query weight matrix to the same preset value.

[0247] The first module length adjustment module 1102 is used to adjust the module length of the normalized query weight matrix using the first module length weight vector to obtain the adjusted query weight matrix; one element of the first module length weight vector is used to adjust the module length of one column vector of the normalized query weight matrix.

[0248] The second module length normalization module 1103 is used to perform second module length normalization on the keyword weight matrix after the training of the large language model is completed, so as to obtain a normalized keyword weight matrix; the second module length normalization is used to normalize the module lengths of different column vectors in the keyword weight matrix to the same preset value.

[0249] The second module length adjustment module 1104 is used to adjust the module length of the normalized keyword weight matrix using the second module length weight vector to obtain the adjusted keyword weight matrix; one element of the second module length weight vector is used to adjust the module length of one column vector of the normalized keyword weight matrix.

[0250] The attention input module 1105 is used to obtain the second attention input matrix corresponding to the attention processing module for the text to be processed;

[0251] Attention processing module 1106 is used to perform attention processing on the second attention input matrix according to the above-mentioned adjusted query weight matrix and adjusted keyword weight matrix to obtain the second attention processing result;

[0252] The text generation module 1107 is used to determine the target generated text output by the large language model based on the results of the second attention processing.

[0253] In summary, the text generation apparatus of this application improves the accuracy and other performance of the large language model by enhancing its training effect. Therefore, the text generation apparatus of this application can improve the accuracy of the target generated text.

[0254] This application also provides a non-volatile readable storage medium storing one or more modules (programs). When these modules are applied to a device, they enable the device to execute the instructions for the method steps in this application.

[0255] This application provides one or more machine-readable media storing instructions that, when executed by one or more processors, cause an electronic device to perform one or more of the methods described in the above embodiments. In this application, the electronic device includes various types of devices such as terminal devices and servers (clusters).

[0256] The embodiments of this disclosure can be implemented as an apparatus configured as desired using any suitable hardware, firmware, software, or any combination thereof, including electronic devices such as terminal devices and servers (clusters). Figure 12 An exemplary apparatus 2100 that can be used to implement the various embodiments described in this application is shown.

[0257] In one embodiment, Figure 12An exemplary device 2100 is shown, which includes one or more processors 2102, a control module (chipset) 2104 coupled to at least one of the processors 2102, a memory 2106 coupled to the control module 2104, an NVM / storage device 2108 coupled to the control module 2104, one or more input / output devices 2110 coupled to the control module 2104, and a network interface 2112 coupled to the control module 2104. Here, NVM (non-volatile memory) refers to memory that stores information that can persist for a long time after the power is turned off, making it less prone to loss.

[0258] Processor 2102 may include one or more single-core or multi-core processors, and processor 2102 may include any combination of general-purpose processors or special-purpose processors (e.g., graphics processors, application processors, baseband processors, etc.). In some embodiments, device 2100 can serve as a terminal device, server (cluster), or other device as described in the embodiments of this application.

[0259] In some embodiments, apparatus 2100 may include one or more computer-readable media (e.g., memory 2106 or non-volatile memory / storage device 2108) having instructions 2114 and one or more processors 2102 that are combined with the one or more computer-readable media and configured to execute the instructions 2114 to implement the module and thus perform the actions described in this disclosure.

[0260] In one embodiment, the control module 2104 may include any suitable interface controller to provide any suitable interface to at least one of the processors 2102 and / or any suitable device or component communicating with the control module 2104.

[0261] The control module 2104 may include a memory controller module to provide an interface to the memory 2106. The memory controller module may be a hardware module, a software module, and / or a firmware module.

[0262] Memory 2106 may be used, for example, to load and store data and / or instructions 2114 for device 2100. In one embodiment, memory 2106 may include any suitable volatile memory, such as suitable DRAM (Dynamic Random Access Memory). In some embodiments, memory 2106 may include double data rate type quad synchronous dynamic random access memory.

[0263] In one embodiment, the control module 2104 may include one or more input / output controllers to provide an interface to the non-volatile memory / storage device 2108 and (one or more) input / output devices 2110.

[0264] For example, non-volatile memory / storage device 2108 may be used to store data and / or instructions 2114. Non-volatile memory / storage device 2108 may include any suitable non-volatile memory (e.g., flash memory) and / or may include any suitable (one or more) non-volatile storage devices (e.g., one or more hard disk drives, one or more optical disk drives, and / or one or more digital universal optical disk drives).

[0265] The non-volatile memory / storage device 2108 may include storage resources that are physically part of a device on which the device 2100 is mounted, or that can be accessed by the device without being part of the device. For example, the non-volatile memory / storage device 2108 may be accessed via a network via one or more input / output devices 2110.

[0266] One or more input / output devices 2110 may provide an interface for device 2100 to communicate with any other suitable device. Input / output devices 2110 may include communication components, audio components, sensor components, etc. A network interface 2112 may provide an interface for device 2100 to communicate via one or more networks. Device 2100 may wirelessly communicate with one or more components of a wireless network according to any of one or more wireless network standards and / or protocols, such as accessing wireless networks based on communication standards, such as WiFi (Wireless Fidelity), 2G (2-Generation wireless telephone technology), 3G (3-Generation wireless telephone technology), 4G (4-Generation wireless telephone technology), 5G (5-Generation wireless telephone technology), etc., or combinations thereof.

[0267] In one embodiment, at least one of the processors 2102 may be logically packaged with one or more controllers (e.g., memory controller modules) of the control module 2104. In one embodiment, at least one of the processors 2102 may be logically packaged with one or more controllers of the control module 2104 to form a system-in-package. In one embodiment, at least one of the processors 2102 may be integrated with the logic of one or more controllers of the control module 2104 on the same die. In one embodiment, at least one of the processors 2102 may be integrated with the logic of one or more controllers of the control module 2104 on the same die to form a system-on-a-chip.

[0268] In various embodiments, device 2100 may be, but is not limited to, a server, desktop computing device, or mobile computing device (e.g., laptop computing device, handheld computing device, tablet computer, netbook, etc.). In various embodiments, device 2100 may have more or fewer components and / or different architectures. For example, in some embodiments, device 2100 includes one or more cameras, a keyboard, a liquid crystal display screen (including a touchscreen display), a non-volatile memory port, multiple antennas, a graphics chip, an application-specific integrated circuit (ASIC), and a speaker.

[0269] The detection device may use a main control chip as a processor or control module, and sensor data, position information, etc. may be stored in a memory or non-volatile memory / storage device. The sensor group may be used as an input / output device, and the communication interface may include a network interface.

[0270] As the device embodiment is basically similar to the method embodiment, the description is relatively simple, and relevant parts can be found in the description of the method embodiment.

[0271] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other.

[0272] This application describes embodiments with reference to flowchart illustrations and / or block diagrams of methods, terminal devices (systems), and computer program products according to embodiments of this application. It should be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing terminal device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing terminal device, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.

[0273] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing terminal device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.

[0274] These computer program instructions may also be loaded onto a computer or other programmable data processing terminal equipment to cause a series of operational steps to be performed on the computer or other programmable terminal equipment to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable terminal equipment, provide steps for implementing the functions specified in one or more flowcharts and / or one or more blocks of a block diagram.

[0275] Although preferred embodiments of the present application have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments as well as all changes and modifications falling within the scope of the embodiments of the present application.

[0276] Finally, it should be noted that in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or terminal device that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or terminal device. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or terminal device that includes said element.

[0277] The above provides a detailed description of a model training method and apparatus, a text generation method and apparatus, an electronic device, and a machine-readable medium provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A model training method, characterized in that, The method is used to train a large language model, wherein the parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector; the method includes: The query weight matrix is ​​subjected to a first modulus normalization to obtain a normalized query weight matrix; the first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value; The normalized query weight matrix is ​​adjusted using the first modulus weight vector to obtain the adjusted query weight matrix; an element of the first modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix. The keyword weight matrix is ​​subjected to a second modulus normalization to obtain a normalized keyword weight matrix; the second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value; The normalized keyword weight matrix is ​​adjusted using the second modulus weight vector to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector in the normalized keyword weight matrix. For text data, obtain the first attention input matrix corresponding to the attention processing module; Based on the adjusted query weight matrix and the adjusted keyword weight matrix, attention processing is performed on the first attention input matrix to obtain the first attention processing result; Based on the first attention processing result, the text information of the large language model is determined; Based on the text information, determine the loss information; The parameters of the large language model are updated based on the loss information.

2. The method according to claim 1, characterized in that, The step of adjusting the modulus of the normalized query weight matrix using the first modulus-length weight vector includes: Using an automatic broadcasting mechanism, the first modulus weight vector is expanded into a first modulus weight matrix; the size of the first modulus weight matrix is ​​the same as the size of the normalized query weight matrix. Perform a dot product operation between the normalized query weight matrix and the first modulus weight matrix to obtain the adjusted query weight matrix.

3. The method according to claim 1, characterized in that, The step of adjusting the modulus of the normalized keyword weight matrix using the second modulus-length weight vector includes: Using an automatic broadcasting mechanism, the second modulus weight vector is expanded into a second modulus weight matrix; the size of the second modulus weight matrix is ​​the same as the size of the normalized keyword weight matrix. Perform a dot product operation between the normalized keyword weight matrix and the second modulus weight matrix to obtain the adjusted keyword weight matrix.

4. The method according to claim 1, characterized in that, The training includes: pre-training; the text data includes: N word units; the text information represents: the prediction information of the i-th word unit; N and i are both positive integers; i is not greater than N; The step of determining the loss information based on the text information includes: The first cross-entropy loss information is determined based on the conditional probability corresponding to the predicted information of the i-th word and the actual information of the i-th word in the text data; the conditional probability is the probability of the predicted information of the i-th word appearing under the condition of the first i-1 words in the text data.

5. The method according to claim 1, characterized in that, The training includes fine-tuning training, and the text data includes: the input sequence and the label values ​​of the output sequence; the text information includes: the predicted values ​​of the output sequence; The step of determining the loss information based on the text information includes: Based on the predicted value and the label value, determine the second cross-entropy loss information or the mean squared error loss information.

6. A text generation method, characterized in that, The method is applied to a large language model, wherein the parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector; the method includes: After training the large language model, the query weight matrix is ​​normalized by the first modulus to obtain the normalized query weight matrix. The first modulus normalization is used to normalize the modulus of different column vectors in the query weight matrix to the same preset value. The normalized query weight matrix is ​​adjusted using the first modulus weight vector to obtain the adjusted query weight matrix; an element of the first modulus weight vector is used to adjust the modulus of a column vector in the normalized query weight matrix. After training the large language model, the keyword weight matrix is ​​subjected to a second modulus normalization to obtain a normalized keyword weight matrix. The second modulus normalization is used to normalize the modulus of different column vectors in the keyword weight matrix to the same preset value. The normalized keyword weight matrix is ​​adjusted using the second modulus weight vector to obtain the adjusted keyword weight matrix; one element of the second modulus weight vector is used to adjust the modulus of one column vector in the normalized keyword weight matrix. For the text to be processed, obtain the second attention input matrix corresponding to the attention processing module; Based on the adjusted query weight matrix and the adjusted keyword weight matrix, attention processing is performed on the second attention input matrix to obtain the second attention processing result; Based on the result of the second attention processing, the target generated text output by the large language model is determined.

7. A model training device, characterized in that, The apparatus is used to train a large language model, wherein the parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector; the apparatus includes: The first module normalization module is used to perform first module normalization on the query weight matrix to obtain a normalized query weight matrix; the first module normalization is used to normalize the module lengths of different column vectors in the query weight matrix to the same preset value. The first module length adjustment module is used to adjust the normalized query weight matrix using the first module length weight vector to obtain the adjusted query weight matrix; an element of the first module length weight vector is used to adjust the module length of a column vector in the normalized query weight matrix. The second module normalization module is used to perform second module normalization on the keyword weight matrix to obtain a normalized keyword weight matrix; the second module normalization is used to normalize the module lengths of different column vectors in the keyword weight matrix to the same preset value. The second module length adjustment module is used to adjust the normalized keyword weight matrix using the second module length weight vector to obtain the adjusted keyword weight matrix; an element of the second module length weight vector is used to adjust the module length of a column vector in the normalized keyword weight matrix. The attention input module is used to obtain the first attention input matrix corresponding to the attention processing module for text data; An attention processing module is used to perform attention processing on the first attention input matrix according to the adjusted query weight matrix and the adjusted keyword weight matrix to obtain a first attention processing result; The output determination module is used to determine the text information of the large language model based on the first attention processing result; The loss determination module is used to determine loss information based on the text information; The parameter update module is used to update the parameters of the large language model based on the loss information.

8. A text generation device, characterized in that, The device is applied to a large language model, wherein the parameters of the attention processing module in the large language model include: a query weight matrix, a first-order-length weight vector, a keyword weight matrix, and a second-order-length weight vector; the device includes: The first module normalization module is used to perform first module normalization on the query weight matrix after the training of the large language model is completed, so as to obtain a normalized query weight matrix; the first module normalization is used to normalize the module lengths of different column vectors in the query weight matrix to the same preset value. The first module length adjustment module is used to adjust the normalized query weight matrix using the first module length weight vector to obtain the adjusted query weight matrix; an element of the first module length weight vector is used to adjust the module length of a column vector in the normalized query weight matrix. The second module normalization module is used to perform second module normalization on the keyword weight matrix after the training of the large language model is completed, so as to obtain a normalized keyword weight matrix. The second module normalization is used to normalize the module length of different column vectors in the keyword weight matrix to the same preset value. The second module length adjustment module is used to adjust the normalized keyword weight matrix using the second module length weight vector to obtain the adjusted keyword weight matrix; an element of the second module length weight vector is used to adjust the module length of a column vector in the normalized keyword weight matrix. The attention input module is used to obtain the second attention input matrix corresponding to the attention processing module for the text to be processed; The attention processing module is used to perform attention processing on the second attention input matrix according to the adjusted query weight matrix and the adjusted keyword weight matrix to obtain the second attention processing result; The text generation module is used to determine the target generated text output by the large language model based on the result of the second attention processing.

9. An electronic device, characterized in that, include: processor; and A memory having executable code stored thereon, which, when executed, causes the processor to perform the method as described in any one of claims 1-6.

10. A machine-readable medium having executable code stored thereon, which, when executed, causes a processor to perform the method as claimed in any one of claims 1-6.

Citation Information

Patent Citations

  • Word weight determination method of text and training method and device of neural network model

    CN115081427A

  • Model training method and device, electronic equipment and storage medium

    CN117786104A