A Model Training Method and Device Based on Re-Attention Mechanism
By introducing a reattention mechanism into the pre-trained language model, the multi-head attention layer and sparse gated network dynamically adjusting the weight of the attention head, the problem of low prediction accuracy in complex contexts is solved, and higher prediction accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202510562334.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Existing pretrained language models are difficult to adaptively focus on key attention heads based on the input context, resulting in lower prediction accuracy of the model in complex contexts.
The model training method based on the reattention mechanism is adopted, and the global semantic features are extracted through the multi-head attention layer, and the dynamic weight is calculated in combination with the sparse gated network to suppress the weight of the non-critical attention head, and the model parameters are adjusted through the preset loss function to achieve dynamic allocation of the weight of the attention head.
It significantly improves the prediction accuracy and generalization ability of the pre-trained language model in complex contexts, solves the problem that the model is difficult to adaptively focus on key attention heads, and improves the prediction accuracy.
Smart Images

Figure CN120087412B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and particularly to a model training method and device based on a re-attention mechanism. Background Art
[0002] With the development of natural language processing (NLP, short for Natural Language Processing) technology, pre-trained language models such as BERT (short for Bidirectional Encoder Representations from Transformers), GPT (short for Generative Pre-trained Transformer), etc., which are based on the Transformer architecture, have achieved remarkable results in various tasks.
[0003] In the traditional pre-trained language model architecture, the multi-head attention (MHA) mechanism is usually adopted. Each attention head in the model independently focuses on different aspects of the input sequence, and linearly combines the outputs of all attention heads through a fixed weight matrix. However, the adopted multi-head attention mechanism has certain limitations. Especially when dealing with texts with complex context changes, it is difficult for the model to adaptively focus on the key attention heads according to the input context. For example, when the sequence corresponding to the sentence "The bank can hold investments or river water" is input into the pre-trained language model, for the token "bank", it is difficult for the model to adaptively focus on the key attention heads "geographical entity" or "financial entity" according to the context of the input sequence, thus affecting the prediction accuracy of the model.
[0004] Based on this, there is an urgent need for a high-precision model training method to achieve adaptive focusing on key attention heads according to the input context, thereby improving the prediction accuracy of the pre-trained language model. Summary of the Invention
[0005] Based on the above problems, this application provides a model training method and device based on a re-attention mechanism, aiming to solve the technical problem that the existing pre-trained language model has a low prediction accuracy because it is difficult to adaptively focus on the key attention heads according to the input context, so as to improve the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0006] The embodiments of this application disclose the following technical solutions:
[0007] In the first aspect of the present application, a model training method based on a re-attention mechanism is provided, which is applied to a pre-trained language model. The pre-trained language model includes a multi-head attention layer and a sparse gating network. The method includes:
[0008] Input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features. The embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text.
[0009] The sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors, and constructs an attention weight vector based on the dynamic weights. The sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads.
[0010] The multi-head attention layer performs weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors to obtain a fused context representation, and performs a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text.
[0011] Adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model.
[0012] In an optional implementation manner, the step of the multi-head attention layer extracting global semantic features from the embedding vector sequences and generating global context vectors based on the global semantic features includes:
[0013] Input each of the embedding vector sequences into the multi-head attention heads respectively to obtain context vectors output by each attention head, and generate the global semantic features based on the context vectors output by each attention head. The attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequences.
[0014] Input the global semantic features into an adaptive pooling layer. The adaptive pooling layer performs pooling processing on the global semantic features based on a low-rank projection matrix to generate global context vectors corresponding to the embedding vector sequences. The low-rank projection matrix is used to map the global semantic features from a high-dimensional feature space to a low-dimensional feature space.
[0015] In an alternative implementation, the formula for calculating the dynamic weights corresponding to each of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors is as follows:
[0016]
[0017] where g represents the dynamic weight, c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix for mapping the global context vector to the gating score space; b g represents the gating bias vector.
[0018] In an alternative implementation, the preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of the non-critical attention heads.
[0019] In an alternative implementation, the method of adjusting the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences by using the preset loss function to obtain the target pre-trained language model includes:
[0020] Calculating a first loss value based on the self-attention values corresponding to each of the embedding vector sequences by using the language model loss function;
[0021] Calculating a second loss value based on the attention weight vectors corresponding to each of the embedding vector sequences by using the weight sparsity regularization term;
[0022] Determining a total loss value based on each of the first loss values and the corresponding second loss values;
[0023] Performing joint model training on the pre-trained language model and the sparse gating network based on each of the total loss values, and stopping the model training until the parameters meet the training cut-off condition, to obtain the target pre-trained language model.
[0024] In an alternative implementation, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0025] In the second aspect of the present application, a prediction method based on a re-attention mechanism is provided, which is applied to a target pre-trained language model. The target pre-trained language model is a model trained by the model training method based on the re-attention mechanism introduced in any one of the implementation manners in the first aspect of the embodiments of the present application. The target pre-trained language model includes a multi-head attention layer and a sparse gating network. The method includes:
[0026] Perform embedding vector encoding on the target natural language text input by the user to obtain a target embedding vector sequence; the target embedding vector sequence is composed of hidden state vectors corresponding to each token in the target natural language text;
[0027] Input the target embedding vector sequence into the multi-head attention layer to obtain a target global context vector output by the multi-head attention layer;
[0028] Input the target global context vector into the sparse gating network to obtain a target attention weight vector output by the sparse gating network; the target attention weight vector is composed of target dynamic weights corresponding to each multi-head attention head in the multi-head attention layer;
[0029] The multi-head attention layer performs weighted fusion on the target context vectors output by each attention head and the target attention weight vector to obtain a target fused context representation;
[0030] The multi-head attention layer performs a linear transformation on the target fused context representation based on a fixed weight matrix to obtain a target self-attention value; the target self-attention value is used to predict the next token corresponding to the target natural language text.
[0031] In the third aspect of the present application, a model training device based on a re-attention mechanism is provided, which is applied to a pre-trained language model. The pre-trained language model includes a multi-head attention layer and a sparse gating network. The device includes:
[0032] A feature extraction module, configured to input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features; the embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text;
[0033] A dynamic weight calculation module, configured to calculate, by the sparse gating network based on the sparse Softmax function and each of the global context vectors, the respective dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer, and construct an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of the non-critical attention heads among the multi-head attention heads;
[0034] A weighted fusion module, configured to perform weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation; and perform a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text;
[0035] A parameter adjustment module, configured to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model.
[0036] In a third aspect of the present application, there is provided a computer-readable storage medium storing a computer program, which when run by a processor, implements the above-mentioned model training method based on the re-attention mechanism.
[0037] In a fourth aspect of the present application, there is provided a processor for running a computer program, and when the computer program runs, it executes the above-mentioned model training method based on the re-attention mechanism.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] In the technical solution of the present application, first, multiple embedding vector sequences for training a pre-trained language model are respectively input into the multi-head attention layer, and the multi-head attention layer can accurately extract global semantic features from the embedding vector sequences, so that the global context vectors generated based on the accurate global semantic features can accurately capture the overall semantics and context of the entire input text (i.e., the natural language text); second, the sparse gating network calculates the respective dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector, and constructs an attention weight vector based on the dynamic weights. Since the sparse Softmax function is used to suppress the weights of the non-critical attention heads among the multi-head attention heads, the weights of each attention head are dynamically allocated by the sparse gating mechanism to sense the input semantic features in real time, which can reduce redundant calculations while accurately focusing on the key attention heads;
[0040] Then, the context vectors output by each attention head are weighted and fused with the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation, and a linear transformation is performed on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text, avoiding the problem that static weights cannot adapt to the input context. Furthermore, according to the self-attention values corresponding to each embedding vector sequence, the parameters of the pre-trained language model are adjusted through a preset loss function to obtain the target pre-trained language model, solving the technical problem that the existing pre-trained language models have low prediction accuracy because it is difficult to adaptively focus on key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0042] Figure 1 It is a flowchart of a model training method based on a re-attention mechanism provided by an embodiment of the present application;
[0043] Figure 2 It is a flowchart of a generation process of a global context vector provided by an embodiment of the present application;
[0044] Figure 3 It is a flowchart of a parameter adjustment process of a pre-trained language model provided by an embodiment of the present application;
[0045] Figure 4 It is a flowchart of a prediction method based on a re-attention mechanism provided by an embodiment of the present application;
[0046] Figure 5 It is a schematic structural diagram of a model training device based on a re-attention mechanism provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0047] As described above, currently, in traditional pre-trained language model architectures, the multi-head attention (MHA) mechanism is usually adopted. Each attention head in the model independently focuses on different aspects of the input sequence and linearly combines the outputs of all attention heads through a fixed weight matrix. However, the adopted multi-head attention mechanism has certain limitations. Especially when dealing with texts with complex context changes, it is difficult for the model to adaptively focus on key attention heads according to the input context. For example, when the sequence corresponding to the sentence "The bank can hold investments or riverwater" is input into the pre-trained language model, for the token "bank", it is difficult for the model to adaptively focus on the key attention heads "geographical entity" or "financial entity" according to the context of the input sequence, thus affecting the prediction accuracy of the model. Based on this, there is an urgent need for a high-precision model training method to achieve adaptive focusing on key attention heads according to the input context, thereby improving the prediction accuracy of the pre-trained language model.
[0048] The inventors have proposed a model training method based on the re-attention mechanism through research. In this solution, first, multiple embedded vector sequences used to train the pre-trained language model are respectively input into the multi-head attention layer. The multi-head attention layer can accurately extract global semantic features from the embedded vector sequences, and thus the global context vector generated based on the accurate global semantic features can accurately capture the overall semantics and context of the entire input text (i.e., natural language text). Secondly, the sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector, and constructs an attention weight vector based on the dynamic weights. Since the sparse Softmax function is used to suppress the weights of non-key attention heads among the multi-head attention heads, the dynamic allocation of the weights of each attention head can be realized by the sparse gating mechanism to perceive the input semantic features in real time, which can accurately focus on the key attention heads while reducing redundant calculations.
[0049] Then, the context vectors output by each attention head are weighted and fused with the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation, and a linear transformation is performed on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text, avoiding the problem that static weights cannot adapt to the input context; furthermore, according to the self-attention values corresponding to each embedding vector sequence, the parameters of the pre-trained language model are adjusted through a preset loss function to obtain a target pre-trained language model, solving the technical problem that the existing pre-trained language models have low prediction accuracy because it is difficult to adaptively focus on key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0050] It should be noted that the application scenarios of a model training method and device based on the re-attention mechanism provided in this application include, but are not limited to, natural language processing (such as machine translation, sentiment analysis), multi-modal models (such as image-text understanding), dialogue systems, etc.
[0051] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0052] Keyword definitions:
[0053] token: In the context of Natural Language Processing (NLP), "token" refers to the smallest analysis unit in text data, which can be a word, a punctuation mark, a number, etc.
[0054] Method Embodiment 1
[0055] An embodiment of a model training method based on the re-attention mechanism is provided in an embodiment of this application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0056] See Figure 1, which is a flowchart of a model training method based on a re-attention mechanism provided in an embodiment of the present application, which is applied to a pre-trained language model (such as a Transformer), wherein the pre-trained language model includes a multi-head attention layer and a sparse gating network. Figure 1 As shown, the method comprises the following steps:
[0057] In step S101, multiple embedding vector sequences used to train the pre-trained language model are respectively input into the multi-head attention layer, and the multi-head attention layer extracts global semantic features from the embedding vector sequences and generates a global context vector based on the global semantic features.
[0058] In step S101, the embedding vector sequence is composed of hidden state vectors corresponding to each token in the natural language text. In an embodiment of the present application, the embedding vector sequence is a sequence formed by converting each word or symbol in the text into a corresponding vector representation by the pre-trained language model through embedding vector encoding of the input natural language text. For example, the input natural language text is "Where did the apple fly to?", and the tokens corresponding to the text are "this", "one", "apple", "fruit", "fly", "to", "where", "in", "go", "have" and "?". After the above text is embedded vector encoded by the pre-trained language model, the hidden state vectors corresponding to the above tokens can be respectively, and the embedding vector sequence can be constructed based on the above hidden state vectors.
[0059] In the embodiment of the present application, the embedding vector sequence H can be expressed as:
[0060]
[0061] Where L represents the length of the embedding vector sequence; represents the set of real numbers; d model Represents the dimension of the embedding vector sequence.
[0062] In the embodiment of the present application, after the embedding vector sequence is input into the multi-head attention layer, the pre-trained language model can obtain accurate global semantic features through the context vector output by the multi-head attention head, and pool the global semantic features based on the low-rank projection matrix to obtain a global context vector that can accurately capture the overall semantics and context of the entire input text. Figure 2 , Figure 2 A flowchart of a process for generating a global context vector provided in an embodiment of the present application, the process comprising the following steps:
[0063] In step S1021, each embedded vector sequence is respectively input into the multi-head attention heads to obtain the context vectors output by each attention head, and global semantic features are generated based on the context vectors output by each attention head.
[0064] In step S1021, the attention heads are used to capture the similarity relationships between different tokens in the embedded vector sequence.
[0065] In step S1022, the global semantic features are input into the adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic features based on the low-rank projection matrix to generate global context vectors corresponding to the embedded vector sequences.
[0066] In step S1022, the low-rank projection matrix is used to map the global semantic features from the high-dimensional feature space to the low-dimensional feature space.
[0067] In the embodiment of the present application, the formula for the adaptive pooling layer to perform pooling processing on the global semantic features based on the low-rank projection matrix can be shown as formula (1):
[0068] (1)
[0069] Where c represents the global context vector; h i represents the feature representation of the i-th element of the input sequence (i.e., the global semantic features); α i represents the attention weight of the i-th element, which is used to measure the importance of this element to the global context; W a represents the low-rank projection matrix; U represents the output dimension of the low-rank projection matrix (i.e., the compressed feature dimension); V represents the input dimension of the low-rank projection matrix.
[0070] Optionally, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedded vector sequence.
[0071] It should be noted that by introducing the low-rank projection matrix to perform pooling processing on the global semantic features and constraining the range of the rank of the low-rank projection matrix to r ≤ 0.2×d model , it is possible to map the global semantic features from the high-dimensional feature space to the low-dimensional feature space, thereby reducing the number of parameters (the parameter compression rate can reach 80%) while enhancing the generalization ability of the pre-trained language model.
[0072] In step S102, the sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector, and constructs an attention weight vector based on the dynamic weights.
[0073] In step S102, the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads. Among them, the non-critical attention head is an attention head with a dynamic weight less than a preset threshold.
[0074] In the embodiment of the present application, based on the sparse Softmax function and each global context vector, the formula for calculating the respective dynamic weights of the multi-head attention heads in the multi-head attention layer is:
[0075]
[0076] Among them, g represents the dynamic weight, , N is the number of attention heads; c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector. Among them, the calculation rule of the sparse Softmax function is as follows:
[0077]
[0078] Among them, x i , x j are respectively the i-th and j-th elements of the input vector (i.e., the global context vector), which are used to represent the unnormalized gating scores.
[0079] In the embodiment of the present application, the pre-trained language model calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer through a sparse gating network based on the sparse Softmax function and each global context vector, which can suppress the weights of non-critical attention heads (such as attention heads with dynamic weights less than a preset threshold) in the multi-head attention heads, and realizes accurate focusing on critical attention heads while reducing redundant calculations.
[0080] Optionally, the sparsity of the sparse Softmax function can be adjusted according to actual needs. For example, the sparsity can be adjusted to 30% according to actual needs.
[0081] Step S103, the multi-head attention layer performs weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors to obtain a fused context representation; and performs a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text.
[0082] In the embodiment of the present application, the multi-head attention layer of the pre-trained language model weights and fuses the context vectors output by each attention head with the corresponding attention weight vectors through formula (2) to obtain the fused context representation Fused_Output.
[0083] (2)
[0084] Where, h i represents the context vector output by the i-th attention head; g i is the dynamic weight corresponding to the i-th attention head in the corresponding attention weight vector; N is the number of attention heads. The multi-head attention layer can also perform a linear transformation on the fused context representation based on a fixed weight matrix through formula (3) to obtain the self-attention value Final_Output for predicting the next token corresponding to the natural language text.
[0085] (3)
[0086] Where, W0 represents the fixed weight matrix corresponding to the multi-head attention layer, which is an immutable constant and will not be updated together with the parameters of the model; d out represents the dimension of the fused context representation.
[0087] In the embodiment of the present application, taking the input natural language text "Apple released a new product. Where did it fly?" as an example, if it is determined through the output self-attention value that the high weight is on "Apple" (company name), the pre-trained language model can accurately predict that the next token "where" may refer to the product release scope; if it is determined through the output self-attention value that the high weight is on "fly" (action), the pre-trained language model can accurately predict that the next token "where" may refer to the physical movement trajectory.
[0088] Step S104, adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence through a preset loss function to obtain the target pre-trained language model.
[0089] In the embodiment of the present application, the preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of non-critical attention heads. Among them, the preset loss function L total can be specifically shown as follows:
[0090]
[0091] Where, L MLMis the loss function of the language model (e.g., the Cross-Entropy Loss function); is the weight sparsity regularization term, is the attention weight vector; λ is used to control the sparsity intensity. For example, when λ = 0.1, the sparsification effect of the sparse gating network is optimal.
[0092] To improve the model prediction accuracy, in the embodiments of this application, the pre-trained language model can determine the total loss value based on the self-attention values and attention weight vectors corresponding to each embedding vector sequence through a preset loss function, and adjust the parameters of the pre-trained language model based on the total loss value to obtain the target pre-trained language model. Specifically, refer to Figure 3 This figure is a flowchart of the parameter adjustment process of a pre-trained language model provided by the embodiments of this application. This process includes the following steps:
[0093] Step S301, calculate the first loss value based on the self-attention values corresponding to each embedding vector sequence through the language model loss function.
[0094] For example, the pre-trained language model can calculate the first loss value based on the self-attention values corresponding to each embedding vector sequence through the Cross-Entropy Loss function.
[0095] Step S302, calculate the second loss value based on the attention weight vectors corresponding to each embedding vector sequence through the weight sparsity regularization term.
[0096] In the embodiments of this application, the weight sparsity regularization term can force the dynamic weights g of some attention heads (e.g., non-critical attention heads) to approach zero. For example, the initial attention weight vector is: g = [0.3, 0.6, 0.1], and after parameter adjustment: g = [0.2, 0.8, 0.0] (the third attention head is suppressed)
[0097] It should be noted that the role of the weight sparsity regularization term is to make gradually become smaller during the training process of the sparse gating network, so that the model selects the adjusted weights through the backpropagation mechanism during training, thereby improving the prediction accuracy of the model.
[0098] Step S303, determine the total loss value based on each first loss value and the corresponding second loss value.
[0099] Step S304, perform joint model training on the pre-trained language model and the sparse gating network based on each total loss value until the parameters meet the training termination condition, and then stop the model training to obtain the target pre-trained language model.
[0100] In the embodiments of the present application, when jointly training the pre-trained language model and the sparse gating network based on each total loss value, the parameters of the sparse gating network are gradually updated through the gradient clipping technique. The clipped gradient limits the step size of parameter update, making the adjustment of dynamic weights smoother, thereby preventing excessive fluctuations in dynamic weight allocation. For example, the dynamic weight may gradually adjust from 0.80.8 to 0.70.7 instead of mutating, which helps the model gradually adapt to different contexts.
[0101] In an alternative embodiment, the present application is significantly superior to the prior art in semantic tasks, especially outstanding in polysemy disambiguation (+4.2%) and short text classification (+1.8%). The specific data is shown in Table 1:
[0102] Table 1
[0103]
[0104] Taking the sentence "The bank can hold investments or river water" as an example, the head activation differences of "bank" under static and dynamic weights can be shown in Table 2, where the geographical entity head weight is increased from 0.3 to 0.8.
[0105] Table 2
[0106]
[0107] Through the model training method based on the re-attention mechanism provided by the embodiments of the present application, the weights of each attention head are dynamically allocated by the sparse gating mechanism to sense the input semantic features in real time, which can accurately focus on the key attention heads while reducing redundant calculations; the weights of non-key attention heads in the multi-head attention are suppressed by the sparse Softmax function, achieving accurate focusing on the key attention heads while reducing redundant calculations; by weighted fusion of the context vectors output by each attention head with the corresponding attention weight vectors, and performing a linear transformation on the fused context representation based on a fixed weight matrix, the self-attention value for predicting the next token corresponding to the natural language text is obtained, avoiding the problem that static weights cannot adapt to the input context; furthermore, by presetting a loss function to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence, the target pre-trained language model is obtained, solving the technical problem that the existing pre-trained language model has a low prediction accuracy because it is difficult to adaptively focus on the key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0108] In addition, the pre-trained language model trained by the model training method based on the re-attention mechanism provided by the embodiments of the present application is applicable to both the pre-training and fine-tuning stages, can be seamlessly integrated into existing models (such as BERT, GPT), and improves the multi-task compatibility of the pre-trained language model.
[0109] Method Embodiment Two
[0110] The embodiments of the present application provide an embodiment of a prediction method based on the re-attention mechanism. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0111] See Figure 4 Figure [FIGURE NUMBER], which is a flowchart of a prediction method based on the re-attention mechanism provided by the embodiments of the present application. This method is applied to a target pre-trained language model, and the target pre-trained language model includes a multi-head attention layer and a sparse gating network. Among them, the target pre-trained language model is a model trained by the model training method based on the re-attention mechanism shown in Embodiment One of the above method. As Figure 4 shown, the method includes the following steps:
[0112] Step S401: Perform embedding vector encoding on the target natural language text input by the user to obtain a target embedding vector sequence.
[0113] In step S401, the target embedding vector sequence is composed of the hidden state vectors corresponding to each token in the target natural language text.
[0114] Step S402: Input the target embedding vector sequence into the multi-head attention layer to obtain a target global context vector output by the multi-head attention layer.
[0115] In the embodiments of the present application, the multi-head attention heads in the multi-head attention layer can generate target context vectors based on the target embedding vector sequence, and extract target global semantic features from all the generated target context vectors; then input the target global semantic features into the adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic features based on the low-rank projection matrix to generate a target global context vector. Among them, the low-rank projection matrix is used to map the global semantic features from the high-dimensional feature space to the low-dimensional feature space.
[0116] Optionally, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0117] Step S403: Input the target global context vector into the sparse gating network to obtain the target attention weight vector output by the sparse gating network.
[0118] In step S403, the target attention weight vector is composed of the target dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer.
[0119] In the embodiment of the present application, after inputting the target global context vector into the sparse gating network, the sparse gating network can calculate the target dynamic weights corresponding to each attention head based on the sparse Softmax function and the target global context vector, and construct the target attention weight vector based on the target dynamic weights. Among them, the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads.
[0120] In the embodiment of the present application, based on the sparse Softmax function and the target global context vector, the formula for calculating the target dynamic weights corresponding to each attention head is:
[0121]
[0122] where g represents the target dynamic weight, , N is the number of attention heads; c represents the target global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector. Among them, the calculation rule of the sparse Softmax function is as follows:
[0123]
[0124] where x i , x j are the i-th and j-th elements of the input vector (i.e., the target global context vector), respectively, which are used to represent the unnormalized gating scores.
[0125] Step S404: The multi-head attention layer performs weighted fusion on the target context vectors output by each attention head and the target attention weight vector to obtain the target fused context representation.
[0126] In the embodiment of the present application, the multi-head attention layer can perform weighted fusion on the target context vectors output by each attention head and the target attention weight vector through formula (2) in Embodiment 1 of the above method to obtain the target fused context representation.
[0127] Step S405: The multi-head attention layer performs a linear transformation on the target fused context representation based on a fixed weight matrix to obtain the target self-attention value.
[0128] In step S405, the target self-attention value is used to predict the next token corresponding to the target natural language text.
[0129] In the embodiment of the present application, the multi-head attention layer can perform a linear transformation on the fused context representation based on the fixed weight matrix through formula (3) in the first method embodiment above to obtain the target self-attention value for predicting the next token corresponding to the target natural language text. Among them, the fixed weight matrix is an immutable constant and will not be updated together with the parameters of the model.
[0130] Through the prediction method based on the re-attention mechanism provided by the embodiment of the present application, it is realized to dynamically allocate the weights of each attention head by sparsely gating the mechanism to sense the input semantic features in real time, which can accurately focus on the key attention heads while reducing redundant calculations; by suppressing the weights of non-key attention heads in the multi-head attention heads through the sparse Softmax function, it is realized to accurately focus on the key attention heads while reducing redundant calculations, solving the technical problem that the existing pre-trained language model has a low prediction accuracy because it is difficult to adaptively focus on the key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0131] Device embodiment
[0132] The embodiment of the present application provides a model training device based on the re-attention mechanism, which is applied to a pre-trained language model. The pre-trained language model includes a multi-head attention layer and a sparse gating network. Among them, Figure 5 FIG. is a schematic structural diagram of a model training device based on the re-attention mechanism provided by the embodiment of the present application. As Figure 5 shown, the device includes: a feature extraction module 51, a dynamic weight calculation module 52, a weighted fusion module 53, and a parameter adjustment module 54. From Figure 5 it can be seen the connection relationship between several modules.
[0133] Among them, the feature extraction module 51 is used to input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features; the embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text.
[0134] The dynamic weight calculation module 52 is used to calculate the respective dynamic weights of the multi-head attention heads in the multi-head attention layer by the sparse gating network based on the sparse Softmax function and each global context vector, and construct an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of the non-critical attention heads in the multi-head attention heads;
[0135] The weighted fusion module 53 is used to perform weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation; and perform a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text;
[0136] The parameter adjustment module 54 is used to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence through a preset loss function to obtain a target pre-trained language model.
[0137] Optionally, the feature extraction module further includes: a generation unit and a pooling processing unit.
[0138] Among them, the generation unit is used to input each embedding vector sequence into the multi-head attention heads respectively to obtain the context vectors output by each attention head, and generate a global semantic feature based on the context vectors output by each attention head; the attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequence;
[0139] The pooling processing unit is used to input the global semantic feature into the adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic feature based on a low-rank projection matrix to generate a global context vector corresponding to the embedding vector sequence; the low-rank projection matrix is used to map the global semantic feature from a high-dimensional feature space to a low-dimensional feature space.
[0140] Optionally, the formula for calculating the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector is:
[0141]
[0142] Among them, g represents the dynamic weight, c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector.
[0143] Optionally, the preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of the non-critical attention heads.
[0144] Optionally, the parameter adjustment module includes: a first calculation unit, a second calculation unit, a determination unit, and a joint training unit.
[0145] Among them, the first calculation unit is used to calculate a first loss value based on the self-attention values corresponding to each embedding vector sequence through the language model loss function;
[0146] The second calculation unit is used to calculate a second loss value based on the attention weight vectors corresponding to each embedding vector sequence through the weight sparsity regularization term;
[0147] The determination unit is used to determine the total loss value based on each first loss value and the corresponding second loss value;
[0148] The joint training unit is used to perform joint model training on the pre-trained language model and the sparse gating network based on each total loss value, and stop the model training until the parameters meet the training cut-off condition, so as to obtain the target pre-trained language model.
[0149] Optionally, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0150] Storage medium embodiment
[0151] An embodiment of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements some or all of the steps of the model training method based on the re-attention mechanism introduced in the foregoing method embodiments of the present application. The storage medium can be various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0152] Processor embodiment
[0153] An embodiment of the present application provides a processor, which is used to run a program. When the program runs, it executes some or all of the steps of the model training method based on the re-attention mechanism introduced in the foregoing method embodiments.
[0154] It should be noted that the embodiments in this specification are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content. The system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0155] The above is only a specific implementation manner of this application, but the protection scope of this application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in this application should be covered by the protection scope of this application. Therefore, the protection scope of this application should be subject to the protection scope of the claims.
Claims
1. A model training method based on a re-attention mechanism, characterized in that Applied to a pre-trained language model, which includes a multi-head attention layer and a sparse gating network, the method includes: Inputting multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features. The embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text; The sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors, and constructs an attention weight vector based on the dynamic weights. The sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads; The multi-head attention layer performs weighted fusion on the context vectors output by each attention head and the corresponding attention weight vector to obtain a fused context representation, and performs a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text; Adjusting the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model; The step of the multi-head attention layer extracting global semantic features from the embedding vector sequences and generating global context vectors based on the global semantic features includes: Inputting each of the embedding vector sequences into the multi-head attention heads respectively to obtain context vectors output by each attention head, and generating the global semantic features based on the context vectors output by each attention head. The attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequences; Inputting the global semantic features into an adaptive pooling layer. The adaptive pooling layer performs pooling processing on the global semantic features based on a low-rank projection matrix to generate global context vectors corresponding to the embedding vector sequences. The low-rank projection matrix is used to map the global semantic features from a high-dimensional feature space to a low-dimensional feature space; 2. The method according to claim 1, wherein The formula for calculating the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors is: ; Among them, g represents the dynamic weight, and c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector.
3. The method according to claim 1, wherein The preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of the non-critical attention heads; 4. The method according to claim 3, characterized in that The step of adjusting the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model includes: Calculating a first loss value based on the self-attention values corresponding to each of the embedding vector sequences through the language model loss function; Calculate a second loss value based on the attention weight vectors corresponding to each of the embedding vector sequences through the weight sparsity regularization term; Determine the total loss value based on each of the first loss values and the corresponding second loss values; Perform joint model training on the pre-trained language model and the sparse gating network based on each of the total loss values, and stop the model training until the parameters meet the training termination condition, to obtain the target pre-trained language model.
5. The method according to claim 1, wherein The constraint condition of the low-rank projection matrix is r ≤ 0.2 × d model , where r represents the rank of the low-rank projection matrix, d model represents the dimension of the embedded vector sequence.
6. A prediction method based on a re-attention mechanism, characterized in that, Applied to a target pre-trained language model, the target pre-trained language model is a model trained by the method according to any one of claims 1-5 above, the target pre-trained language model includes a multi-head attention layer and a sparse gating network, and the method includes: Perform embedding vector encoding on the target natural language text input by the user to obtain a target embedding vector sequence; the target embedding vector sequence is composed of hidden state vectors corresponding to each token in the target natural language text; Input the target embedding vector sequence into the multi-head attention layer to obtain a target global context vector output by the multi-head attention layer; Input the target global context vector into the sparse gating network to obtain a target attention weight vector output by the sparse gating network; the target attention weight vector is composed of target dynamic weights corresponding to each of the multi-head attention heads in the multi-head attention layer; The multi-head attention layer performs weighted fusion on the target context vectors output by each attention head and the target attention weight vector to obtain a target fused context representation; The multi-head attention layer performs a linear transformation on the target fused context representation based on a fixed weight matrix to obtain a target self-attention value; the target self-attention value is used to predict the next token corresponding to the target natural language text.
7. A model training device based on a re-attention mechanism, characterized in that, Applied to a pre-trained language model, the pre-trained language model includes a multi-head attention layer and a sparse gating network, and the apparatus includes: A feature extraction module, configured to respectively input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer, and the multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features; the embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text; A dynamic weight calculation module, configured to calculate, by the sparse gating network, the dynamic weights corresponding to each of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors, and construct an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads; A weighted fusion module, configured to perform weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation; and perform a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text. A parameter adjustment module, configured to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model. The feature extraction module includes: a generation unit and a pooling processing unit. Among them, the generation unit is configured to input each of the embedding vector sequences into the multi-head attention heads respectively to obtain context vectors output by each attention head, and generate the global semantic feature based on the context vectors output by each attention head; the attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequences. The pooling processing unit is configured to input the global semantic feature into an adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic feature based on a low-rank projection matrix to generate a global context vector corresponding to the embedding vector sequence; the low-rank projection matrix is used to map the global semantic feature from a high-dimensional feature space to a low-dimensional feature space.
8. A computer-readable storage medium, characterized in that, A computer-readable storage medium stores a computer program, and when the computer program is run by a processor, it implements the model training method based on the re-attention mechanism according to any one of claims 1-5.
9. A processor, characterized in that, For running a computer program, the computer program, when running, executes the model training method based on the re-attention mechanism according to any one of claims 1-5.
Citation Information
Patent Citations
Self-adaptive relation modeling method for structured data
CN113191441A
Large language model training method and device
CN118445379A