Model training method and device based on reattention mechanism
By introducing a reattention mechanism and sparse gated network into the pre-trained language model, dynamically adjusting the weight of the attention head, the problem that existing models are difficult to adaptively focus on key attention heads is solved, and the prediction accuracy and generalization ability of the model are significantly improved.
Patent Information
- Application Number
- CN202510562334.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2045-04-30
AI Technical Summary
Existing pretrained language models are difficult to adaptively focus on key attention heads based on the input context, resulting in low model prediction accuracy.
Using a model training method based on the reattention mechanism, global semantic features are extracted through the multi-head attention layer and global context vectors are generated. The dynamic weight is calculated in combination with the sparse gated network, and the weight of the attention head is dynamically adjusted to achieve adaptive focus.
It significantly improves the prediction accuracy and generalization capabilities of pre-trained language models in complex contexts, and solves the problem that static weights cannot adapt to the input context.
Smart Images

Figure CN120087412A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of artificial intelligence, and in particular, to a model training method and device based on a re-attention mechanism. Background Art
[0002] With the development of natural language processing (NLP, short for Natural Language Processing) technology, pre-trained language models such as BERT (short for Bidirectional Encoder Representations from Transformers), GPT (short for Generative Pre-trained Transformer), etc., which are based on the Transformer architecture, have achieved remarkable results in various tasks.
[0003] In the traditional pre-trained language model architecture, the multi-head attention (MHA) mechanism is usually adopted. Each attention head in the model independently focuses on different aspects of the input sequence, and linearly combines the outputs of all attention heads through a fixed weight matrix. However, the adopted multi-head attention mechanism has certain limitations. Especially when dealing with texts with complex context changes, it is difficult for the model to adaptively focus on the key attention heads according to the input context. For example, when the sequence corresponding to the sentence "The bank can hold investments or river water" is input into the pre-trained language model, for the token "bank", it is difficult for the model to adaptively focus on the key attention heads "geographical entity" or "financial entity" according to the context of the input sequence, thus affecting the prediction accuracy of the model.
[0004] Based on this, there is an urgent need for a high-precision model training method to achieve adaptive focusing on key attention heads according to the input context, so as to improve the prediction accuracy of the pre-trained language model. Summary of the Invention
[0005] Based on the above problems, the present application provides a model training method and device based on a re-attention mechanism, aiming to solve the technical problem that the existing pre-trained language model has a low model prediction accuracy because it is difficult to adaptively focus on key attention heads according to the input context, so as to improve the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0006] The embodiments of the present application disclose the following technical solutions:
[0007] In the first aspect of the present application, a model training method based on a re-attention mechanism is provided, which is applied to a pre-trained language model. The pre-trained language model includes a multi-head attention layer and a sparse gating network. The method includes:
[0008] Input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features. The embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text.
[0009] The sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors, and constructs an attention weight vector based on the dynamic weights. The sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads.
[0010] The multi-head attention layer performs weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors to obtain a fused context representation, and performs a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text.
[0011] Adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model.
[0012] In an optional implementation manner, the step of the multi-head attention layer extracting global semantic features from the embedding vector sequences and generating global context vectors based on the global semantic features includes:
[0013] Input each of the embedding vector sequences into the multi-head attention heads respectively to obtain context vectors output by each attention head, and generate the global semantic features based on the context vectors output by each attention head. The attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequences.
[0014] Input the global semantic features into an adaptive pooling layer. The adaptive pooling layer performs pooling processing on the global semantic features based on a low-rank projection matrix to generate global context vectors corresponding to the embedding vector sequences. The low-rank projection matrix is used to map the global semantic features from a high-dimensional feature space to a low-dimensional feature space.
[0015] In an alternative implementation, the formula for calculating the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors is as follows:
[0016]
[0017] where g represents the dynamic weight, c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix for mapping the global context vector to the gating score space; b g represents the gating bias vector.
[0018] In an alternative implementation, the preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of the non-critical attention heads.
[0019] In an alternative implementation, the process of adjusting the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence through the preset loss function to obtain the target pre-trained language model includes:
[0020] Calculating a first loss value based on the self-attention values corresponding to each embedding vector sequence through the language model loss function;
[0021] Calculating a second loss value based on the attention weight vectors corresponding to each embedding vector sequence through the weight sparsity regularization term;
[0022] Determining a total loss value based on each of the first loss values and the corresponding second loss values;
[0023] Performing joint model training on the pre-trained language model and the sparse gating network based on each of the total loss values until the model training stops when the parameters meet the training termination condition, and obtaining the target pre-trained language model.
[0024] In an alternative implementation, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0025] In the second aspect of the present application, a prediction method based on a re-attention mechanism is provided, which is applied to a target pre-trained language model. The target pre-trained language model is a model trained by the model training method based on the re-attention mechanism introduced in any one of the implementation manners in the first aspect of the embodiments of the present application. The target pre-trained language model includes a multi-head attention layer and a sparse gating network. The method includes:
[0026] Perform embedding vector encoding on the target natural language text input by the user to obtain a target embedding vector sequence; the target embedding vector sequence is composed of hidden state vectors corresponding to each token in the target natural language text;
[0027] Input the target embedding vector sequence into the multi-head attention layer to obtain a target global context vector output by the multi-head attention layer;
[0028] Input the target global context vector into the sparse gating network to obtain a target attention weight vector output by the sparse gating network; the target attention weight vector is composed of target dynamic weights corresponding to each multi-head attention head in the multi-head attention layer;
[0029] The multi-head attention layer performs weighted fusion of the target context vectors output by each attention head and the target attention weight vector to obtain a target fused context representation;
[0030] The multi-head attention layer performs a linear transformation on the target fused context representation based on a fixed weight matrix to obtain a target self-attention value; the target self-attention value is used to predict the next token corresponding to the target natural language text.
[0031] In the third aspect of the present application, a model training device based on a re-attention mechanism is provided, which is applied to a pre-trained language model. The pre-trained language model includes a multi-head attention layer and a sparse gating network. The device includes:
[0032] A feature extraction module, configured to input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features; the embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text;
[0033] A dynamic weight calculation module, configured to calculate, by the sparse gating network based on the sparse Softmax function and each of the global context vectors, the respective dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer, and construct an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of the non-critical attention heads among the multi-head attention heads;
[0034] A weighted fusion module, configured to perform weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation; and perform a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text;
[0035] A parameter adjustment module, configured to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model.
[0036] In a third aspect of the present application, there is provided a computer-readable storage medium storing a computer program, which when run by a processor, implements the above-mentioned model training method based on the re-attention mechanism.
[0037] In a fourth aspect of the present application, there is provided a processor for running a computer program, and when the computer program runs, it executes the above-mentioned model training method based on the re-attention mechanism.
[0038] Compared with the prior art, the present application has the following beneficial effects:
[0039] In the technical solution of the present application, first, multiple embedding vector sequences for training the pre-trained language model are respectively input into the multi-head attention layer, and the multi-head attention layer can accurately extract global semantic features from the embedding vector sequences, so that the global context vectors generated based on the accurate global semantic features can accurately capture the overall semantics and context of the entire input text (i.e., the natural language text); secondly, the sparse gating network calculates the respective dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector, and constructs an attention weight vector based on the dynamic weights. Since the sparse Softmax function is used to suppress the weights of the non-critical attention heads among the multi-head attention heads, the dynamic allocation of the weights of each attention head is realized by the sparse gating mechanism to sense the input semantic features in real time, which can reduce redundant calculations while accurately focusing on the key attention heads;
[0040] Then, the context vectors output by each attention head are weighted and fused with the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation, and a linear transformation is performed on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text, avoiding the problem that static weights cannot adapt to the input context; furthermore, the parameters of the pre-trained language model are adjusted according to the self-attention values corresponding to each embedding vector sequence by a preset loss function to obtain a target pre-trained language model, solving the technical problem that the existing pre-trained language models have low prediction accuracy because it is difficult to adaptively focus on key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 It is a flowchart of a model training method based on a re-attention mechanism provided by an embodiment of the present application; Figure 2 It is a flowchart of a generation process of a global context vector provided by an embodiment of the present application; Figure 3 It is a flowchart of a parameter adjustment process of a pre-trained language model provided by an embodiment of the present application; Figure 4 It is a flowchart of a prediction method based on a re-attention mechanism provided by an embodiment of the present application; Figure 5 It is a schematic structural diagram of a model training device based on a re-attention mechanism provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0043] As described above, currently, in traditional pre-trained language model architectures, the multi-head attention (MHA) mechanism is usually adopted. Each attention head in the model independently focuses on different aspects of the input sequence and linearly combines the outputs of all attention heads through a fixed weight matrix. However, the adopted multi-head attention mechanism has certain limitations. Especially when dealing with texts with complex context changes, it is difficult for the model to adaptively focus on key attention heads according to the input context. For example, when the sequence corresponding to the sentence "The bank can hold investments or riverwater" is input into the pre-trained language model, for the token "bank", it is difficult for the model to adaptively focus on the key attention heads "geographical entity" or "financial entity" according to the context of the input sequence, thus affecting the prediction accuracy of the model. Based on this, there is an urgent need for a high-precision model training method to achieve adaptive focusing on key attention heads according to the input context, thereby improving the prediction accuracy of the pre-trained language model.
[0044] The inventors have proposed a model training method based on the re-attention mechanism through research. In this scheme, first, multiple embedded vector sequences used to train the pre-trained language model are respectively input into the multi-head attention layer. The multi-head attention layer can accurately extract global semantic features from the embedded vector sequences, so that the global context vectors generated based on the accurate global semantic features can accurately capture the overall semantics and context of the entire input text (i.e., natural language text). Secondly, the sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector, and constructs an attention weight vector based on the dynamic weights. Since the sparse Softmax function is used to suppress the weights of non-key attention heads among the multi-head attention heads, the dynamic allocation of the weights of each attention head can be realized by the sparse gating mechanism to sense the input semantic features in real time, which can accurately focus on the key attention heads while reducing redundant calculations.
[0045] Then, the context vectors output by each attention head are weighted and fused with the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation, and a linear transformation is performed on the fused context representation based on a fixed weight matrix to obtain the self-attention value for predicting the next token corresponding to the natural language text, avoiding the problem that static weights cannot adapt to the input context; furthermore, the parameters of the pre-trained language model are adjusted according to the self-attention values corresponding to each embedding vector sequence by a preset loss function to obtain the target pre-trained language model, solving the technical problem that the existing pre-trained language models have low prediction accuracy because it is difficult to adaptively focus on key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0046] It should be noted that the application scenarios of a model training method and device based on the re-attention mechanism provided in this application include but are not limited to natural language processing (such as machine translation, sentiment analysis), multi-modal models (such as image-text understanding), dialogue systems, etc.
[0047] In order to enable those skilled in the art to better understand the solution of this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of this application.
[0048] Keyword definition:
[0049] token: In the context of Natural Language Processing (NLP), "token" refers to the smallest analysis unit in text data, and this unit can be a word, a punctuation mark, a number, etc.
[0050] Method Embodiment 1
[0051] An embodiment of a model training method based on the re-attention mechanism is provided in an embodiment of this application. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0052] See Figure 1, which is a flow chart of a model training method based on a re-attention mechanism provided in an embodiment of the present application, which is applied to a pre-trained language model (such as a Transformer), wherein the pre-trained language model includes a multi-head attention layer and a sparse gating network. Figure 1 As shown, the method comprises the following steps:
[0053] In step S101, multiple embedding vector sequences used to train the pre-trained language model are respectively input into the multi-head attention layer, and the multi-head attention layer extracts global semantic features from the embedding vector sequences and generates a global context vector based on the global semantic features.
[0054] In step S101, the embedding vector sequence is composed of hidden state vectors corresponding to each token in the natural language text. In an embodiment of the present application, the embedding vector sequence is a sequence formed by converting each word or symbol in the text into a corresponding vector representation by the pre-trained language model through embedding vector encoding of the input natural language text. For example, the input natural language text is "Where did the apple fly to?", and the tokens corresponding to the text are "this", "one", "apple", "fruit", "fly", "to", "where", "in", "go", "have" and "?". After the above text is embedded vector encoded by the pre-trained language model, the hidden state vectors corresponding to the above tokens can be respectively, and the embedding vector sequence can be constructed based on the above hidden state vectors.
[0055] In the embodiment of the present application, the embedding vector sequence H can be expressed as:
[0056]
[0057] Where L represents the length of the embedding vector sequence; represents the set of real numbers; d model Represents the dimension of the embedding vector sequence.
[0058] In the embodiment of the present application, after the embedding vector sequence is input into the multi-head attention layer, the pre-trained language model can obtain accurate global semantic features through the context vector output by the multi-head attention head, and pool the global semantic features based on the low-rank projection matrix to obtain a global context vector that can accurately capture the overall semantics and context of the entire input text. Figure 2 , Figure 2 A flowchart of a process for generating a global context vector provided in an embodiment of the present application, the process comprising the following steps:
[0059] Step S1021: Input each embedding vector sequence into the multi-head attention heads respectively to obtain the context vectors output by each attention head, and generate global semantic features based on the context vectors output by each attention head.
[0060] In step S1021, the attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequence.
[0061] Step S1022: Input the global semantic features into the adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic features based on the low-rank projection matrix to generate global context vectors corresponding to the embedding vector sequence.
[0062] In step S1022, the low-rank projection matrix is used to map the global semantic features from the high-dimensional feature space to the low-dimensional feature space.
[0063] In the embodiment of the present application, the formula for the adaptive pooling layer to perform pooling processing on the global semantic features based on the low-rank projection matrix can be shown as formula (1):
[0064] (1)
[0065] Where, c represents the global context vector; h i represents the feature representation of the i-th element of the input sequence (i.e., the global semantic feature); α i represents the attention weight of the i-th element, which is used to measure the importance of this element to the global context; W a represents the low-rank projection matrix; U represents the output dimension of the low-rank projection matrix (i.e., the compressed feature dimension); V represents the input dimension of the low-rank projection matrix.
[0066] Optionally, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where, r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0067] It should be noted that by introducing the low-rank projection matrix to perform pooling processing on the global semantic features and constraining the range of the rank of the low-rank projection matrix to r ≤ 0.2×d model , it is possible to map the global semantic features from the high-dimensional feature space to the low-dimensional feature space, thereby achieving the reduction of the number of parameters (the parameter compression rate can reach 80%) while enhancing the generalization ability of the pre-trained language model.
[0068] Step S102: The sparse gating network calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector, and constructs an attention weight vector based on the dynamic weights.
[0069] In step S102, the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads. Among them, the non-critical attention heads are attention heads with dynamic weights less than a preset threshold.
[0070] In the embodiment of the present application, based on the sparse Softmax function and each global context vector, the formula for calculating the respective dynamic weights of the multi-head attention heads in the multi-head attention layer is:
[0071]
[0072] where g represents the dynamic weight, , N is the number of attention heads; c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector. Among them, the calculation rule of the sparse Softmax function is as follows:
[0073]
[0074] where x i , x j are the i-th and j-th elements of the input vector (i.e., the global context vector) respectively, and are used to represent the unnormalized gating scores.
[0075] In the embodiment of the present application, the pre-trained language model calculates the respective dynamic weights of the multi-head attention heads in the multi-head attention layer through a sparse gating network based on the sparse Softmax function and each global context vector, and can suppress the weights of non-critical attention heads (such as attention heads with dynamic weights less than a preset threshold) in the multi-head attention heads, achieving precise focus on critical attention heads while reducing redundant calculations.
[0076] Optionally, the sparsity of the sparse Softmax function can be adjusted according to actual needs. For example, the sparsity can be adjusted to 30% according to actual needs.
[0077] In step S103, the multi-head attention layer performs weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors to obtain a fused context representation; and performs a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text.
[0078] In the embodiments of the present application, the multi-head attention layer of the pre-trained language model weights and fuses the context vectors output by each attention head with the corresponding attention weight vectors through formula (2) to obtain the fused context representation Fused_Output.
[0079] (2)
[0080] Where, h i represents the context vector output by the i-th attention head; g i is the dynamic weight corresponding to the i-th attention head in the corresponding attention weight vector; N is the number of attention heads. The multi-head attention layer can also linearly transform the fused context representation based on a fixed weight matrix through formula (3) to obtain the self-attention value Final_Output for predicting the next token corresponding to the natural language text.
[0081] (3)
[0082] Where, W 0 represents the fixed weight matrix corresponding to the multi-head attention layer, which is an immutable constant and will not be updated along with the parameters of the model; d out represents the dimension of the fused context representation.
[0083] In the embodiments of the present application, taking the input natural language text "Apple released a new product. Where did it fly?" as an example, if it is determined through the output self-attention value that the high weight is on "Apple" (company name), then the pre-trained language model can accurately predict that the next token "where" may refer to the product release scope; if it is determined through the output self-attention value that the high weight is on "fly" (action), then the pre-trained language model can accurately predict that the next token "where" may refer to the physical movement trajectory.
[0084] Step S104, adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence through a preset loss function to obtain the target pre-trained language model.
[0085] In the embodiments of the present application, the preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of non-critical attention heads. Among them, the preset loss function L total can be specifically as follows:
[0086]
[0087] Where, L MLMis the loss function of the language model (e.g., the Cross-Entropy Loss function); is the weight sparsity regularization term, is the attention weight vector; λ is used to control the sparsity intensity. For example, when λ = 0.1, the sparsification effect of the sparse gating network is optimal.
[0088] To improve the model prediction accuracy, in the embodiments of this application, the pre-trained language model can determine the total loss value based on the self-attention values and attention weight vectors corresponding to each embedding vector sequence through a preset loss function, and adjust the parameters of the pre-trained language model based on the total loss value to obtain the target pre-trained language model. Specifically, referring to Figure 3 This figure is a flowchart of the parameter adjustment process of a pre-trained language model provided by the embodiments of this application. This process includes the following steps:
[0089] Step S301, calculate the first loss value based on the self-attention values corresponding to each embedding vector sequence through the language model loss function.
[0090] For example, the pre-trained language model can calculate the first loss value based on the self-attention values corresponding to each embedding vector sequence through the Cross-Entropy Loss function.
[0091] Step S302, calculate the second loss value based on the attention weight vectors corresponding to each embedding vector sequence through the weight sparsity regularization term.
[0092] In the embodiments of this application, the weight sparsity regularization term can force the dynamic weights g of some attention heads (e.g., non-critical attention heads) to approach zero. For example, the initial attention weight vector is: g = [0.3, 0.6, 0.1], and after parameter adjustment: g = [0.2, 0.8, 0.0] (the third attention head is suppressed)
[0093] It should be noted that the role of the weight sparsity regularization term is to make gradually become smaller during the training process of the sparse gating network, so that the model selects the adjusted weights through the backpropagation mechanism during training, thereby improving the prediction accuracy of the model.
[0094] Step S303, determine the total loss value based on each first loss value and the corresponding second loss value.
[0095] Step S304, perform joint model training on the pre-trained language model and the sparse gating network based on each total loss value, and stop the model training until the parameters meet the training cut-off condition to obtain the target pre-trained language model.
[0096] In the embodiments of the present application, when jointly training the pre-trained language model and the sparse gating network based on each total loss value, the parameters of the sparse gating network are gradually updated through the gradient clipping technique. The clipped gradient restricts the step size of parameter update, making the adjustment of dynamic weights smoother, thereby preventing excessive fluctuations in dynamic weight allocation. For example, the dynamic weight may gradually adjust from 0.80.8 to 0.70.7 instead of mutating, which helps the model gradually adapt to different contexts.
[0097] In an alternative embodiment, the present application significantly outperforms the prior art in semantic tasks, especially in polysemy disambiguation (+4.2%) and short text classification (+1.8%). The specific data is shown in Table 1 as follows:
[0098] Table 1
[0099]
[0100] Taking the sentence "The bank can hold investments or river water" as an example, the head activation differences of "bank" under static and dynamic weights can be shown in Table 2, where the head weight of geographical entities increases from 0.3 to 0.8.
[0101] Table 2
[0102]
[0103] Through the model training method based on the re-attention mechanism provided by the embodiments of the present application, the dynamic allocation of weights for each attention head is realized by the sparse gating mechanism to perceive input semantic features in real time, which can accurately focus on key attention heads while reducing redundant calculations; the weights of non-key attention heads in the multi-head attention are suppressed by the sparse Softmax function, achieving accurate focus on key attention heads while reducing redundant calculations; by weighted fusion of the context vectors output by each attention head with the corresponding attention weight vectors, and performing a linear transformation on the fused context representation based on a fixed weight matrix, the self-attention value for predicting the next token corresponding to the natural language text is obtained, avoiding the problem that static weights cannot adapt to the input context; furthermore, by presetting a loss function to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence, a target pre-trained language model is obtained, solving the technical problem that the existing pre-trained language model has low prediction accuracy due to its difficulty in adaptively focusing on key attention heads according to the input context, and significantly improving the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0104] In addition, the pre-trained language model trained by the model training method based on the re-attention mechanism provided in the embodiments of the present application is applicable to the pre-training and fine-tuning stages, can be seamlessly integrated into existing models (such as BERT, GPT), and improves the multi-task compatibility of the pre-trained language model.
[0105] Method Embodiment Two
[0106] The embodiments of the present application provide an embodiment of a prediction method based on the re-attention mechanism. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0107] See Figure 4 , which is a flowchart of a prediction method based on the re-attention mechanism provided by the embodiments of the present application. This method is applied to a target pre-trained language model, and the target pre-trained language model includes a multi-head attention layer and a sparse gating network, where the target pre-trained language model is a model trained by the model training method based on the re-attention mechanism shown in the above Method Embodiment One. As Figure 4 shown, the method includes the following steps:
[0108] Step S401: Perform embedding vector encoding on the target natural language text input by the user to obtain a target embedding vector sequence.
[0109] In step S401, the target embedding vector sequence is composed of the hidden state vectors corresponding to each token in the target natural language text.
[0110] Step S402: Input the target embedding vector sequence into the multi-head attention layer to obtain a target global context vector output by the multi-head attention layer.
[0111] In the embodiments of the present application, the multi-head attention heads in the multi-head attention layer can generate target context vectors based on the target embedding vector sequence, and extract target global semantic features from all the generated target context vectors; then input the target global semantic features into the adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic features based on the low-rank projection matrix to generate a target global context vector. Among them, the low-rank projection matrix is used to map the global semantic features from the high-dimensional feature space to the low-dimensional feature space.
[0112] Optionally, the constraint condition of the low-rank projection matrix is r ≤ 0.2 × d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0113] In step S403, input the target global context vector into the sparse gating network to obtain the target attention weight vector output by the sparse gating network.
[0114] In step S403, the target attention weight vector is composed of the target dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer.
[0115] In the embodiment of the present application, after inputting the target global context vector into the sparse gating network, the sparse gating network can calculate the target dynamic weights corresponding to each attention head based on the sparse Softmax function and the target global context vector, and construct the target attention weight vector based on the target dynamic weights. Among them, the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads.
[0116] In the embodiment of the present application, based on the sparse Softmax function and the target global context vector, the formula for calculating the target dynamic weights corresponding to each attention head is:
[0117]
[0118] where g represents the target dynamic weight, , N is the number of attention heads; c represents the target global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector. Among them, the calculation rule of the sparse Softmax function is as follows:
[0119]
[0120] where x i and x j are the i-th and j-th elements of the input vector (i.e., the target global context vector), respectively, and are used to represent the unnormalized gating scores.
[0121] In step S404, the multi-head attention layer performs weighted fusion on the target context vectors output by each attention head and the target attention weight vector to obtain the target fused context representation.
[0122] In the embodiment of the present application, the multi-head attention layer can perform weighted fusion on the target context vectors output by each attention head and the target attention weight vector through formula (2) in Embodiment 1 of the above method to obtain the target fused context representation.
[0123] In step S405, the multi-head attention layer performs a linear transformation on the target fused context representation based on the fixed weight matrix to obtain the target self-attention value.
[0124] In step S405, the target self-attention value is used to predict the next token corresponding to the target natural language text.
[0125] In the embodiment of the present application, the multi-head attention layer can perform a linear transformation on the fused context representation based on the fixed weight matrix through formula (3) in Embodiment 1 of the above method to obtain the target self-attention value for predicting the next token corresponding to the target natural language text. Among them, the fixed weight matrix is an immutable constant and will not be updated together with the parameters of the model.
[0126] Through the prediction method based on the re-attention mechanism provided by the embodiment of the present application, it is realized to dynamically allocate the weights of each attention head by the sparse gating mechanism to sense the input semantic features in real time, which can accurately focus on the key attention heads while reducing redundant calculations; the weights of non-key attention heads in the multi-head attention heads are suppressed by the sparse Softmax function, which realizes accurately focusing on the key attention heads while reducing redundant calculations, solves the technical problem that the existing pre-trained language model has low prediction accuracy because it is difficult to adaptively focus on the key attention heads according to the input context, and significantly improves the prediction accuracy and generalization ability of the pre-trained language model in complex contexts.
[0127] Device embodiment
[0128] The embodiment of the present application provides a model training device based on the re-attention mechanism, which is applied to a pre-trained language model. The pre-trained language model includes a multi-head attention layer and a sparse gating network. Among them, Figure 5 FIG. is a schematic structural diagram of a model training device based on the re-attention mechanism provided by the embodiment of the present application, as Figure 5 shown, the device includes: a feature extraction module 51, a dynamic weight calculation module 52, a weighted fusion module 53, and a parameter adjustment module 54. The connection relationship between several modules can be seen from Figure 5 this.
[0129] Among them, the feature extraction module 51 is used to input multiple embedding vector sequences for training the pre-trained language model into the multi-head attention layer respectively. The multi-head attention layer extracts global semantic features from the embedding vector sequences and generates global context vectors based on the global semantic features; the embedding vector sequences are composed of hidden state vectors corresponding to each token in the natural language text.
[0130] The dynamic weight calculation module 52 is used to calculate the respective dynamic weights of the multi-head attention heads in the multi-head attention layer by the sparse gating network based on the sparse Softmax function and each global context vector, and construct an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads;
[0131] The weighted fusion module 53 is used to perform weighted fusion on the context vectors output by each attention head and the corresponding attention weight vectors by the multi-head attention layer to obtain a fused context representation; and perform a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text;
[0132] The parameter adjustment module 54 is used to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each embedding vector sequence through a preset loss function to obtain a target pre-trained language model.
[0133] Optionally, the feature extraction module further includes: a generation unit and a pooling processing unit.
[0134] Among them, the generation unit is used to input each embedding vector sequence into the multi-head attention heads respectively to obtain the context vectors output by each attention head, and generate global semantic features based on the context vectors output by each attention head; the attention heads are used to capture the similarity relationships between different tokens in the embedding vector sequence;
[0135] The pooling processing unit is used to input the global semantic features into the adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic features based on a low-rank projection matrix to generate a global context vector corresponding to the embedding vector sequence; the low-rank projection matrix is used to map the global semantic features from a high-dimensional feature space to a low-dimensional feature space.
[0136] Optionally, the formula for calculating the respective dynamic weights of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each global context vector is:
[0137]
[0138] Among them, g represents the dynamic weight, c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents the gating weight matrix, which is used to map the global context vector to the gating score space; b g represents the gating bias vector.
[0139] Optionally, the preset loss function includes a language model loss function and a weight sparsity regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparsity regularization term is used to sparsify the weights of non-critical attention heads.
[0140] Optionally, the parameter adjustment module includes: a first calculation unit, a second calculation unit, a determination unit, and a joint training unit.
[0141] The first calculation unit is configured to calculate a first loss value based on the self-attention values corresponding to each embedding vector sequence through the language model loss function.
[0142] The second calculation unit is configured to calculate a second loss value based on the attention weight vectors corresponding to each embedding vector sequence through the weight sparsity regularization term.
[0143] The determination unit is configured to determine a total loss value based on each first loss value and the corresponding second loss value.
[0144] The joint training unit is configured to perform joint model training on the pre-trained language model and the sparse gating network based on each total loss value, and stop the model training until the parameters meet the training termination condition, so as to obtain a target pre-trained language model.
[0145] Optionally, the constraint condition of the low-rank projection matrix is r ≤ 0.2×d model , where r represents the rank of the low-rank projection matrix, and d model represents the dimension of the embedding vector sequence.
[0146] Storage medium embodiment
[0147] An embodiment of the present application provides a computer-readable storage medium, on which a program is stored. When the program is executed by a processor, it implements some or all of the steps of the model training method based on the re-attention mechanism introduced in the foregoing method embodiments of the present application. The storage medium can be various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc.
[0148] Processor embodiment
[0149] An embodiment of the present application provides a processor for running a program. When the program is running, it executes some or all of the steps of the model training method based on the re-attention mechanism introduced in the foregoing method embodiments.
[0150] It should be noted that the various embodiments in this specification are described in a progressive manner. For the identical or similar parts among the embodiments, reference can be made to each other, and the key point of each embodiment is to illustrate the differences from other embodiments. In particular, for the system embodiments, since they are basically similar to the method embodiments, the description is relatively simple, and reference can be made to the relevant parts of the method embodiments for the relevant content. The system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.
[0151] The above is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed in the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.
Claims
1. A model training method based on re-attention mechanism, characterized in that: Applied to a pre-trained language model, the pre-trained language model includes a multi-head attention layer and a sparse gating network, the method includes: Inputting multiple embedding vector sequences used to train the pre-trained language model into the multi-head attention layer respectively, extracting global semantic features from the embedding vector sequences by the multi-head attention layer, and generating a global context vector based on the global semantic features; the embedding vector sequence is composed of hidden state vectors corresponding to each token in the natural language text; The sparse gating network calculates the dynamic weights corresponding to each of the multiple attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors, and constructs an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads; The multi-head attention layer performs weighted fusion on the context vector output by each attention head and the corresponding attention weight vector to obtain a fused context representation; and performs a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text; The parameters of the pre-trained language model are adjusted according to the self-attention values corresponding to each of the embedded vector sequences through a preset loss function to obtain a target pre-trained language model.
2. The method according to claim 1, characterized in that: The extracting global semantic features from the embedding vector sequence by the multi-head attention layer and generating a global context vector based on the global semantic features includes: Input each of the embedding vector sequences into the multi-head attention head, obtain the context vector output by each attention head, and generate the global semantic feature based on the context vector output by each attention head; the attention head is used to capture the similarity relationship between different tokens in the embedding vector sequence; The global semantic features are input into an adaptive pooling layer, and the adaptive pooling layer performs pooling processing on the global semantic features based on a low-rank projection matrix to generate a global context vector corresponding to the embedded vector sequence; the low-rank projection matrix is used to map the global semantic features from a high-dimensional feature space to a low-dimensional feature space.
3. The method according to claim 1, characterized in that The formula for calculating the dynamic weights corresponding to each of the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors is: ; Wherein, g represents the dynamic weight, c represents the global context vector; SparseSoftmax() represents the sparse Softmax function; W g represents a gating weight matrix, which is used to map the global context vector to the gating score space; b g Represents the gate bias vector.
4. The method according to claim 1, characterized in that The preset loss function includes a language model loss function and a weight sparse regularization term. The language model loss function is used to measure the difference between the predicted value and the true value of the pre-trained language model, and the weight sparse regularization term is used to perform sparse processing on the weights of the non-critical attention heads.
5. The method according to claim 4, characterized in that The method of adjusting the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedding vector sequences through a preset loss function to obtain a target pre-trained language model includes: Calculating a first loss value based on the self-attention values corresponding to each of the embedding vector sequences by using the language model loss function; Calculating a second loss value based on the attention weight vectors corresponding to each of the embedding vector sequences through the weight sparse regularization term; determining a total loss value based on each of the first loss values and the corresponding second loss values; Based on each of the total loss values, the pre-trained language model and the sparse gating network are jointly trained until the model training is stopped when the parameters meet the training cutoff condition, so as to obtain the target pre-trained language model.
6. The method according to claim 2, characterized in that The constraint condition of the low-rank projection matrix is r≤0.2×d model , where r represents the rank of the low-rank projection matrix, d model represents the dimension of the embedding vector sequence.
7. A prediction method based on re-attention mechanism, characterized in that: Applied to a target pre-trained language model, the target pre-trained language model is a model trained by any one of the methods of claims 1 to 6 above, the target pre-trained language model includes a multi-head attention layer and a sparse gating network, and the method includes: Performing embedding vector encoding on the target natural language text input by the user to obtain a target embedding vector sequence; the target embedding vector sequence is composed of hidden state vectors corresponding to each token in the target natural language text; Inputting the target embedding vector sequence into the multi-head attention layer to obtain the target global context vector output by the multi-head attention layer; Inputting the target global context vector into the sparse gating network to obtain a target attention weight vector output by the sparse gating network; the target attention weight vector is composed of target dynamic weights corresponding to each of the multiple attention heads in the multi-head attention layer; The multi-head attention layer performs weighted fusion on the target context vector output by each attention head and the target attention weight vector to obtain a target fused context representation; The multi-head attention layer performs a linear transformation on the target fusion context representation based on a fixed weight matrix to obtain a target self-attention value; the target self-attention value is used to predict the next token corresponding to the target natural language text.
8. A model training device based on re-attention mechanism, characterized in that: Applied to a pre-trained language model, the pre-trained language model includes a multi-head attention layer and a sparse gating network, and the device includes: A feature extraction module, used to input multiple embedding vector sequences used to train the pre-trained language model into the multi-head attention layer, respectively, and the multi-head attention layer extracts global semantic features from the embedding vector sequences, and generates a global context vector based on the global semantic features; the embedding vector sequence is composed of hidden state vectors corresponding to each token in the natural language text; A dynamic weight calculation module, configured to calculate the dynamic weights corresponding to the multi-head attention heads in the multi-head attention layer based on the sparse Softmax function and each of the global context vectors by the sparse gating network, and construct an attention weight vector based on the dynamic weights; the sparse Softmax function is used to suppress the weights of non-critical attention heads in the multi-head attention heads; A weighted fusion module is used to perform weighted fusion of the context vector output by each attention head and the corresponding attention weight vector by the multi-head attention layer to obtain a fused context representation; and to perform a linear transformation on the fused context representation based on a fixed weight matrix to obtain a self-attention value for predicting the next token corresponding to the natural language text; The parameter adjustment module is used to adjust the parameters of the pre-trained language model according to the self-attention values corresponding to each of the embedded vector sequences through a preset loss function to obtain a target pre-trained language model.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the model training method based on the re-attention mechanism as described in any one of claims 1 to 6 is implemented.
10. A processor, characterized in that: Used to run a computer program, which, when running, executes the model training method based on the re-attention mechanism as described in any one of claims 1-6.
Citation Information
Patent Citations
Self-adaptive relation modeling method for structured data
CN113191441A
UNILM abstract generation method based on improvement
CN114691858A
Chinese multi-label classification method fused with named entity recognition
CN118312612A
Large language model training method and device
CN118445379A
Fine-grained image classification method based on feature fusion and semantic enhancement
CN118799646A
Cited By
Word sequence prediction method for large-model mixing precision quantification driven by multiple weight saliency
CN120597871A
A large-scale word sequence prediction method driven by multiple weights and saliency with mixed precision quantization
CN120597871B
Intelligent weight prediction method for dynamic scene
CN120805990A
Transverse mixed attention mechanism model training method, medium, device and program product
CN121031665A