Text generation method, text generation device, electronic equipment and storage medium
By dynamically combining attention heads and optimizing attention score calculations, the information loss, repetition generation and dead cycle problems of traditional multi-head attention mechanisms during text generation is solved, which significantly improves the text generation performance and diversity of the model.
Patent Information
- Application Number
- CN202411448797.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-16
- Publication Date
- 2025-06-13
AI Technical Summary
The traditional multi-head attention mechanism has problems such as information loss, repeated generation and dead loop when text generation, resulting in insufficient expression ability and poor generation effect of the model.
By generating dynamic input parameters, dynamically combining attention heads, the interaction between heads is realized, and the calculation of attention scores is optimized, thereby improving the representation ability of text.
It significantly improves the text generation performance of the model, reduces information redundancy and duplicate generation, reduces model deviation, and enhances the diversity and accuracy of text generation.
Smart Images

Figure CN120144731A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technologies, and in particular, to a text generation method, a text generation device, an electronic device, and a storage medium. Background Art
[0002] In the fields of natural language processing (NLP) and machine learning, the Transformer model has become the mainstream architecture for many tasks, including but not limited to machine translation, text summarization, question answering systems, etc., for text processing and generation. The core component of the Transformer model is the multi-head attention mechanism (MHA), which allows the model to simultaneously focus on information in different representation subspaces, thereby capturing long-range dependencies in sequential data.
[0003] The traditional MHA extends the attention mechanism to multiple heads, thereby increasing the model's attention to different features. Since the individual heads of the traditional MHA work independently, there are some deficiencies in text generation as follows:
[0004] (1) When inputting the text "Shanghai is the economic center of our country and is located in the _____ of our country.", the traditional MHA method will wrongly focus on "center" and thus generate "center". Then a sentence like "Shanghai is the economic center of our country and is located in the center of our country" will be obtained. This is mainly likely to cause the problem of a low-rank bottleneck in the attention score matrix, which is manifested as information loss of certain words in the text. Since "center" appears in the previous text, the subsequent generated text focuses on the previous text. If the model's expression ability is insufficient, it will largely lead to copying the previous text, resulting in insufficient accuracy of text generation.
[0005] (2) The traditional MHA method will overly focus on some words in the previous text, greatly increasing the probability of generating the previous words and easily resulting in repetition. For example, when the input text is: "Geese, geese, geese, curving their necks towards the sky. White feathers floating on green water, red palms paddling clear waves.", it is very easy to generate "Geese, geese, geese, geese, geese, geese, geese...". This is because of multiple attention heads, which may learn similar representations, thus greatly wasting the model's representation space for the text, overly focusing on some useless information, and resulting in redundant text information representation.
[0006] (3) Generally, the performance evaluation metrics for a model are loss or perplexity (ppl). The lower the evaluation metric, the better the text generation effect. Therefore, the evaluation metric can be understood as the deviation of the text, which accumulates gradually during the text generation process. Therefore, the existing MHA method will cause a relatively large deviation in the model. When it accumulates to a certain extent, it will lead the model to fall into a dead loop of text generation. For example, the model will continuously repeat generating certain words or sentences.
[0007] Therefore, the traditional MHA has certain limitations for text generation, which greatly restricts the model's ability to express text, thus affecting the text generation effect. Summary of the Invention
[0008] The object of the present invention aims to solve at least one of the above technical problems to a certain extent.
[0009] To this end, the first object of the present invention is to propose a text generation method. By generating dynamic input parameters, this method can dynamically combine attention heads, combine the representations of different heads for the text, thereby realizing the interaction between heads, so as to be able to transform the attention scores according to different input texts, thereby optimizing the text representation ability. This method can be applied to any Transformer model architecture, significantly improving the text generation performance of the model.
[0010] To achieve the above object, the first aspect embodiment of the present invention proposes a text generation method, which includes: obtaining an input text, and fusing the vector linear transformation and position encoding of the input text to generate a position vector carrying position information of the input text, where the position vector includes a position query vector, a position key vector, and a position value vector; generating dynamic input parameters for multi-head attention transformation based on the vector of the input text, preset processing parameters, and activation functions; obtaining a basic attention vector of the input text based on the position query vector and position key vector of the input text and the dimension of the head; obtaining an attention score of the input text based on the dynamic input parameters and the basic attention vector; generating an attention residual vector associated with the target text based on the attention score of the input text, the position value vector of the input text, and preset attention output parameters; generating the target text to be output based on the attention residual vector and a preset vocabulary.
[0011] According to an embodiment of the present invention, the dynamic input parameters include dynamic attention parameters and gated dynamic attention parameters.
[0012] According to an embodiment of the present invention, based on the vector of the input text, preset processing parameters, and activation function, dynamic attention parameters are generated, including: performing dimensionality reduction and non-linear activation on the query vector of the input text based on the preset dimensionality reduction parameters and activation function to obtain a corresponding non-linear activation vector; performing linear transformation and decomposition on the non-linear activation vector based on the preset attention head generation parameters to obtain the dynamic attention parameters.
[0013] According to an embodiment of the present invention, based on the vector of the input text, preset processing parameters, and activation function, gated dynamic attention parameters are generated, including: processing the query vector of the input text based on the preset gating parameters and activation function to obtain a corresponding intermediate gating parameter; decomposing the intermediate gating parameter to obtain the gated dynamic attention parameters.
[0014] According to an embodiment of the present invention, obtaining the attention score of the input text based on the dynamic input parameters and the basic attention vector includes: obtaining the fused attention vector of the input text based on the dynamic input parameters and the basic attention vector; performing normalization processing on the fused attention vector to obtain the preprocessed attention score of the input text; optimizing and adjusting the preprocessed attention score based on the dynamic input parameters to obtain the postprocessed attention score of the input text, and using the postprocessed attention score as the attention score of the input text.
[0015] According to an embodiment of the present invention, obtaining the fused attention vector of the input text based on the dynamic input parameters and the basic attention vector includes: performing linear transformation on the basic attention vector based on the preset head transformation parameters to obtain a corresponding attention transformation vector; fusing the basic attention vector based on the dynamic input parameters to obtain a head-fused attention vector and a gated attention vector; obtaining the fused attention vector of the input text based on the attention transformation vector, the head-fused attention vector, and the gated attention vector.
[0016] According to an embodiment of the present invention, generating an attention residual vector based on the attention score of the input text, the position value vector of the input text, and preset attention output parameters includes: performing feature extraction and merging on the position value vector of the input text based on the attention score to obtain a head-fused attention sentence vector; performing linear transformation on the head-fused attention vector based on the preset attention output parameters to generate the attention residual vector.
[0017] To achieve the above object, an embodiment of the second aspect of the present invention provides a text generation device, which includes: a position vector unit configured to obtain an input text, perform vector linear transformation and position encoding on the input text, and generate a position vector carrying position information of the input text, where the position vector includes a position query vector, a position key vector, and a position value vector; a dynamic parameter unit configured to generate dynamic input parameters for multiple head attention transformations based on the vector of the input text, preset processing parameters, and activation functions; a basic vector unit configured to obtain a basic attention vector of the input text based on the position query vector and position key vector of the input text and the dimension of the head; an attention score unit configured to obtain an attention score of the input text based on the dynamic input parameters and the basic attention vector; an attention residual unit configured to generate an attention residual vector associated with the target text based on the attention score of the input text, the position value vector of the input text, and preset attention output parameters; and a target text unit configured to generate the target text to be output based on the attention residual vector and a preset vocabulary table.
[0018] To achieve the above object, an embodiment of the third aspect of the present invention provides an electronic device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the text generation method according to the embodiment of the first aspect of the present invention is implemented.
[0019] To achieve the above object, a computer-readable storage medium provided by an embodiment of the fourth aspect of the present invention, when the computer program is executed by a processor, implements the text generation method according to the embodiment of the first aspect of the present invention.
[0020] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the practice of the present invention. Description of the Drawings
[0021] The above and / or additional aspects and advantages of the present invention will become apparent and be readily understood from the following description of the embodiments in conjunction with the drawings, where:
[0022] Figure 1 is a flowchart of a text generation method shown according to an exemplary embodiment;
[0023] Figure 2 is a flowchart of a method for generating dynamic attention parameters shown according to an exemplary embodiment;
[0024] Figure 3 is a flowchart of a method for generating gated dynamic attention parameters shown according to an exemplary embodiment;
[0025] Figure 4 It is a flowchart of a method for generating attention scores of input text shown according to an exemplary embodiment;
[0026] Figure 5 It is a flowchart of a method for generating a fused attention vector of input text shown according to an exemplary embodiment;
[0027] Figure 6 It is a flowchart of a method for generating an attention residual vector shown according to an exemplary embodiment;
[0028] Figure 7 It is a schematic diagram of a loss curve shown according to an exemplary embodiment;
[0029] Figure 8 It is a schematic block diagram of a device shown according to an exemplary embodiment; and
[0030] Figure 9 It is a schematic structural diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0031] The embodiments of the present invention will be described in detail below. The examples of the embodiments are shown in the drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the drawings are exemplary and are intended to explain the present invention and should not be construed as limiting the present invention.
[0032] In view of the deficiencies of the existing MHA in text generation, the present invention proposes a text generation method, a text generation method and device, an electronic device and a storage medium to improve the expression ability of text and thus enhance the text generation performance of the model.
[0033] In order to clearly describe the text generation method process of the present invention, it is assumed in the following embodiments that the dimension of the Transformer model is D m , the input length is T, the batch size is B, the number of groups of heads is G, the number of heads in each group is N, the dimension of the head is H (H = Dm / (N * G)), the rank of the dynamically combined head is R (note: R / 2 must be an integer), and the number of parameters of the dynamically combined head is C. In the embodiments of the present invention, C = 4 is mainly taken as an example.
[0034] Figure 1 It is a schematic diagram of a text generation method process shown according to an exemplary embodiment, as Figure 1 shown, the text generation method includes:
[0035] Step S110: Obtain the vector of the input text, and fuse the linear transformation and positional encoding of the vector of the input text to generate a positional vector carrying positional information of the input text, where the positional vector includes a positional query vector, a positional key vector, and a positional value vector.
[0036] Specifically, assume the input text is: "Shanghai is the economic center of our country and is located in the ___ of our country." Then, tokenize the input text. For example, after Tokenize, the input text becomes tokens -> ["Shanghai", "is the", "of", "economic", "center", ",", "located", "in the", "of"], and number the tokens according to the vocabulary to get ids -> [100633, 108659, 9370, 99346, 99488, 3837, 103987, 101055, 9370]. Finally, after inputting the ids into the model, obtain the vectors Q, K, and V of all ids (i.e., the sentence). The vectors Q, K, and V are the vectors of the input text. Among them, the vector Q (Query) represents the query vector, which is used to query information at other positions in the attention mechanism, and its dimension is BxTxD m The vector K (Key) represents the key vector, which is used to be queried in the attention mechanism to obtain relevant information, and its dimension is BxSxD m The vector V (Value) represents the value vector, which contains the actual information content and is used to perform weighted combination according to the matching degree of Q and K in the attention mechanism, and its dimension is BxSxD m .
[0037] In the embodiments of the present invention, based on the transformation parameters of the vector, a linear transformation is performed on the vector, so that it can be mapped to another vector space. The transformation parameters here refer to the weight matrices of Q, K, and V respectively. By multiplying Q, K, and V with their respective weight matrices W q , W k , W v respectively, new vectors with different representations are obtained accordingly. This process can focus on and combine information according to different parts of the input text.
[0038] Considering that the Transformer model itself does not contain a recurrent or convolutional structure and cannot directly capture the sequential information in the input text, the text generation method in the embodiments of the present invention adopts Rotary positional encoding. A unique vector is generated for each position in the form of a rotation matrix, and then this unique vector is added to the new vectors with different representations obtained above to obtain a positional vector carrying positional information, namely the positional query vector Q proj , the positional key vector K proj and the positional value vector V projIn this way, through this process, the positions and mutual relationships of words in the text can be obtained, so as to achieve a deeper understanding of the input text.
[0039] For the above processing process, the following formula can be used for representation:
[0040]
[0041] Step S120: Based on the vector of the input text, the preset processing parameters, and the activation function, generate dynamic input parameters for multiple head attention transformations.
[0042] For example, this step mainly uses the parameter matrix to perform a series of linear transformations, activations, and decompositions on the vector of the input text obtained in step S110 to generate dynamic input parameters for subsequent multiple head attention transformations. For the specific operations of this step, reference can be made to the subsequent embodiments, and no further elaboration will be made here.
[0043] Step S130: Based on the position query vector and position key vector of the input text and the dimension of the head, obtain the basic attention vector of the input text.
[0044] For example, in the attention mechanism of the Transformer model, calculating the attention vector of the text is crucial. Specifically, by performing a matrix multiplication operation on Q proj and K proj obtained in step S110, and using Reshape and combining the dimension of the attention head, finally obtain the basic attention vector of the input text, which is specifically represented by the following formula:
[0045]
[0046] A = Reshape(A′)
[0047] where the dimension of the basic attention vector A is BxGxNxTxS
[0048] Step S140: Based on the dynamic input parameters and the basic attention vector, obtain the attention score of the input text.
[0049] For example, after obtaining the dynamic input parameters and the basic attention vector, the present invention needs to adjust the combined interaction between different heads according to the dynamic input parameters to capture the relationships between various words in the input text, or adjust the weights between different attention heads, so as to obtain a reasonable attention distribution, that is, the attention score. This step can be regarded as the feature fusion of dynamic attention. For the specific implementation content of this step, reference can be made to the subsequent embodiments, and no further elaboration will be made here.
[0050] Step S150: Generate an attention residual vector for outputting the target text based on the attention scores of the input text, the position value vector of the input text, and preset attention output parameters.
[0051] For example, after obtaining the attention scores, in the embodiments of the present invention, by performing feature extraction and merging operations on the position vectors of the text, fusing the features of different attention heads to form a more comprehensive semantic representation, and finally performing a linear transformation to obtain the final output of this attention mechanism, that is, generating an attention residual vector rich in semantic information.
[0052] Step S160: Generate the target text to be output based on the attention residual vector and a preset vocabulary.
[0053] For example, when the attention residual vector is obtained, in the embodiments of the present invention, after processing through MLP, Lm_head, etc., a word probability distribution matrix Prob can be output. Based on this Prob matrix, the id at the maximum probability can be obtained, and then based on the id, the corresponding word can be obtained from the preset vocabulary, thereby finally generating the target text to be output.
[0054] The text generation method proposed by the embodiments of the present invention combines multiple heads of attention dynamically, combines the representations of different heads for the text, thereby realizing the interaction between heads. In addition, the embodiments of the present invention can change the attention scores according to different input texts, thereby improving the text representation ability of the model.
[0055] The text generation method of the present invention will be further described and introduced below with reference to the accompanying drawings.
[0056] In a preferred embodiment, the dynamic input parameters include dynamic attention parameters and gated dynamic attention parameters.
[0057] The generation processes of the dynamic attention parameters and the gated dynamic attention parameters will be further described below.
[0058] In a preferred embodiment, as Figure 2 shown, generate dynamic attention parameters based on the vector of the input text, preset processing parameters, and activation functions, including:
[0059] Step S210: Perform dimensionality reduction and non-linear activation on the query vector of the input text based on preset dimensionality reduction parameters and activation functions to obtain a corresponding non-linear activation vector;
[0060] Step S220: Perform a linear transformation and decomposition on the non-linear activation vector based on preset attention head generation parameters to obtain the dynamic attention parameters.
[0061] For example, first, the query vector Q is dimensionally reduced via the dimensionality reduction parameter dw1, reducing D m to H, and a non-linear activation such as the GELU activation function is used to increase its non-linear ability, obtaining the non-linearly activated vector dw_hidden, which is specifically represented by the following formula:
[0062]
[0063] where the dimension of dw1 is D m xGxCxH, and the dimension of dw_hidden is BxTxGxCxH. This process enables the model to learn more complex feature representations.
[0064] Next, a linear transformation is performed on dw_hidden and it is split through the RmsNorm function and the Split function to obtain the dynamic attention parameters. Specifically, the following formula is used to multiply the attention head parameter qkw by the non-linearly activated vector dw_hidden and perform the decomposition:
[0065]
[0066] w 1 = RmsNorm(w′ 1 )
[0067] where w 1 , w 2 are intermediate decomposition parameters, 2 indicates splitting into 2 parts, and the intermediate decomposition parameters w 1 , w 2 both have dimensions of BxTxGxCx(R / 2)xN, and the dimension of the attention head generation parameter qkw is GxCxHxRxN.
[0068] Furthermore, based on the intermediate decomposition parameters w 1 , w 2 , further decomposition is performed on dimension C to obtain the final dynamic attention parameters, which are represented by the following formula:
[0069] pre_qw1, pre_kw1, post_qw1, post_kw1 = Split(w1, C)
[0070] pre_qw2, pre_kw2, post_qw2, post_kw2 = Split(w2, C)
[0071] Among them, pre_qw1, pre_kw1, pre_qw2, and pre_kw2 are preprocessing dynamic attention parameters, and post_qw1, post_kw1, post_qw2, and post_kw2 are postprocessing dynamic attention parameters. The dimensions of these two sets of dynamic attention parameters are both BxTxGx1x(R / 2)xN.
[0072] In a preferred embodiment, as Figure 3 shown, based on the vector of the input text, preset processing parameters, and activation function, gated dynamic attention parameters are generated, including:
[0073] Step S310: Based on preset gating parameters and activation function, process the query vector of the input text to obtain corresponding intermediate gating parameters;
[0074] Step S320: Decompose the intermediate gating parameters to obtain the gated dynamic attention parameters.
[0075] For example, the dynamic gating parameters mainly play a gating role, that is, retaining useful attention scores and removing useless attention scores.
[0076] Specifically, first multiply the preset gating parameter dd (gating parameter generation matrix) with the query vector Q of the input text, and map the eigenvalue to the range of -1 to 1 through, for example, the Tanh activation function. The main role of the Tanh activation function here is to prevent the value of the generated dynamic gating attention parameter from being too large or too small. The processing process is represented by the following formula:
[0077]
[0078] where qdd is the intermediate gating decomposition parameter.
[0079] Then, decompose the intermediate gating decomposition parameter qdd into C parts in dimension H to obtain C preprocessing / postprocessing dynamic gating attention parameters respectively. The specific formula is as follows:
[0080] pre_qdd, pre_kdd, post_qdd, post_kdd = Split(qdd, C)
[0081] where the dimension of the gating parameter dd is D mxGxH, the dimension of the intermediate gated decomposition parameter qdd is BxTxGxH, and the dimensions of the pre-processing / post-processing dynamic gated attention parameters pre_qdd, pre_kdd, post_qdd, post_kdd are all BxTxGx(H / C). Generally speaking, H / C=N, so it can be considered that the dimensions of *dd are all BxTxGxN.
[0082] The dynamic gated attention parameters obtained above can control the weights in subsequent processing, thereby achieving more flexible feature combination and transformation.
[0083] It should be noted that the dimensions of the parameters dw1, qkw, etc. mentioned above for generating dynamic parameters can be flexibly adjusted.
[0084] In a preferred embodiment, Figure 4 As shown, the attention score of the input text is obtained based on the dynamic input parameter and the basic attention vector, including:
[0085] Step S410, obtaining a fused attention vector of the input text based on the dynamic input parameters and the basic attention vector.
[0086] For example, this step can be regarded as a preprocessing process, which can dynamically adjust the interaction between different heads according to the input data, helping the model to better capture the complex relationships in the data and improve the model's expressiveness. In addition, it can also reduce the redundancy between heads, so that different heads can focus more on their best feature representation, rather than each head having to learn similar information independently.
[0087] More preferably, Figure 5 As shown, the fused attention vector of the embodiment of the present invention is obtained based on the following steps:
[0088] Step S510, linearly transform the basic attention vector based on preset head transformation parameters to obtain a corresponding attention transformation vector.
[0089] Specifically, the basic attention vector is first linearly transformed to prepare for the subsequent fusion between heads, which is specifically expressed by the following formula:
[0090]
[0091] Among them, O b1 is the attention conversion vector, A is the basic attention vector, and W b is the header conversion parameter. Header conversion parameter W b The dimension is NxN, O b1The dimension is BxGxNxTxS. Among them, N represents N attention heads. It is equivalent that each head has an attention score from T -> S. Generally, T = S. Therefore, this vector can also be understood as each head having an attention score between words.
[0092] Step S520, fuse the basic attention vector based on the dynamic input parameter to obtain a head-fused attention vector and a gated attention vector.
[0093] Specifically, the head-fused attention vector in the embodiment of the present invention includes a head-fused attention vector based on the query vector (information flow based on the query vector) and a head-fused attention vector based on the key vector (information flow based on the key vector). Similarly, the gated attention vector includes a gated attention vector based on the query vector (information flow gating based on the query vector) and a gated attention vector based on the key vector (information flow gating based on the key vector). The specific processing process is as follows:
[0094] For example, let the input of the model be W q1 = pre_qw1, W q2 = pre_qw2, W k1 = pre_kw1, W k2 = pre_kw2, W qg = pre_qdd, W kg = pre_kdd, and the attention vector matrix A.
[0095] Information flow based on the query vector (query-wise):
[0096] This process is a process of fusing the basic attention vector based on the parameters related to the query vector in the preprocessed dynamic attention parameters to generate a head-fused attention vector based on the query vector. It has been introduced before that the dimension of the dynamic attention parameter is BxTxGx1x(R / 2)xN, where the R / 2 dimension is the dimension of the rank, which represents the number after head fusion, that is, the scores of N heads of the basic attention vector A are fused into the scores of R / 2 heads. Then, based on the scores of R / 2 heads, they are redistributed to N heads. This is the fusion and recombination of head attention scores.
[0097] Specifically, first, decompose in the dimension of rank R to obtain the i-th dynamic parameter W q1i and W q2i . W q1i represents the head attention fusion parameter with the query-wise serial number i, and W q2i represents the head attention recombination parameter with the query-wise serial number i. Among them, W q1i and W q2iIt is represented by the following formula:
[0098] W q1i = Split(W q1 , R / 2)
[0099] W q2i = Split(W q2 , R / 2)
[0100] Furthermore, based on W q1i , the scores of each head of the attention vector are fused to obtain the intermediate attention vector Q′ of the i-th head fusion b2i :
[0101]
[0102] Furthermore, based on W q2i , the intermediate attention vector Q′ of the head fusion b2i is recombined to obtain the attention vector O of the i-th head fusion b2i :
[0103]
[0104] Finally, the attention vectors of all parameters after rank decomposition are summed to obtain the query head fusion attention vector O b2 :
[0105]
[0106] Among them, the dimension of Q′ b2i is BxGx1xTxS. The dimensions of O b2i and O b2 are both BxGxNxTxS.
[0107] After the heads are fused in this process, the attention score information between the heads is interacted. This can reduce head information redundancy. Thus, it increases the diversity of text generation and reduces repetition.
[0108] Information flow based on key vectors (key-wise):
[0109] This process is the process of fusing the basic attention vector based on the parameters related to the key vector in the preprocessed dynamic attention parameters to generate the head fusion attention vector based on the key vector. The processing of this process is the same as the information flow based on the query vector mentioned above.
[0110] Specifically, it is decomposed in the dimension of rank R to obtain the i-th dynamic parameter W k1i . W k1i and W k2iDenote the head attention fusion parameter in a key-wise manner, \(W\) k2i Denote the head attention recombination parameter in a key-wise manner.
[0111] \(W\) k1i =\(Split(W\) k1 , R / 2)
[0112] \(W\) k2i =\(Split(W\) k2 , R / 2)
[0113] Based on \(W\) k1i , fuse the scores of each head of the attention vector to obtain the intermediate head-fused attention vector \(Q'\) b4i :
[0114]
[0115] Based on \(W\) k2i , recombine the intermediate head-fused attention vector \(Q'\) b4i to obtain the head-fused attention vector \(O\) of the \(i\)-th parameter b4i :
[0116]
[0117] Finally, sum up the head-fused attention vectors of all parameters after rank decomposition to obtain the key head-fused attention vector \(O\) b4 :
[0118]
[0119] Among them, the dimension of \(Q'\) b4i is \(BxGx1xTxS\). The dimensions of \(O\) b4i and \(O\) b4 are both \(BxGxNxTxS\).
[0120] Similarly, after this process fuses the heads, the attention score information between the heads is interacted. This can reduce the head information redundancy. Thus, it increases the diversity of text generation and reduces repetition.
[0121] Query-wise gating based on the query vector
[0122] This process mainly performs feature transformation based on the dynamic gating attention parameter, extracts the base attention vector \(A\) twice to obtain the query gating attention vector, and further extracts important text information from the base attention vector \(A\), thereby improving the accuracy of text generation. It is specifically represented by the following formula:
[0123]
[0124] Among them, O b3 Query gating attention vector, A is the base attention vector, O b3 has a dimension of BxGxNxTxS.
[0125] Key-wise gating of the information flow based on the key vector
[0126] This process is the same as the process of information flow gating based on the query vector, and is specifically represented by the following formula:
[0127]
[0128] Among them, O b5 is the key gating attention vector, A is the base attention vector, O b5 has a dimension of BxGxNxTxS.
[0129] Step S530, based on the attention conversion vector, the head fusion attention vector, and the gating attention vector, obtain the fusion attention vector of the input text.
[0130] Specifically, the fusion attention vector is represented by the following formula:
[0131] A new = O b1 + O b2 + O b3 + O b4 + O b5
[0132] Among them, among them, A new is the fusion attention vector, and its dimension is BxGxNxTxS.
[0133] It should be noted that the operation processes of the above information flow based on the query vector, the information flow based on the key vector, the information flow gating based on the query vector, and the information flow gating based on the key vector can be replaced by other operations, such as convolution, linear transformation, MLP, etc.
[0134] Step S420, normalize the fusion attention vector to obtain the preprocessed attention score of the input text.
[0135] For example, normalize the attention vector to obtain the pre-attention score
[0136] Pre_score = Softmax(A new )
[0137] Among them, Pre_score is the pre-attention score, and its dimension is BxGxNxTxS.
[0138] Step S430: Optimize and adjust the preprocessed attention scores based on the dynamic input parameters to obtain the postprocessed attention scores of the input text.
[0139] For example, this step further adjusts the attention weights after the above normalization stage, so as to use the obtained postprocessed attention scores as the attention scores of the input text. This allows the model to dynamically adjust the weights of different attention heads according to the input data when the attention distribution has been determined. Through postprocessing, the model can more finely control which information in the text is important and how to combine this information, thereby improving the expressive ability of the attention distribution. In addition, by dynamically adjusting the attention weights, the model can better adapt to different tasks and data characteristics and improve its generalization ability on multiple tasks.
[0140] Specifically, similar to Step S410, the processing procedure of this step only has different inputs. Among them, let W q1 = post_qw1, W q2 = post_qw2, W k1 = post_kw1, W k2 = post_kw2, W qg = post_qdd, W kg = post_kdd. The attention vector matrix changes from A to Pre_score obtained by the above normalization.
[0141] Based on these inputs, the postprocessed attention scores can be correspondingly obtained, and their dimension is also BxGxNxTxS. Then the obtained postprocessed attention scores are an attention score that fuses the weights between heads.
[0142] It should be noted that the process of obtaining the attention scores in Steps S410 to S430 above is only one preferred example. The present invention is not limited to this example, and those skilled in the art can adjust the above steps according to actual application requirements. For example, in some examples, one of Step S410 and Step S430 can be removed, or any one of these two steps can be replaced with other operations, such as convolution, linear transformation, Multilayer Perceptron (MLP), etc.
[0143] In a preferred embodiment, as Figure 6 shown, generating an attention residual vector associated with the target text based on the attention scores of the input text, the position vector of the input text, and preset attention output parameters includes:
[0144] Step S610: Based on the attention scores, perform feature extraction and merging on the position vectors of the input text to obtain the head-fused attention sentence vectors;
[0145] Step S620: Based on the preset attention output parameters, perform linear transformation on the head-fused attention vectors to generate the attention residual vectors.
[0146] For example, in this embodiment, the attention scores are equivalent to feature selection scores, which are scores obtained by fusing all attention heads and performing secondary extraction. Therefore, the accuracy of these attention scores is very high, and the extracted features are correspondingly more accurate. Specifically, the head-fused attention intermediate sentence vectors are obtained using the following formula:
[0147]
[0148] where V o ′ is the head-fused attention intermediate sentence vector, whose dimension is BxGxTxNxH, Post_score is the attention score, and V proj is the position value vector.
[0149] Furthermore, based on the obtained head-fused attention intermediate sentence vector V o ′, the head-fused attention sentence vectors are represented using the following formula:
[0150] V o = Merge(V o ′)
[0151] where V o is the head-fused attention sentence vector, whose dimension is BxTx(G*N*H), that is, BxTxDm.
[0152] Subsequently, based on the attention output parameters, perform linear transformation on the head-fused attention sentence vector V o . The final output of the multi-head attention in the present invention - the attention residual vector - is represented using the following formula:
[0153]
[0154] where W P is the attention output parameter, whose dimension is D m xD m , and O attn is the attention residual vector, whose dimension is BxTxD m .
[0155] Taking the example given above, based on the traditional MHA, the final model output is that Shanghai is the economic center of our country and is located in the center of our country. While based on the method of the present invention, the final output is that Shanghai is the economic center of our country and is located in the east of our country.
[0156] In addition, based on the text generation method of the present invention, the present invention trained a Transformer model with a size of 405 megabytes using a dataset and compared it with a model based on the traditional MHA. As Figure 7 shown, where the red line is the traditional MHA and the blue line is the MHA of the present invention. From Figure 4 it can be seen that after adopting the MHA of the present invention, the loss of the model is significantly lower and the error deviation is smaller.
[0157] In summary, the text generation method of the present invention has the following advantages:
[0158] 1) By combining attention heads, the information fusion between heads is increased, and the text expression ability is improved. For example, the previous 8 heads were all independent, and now the 8 heads interact with each other, and the information combination and communication method changes from 8 -> 8 * 8 = 64. Therefore, the present invention enhances the model's ability to process complex relationships, thereby improving the accuracy of model text generation.
[0159] 2) Obtaining dynamic input parameters based on the input text makes the fusion between heads dynamic. That is to say, the fusion between the heads of the model is determined according to the input. This greatly increases the model's representation ability for text and further improves the diversity of the text generated by the model.
[0160] 3) Every time the heads are fused, it is actually a correction of the attention scores. The text generation method of the present invention can make the model more precisely adjust the attention scores through multiple correction operations, reduce the redundancy of scores between heads, and while reducing redundancy, also reduce the deviation of text generation and increase the diversity of the text.
[0161] 4) It can be applied to any type of Transformer model, including models with Transformer's Encoder and Decoder structures, such as BERT, GPT, Llama, etc.
[0162] Correspondingly, the present invention also provides a text generation device, as Figure 8As shown, the multi-head attention device 800 includes: a position vector unit 810, configured to obtain an input text, perform a linear transformation and position encoding on the vector of the input text, and generate a position vector carrying position information of the input text, where the position vector includes a position query vector, a position key vector, and a position value vector; a dynamic parameter unit 820, configured to generate dynamic input parameters for multi-head attention transformation based on the vector of the input text, a preset processing parameter, and an activation function; a basic vector unit 830, configured to obtain a basic attention vector of the input text based on the position query vector and the position key vector of the input text and the dimension of the head; an attention score unit 840, configured to obtain an attention score of the input text based on the dynamic input parameters and the basic attention vector; an attention residual unit 850, configured to generate an attention residual vector associated with a target text based on the attention score of the input text, the position value vector of the input text, and a preset attention output parameter; and a target text unit 860, configured to generate a target text to be output based on the attention residual vector and a preset vocabulary table.
[0163] For the specific implementation content and advantages of the text generation device of the present invention, reference can be made to the embodiments of the above-mentioned text generation method, and details are not described herein again.
[0164] It should be noted that the text generation device of the present invention can be understood as a natural language model for text generation, such as the Transformer model mentioned in the previous example.
[0165] Correspondingly, the present invention also provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method according to any one of the above embodiments.
[0166] Correspondingly, the present invention also provides a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the text generation method according to any one of the above embodiments.
[0167] Correspondingly, an embodiment of the present invention also provides an electronic device, such as Figure 9As shown, it shows a schematic structural diagram of an electronic device suitable for implementing the embodiments of the present invention. The electronic device in the embodiments of the present invention may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), in-vehicle terminals (such as in-vehicle navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 9 The electronic device shown is merely an example and should not impose any limitations on the functions and scope of use of the embodiments of the present invention.
[0168] As Figure 9 shown, the electronic device may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 901, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 902 or the program loaded from the storage device 908 into the random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device are also stored. The processing device 901, the ROM 902, and the RAM 903 are connected to each other through a bus 904. The input / output (I / O) interface 905 is also connected to the bus 904.
[0169] Generally, the following devices may be connected to the I / O interface 905: an input device 906 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 907 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 908 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 909. The communication device 909 may allow the electronic device to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 9 it shows an electronic device having various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had.
[0170] Specifically, according to the embodiments of the present invention, the process described above with reference to the flowchart may be implemented as a computer software program. For example, the embodiments of the present invention provide a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program contains program codes for performing the method shown in the flowchart. In such an embodiment, the computer program may be downloaded and installed from the network through the communication device 909, or installed from the storage device 908, or installed from the ROM 902. When the computer program is executed by the processing device 901, the above functions defined in the method of the embodiments of the present invention are executed.
[0171] It should be noted that the computer-readable medium described above in the present invention can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, the computer-readable storage medium can be any tangible medium that contains or stores a program, and this program can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present invention, the computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium other than the computer-readable storage medium, and this computer-readable signal medium can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.
[0172] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (HyperText Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0173] The above computer-readable medium can be included in the above electronic device; it can also exist separately and not be assembled into the electronic device.
[0174] The above computer-readable medium carries one or more programs, which, when executed by the electronic device, cause the electronic device to: obtain a vector of the input text, and fuse a linear transformation and positional encoding of the vector of the input text to generate a positional vector carrying positional information of the input text; generate dynamic input parameters for multiple head attention transformations based on the vector of the input text, where the dynamic input parameters include dynamic attention parameters and gated dynamic attention parameters; obtain a basic attention vector of the input text based on the positional vector of the input text; obtain an attention score of the input text based on the dynamic input parameters and the basic attention vector; generate an attention residual vector based on the attention score of the input text and the positional vector of the input text; and generate a target text matching the input text based on the attention residual vector.
[0175] Computer program code for performing the operations of the present invention may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may execute entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on the remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., through the Internet using an Internet service provider).
[0176] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present invention. In this regard, each block in the flowchart or block diagram may represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that, in some alternative implementations, the functions marked in the blocks may occur in a different order than marked in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, may be implemented by a dedicated hardware-based system for performing the specified functions or operations, or may be implemented by a combination of dedicated hardware and computer instructions.
[0177] The units involved in the embodiments of the present invention can be implemented in software or in hardware. Among them, the name of the unit does not constitute a limitation to the unit itself in some cases. For example, the first acquisition unit can also be described as "the unit for acquiring at least two Internet protocol addresses".
[0178] The functions described above in this article can be performed, at least in part, by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: Field Programmable Gate Array (FPGA), Application Specific Integrated Circuit (ASIC), Application Specific Standard Product (ASSP), System on Chip (SOC), Complex Programmable Logic Device (CPLD), and so on.
[0179] In the context of the present invention, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media would include electrical connections based on one or more wires, portable computer disks, hard disks, Random Access Memory (RAM), Read Only Memory (ROM), Erasable Programmable Read Only Memory (EPROM or Flash Memory), optical fibers, portable compact disc read only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0180] The above description is only for the preferred embodiments of the present invention and the description of the applied technical principles. Those skilled in the art should understand that the scope of disclosure involved in the present invention is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above disclosure concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the present invention.
[0181] Moreover, although the operations are depicted in a particular order, this should not be construed as requiring that the operations be performed in the particular order shown or in sequential order. In certain environments, multitasking and parallel processing may be advantageous. Similarly, although several specific implementation details are included in the foregoing discussion, these should not be construed as limitations on the scope of the invention. Certain features that are described in the context of separate embodiments may also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment may also be implemented separately or in any suitable sub-combination in multiple embodiments.
[0182] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.
Claims
1. A text generation method, characterized in that: The text generation method comprises: Acquire input text, and perform linear transformation and position encoding fusion on the vector of the input text to generate a position vector carrying position information of the input text, wherein the position vector includes a position query vector, a position key vector, and a position value vector; Based on the vector of the input text and the preset processing parameters and activation function, generating dynamic input parameters for transforming multiple head attentions; Obtaining a basic attention vector of the input text based on the position query vector and the position key vector of the input text and the dimension of the head; Obtaining an attention score of the input text based on the dynamic input parameter and the basic attention vector; Generate an attention residual vector associated with the target text based on the attention score of the input text, the position value vector of the input text and a preset attention output parameter; Based on the attention residual vector and a preset vocabulary, a target text to be output is generated.
2. The text generation method according to claim 1, characterized in that: The dynamic input parameters include dynamic attention parameters and gated dynamic attention parameters.
3. The text generation method according to claim 1, characterized in that: Based on the vector of the input text and the preset processing parameters and activation function, dynamic attention parameters are generated, including: Performing dimension reduction and non-linear activation on the query vector of the input text based on preset dimension reduction parameters and activation functions to obtain a corresponding non-linear activation vector; The nonlinear activation vector is linearly transformed and decomposed based on preset attention head generation parameters to obtain the dynamic attention parameters.
4. The text generation method according to claim 1, characterized in that: Based on the vector of the input text and the preset processing parameters and activation function, a gated dynamic attention parameter is generated, including: Based on preset gating parameters and activation functions, the query vector of the input text is processed to obtain corresponding intermediate gating parameters; The intermediate gating parameters are decomposed to obtain the gated dynamic attention parameters.
5. The text generation method according to claim 1, characterized in that: The obtaining, based on the dynamic input parameter and the basic attention vector, an attention score of the input text comprises: Based on the dynamic input parameter and the basic attention vector, obtaining a fused attention vector of the input text; Normalizing the fused attention vector to obtain a preprocessed attention score of the input text; The pre-processing attention score is optimized and adjusted based on the dynamic input parameters to obtain a post-processing attention score of the input text, so as to use the post-processing attention score as the attention score of the input text.
6. The text generation method according to claim 5, characterized in that: Based on the dynamic input parameter and the basic attention vector, a fused attention vector of the input text is obtained, including: Linearly transform the basic attention vector based on a preset head transformation parameter to obtain a corresponding attention transformation vector; The basic attention vectors are fused based on the dynamic input parameters to obtain a head fused attention vector and a gated attention vector; Based on the attention conversion vector, the head fused attention vector and the gated attention vector, a fused attention vector of the input text is obtained.
7. The text generation method according to claim 1, characterized in that: The step of generating an attention residual vector associated with a target text based on the attention score of the input text, the position value vector of the input text, and a preset attention output parameter comprises: Based on the attention score, feature extraction and merging are performed on the position value vector of the input text to obtain a head fusion attention sentence vector; The head fusion attention vector is linearly transformed based on preset attention output parameters to generate the attention residual vector.
8. A text generation device, characterized in that: The text generating device comprises: A position vector unit, used to obtain input text, and linearly transform and position encode the vector of the input text to generate a position vector carrying position information of the input text, wherein the position vector includes a position query vector, a position reference vector and a position value vector; A dynamic parameter unit, for generating dynamic input parameters for transforming multiple head attentions based on the vector of the input text and preset processing parameters and activation functions; A basic vector unit, configured to obtain a basic attention vector of the input text based on a position query vector and a position key vector of the input text and a dimension of the head; an attention score unit, configured to obtain an attention score of the input text based on the dynamic input parameter and the basic attention vector; An attention residual unit, configured to generate an attention residual vector associated with a target text based on an attention score of the input text, a position value vector of the input text, and a preset attention output parameter; The target text unit is used to generate the target text to be output based on the attention residual vector and a preset vocabulary.
9. An electronic device, comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the text generation method according to any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable the computer to execute the text generation method according to any one of claims 1-7.