An Improved Method for Emotion-Enhanced Dialogue Generation with Transformer

By adding emotional attention modules to the self-attention layer of Transformer, a TEECG model was generated, which solved the problem of low quality of emotional annotation in existing human-computer dialogues, and achieved more emotional human-computer dialogue generation.

CN115048943BActive Publication Date: 2025-07-04NORTHEASTERN UNIV CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202210652231.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-10
Publication Date
2025-07-04
Estimated Expiration
2042-06-10

AI Technical Summary

Technical Problem

The quality of emotional labeling in existing human-computer dialogues is not high, and the machine cannot accurately perceive the user's emotions and generate emotional replies.

Method used

Add emotional attention modules to Transformer's self-attention layer, generate TEECG models, including the Encoder layer and the Decoder layer, and improve Transformer's emotion-enhanced dialogue generation method through multi-head attention layer and feedforward neural network.

Benefits of technology

It improves the emotional investment in human-computer dialogue, and the generated emotional replies are better, and can more accurately capture the semantics and emotions between dialogue sentences, and the generated replies are more emotional.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115048943B_ABST
    Figure CN115048943B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of human-computer dialogue, and particularly to an emotion-enhanced dialogue generation method for improving Transformer. Aiming at the problems of low quality of existing human-computer dialogue emotion annotation, the machine being unable to accurately perceive the user's emotion and unable to generate emotional responses, the following technical solutions are proposed: It includes an improvement on the self-attention layer of the Transformer encoding based on Transformer. The improvement lies in adding an emotion attention module to the original self-attention layer of Transformer to generate a TEECG model. The TEECG model includes an Encoder layer and a Decoder layer. The present invention improves the original self-attention layer of Transformer encoding, capable of capturing the semantics and emotions between dialogue sentences, making the human-computer dialogue more emotional. The decoder corrects the emotional responses in the dialogue, and the generated emotional response effect is better. It is mainly applied to the improvement of sentence emotions in human-computer dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of human-computer dialogue, and particularly to an emotion-enhanced dialogue generation method for improving Transformer. Background Art

[0002] A variety of dialogue robots have been active in people's daily lives, with a wide range of application scenarios, saving a large amount of labor costs and improving work efficiency. Therefore, technology companies at home and abroad have successively formulated their respective development strategies in the direction of human-computer dialogue and released related products. The application scope covers a variety of scenarios from daily life to transportation. Among them, the more common applications include intelligent assistants, intelligent customer service, in-vehicle systems, smart homes, etc. The main purpose is to help complete the tasks specified by users.

[0003] At the same time, human-computer dialogue also has the following challenges: there is a lack of Chinese dialogue corpus and the quality of emotion annotation is not high; in order to improve the emotional investment of human-computer dialogue, the solution aims to improve the emotion-enhanced dialogue generation model of Transformer, add an emotion attention mechanism to the Transformer encoder, input the entire dialogue sentence into the encoder, learn the emotion attention between sentences, and then correct the emotion response in the dialogue through the Transformer decoder, aiming to generate a better emotion response. Summary of the Invention

[0004] The purpose of the present invention is to address the problem in the background art that the existing human-computer dialogue emotion annotation quality is not high, the machine cannot accurately perceive the user's emotion and cannot generate an emotional response, and to propose an emotion-enhanced dialogue generation method for improving Transformer.

[0005] The technical solution of the present invention: an emotion-enhanced dialogue generation method for improving Transformer, including an improvement on the self-attention layer of the Transformer encoding based on Transformer. The improvement lies in adding an emotion attention module to the original self-attention layer of Transformer to generate a TEECG model, and the TEECG model includes an Encoder layer and a Decoder layer; the specific improvements include the following features:

[0006] The input of the TEECG model is a dialogue (U1, U1) pair represented by VADBW word vectors. The dialogue is a Chinese dialogue, and the dialogue pair consists of n + m words, that is, U1 = (W1 + W2 +..., W n ), U2 = (W1 + W2 +..., W m ), and the output of the TEECG model is the probability distribution of generating the i-th word at the i-th moment;

[0007] The Encoder layer is stacked by L identical Encoders. The output of the previous Encoder is used as the input of the next Encoder, and the input of the first Encoder is the dialogue pair (U1, U1) that meets the requirements of the problem description.

[0008] Chinese word segmentation W i Adopt the multi-feature representation method of words VADBW, and the specific formula is: I(W i ) = WE(W i ) + VADE(W i ) + BE(W i ). The word embedding representation of the entire dialogue pair is divided into two parts: E1 (0) , E2 (0) . E1 (0) = [I(W1), I(W2),... I(W n )], E2 (0) = [I(W1), I(W2),... I(W m )];

[0009] The Encoder consists of two sub-layers. One is the multi-headed attention layer (Multi-Headed Attention), and the other is the FFN. There are two multi-headed attention layers, namely the general semantic self-attention layer and the sentiment attention layer. The calculation formula is: E M (l) = MultiHead(E1 (l-1) , E1 (l-1) , E1 (l-1) ),

[0010] The input of the Encoder feed-forward neural network is the output of the multi-headed attention layer and The specific calculation formula is: where E (l) is the connection of the output vectors of this sub-layer, and E (L) is the output of the last sub-layer;

[0011] The Decoder layer is composed of L Decoders stacked. Each Decoder includes three sub-layers. The first sub-layer is a masked multi-headed self-attention layer, and the formula is: M(D) l = MultiHead(D (l-1) , D (l-1) , D (l-1) ), D (0) = R, where R is the output reply. The second sub-layer is composed of encoder-decoder attention, and the calculation formula is: C(E) l= MultiHead(M(D) (l) , E (L) , E (L) ) The third sub - layer is a position - based FNN, and its calculation formula is: D l = FNN(C(E) l ).

[0012] Preferably, the TEECG model further includes a calculation algorithm. The input of the algorithm is: the dialogue pair (U1, U1) represented by the VADBW word vectors = {x1, x2,... x i ,... x n+m}, and the output reply of the algorithm is R = {y1, y2,... y i ,... y Ty}. The specific calculation algorithm is:

[0013] Row1: Begin

[0014] Row2: First, convert each word to an id, then add Start at the beginning position and End identifier at the end of the sentence to form a batch. Each batch includes a complete sentence pair

[0015] Row3: Initialize the model parameters, L = 2

[0016] Row4: Represent U1 as E1 using the formula E1 (0) = [I(W1), I(W2),... I(W n )] and input it into the formula (0) Input to the formula

[0017] E M (l) = MultiHead(E1 (l-1) , E1 (l-1) , E1 (l-1) ) to obtain E M (0)

[0018] Row5: Represent U2 as E2 using the formula E2 (0) = [I(W1), I(W2),... I(W m )] and input it into the formula (0) Input to the formula to obtain

[0019] Row6: Encoder Repeat

[0020] Row7: For l in L

[0021] Row8: Input into the formula EM (l) = MultiHead(E1 (l-1) , E1 (l-1) , E1 (l-1) ) to obtain

[0022] Row9: Input into the formula to obtain

[0023] Row10: Use the formula to obtain E (l)

[0024] Row11: l++

[0025] Row12: End For

[0026] Row13: Represent R as VADBW to obtain D (0)

[0027] Row14: Decoder Repeat

[0028] Row15: For l in L

[0029] Row16: Input D (l-1) into the formula M(D) l = MultiHead(D (l-1) , D (l-1) , D (l-1) ) to obtain M(D) l

[0030] Row17: Input M(D) l and E (L) into the formula C(E) l = MultiHead(M(D) (l) , E (L) , E (L) ) to obtain C(E) (l)

[0031] Row18: Input C(E) (l) into the formula D (l) = FNN(C(E) l ) to obtain D (l)

[0032] Row19: l++

[0033] Row20: End For

[0034] Row21: Input D (L)Perform softmax and log operations to calculate the probability distribution score of the current word: P(a, b) = P(a / b)P(b); a is the word to be selected currently, and b is the word selected at the previous moment; log is expressed as: logP(a, b) = logP(a / b) + logP(b)

[0035] Row22: Record the id and log probability score of the word selected in the current step. Eventually, all the ids form R

[0036] Row23: Input the final R into the BERT sentiment classifier to verify the accuracy of the sentiment generated by the model

[0037] Row24: End.

[0038] Preferably, Row2 is to read all the dialogue pairs in the sentiment dialogue dataset; Row3 is to set the hyperparameters of the training process; Row4 - 12 are the calculation processes of the improved Transformer encoder; Row13 - 20 are the calculation processes of the decoder of this sentiment dialogue generation model; Row21 - 22 are the processes of calculating the probability distribution of the output words; Row23 is the process of verifying the sentiment accuracy of the generated sentiment response

[0039] Preferably, there is also a soft control gate λ during the operation of the Encoder, which is used to obtain the decoding concept. The generation of the response is completed within t time steps during the calculation algorithm operation, and the decoder state at each time step is Obtain a context vector a by combining each output E of the encoder through the attention mechanism (L) , obtaining a context vector a t , a t Connect with and feed it through two linear layers to generate the vocabulary distribution

[0040] Preferably, use λ to obtain the final decoding probability, and the specific calculation formula is: a = Attention(E (L) , E (L) , E (L) );

[0041] where V, V', b, b' are the parameters learned during model training

[0042] σ(.) is the sigmoid function used to generate values between 0 and 1

[0043] W d ,W c(E) and W E are the parameters learned during model training and are the transposes of W d , W c(E) and W E respectively,

[0044] Attention(.) is the attention distribution function, which is used to calculate the attention weights between the decoder and the encoder when generating the sentence R.

[0045] Compared with the prior art, the present invention has the following beneficial technical effects:

[0046] 1. The TEECG model of the present invention improves the self-attention layer of the original Transformer encoding. The original self-attention can only learn the semantic and emotional information of each word in a single sentence, but cannot learn the semantic and emotional information between dialogue sentences. The Encoder of the TEECG model adds emotional attention, which will pay attention to the emotional information between dialogue sentences. The Decoder structure remains unchanged. The output of the TEECG model is the probability distribution of generating words at a certain moment. The Encoder-Decoder attention mechanism is used to rewrite the reply sentence, and more emotional dialogue replies are obtained;

[0047] 2. In summary, the present invention improves the self-attention layer of the original Transformer encoding, has the ability to capture the semantics and emotions between dialogue sentences, makes the human-computer dialogue more emotional, and the decoder corrects the emotional replies in the dialogue, and the generated emotional reply effect is better. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is the framework diagram of the TEECG model in the method for generating emotion-enhanced dialogue by improving Transformer. DETAILED DESCRIPTION OF THE INVENTION

[0049] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and specific embodiments.

[0050] Embodiment

[0051] As Figure 1 shown, the method for generating emotion-enhanced dialogue by improving Transformer proposed by the present invention includes improving the self-attention layer of Transformer encoding on the basis of Transformer. The improvement lies in adding an emotional attention module to the original self-attention layer of Transformer to generate the TEECG model. The TEECG model includes an Encoder layer and a Decoder layer; the specific improvements include the following features:

[0052] The input of the TEECG model is the dialogue (U1, U1) pair represented by VADBW word vectors. The dialogue is a Chinese dialogue, and the dialogue pair consists of n + m words, that is, U1 = (W1 + W2 +..., W n ), U2 = (W1 + W2 +..., W m ). The output of the TEECG model is the probability distribution of generating the i-th word at the i-th moment;

[0053] The Encoder layer is stacked by L identical Encoders. The output of the previous Encoder is used as the input of the next Encoder. The input of the first Encoder is the dialogue pair (U1, U1) that meets the requirements of the problem description;

[0054] Chinese word segmentation W i Adopts the multi-feature representation VADBW method of words. The specific formula is: I(W i ) = WE(W i ) + VADE(W i ) + BE(W i ). The word embedding representation of the entire dialogue pair is divided into two parts: E1 (0) , E2 (0) . E1 (0) = [I(W1), I(W2),... I(W n )], E2 (0) = [I(W1), I(W2),... I(W m )];

[0055] The Encoder consists of two sub-layers. One is the multi-headed attention layer (Multi-Headed Attention), and the other is the FFN. There are two multi-headed attention layers, namely the general semantic self-attention layer and the emotion attention layer. The calculation formula is: E M (l) = MultiHead(E1 (l-1) , E1 (l-1) , E1 (l-1) );

[0056] The input of the Encoder feed-forward neural network is the output of the multi-headed attention layer and The specific calculation formula is: where E (l) is the concatenation of the output vectors of this sub-layer, and E (L) is the output of the last sub-layer;

[0057] The Decoder layer is composed of L stacked Decoders. Each Decoder includes three sub-layers. The first sub-layer is a masked multi-head self-attention layer, with the formula: M(D) l = MultiHead(D (l-1) , D (l-1) , D (l-1) ), D (0) = R, where R is the output response. The second sub-layer is composed of encoder-decoder attention, with the calculation formula: C(E) l = MultiHead(M(D) (l) , E (L) , E (L) ). The third sub-layer is a position-based FNN, with the calculation formula: D l = FNN(C(E) l ).

[0058] The TEECG model also includes a calculation algorithm. The input of the algorithm is: the dialogue pair (U1, U1) represented by the VADBW word vector = {x1, x2,... x i ,... x n+m}), and the output response of the algorithm is R = {y1, y2,... y i ,... y Ty}. The specific calculation algorithm is:

[0059] Row1: Begin

[0060] Row2: First, convert each word to an id, then add Start at the beginning position and End identifier at the end of the sentence to form a batch. Each batch includes a complete sentence pair

[0061] Row3: Initialize the model parameters, L = 2

[0062] Row4: Represent U1 as E1 (0) = [I(W1), I(W2),... I(W n )] and input it into the formula E (0) M (l) (l-1) = MultiHead(E1 (l-1) , E1 (l-1) , E1 M ) to get E M (0)

[0063] Row5: Represent U2 as E2 (0) = [I(W1), I(W2),... I(W m )] and represent it as E2 (0)Input to formula Get

[0064] Row6: Encoder Repeat

[0065] Row7: For l in L

[0066] Row8: Input to formula E M (l) = MultiHead(E1 (l-1) , E1 (l-1) , E1 (l-1) ) to get

[0067] Row9: Input to formula to get

[0068] Row10: Use formula to get E (l)

[0069] Row11: l++

[0070] Row12: End For

[0071] Row13: Represent R as VADBW to get D (0)

[0072] Row14: Decoder Repeat

[0073] Row15: For l in L

[0074] Row16: Input D (l-1) to formula M(D) l = MultiHead(D (l-1) , D (l-1) , D (l-1) ) to get M(D) l

[0075] Row17: Input M(D) l and E (L) to formula C(E) l = MultiHead(M(D) (l) , E (L) , E (L) ) to get C(E) (l)

[0076] Row18: Input C(E) (l) to formula D(l) = FNN(C(E) l ) to obtain D (l)

[0077] Row19: l++

[0078] Row20: End For

[0079] Row21: Apply softmax and log operations to D (L) to calculate the probability distribution score of the current word: P(a, b) = P(a / b)P(b); a is the word to be selected currently, b is the word selected at the previous moment; log is expressed as: logP(a, b) = logP(a / b) + logP(b)

[0080] Row22: Record the id and log probability score of the word selected in the current step. Eventually, all ids form R

[0081] Row23: Input the final R into the BERT sentiment classifier to verify the accuracy of the sentiment generated by the model

[0082] Row24: End

[0083] Row2 is to read all the dialogue pairs in the sentiment dialogue dataset; Row3 is to set the hyperparameters of the training process; Rows 4 - 12 are the calculation process of the improved Transformer encoder; Rows 13 - 20 are the calculation process of the decoder of this sentiment dialogue generation model; Rows 21 - 22 are the process of calculating the probability distribution of the output words; Row23 is the process of verifying the sentiment accuracy of the generated sentiment response

[0084] During the operation of the Encoder, there is also a soft control gate λ for obtaining the decoding concept. The generation of the response is completed within t time steps during the calculation algorithm operation, and the decoder state at each time step is Combine each output E of the Encoder through the attention mechanism (L) to obtain a context vector a t , a t Connect with and feed it through two linear layers to generate the vocabulary distribution Use λ to obtain the final decoding probability, and the specific calculation formula is: a = Attention(E (L) , E (L) , E (L) );

[0085] where V, V', b, b' are the parameters learned during the model training

[0086] σ(.) is the sigmoid function used to generate values between 0 and 1.

[0087] W d ,W c(E) and W E are parameters learned during model training. and are the transposes of W d ,W c(E) and W E respectively.

[0088] Attention(.) is the attention distribution function used to calculate the attention weights between the decoder and the encoder when generating sentence R.

[0089] To verify the effectiveness of the improved Transformer-based sentiment dialogue generation model proposed in this scheme, this scheme conducts comparative experiments with baseline models, specifically including the following baseline models: TEECG + Word2vec: Using Word2vec as the word embedding to input the improved Transformer-based sentiment enhanced dialogue generation model proposed in this paper; TEECG + VADBW: Using VADBW as the word embedding to input the improved Transformer-based sentiment enhanced dialogue generation model proposed in this paper.

[0090] The hardware environment of the experiment is shown in the following table:

[0091]

[0092] The software environment of the experiment is shown in the following table:

[0093]

[0094] Experimental evaluation metrics (automatic evaluation + manual evaluation):

[0095] For automatic evaluation, three metrics, PPL, BLUE, and Distinct, are used, and the emotion accuracy (Accuracy) is used as an indicator of the consistency between the emotion of the generated sentence and the emotion category of the target response.

[0096] PPL (perplexity) is used to estimate whether the response text is relevant to the input sentence in content or whether the grammar conforms to the grammar rules. It estimates the probability of a sentence based on each word and normalizes it by the sentence length. Its calculation formula is: where s represents the output sentence; N represents the length of the output sentence; p(w i ) represents the probability of generating the i-th word in the output sentence; W0 represents the START identifier, marking the starting point of the output sentence.

[0097] BLEU is a commonly used evaluation metric in machine translation tasks. It mainly estimates the quality of a model by comparing the number of n-gram matches between the output sequence and the reference answer. The matching method is relatively rough. The matching process ignores the positions of the matching segments and only focuses on the number of successful matches. The more n-grams that can be matched, the closer the output sequence is considered to be to the reference answer. Its calculation formula

[0098]

[0099] is as follows: where: N represents the maximum value of n-grams after selection

[0100] of n-grams; represents the accuracy of the n-gram phrases in the entire dataset; h(k,r) represents the sum of the number of times the n-gram phrases appear in the reference response sentences; h(k,r i ) represents the number of times the n-gram phrases appear in the reference response sentences; β n represents the weight of each n-gram; BP represents the length penalty factor; c represents the candidate sentence of the expected response; r represents the reference response sentence;

[0101] Distinct is used to judge the diversity of the generated responses and determine whether there are a large number of general and repetitive responses. Its calculation formula is where Count(unqique_ngram) represents the number of non-repeating n-grams in the response; Count(word) represents the total number of words in the response;

[0102] Sentiment accuracy (Accuracy) is used to evaluate whether the sentiment of the responses generated by the dialogue generation model is accurate. In this solution, the sentiment accuracy is used as an evaluation metric for the consistency classification between the expected emotion category and the predicted emotion category in the sentiment generation response. Its calculation formula is: where TP represents the number of sentences that actually belong to a certain sentiment category and are predicted to belong to a certain sentiment category; FP represents the number of sentences that actually do not belong to a certain sentiment category but are predicted to belong to a certain sentiment category; FN represents the number of sentences that actually belong to a certain sentiment category but are predicted not to belong to a certain sentiment category; TN represents the number of sentences that actually do not belong to a certain sentiment category and are predicted not to belong to a certain sentiment category.

[0103] For manual evaluation, three students with a bachelor's degree or above are selected as users to evaluate the quality of the generated responses. Three rating metrics are formulated for manual evaluation, namely:

[0104] "Poor", indicating that the semantics are unclear, the question-and-answer intentions are different, and the response sentences are unfriendly;

[0105] "General" indicates that the semantics is relatively smooth, the Q&A intention is relevant, and the response statement is friendly;

[0106] "Excellent": indicates that the semantics is smooth, the Q&A intention is closely related, and the response statement is friendly.

[0107] The evaluation rule is as follows: 20 common sentences are used as input sentences, which are respectively input into the baseline model and the model proposed in this chapter, and the responses of different models to the same sentence are obtained respectively;

[0108] Three evaluators, through a non-interfering method, respectively conduct manual evaluations on 20 groups of dialogue responses obtained by different models, and then obtain three different evaluation results;

[0109] After that, the evaluation results of the three people are made public. If the three groups of evaluation results are consistent, then this evaluation is used as the final evaluation of this sentence; if the three groups of evaluation results are inconsistent, then the majority rule is followed, and the evaluation result of the majority is used as the final evaluation of this sentence; if all three groups of evaluations are inconsistent, then the three people discuss again, and the result after discussion is used as the final evaluation of this sentence.

[0110] The automatic evaluation results are shown in the following table:

[0111]

[0112] It can be seen from the above table that the Transformer model with the Emotion-Attention mechanism can improve the emotion perception ability of the generation model, but the improvement of Distinct-2 is not significant. Therefore, it can be known that the emotion generation model tends to generate relatively safe sentences;

[0113] In this solution, the emotional responses in the test set are also counted, and the specific results are shown in the following table:

[0114]

[0115] The statistical results in the table show that there are 450 emotional responses in the test set. The emotional responses generated by the TEECG model of the solution have increased by 12.7% compared with the emotional sentences in the test set, and the number of emotional responses has increased from 450 to 540; there are 260 non-emotional responses in the test set. After being processed by the TEECG model of this solution, the proportion of non-emotional sentences has decreased from 36.7% to 24%, specifically from 260 to 170. It can be known from this that the improved TEECG model of this solution is effective in improving the emotional information in the generated responses.

[0116] The above specific embodiments are merely a preferred embodiment of the present invention. Based on the technical solution of the present invention and the relevant revelations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. An improved method for sentiment-enhanced dialogue generation using Transformer, characterized in that, It includes improvements to the self-attention layer encoded by Transformer based on Transformer. The improvement lies in adding an emotion attention module to the original self-attention layer of Transformer to generate the TEECG model, and the TEECG model includes an Encoder layer and a Decoder layer; specific improvements include the following features: The input of the TEECG model is the dialogue pair (U1, U2) represented by VADBW word vectors. The dialogue is in Chinese and the dialogue pair consists of n + m words, that is, U1 = (W1 + W2 +..., W n ), U2 = (W1 + W2 +..., W m ). The output of the TEECG model is the probability distribution of generating the i-th word at the i-th moment; The Encoder layer is stacked by L identical Encoders, and the output of the previous Encoder is used as the input of the next Encoder. The input of the first Encoder is the dialogue pair (U1, U1) that meets the requirements of the problem description. Chinese word segmentation W i Adopt the multi-feature representation VADBW representation method of words. The specific formula is: I(W i ) = WE(W i ) + VADE(W i ) + BE(W i ). The word embedding representation of the entire dialogue pair is divided into two parts: E1 (0) , E2 (0) . E2 (0) = [I(W1), I(W2),... I(W m )]; The Encoder consists of two sub-layers, one is the Multi-Headed Attention layer, and the other is the FFN. There are two Multi-Headed Attention layers, namely the general semantic self-attention layer and the sentiment attention layer respectively. The calculation formula is as follows: The input of the encoder feed-forward neural network is the output of the multi-head attention layer and The specific calculation formula is as follows: where E (l) is the concatenation of the output vectors of this sub-layer, and E (L) is the output of the last sub-layer; The Decoder layer is composed of L Decoder stacks. Each Decoder includes three sub-layers. The first sub-layer is a masked multi-head self-attention layer, with the formula: M(D) l = MultiHead(D (l-1) , D (l-1) , D (l-1) ), D (0) = R, where R is the output response. The second sub-layer is composed of encoder-decoder attention, with the calculation formula: C(E) l = MultiHead(M(D) (l) , E (L) , E (L) ) The third sub-layer is a position-based FNN, with the calculation formula: D l = FNN(C(E) l ).

2. The method for generating an emotion-enhanced dialogue by improving a Transformer according to claim 1, characterized in that, The TEECG model further includes a calculation algorithm. The inputs of the algorithm are: the dialogue pair (U1, U2) represented by the VADBW word vectors = {x1, x2,... x i ,... x n+m}, and the output reply of the algorithm is R = {y1, y2,... y i ,... y Ty}. The specific calculation algorithm is as follows: Row1:Begin Row2: First, convert each word into an id, then add Start at the beginning position and End identifier at the end of the sentence to form a batch, and each batch includes a complete sentence pair. Row3: Initialize the model parameters, L = 2 Row 4: Express U1 using the formula E1 (0) = [I(W1), I(W2),... I(W n )] as E1 (0) Input into the formula to obtain E M (1) Row 5: Express U2 using the formula E2 (0) = [I(W1), I(W2),...I(W m )] as E2 (0) Input into the formula Obtain Row6:EncoderRepeat Row7:For l in L Row 8: Input into the formula to obtain Row 9: Input into the formula to obtain Row 10: Using the formula to obtain E (l) Row11:l++ Row12:EndFor Row 13: Represent R as VADBW to obtain D (0) Row14:Decoder Repeat Row15:For l in L Row16: Input D (l-1) into the formula M(D) l = MultiHead(D (l-1) , D (l-1) , D (l-1) ) to obtain M(D) l Row17: Input M(D) l and E (l) into the formula C(E) l = MultiHead(M(D) (l) , E (L) , E (L) ) to obtain C(E) (l) Row 18: Input C(E) (l) Input formula D (l) = FNN(C(E) l ) to obtain D (l) Row19:l++ Row20:EndFor Row 21: Take D (L) Perform softmax and log operations to calculate the probability distribution score of the current word: P(a, b) = P(ab) / P(b); a is the word to be selected currently, b is the word selected at the previous moment; log is expressed as: logP(a, b) = logP(a|b) + logP(b) Row22: Record the id and log probability score of the word selected in the current step. Finally, all the ids form R. Row23: Input the final R into the BERT sentiment classifier to verify the accuracy of the sentiment generated by the model. Row24:End.

3. The method for generating sentiment-enhanced conversations by improving Transformer according to claim 2, wherein Row2 is to read all the dialogue pairs in the emotion dialogue dataset; Row3 is to set the hyperparameters of the training process; Rows 4 - 12 are the calculation process of the improved Transformer encoder; Rows 13 - 20 are the calculation process of the decoder of this emotion dialogue generation model; Rows 21 - 22 are the process of calculating the probability distribution of the output words; Row23 is the process of verifying the sentiment accuracy of the generated emotion response.

4. The method for generating an emotion-enhanced dialogue by improving a Transformer according to claim 3, wherein During the operation of the encoder, there is also a soft control gate λ for obtaining the decoding concept. The generation of the reply is completed within t time steps during the operation of the computing algorithm, and the decoder state at each time step is Combining each output E of the encoder through the attention mechanism (L) , obtaining a context vector a t , a t And Connected and fed through two linear layers to generate the vocabulary distribution 5. The method for generating sentiment-enhanced dialogues by improving Transformer according to claim 4, wherein Use λ to obtain the final decoding probability, and the specific calculation formula is: a = Attention(E (L) , E (L) , E (L) ); where V, V', b, b' are parameters learned during model training. σ(.) is the sigmoid function used to generate a value between 0 and 1. W d , W c(E) and W E are parameters learned during model training, and are the transposes of W d , W c(E) and W E respectively, Attention(.) is the attention distribution function used to calculate the attention weights between the decoder and the encoder when generating sentence R.