Multi-round Dialogue Method and System with Dialogue Style Perception and Topic Guidance

By building a multi-task deep learning network model, combining BERT model and data augmentation technology, the challenges of dialogue style recognition and context understanding in multiple rounds of dialogue systems are solved, and the accuracy of dialogue response is improved.

CN116860943BActive Publication Date: 2025-07-08FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310874629.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-07-17
Publication Date
2025-07-08
Estimated Expiration
2043-07-17

AI Technical Summary

Technical Problem

Existing multi-round dialogue systems have challenges in dealing with long-term dependence on dialogue contexts and dialogue style recognition, resulting in insufficient accuracy in response selection.

Method used

Using a multi-round dialogue method of dialogue style perception and theme-guided, we use the multi-task deep learning network model, and use the BERT model for encoding and data enhancement. Combining the comparison learning and attention mechanism, we calculate the loss function of the multi-round dialogue response to improve the accuracy of response selection.

Benefits of technology

Improves the accuracy of multi-round dialogue response selection, allowing better understanding of dialogue context and identifying dialogue styles, thereby generating more appropriate responses.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116860943B_ABST
    Figure CN116860943B_ABST
Patent Text Reader

Abstract

The present invention proposes a multi-turn dialogue method and system for dialogue style perception and topic guidance. The method includes the following steps: Step A: Collect multi-turn dialogue texts in a specific scenario, annotate the category to which the response of each dialogue belongs, and construct a multi-turn dialogue training set DS with positive and negative category labels; Step B: Use the training set DS to train a multi-task deep learning network model for multi-turn dialogue response selection; Step C: Input the multi-turn dialogue context and candidate responses into the trained deep learning network model to obtain the most appropriate response for the corresponding multi-turn dialogue context. The present invention is beneficial to improving the accuracy of multi-turn dialogue response selection in the process of human-computer dialogue.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing, in particular to a multi-turn dialogue method and system for dialogue style perception and topic guidance. Background Art

[0002] With the continuous development of computer technology, the development of artificial intelligence technology has also become increasingly rapid, and artificial intelligence has played an important role in various fields. In recent years, with the development of machine learning and deep learning networks, natural language processing has made great progress in the dialogue between humans and computers, and the human-computer dialogue system has gradually become a key research issue. The human-computer dialogue system is a human-computer interaction system that enables humans and computers to communicate through natural language processing. The human-computer dialogue system has important significance and application value and can be widely applied in multiple fields. The human-computer dialogue system has broad commercial value, and currently many commercial companies have launched products related to the dialogue system, such as Apple's Siri, Microsoft's Cortana and Xiaoice, Google's Google Now, and Alibaba's Alibaba Genie, etc. The current dialogue systems mainly have two methods: generation-based and retrieval-based. The former method aims to generate a suitable response given the context, and the latter method aims to select the correct answer from candidate response answers.

[0003] With the continuous development of machine learning methods, generation-based methods have been continuously proposed. The responses generated by the generation-based dialogue system are diverse, but often lack logic and sometimes fall into the trap of safe responses. The retrieval-based dialogue system retrieves the most suitable response from the corpus according to the specific context, can mine the correlation information between the context and the correct response, and infer a suitable response based on this information. Compared with the generation-based dialogue system model, it is more reliable and practical.

[0004] Currently, retrieval-based methods are constantly evolving. Early information retrieval-based systems were mainly applied to short texts and single-round conversations, using simple text matching and rule methods to solve problems. However, it is difficult to apply these methods to multi-round conversation problems that are closer to actual scenarios. Multi-round conversations need to consider the long-term dependencies of the conversation context and are more challenging than single-round conversations. The rapid development of deep learning in recent years has provided solutions to multi-round conversation problems. Wu et al. proposed the SMN model, which matches the similarity between the candidate answers and each sentence in the conversation context respectively and calculates the final score. An et al. used cross-convolution operations to expand the encoder and improved the representation of rare words by learning different representation vectors from the dataset. Dong and Huang proposed to handle out-of-vocabulary words by pre-training word vectors. With the proposal of the BERT model, the BERT-VFT model proposed by Whang et al. uses a post-training method to continue training the pre-trained language model on the dataset in the dialogue field to master the language style of the dialogue. The SA-BERT model proposed by Gu et al. adds special tokens to the input of the pre-trained model to capture the language styles of different dialogue participants and enables the pre-trained model to better master the dialogue information. TWhang et al. proposed auxiliary tasks to improve the pre-trained model, and the auxiliary strategies they proposed can all improve the performance. The BERT-SL model proposed by Xu et al. designed four auxiliary tasks to enable the model to better learn dialogue information. The BERT-FP model proposed by Han Janghoon et al. introduced a fine-grained post-training task by generating short contexts, further improving the performance of the model. The MDKEMN model proposed by Chen Xiuyi et al. introduced a knowledge graph and achieved SOTA performance on the KdConv dataset. The GLIMM model proposed by Ma Hui et al. interacted with the utterance and response from both local and global perspectives. The model proposed by Zhang Wentao et al. performed data augmentation on the dataset and further improved the performance of the model through supervised contrast learning. Summary of the Invention

[0005] The present invention proposes a multi-round conversation method and system for dialogue style perception and topic guidance, which is beneficial to improving the accuracy of multi-round conversation response selection in the process of human-computer conversation.

[0006] The present invention adopts the following technical solutions.

[0007] A multi-round conversation method for dialogue style perception and topic guidance, the method includes the following steps;

[0008] Step A: Collect multi-round conversation texts in a specific scenario, label the category to which the response of each conversation belongs, and construct a multi-round conversation training set DS with positive and negative category labels;

[0009] Step B: Use the training set DS to train a multi-task deep learning network model G for multi-turn dialogue response selection;

[0010] Step C: Input the multi-turn dialogue context and candidate responses into the trained deep learning network model G to obtain the most appropriate response for the corresponding multi-turn dialogue context.

[0011] The said Step B includes the following steps;

[0012] Step B1: For each training sample in the training set DS, extract the historical dialogue of a single speaker according to its turn number, concatenate the correct candidate response after the historical dialogue to obtain a dialogue history sample H; randomly extract a sentence of another speaker to replace the response to obtain a dialogue history negative sample H - ; Encode each training sample in the training set DS to obtain the representation vectors E of the context and response cr ;

[0013] Step B2: Input the representation vectors E obtained in Step B1 cr into the classification layer to calculate scores, and calculate the loss loss of the response selection main task according to the target loss function main ;

[0014] Step B3: Augment the dialogue history sample H obtained in Step B1 to obtain an augmented dialogue history sample H + , and form a triple (H, H + , H - ) with the dialogue history sample and the dialogue history negative sample. Encode the obtained triple (H, H + , H - ), and calculate the contrast loss loss using contrastive learning for the obtained representations CL ;

[0015] Step B4: Split the representation vectors E obtained in Step B1 cr by position to obtain the representation E of each utterance. Calculate the cosine similarity between the current window and adjacent windows through a variable-size window. Split the context except the last sentence into multiple topic windows W according to the similarity level, and splice the representations of all utterances within each window to obtain the representation of the topic window U Regard the last sentence as the question q, and its representation is E q ;

[0016]

[0016] Step B5: Input the window representations obtained in Step B4 into a convolutional neural network to obtain the representation vectors E of the windows W , for E qCalculate the attention scores for the representation vectors of all windows, and select window W according to the scores q , concatenate W q , q and the response to obtain a window-question sample;

[0017] Step B6: Encode the window-question sample obtained in step B5 to obtain a representation vector E WQ , input the representation vector E WQ into the classification layer to calculate the scores, and calculate the loss loss of the problem-oriented window selection according to the objective loss function WQ ;

[0018] Step B7: Add the loss loss obtained in step B2 main , the loss loss obtained in step B3 CL and the loss loss obtained in step B6 WQ to get the final loss loss, calculate the gradients of the parameters in the deep learning network model using the backpropagation method, and update the parameters using the stochastic gradient descent method;

[0019] Step B8: When the iterative change of the loss value generated by the deep learning network model is less than the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.

[0020] The specific steps of step B1 are as follows;

[0021] Step B11: Extract all the historical conversations h of a single speaker according to the number of turns in the dialogue context, form a dialogue history sample H with the response r, and randomly select a sentence u of another speaker other to form a negative dialogue history sample H - :

[0022]

[0023] H = {h, r} Equation 2;

[0024] H - = {h, u other} Equation 3;

[0025] Step B12: Concatenate the utterances and responses in the dialogue context to obtain the input x of the BERT model:

[0026] x = {[CLS], u1, [EOS], u2, [EOS], …, u n , [EOS], [SEP], r, [SEP]} Equation 4;

[0027] Among them, x is the long text obtained by splicing, [sEP] is the separator, [CLS] is the token used by the BERT model to learn global features, and [EOS] is the special token for distinguishing each utterance;

[0028] Step B13: Input the x obtained in step B12 into the BERT model to obtain the representation vectors E of the context and the response cr :

[0029] E cr = BERT(x) Equation Five.

[0030] The said step B2 includes the following steps;

[0031] Step B21: Input the representation vector E obtained in step B13 cr into the classification layer to calculate the score, and the calculation formula is as follows:

[0032] g main (c, r) = σ(W1E cr + b1) Equation Six;

[0033] Among them, W1 is the trainable parameter matrix, b1 is the bias vector, c is the context, r is the response, and σ(·) is the sigmoid activation function;

[0034] Step B22: The objective function uses the cross-entropy loss function to calculate the loss value loss of the multi-turn dialogue response selection main task main :

[0035]

[0036] Among them, DS is the training set and y is the corresponding label.

[0037] The said step B3 specifically includes the following steps:

[0038] Step B31: Perform data augmentation on h in the dialogue history sample H = {h, r} obtained in step B11 to obtain the augmented historical dialogue h + , and the data augmentation methods include: random shuffling, random swapping, and crossing. Among them, random shuffling is to randomly shuffle all the tokens of an utterance, random swapping is to randomly swap any two tokens in an utterance, and crossing is to divide two utterances into two halves respectively and swap the second halves with each other, h + is expressed as:

[0039]

[0040] Among them The utterance for data augmentation is as follows. The specific method of data augmentation is as follows: If there is only one utterance in h, randomly select one method from random shuffling and random swapping for augmentation. If the number of utterances in h is greater than 1, randomly select one method from three methods for augmentation. The utterance for augmentation is randomly selected from h. The augmented historical dialogue h + and the response r are combined to form an augmented dialogue history sample H + ;

[0041] Step B32: The augmented dialogue history sample H obtained in Step B31 + is combined with the dialogue history sample H and the dialogue history negative sample H - to form a triple (H, H + , H - ). The utterances and responses in the sample are concatenated to obtain the input x of the BERT model h :

[0042]

[0043] Formula Nine;

[0044] Step B33: The x obtained in Step B32 h is input into the BERT model to obtain the representation vectors E of the dialogue history context and the response CL :

[0045] E CL = BERT(x h ) Formula Ten;

[0046]

[0047] The obtained feature vectors are used for contrastive learning to calculate the contrastive loss loss of the language style CL :

[0048]

[0049] where sim(·) represents the cosine similarity function, and τ is the temperature coefficient used to adjust the attention to hard negative samples

[0050] Step B4 specifically includes the following steps

[0051] Step B41: The representation vectors E of the context and the response obtained in Step B13 cr are sliced at the position of [EOS] to obtain the representation E of each utterance U , and the formula is

[0052]

[0053] Step B42: After obtaining the utterance representation E in Step B41, set the maximum window size M and the variable-size window W U ={W change | i = 1, 2,..., M}, and sequentially pass through W i to segment M windows starting from change From At the same time, segment each W and the subsequent window of size M change Concatenate all the utterance representations within the window and input them into a convolutional neural network to obtain the representation for each window where concat(·) represents the concatenation operation; And

[0054] And

[0055]

[0056]

[0057]

[0058]

[0059] Among them, sim(·) represents the cosine similarity function. According to the similarity list, select the position j with the lowest similarity, use a window of size j to segment a topic window from the context, and then continue to repeat segmenting the context except u

[0060] from after this window to obtain multiple topic windows W, and concatenate all the utterance representations within each topic window to obtain the topic window representationStep B43: Calculate the cosine similarity at the corresponding positions for the results obtained in Step B42 to obtain the similarity list similarity: And similarity = [s1, s2,..., s

[0061]

[0062] Equation XIX; M The formula is as follows:

[0063] n Outside to obtain multiple topic windows W, and concatenate all the utterance representations within each topic window to obtain the topic window representation The formula is as follows:

[0064]

[0065]

[0066] ​where concat(·) represents the concatenation operation, and u n is regarded as the problem q, which is represented as E q .

[0067] The specific steps of step B5 include the following steps:

[0068] Step B51: Input the topic window representation obtained in step B43 into the convolutional neural network to obtain the topic window representation vector E W , and calculate the attention score score q for E W and the topic window representation vector E Att :

[0069]

[0070]

[0071]

[0072] where W2, W3, and W4 are trainable parameters. According to the scores, select the topic window corresponding to the highest score in the topic window W as the question-topic window W q :

[0073] W q ={u i , u i+1 , …, u i+l} Formula XXV;

[0074] Step B52: Concatenate the question-topic window W q obtained in step B51 with the question q and response r obtained in step B43, and add special markers to obtain the input x h for the BERT model. The formula is:

[0075] x h ={[CLS], u i , [EOS], u i+1 , [EOS], …, u i+l , [EOS], q, [EOS][SEP], r, [SEP]} Formula XXVI.

[0076] The specific steps of step B6 include the following steps:

[0077] Step B61: Input the x h obtained in step B52 into the BERT model to obtain the window-question representation vector E WQ :

[0078] EWQ = BERT(x h ) Formula 27;

[0079] Step B62: Input the feature vector E obtained in Step B61 WQ into the classification layer to calculate the score. The calculation formula is as follows:

[0080] g WQ (c q , r) = σ(W5E WQ + b2) Formula 28;

[0081] where W5 is a trainable parameter matrix, b2 is a bias vector, c q is the window-question context, r is the response, and σ(·) is the sigmoid activation function;

[0082] Step B63: The objective function uses the cross-entropy loss function to calculate the loss value loss of the multi-turn dialogue response selection auxiliary task WQ :

[0083]

[0084] Formula 28;

[0085] where DS is the training set and y is the corresponding label.

[0086] The specific steps of Step B7 are as follows:

[0087] Step B71: Add the main task loss loss obtained in Step B22 main , the language style auxiliary task loss loss obtained in Step B33 CL and the question-oriented window selection auxiliary task loss loss obtained in Step B63 WQ to obtain the final loss function of the model:

[0088] loss = loss main + αloss cL + βloss WQ Formula 29;

[0089] where α and β are hyperparameters used to control the influence of the two auxiliary tasks on the model;

[0090] Step B72: Use the backpropagation method to calculate the gradients of the parameters in the deep learning network model and use the stochastic gradient descent method to update the parameters.

[0091] A multi-turn dialogue system with dialogue style perception and topic guidance, which is a human-computer dialogue system, adopts the multi-turn method of dialogue style perception and topic guidance described above. The modules included in the dialogue system and their module functions are as follows;

[0092] Data collection module: used to collect multi-turn dialogue texts in specific scenarios, label the category to which the response of each dialogue belongs, and construct a multi-turn dialogue training set DS with positive and negative category labels;

[0093] Pre-trained language model training module: used to input dialogue context, response, and label into the pre-trained language model BERT to obtain the representation vectors of the context and response, calculate the target loss function through the representation vectors, and add the losses of two auxiliary tasks to train the entire model with the goal of minimizing the loss, obtaining a multi-turn dialogue response selection model. The pre-trained language models of the main task and the auxiliary tasks share parameters;

[0094] Language style auxiliary task module: for each sample in the training set DS, extract the historical dialogue of a single speaker to construct a dialogue history sample, and construct a dialogue history negative sample and an enhanced sample to form a triple of original sample, enhanced sample, and negative sample. Use the pre-trained language model to encode the triple, and apply contrastive learning to the obtained representations to calculate the loss of the language style auxiliary task;

[0095] Problem-oriented window selection auxiliary task module: divide the context except the last sentence into multiple topic windows, regard the last sentence of the context as the problem, calculate the attention score between the problem and each window, select the window with the highest score to form a window-problem sample with the problem and the response, use the pre-trained language model to encode the window-problem sample, and calculate the loss of the problem-oriented window selection auxiliary task for the obtained representation vectors;

[0096] Multi-turn dialogue response selection module: use the trained multi-turn dialogue response selection model of language style and problem-oriented window selection to analyze and process the input dialogue text, and output the most appropriate response for the current dialogue context.

[0097] The present invention is beneficial to improving the accuracy of multi-turn dialogue response selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] The present invention will be further described in detail below with reference to the drawings and specific embodiments:

[0099] Attached Figure 1 is a schematic flowchart of the method implementation of the embodiment of the present invention;

[0100] Attached Figure 2 is a schematic diagram of the model architecture in the embodiment of the present invention;

[0101] Attached Figure 3It is a schematic diagram of the system structure of an embodiment of the present invention. Detailed implementation manners

[0102] As shown in the figure, a multi-turn dialogue method for dialogue style perception and topic guidance, the method includes the following steps;

[0103] Step A: Collect multi-turn dialogue texts in a specific scenario, label the category to which the response of each dialogue belongs, and construct a multi-turn dialogue training set DS with positive and negative category labels;

[0104] Step B: Use the training set DS to train a multi-task deep learning network model G for multi-turn dialogue response selection;

[0105] Step C: Input the multi-turn dialogue context and candidate responses into the trained deep learning network model G to obtain the most appropriate response for the corresponding multi-turn dialogue context.

[0106] The said Step B includes the following steps;

[0107] Step B1: Extract the historical dialogue of a single speaker according to the number of turns for each training sample in the training set DS, concatenate the correct candidate response after the historical dialogue to obtain a dialogue history sample H; randomly extract a sentence of another speaker to replace the response to obtain a dialogue history negative sample H - ; Encode each training sample in the training set DS to obtain the representation vectors E of the context and the response cr ;

[0108] Step B2: Input the representation vector E obtained in Step B1 cr into the classification layer to calculate the score, and calculate the loss loss of the response selection main task according to the target loss function main ;

[0109] Step B3: Augment the dialogue history sample H obtained in Step B1 to obtain an augmented dialogue history sample H + , and form a triple (H, H + , H - ) with the dialogue history sample and the dialogue history negative sample, encode the obtained triple (H, H + , H - ), and calculate the contrast loss loss using contrastive learning for the obtained representation CL ;

[0110] Step B4: Cut out the representation of each utterance from the representation vector E obtained in Step B1 cr by position to obtain the representation E of each utterance U, calculate the cosine similarity between the current window and adjacent windows through a window of variable size, segment the context except the last sentence into multiple topic windows W according to the similarity level, and splice the representations of all utterances within each window to obtain the representation of the topic window Regard the last sentence as the question q, and its representation is E q ;

[0111] Step B5: Input the window representation obtained in step B4 into a convolutional neural network to obtain the feature vector E of the window W , calculate the attention scores for E q and the feature vectors of all windows, select the window W according to the scores q , splice W q , q and the response to obtain the window-question sample;

[0112] Step B6: Encode the window-question sample obtained in step B5 to obtain the feature vector E WQ , input the feature vector E WQ into the classification layer to calculate the score, and calculate the loss loss of the question-oriented window selection according to the objective loss function WQ ;

[0113] Step B7: Add the loss loss obtained in step B2 main , the loss loss obtained in step B3 CL and the loss loss obtained in step B6 WQ to get the final loss loss, calculate the gradients of the parameters in the deep learning network model using the backpropagation method, and update the parameters using the stochastic gradient descent method;

[0114] Step B8: When the iterative change of the loss value generated by the deep learning network model is less than the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model.

[0115] The specific steps of step B1 are as follows;

[0116] Step B11: Extract all historical conversations h of a single speaker according to the number of turns in the conversation context, form the conversation history sample H with the response r, and randomly select an utterance u of another speaker other to form the negative conversation history sample H - :

[0117]

[0118] H = {h, r} Formula Two;

[0119] H- = {h, u other} Formula Three;

[0120] Step B12: Concatenate the utterances and responses in the dialogue context to obtain the input x for the BERT model:

[0121] x = {[CLS], u1, [EOS], u2, [EOS],..., u n , [EOS], [SEP], r, [SEP]} Formula Four;

[0122] where x is the long text obtained by concatenation, [SEP] is the separator, [CLS] is the token used by the BERT model to learn global features, and [EOS] is the special token to distinguish each utterance;

[0123] Step B13: Input the x obtained in Step B12 into the BERT model to obtain the representation vectors E of the context and the response cr :

[0124] E cr = BERT(x) Formula Five.

[0125] Step B2 includes the following steps;

[0126] Step B21: Input the representation vector E obtained in Step B13 into the classification layer to calculate the score. The calculation formula is as follows: cr Input into the classification layer to calculate the score, and the calculation formula is as follows:

[0127] g main (c, r) = σ(W1E cr + b1) Formula Six;

[0128] where W1 is a trainable parameter matrix, b1 is a bias vector, c is the context, r is the response, and σ(·) is the sigmoid activation function;

[0129] Step B22: The objective function uses the cross-entropy loss function to calculate the loss value loss of the multi-round dialogue response selection main task main :

[0130]

[0131] where DS is the training set and y is the corresponding label.

[0132] Step B3 specifically includes the following steps:

[0133] Step B31: Augment the h in the dialogue history sample H = {h, r} obtained in Step B11 to obtain the augmented historical dialogue h +, the data augmentation methods include: random shuffling, random swapping, and cross-over. Among them, random shuffling means randomly shuffling all the tokens of a utterance, random swapping means randomly swapping any two tokens in an utterance, and cross-over means splitting two utterances into two halves respectively and swapping the second halves with each other, h + It is expressed as:

[0134]

[0135] where represents the utterance for data augmentation. The specific method of data augmentation is as follows: if there is only one utterance in h, randomly select one method from random shuffling and random swapping for augmentation; if the number of utterances in h is greater than 1, randomly select one method from the three methods for augmentation. The utterance for augmentation is randomly selected from h. The augmented historical dialogue h + and the response r form an augmented dialogue history sample H + ;

[0136] Step B32: The augmented dialogue history sample H obtained in Step B31 + and the dialogue history sample H and the dialogue history negative sample H - constitute a triple (H, H + , H - ). Concatenate the utterances and responses in the sample to obtain the input x of the BERT model h :

[0137] Formula Nine;

[0138] Step B33: Input the x obtained in Step B32 h into the BERT model to obtain the representation vectors E of the dialogue history context and the response CL :

[0139] E CL = BERT(x h ) Formula Ten;

[0140]

[0141] Apply contrastive learning to the obtained feature vectors and calculate the contrastive loss loss of the language style CL :

[0142]

[0143] where sim(·) represents the cosine similarity function, and τ is the temperature coefficient used to adjust the attention to hard negative samples.

[0144] The specific steps of the said Step B4 include the following steps:

[0145] Step B41: For the context and response representation vectors E obtained in Step B13 cr Split the representation E of each utterance at the position of [EOS], and the formula is: U , the formula is:

[0146]

[0147] Step B42: After obtaining the utterance representation E in Step B41 U , set the maximum window size M and the variable-size window W change ={W i | i = 1, 2,..., M}, and sequentially pass through W change From Start splitting out M windows At the same time, split out each W change The subsequent window of size M Concatenate all the utterance representations within the window and input them into the convolutional neural network to obtain the representation representing each window And

[0148] where concat(·) represents the concatenation operation; And

[0149]

[0150]

[0151]

[0152]

[0153] where sim(·) represents the cosine similarity function. According to the similarity list, select the position j with the lowest similarity, use a window of size j to split out a topic window from the context, and then continue to repeat the splitting from after this window except for u

[0154] Step B43: Calculate the cosine similarity at the corresponding positions for the results obtained in Step B42 And to obtain the similarity list similarity:

[0155]

[0156] similarity = [s1, s2,..., s M Formula XIX;

[0157] where sim(·) represents the cosine similarity function. According to the similarity list, select the position j with the lowest similarity, use a window of size j to split out a topic window from the context, and then continue to repeat the splitting from after this window except for u nObtain multiple topic windows W from the context outside, and splice all the utterance representations within each topic window to obtain the topic window representation The formula is as follows:

[0158]

[0159]

[0160] where concat(·) represents the splicing operation, and regard u n as the question q, and its representation is E q .

[0161] The specific steps of step B5 include the following steps:

[0162] Step B51: Input the topic window representation obtained in step B43 into the convolutional neural network to obtain the topic window characterization vector E W , and calculate the attention score score q for E W and the topic window characterization vector E Att :

[0163]

[0164]

[0165] score Att = α(E q , E W ) Formula 24;

[0166] where W2, W3, and W4 are trainable parameters. According to the score, select the topic window corresponding to the highest score in the topic window W as the question-topic window W q :

[0167] W q = {u i , u i+1 ,..., u i+l} Formula 25;

[0168] Step B52: Splice the question-topic window W q obtained in step B51 with the question q and response r obtained in step B43, and add special markers to obtain the input x h to the BERT model. The formula is:

[0169] x h = {[CLS], u i , [EOS], u i+1 , [EOS],..., ui+l , [EOS], q, [EOS][SEP], r, [SEP]} Formula 26.

[0170] Step B6 specifically includes the following steps:

[0171] Step B61: Input the x obtained in step B52 h into the BERT model to obtain the representation vector E of the window-question WQ :

[0172] E WQ = BERT(x h ) Formula 27;

[0173] Step B62: Input the representation vector E obtained in step B61 WQ into the classification layer to calculate the score. The calculation formula is as follows:

[0174] g WQ (c q , r) = σ(W5E WQ + b2) Formula 28;

[0175] where W5 is a trainable parameter matrix, b2 is a bias vector, c q is the window-question context, r is the response, and σ(·) is the sigmoid activation function;

[0176] Step B63: The objective function uses the cross-entropy loss function to calculate the loss value loss of the multi-round dialogue response selection auxiliary task WQ :

[0177]

[0178] where DS is the training set and y is the corresponding label.

[0179] Step B7 specifically includes the following steps:

[0180] Step B71: Add the main task loss loss obtained in step B22 main , the language style auxiliary task loss loss obtained in step B33 CL and the problem-oriented window selection auxiliary task loss loss obtained in step B63 WQ to obtain the final loss function of the model:

[0181] loss = loss main + αloss CL + βloss WQ Formula 29;

[0182] Among them, α and β are hyperparameters used to control the influence of the two auxiliary tasks on the model;

[0183] Step B72: Calculate the gradients of the parameters in the deep learning network model using the backpropagation method, and update the parameters using the stochastic gradient descent method.

[0184] The dialogue style perception and topic-guided multi-round dialogue system is a human-computer dialogue system that adopts the above-mentioned dialogue style perception and topic-guided multi-round method. The modules included in the dialogue system and their module functions are as follows;

[0185] Data collection module: Used to collect multi-round dialogue texts in a specific scenario, annotate the category to which the response of each dialogue belongs, and construct a multi-round dialogue training set DS with positive and negative category labels;

[0186] Pre-trained language model training module: Used to input the dialogue context, response, and label into the pre-trained language model BERT to obtain the representation vectors of the context and response, calculate the target loss function through the representation vectors, and add it to the losses of the two auxiliary tasks. The entire model is trained with the goal of minimizing the loss to obtain a multi-round dialogue response selection model. The pre-trained language models of the main task and the auxiliary tasks share parameters;

[0187] Language style auxiliary task module: Extract the historical dialogue of a single speaker from each sample in the training set DS to construct a dialogue history sample, and construct a dialogue history negative sample and an enhanced sample to form a triple of original sample, enhanced sample, and negative sample. Use the pre-trained language model to encode the triple, and apply contrastive learning to the obtained representations to calculate the loss of the language style auxiliary task;

[0188] Problem-oriented window selection auxiliary task module: Split the context except the last sentence into multiple topic windows, regard the last sentence of the context as the problem, calculate the attention scores between the problem and each window, select the window with the highest score and combine it with the problem and response to form a window-problem sample. Use the pre-trained language model to encode the window-problem sample, and calculate the loss of the problem-oriented window selection auxiliary task for the obtained representation vectors;

[0189] Multi-round dialogue response selection module: Use the trained multi-round dialogue response selection model of language style and problem-oriented window selection to analyze and process the input dialogue text, and output the most suitable response for the current dialogue context.

[0190] As described above, it is only the preferred embodiment of the present invention, and it is not intended to limit the present invention in other forms. Any person skilled in the art may use the technical content disclosed above to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modification, equivalent change, and modification made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still fall within the protection scope of the technical solution of the present invention.

Claims

1. A multi-round dialogue method for dialogue style perception and topic guidance, characterized in that: The method includes the following steps; Step A: Collect multi-turn dialogue texts in a specific scenario, label the category to which the response of each dialogue belongs, and construct a multi-turn dialogue training set DS with positive and negative category labels; Step B: Use the training set DS to train a multi-task deep learning network model G for multi-turn dialogue response selection; Step C: Input the multi-turn dialogue context and candidate responses into the trained deep learning network model G to obtain the most suitable response for the corresponding multi-turn dialogue context; The said Step B includes the following steps; Step B1: For each training sample in the training set DS, extract the historical dialogue of a single speaker according to the training sample round, concatenate the correct candidate response after the historical dialogue to obtain the dialogue history sample H; randomly extract a sentence of another speaker to replace the response to obtain the dialogue history negative sample H - ; Encode each training sample in the training set DS to obtain the representation vectors E of the context and the response cr ; Step B2: Input the feature vector E obtained in Step B1 cr into the classification layer to calculate the score, and calculate the loss of the response selection main task according to the objective loss function main ; Step B3: Augment the dialogue history sample H obtained in Step B1 to obtain the augmented dialogue history sample H + , which together with the dialogue history sample and the dialogue history negative sample form a triple (H, H + , H - ). Encode the obtained triple (H, H + , H - ), and calculate the contrastive loss loss using contrastive learning for the obtained representations CL ; Step B4: Take the characterization vector E obtained in Step B1 cr and split out the representation E of each utterance by position U . Calculate the cosine similarity between the current window and adjacent windows through a window of variable size. According to the similarity level, split the context except the last sentence into multiple topic windows W, and splice the representations of all utterances within each window to obtain the representation of the topic window Regard the last sentence as the question q, and its representation is E q ; Step B5: The window representation obtained in step B4 is input into a convolutional neural network to obtain the representation vector E of the window W . For E q and the representation vectors of all windows, attention scores are calculated, and the window W is selected according to the scores q . W q , q and the response are concatenated to obtain the window-question sample; Step B6: Encode the window-problem samples obtained in Step B5 to obtain a representation vector E WQ , and input the representation vector E WQ into the classification layer to calculate a score, and calculate the loss loss of problem-oriented window selection according to the objective loss function WQ ; Step B7: Take the loss obtained in Step B2 main , the loss obtained in Step B3 CL and the loss obtained in Step B6 WQ add them up to get the final loss, calculate the gradients of each parameter in the deep learning network model using the backpropagation method, and update each parameter using the stochastic gradient descent method; Step B8: When the iterative change of the loss value generated by the deep learning network model is less than the given threshold or reaches the maximum number of iterations, terminate the training process of the deep learning network model; In step B2, the cross-entropy loss function is adopted as the objective function to calculate the loss value loss of the main task of multi-turn dialogue response selection main : where y is the corresponding label; g main (c, r) represents the model's estimated matching probability of the given dialogue context c and the candidate response r; In step B3, contrastive learning is applied to the obtained feature vectors to calculate the contrastive loss of the language style loss CL : Where sim(·) represents the cosine similarity function, and τ is the temperature coefficient used to adjust the attention to hard negative samples; In step B6, the cross-entropy loss function is used as the objective function to calculate the loss value loss of the multi-turn dialogue response selection auxiliary task WQ : g WQ (c q , r) is the representation vector E WQ input to the calculation score in the classification layer, c q is the window-question context, r is the response; In step B7, add the main task loss main , the language style auxiliary task loss CL and the problem-oriented window selection auxiliary task loss WQ to obtain the final loss function of the model: loss = loss main + αloss CL + βloss WQ Formula 29; Where α and β are hyperparameters used to control the influence of the two auxiliary tasks on the model.

2. The multi-turn dialogue method for dialogue style perception and topic guidance according to claim 1, characterized in that: The said Step B1 specifically includes the following steps; Step B11: Extract all the historical conversations h of a single speaker according to the turn number in the dialogue context, form a dialogue history sample H with the response r, and randomly select a sentence u of the other speaker other to form a negative dialogue history sample H with the dialogue history h - : H - = {h, u other} Equation 3; Step B12: Concatenate the utterances and responses in the dialogue context to obtain the input x of the BERT model: x = {[CLS], u1, [EOS], u2, [EOS],..., u n , [EOS], [SEP], r, [SEP]} Equation 4; Where x is the long text obtained by splicing, [SEP] is the separator, [CLS] is the token used by the BERT model to learn global features, and [EOS] is the special token used to distinguish each utterance; Step B13: Input the x obtained in step B12 into the BERT model to obtain the representation vectors E of the context and the response cr : E cr = BERT(x), Equation Five.

3. The multi-turn dialogue method for dialogue style perception and topic guidance according to claim 2, wherein: The step B2 inputs the characterization vector E obtained in step B13 cr into the classification layer to calculate the score, and the calculation formula is as follows: g main (c, r) = σ(W1E cr + b1) Equation 6; Where W1 is the trainable parameter matrix, b1 is the bias vector, c is the context, r is the response, and σ(·) is the sigmoid activation function.

4. The multi-turn dialogue method for dialogue style perception and topic guidance according to claim 3, wherein: The said Step B3 specifically includes the following steps: Step B31: Augment the history dialogue h in the dialogue history sample H obtained in Step B11 to obtain an augmented historical dialogue h + , and the data augmentation methods include: random shuffling, random swapping, and crossing. Among them, random shuffling means randomly shuffling all the tokens of a utterance, random swapping means randomly swapping any two tokens in an utterance, and crossing means dividing two utterances into two halves respectively and swapping the second halves with each other, where h + is expressed as: Among them represents the utterance for data augmentation. The specific method of data augmentation is as follows: if there is only one utterance in h, randomly select one method from random shuffling and random swapping for augmentation; if the number of utterances in h is greater than 1, randomly select one method from three methods for augmentation. The utterance for augmentation is randomly selected from h. The augmented historical dialogue h + and the response r form the augmented dialogue history sample H + ; Step B32: Use the enhanced dialogue history sample H obtained in step B31 + along with the dialogue history sample H and the negative dialogue history sample H - to form a triple (H, H + , H - ). Concatenate the utterances and responses in the sample to obtain the input x h : Step B33: Input the x obtained in step B32 h into the BERT model to obtain the representation vectors E of the dialogue history context and the response CL : E CL = BERT(x h ) Formula Ten; 5. The multi-turn dialogue method for dialogue style perception and topic guidance according to claim 4, wherein: The said Step B4 specifically includes the following steps: Step B41: Take the context and response representation vectors E obtained in step B13 cr and split the representation E of each utterance according to the position of [EOS] U , with the formula: Step B42: After obtaining the utterance representation E in step B41, set the maximum window size M, and calculate the representation for each window through a window of variable size; U ​ Step B43: Calculate the cosine similarity of the representations of each window obtained in Step B42 at the corresponding positions to obtain a similarity list similarity: According to the similarity list, select the position with the lowest similarity, use a window of size j to segment the context into a topic window, and then continue to repeat segmenting the context except u after this window to obtain multiple topic windows W. Concatenate all the utterance representations within each topic window to obtain the topic window representation n Except for the context outside u, multiple topic windows W are obtained. Concatenate all the utterance representations within each topic window to obtain the topic window representation 6. The multi-round dialogue method for dialogue style perception and topic guidance according to claim 5, characterized in that: The said Step B5 specifically includes the following steps: Step B51: Input the topic window representation obtained in step B43 into the convolutional neural network to obtain the topic window representation vector E W , and calculate the attention score score q for E W and the topic window representation vector E Att : score Att = α(E q , E W ) Formula 24; Where W2, W3, and W4 are trainable parameters, and according to the scores, the topic window with the highest score is selected from the topic windows in the topic window W as the question-topic window W q ; Step B52: Concatenate the problem-topic window W obtained in step B51 q with the problem q and response r obtained in step B43, and add special markers to obtain the input x of the BERT model h , and the formula is: x h = {[CLS], u i , [EOS], u i+1 , [EOS],..., u i+l , [EOS], q, [EOS][SEP], r, [SEP]} Formula XXVI.

7. The multi-turn dialogue method for dialogue style perception and topic guidance according to claim 6, wherein: The step B6 inputs the x obtained in the step B52 h into the BERT model to obtain the representation vector E of the window-question WQ : E WQ = BERT(x h ) Formula XXVII; Step B62: Input the feature vector E obtained in step B61 WQ into the classification layer to calculate the score. The calculation formula is as follows: g WQ (c q , r) = σ(W5E WQ + b2) Formula 28; where W5 is a trainable parameter matrix, b2 is a bias vector, c q is the window-question context, r is the response, and σ(·) is the sigmoid activation function.

8. The multi-turn dialogue method for dialogue style perception and topic guidance according to claim 7, characterized in that: The said Step B7 specifically includes the following steps: Step B72: Use the backpropagation method to calculate the gradients of the parameters in the deep learning network model, and use the stochastic gradient descent method to update the parameters.

9. A multi-turn dialogue system with dialogue style perception and topic guidance, which is a human-computer dialogue system, adopts the multi-turn dialogue method with dialogue style perception and topic guidance as described in any one of claims 1-8, and is characterized in that: The modules included in the dialogue system and their module functions are; Data collection module: Used to collect multi-turn dialogue texts in a specific scenario, label the category to which the response of each dialogue belongs, and construct a multi-turn dialogue training set DS with positive and negative category labels; Pre-trained language model training module, which is used to input the dialogue context, response and label into the pre-trained language model BERT to obtain the representation vectors of the context and response, calculate the target loss function through the representation vectors, and add the losses of the two auxiliary tasks, and train the entire model with the goal of minimizing the loss to obtain a multi-turn dialogue response selection model. The pre-trained language models of the main task and the auxiliary tasks share parameters; Language style auxiliary task module, which extracts the historical dialogue of a single speaker from each sample in the training set DS to construct a dialogue history sample, and constructs a dialogue history negative sample and an augmented sample to form a triple of the original sample, augmented sample, and negative sample. Use the pre-trained language model to encode the triple, apply contrastive learning to the obtained representations, and calculate the loss of the language style auxiliary task; The problem-oriented window selection auxiliary task module divides the context except the last sentence into multiple topic windows, regards the last sentence of the context as the problem, calculates the attention score between the problem and each window, selects the window with the highest score to form a window-problem sample with the problem and the response, encodes the window-problem sample using a pre-trained language model, and calculates the loss of the problem-oriented window selection auxiliary task for the obtained representation vectors; the multi-turn dialogue response selection module uses the trained multi-turn dialogue response selection model with the language style and problem-oriented window selection to analyze and process the input dialogue text and outputs the most appropriate response for the current dialogue context.

Citation Information

Patent Citations

  • Local information perception dialogue method and system based on pre-training language model

    CN114443827A

  • Multi-round dialogue method and system for domain data enhancement and multi-granularity semantic understanding

    CN116050432A