Multi-turn Dialogue Method and System with Co-optimization of Response Enhancement and Span Prediction

Through the deep learning network model optimized by reply enhancement and span prediction, combined with data enhancement and pre-trained language model BERT, the problem of insufficient utilization of global semantic information in multiple rounds of dialogue systems is solved, and the accuracy of reply selection is improved.

CN116361437BActive Publication Date: 2025-07-29FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310333768.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-03-31
Publication Date
2025-07-29
Estimated Expiration
2043-03-31

AI Technical Summary

Technical Problem

When the existing multi-round dialogue system handles complex contexts and dialogue history, it is difficult to effectively learn the semantic information contained in the reply discourse, especially when the global semantic information classification, it ignores the valuable semantic information of other locations in the context, resulting in insufficient accuracy of reply selection.

Method used

A deep learning network model that is jointly optimized by reply enhancement and span prediction is adopted. Through data augmentation and local semantic relationship learning, combined with the pre-trained language model BERT, the representation vectors of global and timing information are optimized, and the multi-task learning framework is used to improve the accuracy of reply selection.

Benefits of technology

By enhancing the model's learning of semantic information and local contextual relevance contained in the reply discourse, the accuracy of reply selection in multiple rounds of dialogue systems is significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116361437B_ABST
    Figure CN116361437B_ABST
Patent Text Reader

Abstract

The present invention relates to a multi-turn dialogue method and system for jointly optimizing response enhancement and span prediction. The method includes the following steps: Step A: Extract user conversations, responses involved in the user conversations, and label the tags of relevant response utterances involved in the user conversations. The positive samples are the correct responses in the conversations, and the negative samples are the incorrect responses, and a training set is constructed. UB ; Step B: Use the training set UB to train a deep learning network model G for jointly optimizing response enhancement and span prediction, which is used to learn the local semantic relationships in the user conversations and the responses involved in the user conversations, and at the same time learn the content in the responses involved in the user conversations; Step C: Input the complete user conversation and the responses involved in the user conversation into the trained deep learning network model G to obtain the correct response for the complete user conversation. The method and system can effectively improve the accuracy of response selection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a multi-turn dialogue method and system for jointly optimizing reply enhancement and span prediction. Background Art

[0002] In recent years, with the continuous development and improvement of artificial intelligence technology, the research on machine learning and deep learning technologies has also been in full swing. Dialogue systems have gradually come into view with the research on machine learning and deep learning. The applications of dialogue systems are very extensive, and they have important research value both in the industrial field and in the academic field. Dialogue systems can be divided into two types according to the type of reply generated: generative dialogue and retrieval-based dialogue. Among them, the generative dialogue model can generate replies autonomously according to the dialogue content without having to select from predefined replies. The generative dialogue has high requirements for data, and it is difficult to collect data sets. The generative dialogue model is complex, prone to generating incorrect replies, and is easily trapped in the trap of safe replies. The retrieval-based dialogue model selects the most suitable reply for the current context from a large number of given candidate replies according to a piece of context. Compared with the generative dialogue model, the retrieval-based dialogue model is simpler and more efficient, can reply to users' questions faster, and has higher accuracy, but it cannot generate new replies or answer questions beyond the scope of the knowledge base.

[0003] Retrieval-based dialogue is divided into two types: single-turn dialogue and multi-turn dialogue. Single-turn dialogue is mainly used for short-text question and answer, cannot process dialogue history information, and has insufficient ability to understand long texts. In daily life, users often express their intentions through multiple utterances. Therefore, single-turn dialogue has little value in actual application scenarios. Retrieval-based multi-turn dialogue selects from a large number of candidate replies according to the dialogue history and the current user's question. However, the long-distance dependency relationship, the temporal semantic information contained, and the topic conversion in the dialogue context make it very difficult to select a suitable reply. In the multi-turn dialogue method introduced in this article, the focus is on the multi-turn reply selection problem in retrieval-based dialogue. Given the context in a dialogue, which consists of multiple utterances. The purpose of this task is to select the most matching reply for the current context from a large number of candidate replies.

[0004] Before the rise of deep learning, early research work was carried out based on rule- and template-matching methods. Eliza was one of the earliest dialogue systems, which was implemented based on rule and template matching and could perform simple semantic understanding and response generation. There were also studies such as Term Frequency-Inverse Document Frequency (TF-IDF) and BM25. These studies were all based on the expression and matching of short texts and were difficult to model long conversations. With the development of artificial intelligence technology, researchers began to apply machine learning to dialogue systems. The methods based on traditional machine learning combined topic analysis algorithms such as Latent Dirichlet Allocation (LDA) and Latent Semantic Analysis (LSA) with short text matching algorithms such as cosine similarity to solve the single-round response selection problem. These models may have been very advanced at that time, but they still had quite a few limitations. They lacked flexibility, required a large amount of manual work, and could not handle complex contexts and conversation histories. With the development of deep learning technology, multi-turn dialogue models based on neural networks have gradually become the mainstream. Due to the temporal characteristics of dialogue content, early deep learning algorithms used Recurrent Neural Networks to process the temporal information contained in the context. With the continuous development of deep learning, researchers have applied various deep neural networks to multi-turn dialogue research. For the multi-turn dialogue response selection task, at first, a single network structure or the aggregation of multiple single network structures was used to model the context.

[0005] In 2017, with the emergence of the Transformer model based on the attention mechanism by Vaswani et al., many excellent multi-turn dialogue models based on Transformer began to appear. The Transformer model has powerful semantic understanding ability. It links all words together through the multi-head self-attention mechanism, effectively extracting the long-distance dependencies in the dialogue. Zhou et al. proposed the DAM (Deep Attention Matching Network) model. This model uses the attention structure in Transformer to represent sentences and adopts the cross-attention mechanism to promote the interaction between sentences. Finally, three-dimensional convolution is used for aggregation. Yuan et al. proposed the MSN (Multi-hop Selector Network) model, which combines the advantages of the DUA model and the DAM model. MSN selects the context more relevant to the question through a multi-hop selector. Inspired by the multi-head attention mechanism of Transformer, these works use multi-head attention to construct semantic information of multiple granularities in the matching stage, thus enriching the feature representation of the model. These improvements have achieved significant performance enhancements. However, these models still have the problem of being unable to understand the dialogue content from a global perspective, which may lead to information loss during the model encoding process.

[0006] In recent years, with the emergence of the famous pre-trained model BERT, more and more methods using mature pre-trained models have also been applied to the multi-turn dialogue field. Scholars have found that pre-trained models can be well combined with downstream tasks. BERT-VFT first applied BERT to the multi-turn response selection task. It converted the original matching task into a classification task. By splicing the historical dialogue and the response together and adding a special [CLS] label to represent the global semantic information for classification, the experimental results demonstrated its superior performance. Subsequently, more and more methods of combining pre-trained models with downstream tasks were proposed. The SA-BERT model was proposed by Gu et al. Since a dialogue usually involves multiple participants, speaker information was added when inputting the data, and the speaker embedding representation was learned through the pre-trained model, thus enhancing the model's ability to understand the language styles of different dialogue participants. Lu et al. proposed an enhanced context language model that can obtain the top-level context dialogue representation, which not only distinguishes different speakers but also cuts off the real dialogue at different time points to expand the training corpus, so as to understand the context semantics more deeply. These models can effectively learn the relevant semantic features in the context and responses through pre-training on large-scale datasets, further improving the model's multi-turn response selection ability. However, for the hidden features contained in the dialogue, such as temporal information, dialogue structure information, etc., there is still a lack of sufficient exploration.

[0007] In recent years, multi-task joint training methods based on pre-trained language models have received much attention. Whang et al. proposed three different auxiliary tasks aiming to optimize the pre-trained model to better adapt to the domain dataset. These auxiliary tasks include utterance deletion, utterance insertion, and utterance search, which further utilize semantic structure information by learning the latent structural features in the dialogue to select the best result from multiple candidate responses. On this basis, researchers strengthen the adaptation ability of the pre-trained language model in dialogue tasks by having the pre-trained language model learn the post-training strategy of the domain data before fine-tuning. However, there are many topic transitions and jumping scenarios in the context of multi-turn conversations, and the information contained in each turn of the conversation varies. These complex scenarios require the model to pay more careful attention to the local information in the context and the semantic information contained in the response utterances. However, there are still many problems in the current research methods. For example, the number of utterances in the context far exceeds the number of utterances in the response. Therefore, when using global semantic information for classification, it is difficult to effectively learn the semantic information contained in the response utterance. In addition, directly using global semantic information for classification ignores the valuable semantic information contained in the utterances at other positions in the context. Summary of the Invention

[0008] The purpose of the present invention is to provide a multi-turn dialogue method and system that jointly optimize response enhancement and span prediction, which can effectively improve the accuracy of response selection.

[0009] To achieve the above purpose, the technical solution adopted by the present invention is: A multi-turn dialogue method that jointly optimizes response enhancement and span prediction, including the following steps:

[0010] Step A: Extract the user dialogue, the responses involved in the user dialogue, and label the tags of the relevant response utterances involved in the user dialogue. The positive sample is the correct response in the dialogue, and the negative sample is the incorrect response, to construct the training set UB;

[0011] Step B: Use the training set UB to train the deep learning network model G that jointly optimizes response enhancement and span prediction, for learning the local semantic relationships in the user dialogue and the responses involved in the user dialogue, and at the same time learning the content in the responses involved in the user dialogue;

[0012] Step C: Input the complete user dialogue and the responses involved in the user dialogue into the trained deep learning network model G to obtain the correct response to the complete user dialogue.

[0013] Further, the specific steps of step B include the following steps:

[0014] Step B1: Encode each training sample in the training set UB to obtain the initial representation vector of the context Initial representation vector of the candidate response and the label y;

[0015] Step B2: For each training positive sample in the training set UB, perform data augmentation on the word-level vectors in the response through two operations: random deletion or random scrambling, to obtain the initial representation vector of the response after data augmentation Use the initial representation vector of the context in Step B1 as the context corresponding to the augmented response, to obtain the corresponding context representation vector

[0016] Step B3: For each training sample in the training set UB, randomly intercept a dialogue segment from the context to obtain the initial representation vector of the dialogue segment Use the original candidate response corresponding to the context as the response of the obtained dialogue segment, to obtain the initial representation vector of the candidate response corresponding to the dialogue segment

[0017] Step B4: Input the representation vectors and into the pre-trained language model to obtain the representation vector of the optimized global semantic information

[0018] Step B5: Input into the sigmoid layer. According to the objective loss function loss main , use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters;

[0019] Step B6: Concatenate the initial representation vector of the context and the initial representation vector of the response after data augmentation and input them into the pre-trained language model to obtain the label vector representing the global semantic information Input into the sigmoid layer. According to the objective loss function loss ra , use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters;

[0020] Step B7: Concatenate the initial representation vector of the dialogue segment and the response representation vector corresponding to the dialogue segment and input them into the pre-trained language model to obtain Input into the LSTM network layer to learn the temporal information, to obtain Input into the sigmoid layer, and calculate the gradients of the parameters in the deep learning network model G according to the target loss function loss sp , and update the parameters using the stochastic gradient descent method;

[0021] Step B8: Based on the target loss functions loss main , loss ra , loss sp obtain the total target loss function loss, and calculate the gradients of the parameters in the deep learning network model G according to the total target loss function loss, and update the parameters using the stochastic gradient descent method; when the iterative change of the loss value generated by the deep learning network model G is less than the set threshold and no longer decreases or reaches the maximum number of iterations, terminate the training of the deep learning network model G.

[0022] Furthermore, the specific steps of the step B1 include the following steps:

[0023] Step B11: Traverse the training set UB, and each training sample in UB is represented as D = (c, r, y). Perform word segmentation on c and r in the training sample D, remove stop words and add new words; where c is the context in the dialogue, r is the reply corresponding to this context, and y is the label of this reply, and the labels include {0, 1}, 0 means that r is not the reply corresponding to c, and 1 means that r is the reply corresponding to c;

[0024] After the context c is word-segmented and stop words are removed, add the new word [CLS] at the front of the context, which is expressed as:

[0025]

[0026] where is the i-th word in the remaining words after the context c is word-segmented, stop words are removed and new words are added, i = 1, 2,..., n, and n is the number of remaining words after the context c is word-segmented, stop words are removed and new words are added;

[0027] After the reply r is word-segmented and stop words are removed, it is expressed as:

[0028]

[0029] where represents the i-th word in the remaining words after the reply r is word-segmented, stop words are removed and new words are added, i = 1, 2,..., m, and m is the number of remaining words after the reply r is word-segmented, stop words are removed and new words are added;

[0030] Step B12: The context obtained in step B11 after word segmentation, removal of stop words and addition of new words Encode and get the initial representation vector of context c

[0031] in, It is expressed as:

[0032]

[0033] in, For the i-th word The corresponding word vector is obtained by pre-training the word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

[0034] Step B13: Response to step B11 after word segmentation, removal of stop words, and addition of new words Encode and get the initial representation vector of the reply r

[0035]

[0036] in, Represents the i-th word The corresponding word vector is obtained by pre-training the word vector matrix It can be found by searching in , where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

[0037] Furthermore, the step B2 specifically includes the following steps:

[0038] Step B21: Return the response from step B11 to Data enhancement is performed by using two strategies: shuffling or deleting data with a certain probability:

[0039]

[0040] in, It represents the response after random ordering of word granularity. represents the i-th word in the reply r, i = 1, 2, ..., m, and m is the number of words in the reply r after random shuffling;

[0041]

[0042] in, Represents the response after random deletion of word granularity, Denote the $i$-th word in the response $r$, where $i = 1, 2, \ldots, n$, $n \lt m$, and $n$ is the number of words in the response $r$ after random deletion;

[0043] Obtain the response after data augmentation:

[0044]

[0045] where $\|$ represents data augmentation of the response $r$ by scrambling or deleting with a certain probability;

[0046] Step B22: Concatenate the response obtained in Step B21 after the context to get $c1=\{u1, u2, \ldots, u n , r1\}$, add the word $[CLS]$ in front of $u1$ to get $c1=\{[CLS], u1, u2, \ldots, u n , r1\}$, and then encode $c1$ at the word granularity level to obtain the initial representation vectors of the context and the augmented response

[0047]

[0048] where denotes the word vector corresponding to the $i$-th word in , where $i = 1, 2, \ldots, p$, and $p$ is the number of word vectors in obtained by looking up in the pre-trained word vector matrix

[0049] Furthermore, Step B3 specifically includes the following steps:

[0050] Step B31: Merge the context in Step B11 at the utterance level to get $c2=\{u1, u2, \ldots, u n \}$, where $n$ represents the number of utterances in the context; randomly intercept a segment of the conversation in the context $c2$ as the local context for the span prediction task:

[0051] c2'=\{u i , u i+1 , \ldots, u i+m-1 \}

[0052] where $u i represents the $i$-th sentence in the context $c2$, $i\in[1, n - m + 1]$, $m$ represents the number of sentences in the conversation segment, and $m\in(0, n)$;

[0053] In $c2'=\{u i , u i+1,...,u i+m-1 Add the new word [CLS] at the beginning of {...} to obtain:

[0054]

[0055] Wherein, represents the i-th word in c2' after adding the new word, i = 1, 2,... l, and l is the number of words in c2';

[0056] Step B32: Encode the context c2' of the dialogue segment intercepted in Step B31 to obtain the initial representation vector of the dialogue segment

[0057]

[0058] Wherein, represents in the corresponding word vector, i = 1, 2,..., l, and l is the number of word vectors in; obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the dictionary V;

[0059] Use the reply obtained in Step B11 as the reply r2 of the dialogue segment c2' and encode it to obtain the initial representation vector of the reply corresponding to the dialogue segment

[0060]

[0061] Wherein, represents the word vector corresponding to the i-th word i = 1, 2,..., m, and m represents the number of word vectors in; obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

[0062] Furthermore, Step B4 specifically includes the following steps:

[0063] Step B41: Input the representation vector and the representation vector into the pre-trained language model BERT and output h 1 :

[0064]

[0065] Wherein, is the output of the i-th word vector of the pre-trained language model BERT, The calculation formula is as follows:

[0066]

[0067] Where x is the context representation vector and each word vector in the response representation vector ;

[0068] Step B42: Extract the representation vector of the global semantic information [CLS] optimized by two auxiliary tasks

[0069]

[0070] Furthermore, the specific steps of step B5 are as follows:

[0071] Step B51: Input the result obtained in step B42 into a fully connected layer for dimensionality reduction, and calculate the probability distribution through the sigmoid activation function. The calculation process is as follows:

[0072]

[0073] Where W is the trainable parameter, b is the bias, and σ(·) is the sigmoid activation function;

[0074] Step B52: Input the probability distribution g(c,r) obtained in step B51 into the loss function to calculate the loss and perform iterative updates. Cross-entropy is used as the loss function during the fine-tuning process of the response selection main task:

[0075]

[0076] Use the gradient accumulation strategy to update the gradient after accumulating 5 batches; The model uses the AdamW algorithm as the optimizer for gradient descent.

[0077] Furthermore, the specific steps of step B6 are as follows:

[0078] Step B61: Input the context obtained in step B22 and the response after data augmentation into the pre-trained language model BERT for encoding:

[0079]

[0080] Step B62: Extract the representation vector of the [CLS] label in the vector after BERT output

[0081]

[0082] Among them Indicates taking out The vector of the first dimension in the vector;

[0083] The obtained Is input into the linear layer. According to the target loss function loss ra , the gradients of the parameters in the deep learning network model G are calculated using the backpropagation method, and the parameters are updated using the stochastic gradient descent method;

[0084]

[0085] Among them, W1 is a trainable parameter, b1 is a bias, and σ(·) is a sigmoid activation function; the cross-entropy is used as the loss function to calculate the loss value, the learning rate is updated through the gradient optimization algorithm Adam, and the model parameters are iteratively updated using backpropagation to train the model by minimizing the loss function;

[0086] Among them, the calculation formula for minimizing the loss function Loss is as follows:

[0087]

[0088] Furthermore, the specific steps of step B7 are as follows:

[0089] Step B71: Concatenate the initial representation vector of the dialogue segment obtained in step B32 And the reply representation vector corresponding to the obtained dialogue segment To obtain Where ⊕ represents an operation of concatenation;

[0090] Input Into the pre-trained model to obtain

[0091]

[0092] Then Is input into the LSTM network layer for fusion learning of temporal information, and the output after the final state hidden layer of the LSTM is obtained The update state of the LSTM network layer at the current moment is as follows:

[0093]

[0094]

[0095]

[0096]

[0097] c t = f t ⊙ c t-1 + i t ⊙ g t

[0098]

[0099] Then input into the fully connected layer, and then calculate the final score after passing through the sigmoid function;

[0100]

[0101] where W2 is a trainable parameter, b2 is a bias, and σ(·) is the sigmoid activation function;

[0102] Step B72: Use cross-entropy as the loss function to calculate the loss value, update the learning rate through the gradient optimization algorithm Adam, and use backpropagation to iteratively update the model parameters to train the model by minimizing the loss function;

[0103] Among them, the calculation formula for minimizing the loss function Loss is as follows:

[0104]

[0105] Furthermore, in step B8, while fine-tuning, the pre-trained language model BERT uses a multi-task joint training framework. Export the parameters of the pre-trained language model BERT, and learn the semantic information contained in the response discourse and the correlation between the local context and the response discourse during the optimization process of the two auxiliary tasks; optimize the two auxiliary tasks and a main task simultaneously during the fine-tuning process. Therefore, the total objective loss function of the deep learning network model is as follows:

[0106] Loss = Loss main + αLoss ra + βLoss sp

[0107] where α and β are two hyperparameters, which are respectively used to control the influence of the two auxiliary tasks on the multi-turn dialogue model.

[0108] The present invention also provides a multi-turn dialogue system adopting the above method, including:

[0109] A data collection module, which is used to extract user conversations, responses involved in the user conversations, and label the tags of relevant response discourses in the conversations to construct a training set;

[0110] A preprocessing module for preprocessing training samples in a training set, including word segmentation, stop word removal, and new word addition;

[0111] An encoding module for looking up word vectors in a pre-trained word vector matrix for the preprocessed dialogue context to obtain an initial representation vector of the context and an initial representation vector of the response;

[0112] A network training module for inputting the initial representation vectors of the context and the response obtained by processing different data into a deep learning network. The pre-trained language model in the deep learning network shares parameters to obtain a response after data augmentation and multi-granularity representation vectors of the context and the response, and uses these to train the deep learning network. Using the probability that the representation vector belongs to a certain category and the annotation in the training set as the loss, the entire deep learning network is trained with the goal of minimizing the loss to obtain a deep learning network model that jointly optimizes response enhancement and span prediction;

[0113] A response selection module that uses the trained deep learning network model that jointly optimizes response enhancement and span prediction to analyze and process the input dialogue context and candidate responses, and outputs the response that best matches the current context content among the candidate responses.

[0114] Compared with the prior art, the present invention has the following beneficial effects: It provides a multi-turn dialogue method and system that jointly optimize response enhancement and span prediction. The method and system use a multi-task learning framework and utilize the response enhancement task to rewrite the utterances in the response through two strategies, thereby enhancing the model's learning of the semantic information contained in the response utterances. Aiming at the problem of lack of learning of the correlation between local context and response, a span prediction task is proposed. This task randomly selects a dialogue segment from the context and then predicts the correctness of the dialogue response through this dialogue segment, enabling the model to learn the local semantic information related to the response in the context. Finally, classification is performed using the new [CLS] representation vector that has learned the response utterance information and the temporal information in the local context, which can effectively improve the accuracy of response selection. BRIEF DESCRIPTION OF THE DRAWINGS

[0115] Figure 1 is a flowchart of the method according to an embodiment of the present invention;

[0116] Figure 2 is an architecture diagram of the deep learning network model according to an embodiment of the present invention;

[0117] Figure 3 is a schematic structural diagram of the system according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0118] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.

[0119] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs.

[0120] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0121] As Figure 1 shown, this embodiment provides a multi-round dialogue method for jointly optimizing response enhancement and span prediction, including the following steps:

[0122] Step A: Extract the user dialogue, the responses involved in the user dialogue, and label the tags of the relevant response words involved in the user dialogue. The positive samples are the correct responses in the dialogue, and the negative samples are the incorrect responses, to construct the training set UB.

[0123] Step B: Use the training set UB to train the deep learning network model G for jointly optimizing response enhancement and span prediction, which is used to learn the local semantic relationships in the user dialogue and the responses involved in the user dialogue, and at the same time learn the content in the responses involved in the user dialogue. The architecture of the deep learning network model G is as Figure 2 shown.

[0124] In this embodiment, step B specifically includes the following steps:

[0125] Step B1: Encode each training sample in the training set UB to obtain the initial representation vector of the context the initial representation vector of the candidate response and the label y.

[0126] In this embodiment, step B1 specifically includes the following steps:

[0127] Step B11: Traverse the training set UB. Each training sample in UB is represented as D=(c, r, y). Tokenize c and r in the training sample D, remove stop words and add new words; where c is the context in the dialogue, r is the response corresponding to this context, and y is the label of this response. The labels include {0, 1}, 0 indicating that r is not the response corresponding to c, and 1 indicating that r is the response corresponding to c.

[0128] After the context c is segmented and stop words are removed, the new word [CLS] is added at the beginning of the context, which is expressed as:

[0129]

[0130] Among them, is the i-th word among the remaining words after the context c is segmented, stop words are removed, and new words are added. i = 1, 2,..., n, where n is the number of remaining words after the context c is segmented, stop words are removed, and new words are added.

[0131] After the response r is segmented and stop words are removed, it is expressed as:

[0132]

[0133] Among them, represents the i-th word among the remaining words after the response r is segmented, stop words are removed, and new words are added. i = 1, 2,..., m, where m is the number of remaining words after the response r is segmented, stop words are removed, and new words are added.

[0134] Step B12: Encode the context obtained in Step B11 after being segmented, stop words are removed, and new words are added to obtain the initial representation vector of the context c

[0135] Among them, is expressed as:

[0136]

[0137] Among them, is the word vector corresponding to the i-th word obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

[0138] Step B13: Encode the response obtained in Step B11 after being segmented, stop words are removed, and new words are added to obtain the initial representation vector of the response r

[0139]

[0140] Among them, represents the word vector corresponding to the i-th word obtained by looking up in the pre-trained word vector matrix found in, where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

[0141] Step B2: For each training positive sample in the training set UB, perform data augmentation on the word-level vectors in the reply through two operations: randomly deleting or randomly shuffling, to obtain the initial representation vector of the reply after data augmentation. Use the initial representation vector of the context in Step B1 as the context corresponding to the augmented reply, to obtain the corresponding context representation vector

[0142] In this embodiment, Step B2 specifically includes the following steps:

[0143] Step B21: Perform data augmentation on the reply in Step B11 according to a certain probability using two strategies: shuffling or deleting: where

[0144]

[0145] represents the reply obtained after randomly shuffling at the word level, where represents the i-th word in the reply r, i = 1, 2,..., m, and m is the number of words in the reply r after random shuffling.

[0146]

[0147] where represents the reply obtained after randomly deleting at the word level, where

[0148] represents the i-th word in the reply r, i = 1, 2,..., n, n < m, and n is the number of words in the reply r after random deletion.

[0149]

[0150] where || represents performing data augmentation on the reply r according to a certain probability using two strategies: shuffling or deleting.

[0151] Step B22: Concatenate the reply obtained in Step B21 behind the context to get c1 = {u1, u2,..., u n , r1}, add the word [CLS] in front of u1 to get c1 = {[CLS], u1, u2,..., u n , r1}, and then encode c1 at the word level to obtain the initial representation vector of the context and the augmented reply.

[0152]

[0153] Among them, denotes the word vector corresponding to the i-th word in , where i = 1, 2,.., p, and p is the number of word vectors in

[0154] Step B3: For each training sample in the training set UB, randomly intercept a dialogue segment from the context to obtain the initial representation vector of the dialogue segment Use the candidate response corresponding to the original context as the response of the obtained dialogue segment to obtain the initial representation vector of the candidate response corresponding to the dialogue segment

[0155] In this embodiment, the specific steps of step B3 are as follows:

[0156] Step B31: Merge the context in step B11 at the utterance level to obtain c2 = {u1, u2,..., u n}, where n represents the number of utterances in the context; randomly intercept a segment of dialogue in the context c2 as the local context for the span prediction task:

[0157] c2' = {u i , u i+1 ,..., u i+m-1}

[0158] where u i represents the i-th sentence in the context c2, i ∈ [1, n - m + 1], m represents the number of utterances in the dialogue segment, and m ∈ (0, n).

[0159] Add the new word [CLS] at the front of c2' = {u i , u i+1 ,..., u i+m-1} to obtain:

[0160]

[0161] where denotes the i-th word in c2' after adding the new word, i = 1, 2,…l, and l is the number of words in c2';.

[0162] Step B32: Encode the dialogue segment context c2' intercepted in step B31 to obtain the initial representation vector of the dialogue segment

[0163]

[0164] Among them, denotes in the corresponding word vector, where \(i = 1, 2, \cdots, l\), and \(l\) is the number of word vectors in It is obtained by looking up in the pre-trained word vector matrix

[0165] Take the reply obtained in step B11 as the reply \(r2\) of the dialogue segment \(c2'\) and encode it to obtain the initial representation vector of the reply corresponding to the dialogue segment

[0166]

[0167] Among them, denotes the word vector corresponding to the \(i\)-th word where \(i = 1, 2, \cdots, m\), and \(m\) represents the number of word vectors in It is obtained by looking up in the pre-trained word vector matrix

[0168] Step B4: Input the representation vectors and into the pre-trained language model to obtain the representation vector of the optimized global semantic information

[0169] In this embodiment, step B4 specifically includes the following steps:

[0170] Step B41: Input the representation vector and the representation vector into the pre-trained language model BERT and output \(h\) 1 :

[0171]

[0172] Among them, is the output of the \(i\)-th word vector of the pre-trained language model BERT, The calculation formula of

[0173]

[0174] Among them, \(x\) is the context representation vector and the reply representation vector Each word vector in

[0175] Step B42: Extract the representation vector of the global semantic information [CLS] optimized by two auxiliary tasks

[0176]

[0177] Step B5: Input into the sigmoid layer, and according to the objective loss function loss main , use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters.

[0178] In this embodiment, the specific steps of Step B5 are as follows:

[0179] Step B51: Input the result obtained in Step B42 into a fully connected layer for dimensionality reduction, and calculate the probability distribution through the sigmoid activation function. The calculation process is as follows:

[0180]

[0181] where W is a trainable parameter, b is a bias, and σ(·) is the sigmoid activation function.

[0182] Step B52: Input the probability distribution g(c,r) obtained in Step B51 into the loss function to calculate the loss and perform iterative update. Cross-entropy is used as the loss function during the fine-tuning process of the response selection main task:

[0183]

[0184] Use the gradient accumulation strategy to update the gradient after accumulating 5 batches; the model uses the AdamW algorithm as the optimizer for gradient descent.

[0185] Step B6: Concatenate the initial context representation vector and the initial representation vector of the response after data augmentation and input them into the pre-trained language model to obtain the label vector representing the global semantic information Input into the sigmoid layer, and according to the objective loss function loss ra , use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters.

[0186] In this embodiment, the specific steps of Step B6 are as follows:

[0187] Step B61: Input the context obtained in Step B22 and the reply after data augmentation into the pre-trained language model BERT for encoding:

[0188]

[0189] Step B62: Take out the representation vector of the [CLS] label in the vector

[0190]

[0191] where denotes taking out the vector of the first dimension in the vector.

[0192] Input the obtained into the linear layer. According to the target loss function loss ra , calculate the gradients of the parameters in the deep learning network model G using the backpropagation method, and update the parameters using the stochastic gradient descent method.

[0193]

[0194] where W1 is a trainable parameter, b1 is a bias, σ(·) is the sigmoid activation function; use cross-entropy as the loss function to calculate the loss value, update the learning rate through the gradient optimization algorithm Adam, and iteratively update the model parameters using backpropagation to train the model by minimizing the loss function.

[0195] where the calculation formula for minimizing the loss function Loss is as follows:

[0196]

[0197] Step B7: Concatenate the initial representation vector of the dialogue segment and the reply representation vector corresponding to the dialogue segment and input the concatenated result into the pre-trained language model to obtain Input into the LSTM network layer for learning temporal information to obtain Input into the sigmoid layer. According to the target loss function loss sp , calculate the gradients of the parameters in the deep learning network model G using the backpropagation method, and update the parameters using the stochastic gradient descent method.

[0198] In this embodiment, Step B7 specifically includes the following steps:

[0199] Step B71: The initial representation vector of the dialogue segment obtained in Step B32 and the response representation vector corresponding to the obtained dialogue segment are concatenated to obtain where ⊕ represents a concatenation operation.

[0200] Input into the pre-trained model to obtain

[0201]

[0202] Then input into the LSTM network layer for fusion learning of temporal information to obtain the output after the final state hidden layer of LSTM The updated state of the LSTM network layer at the current moment is as follows:

[0203]

[0204]

[0205]

[0206]

[0207] c t = f t ⊙ c t-1 + i t ⊙ g t

[0208]

[0209] Then input into the fully connected layer and calculate the final score after passing through sigmoid.

[0210]

[0211] where W2 is a trainable parameter, b2 is a bias, and σ(·) is the sigmoid activation function.

[0212] Step B72: Use cross-entropy as the loss function to calculate the loss value, update the learning rate through the gradient optimization algorithm Adam, and iteratively update the model parameters using backpropagation to train the model by minimizing the loss function.

[0213] Among them, the calculation formula for minimizing the loss function Loss is as follows:

[0214]

[0215] Step B8: Based on the target loss function loss main 、loss ra 、loss sp Obtain the total target loss function loss, and according to the total target loss function loss, use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters; when the iterative change of the loss value generated by the deep learning network model G is less than the set threshold and no longer decreases or reaches the maximum number of iterations, terminate the training of the deep learning network model G.

[0216] In the step B8, while fine-tuning, the pre-trained language model BERT uses a multi-task joint training framework, exports the parameters of the pre-trained language model BERT, and learns the semantic information contained in the response discourse and the correlation between the local context and the response discourse during the optimization of the two auxiliary tasks; during the fine-tuning process, optimize the two auxiliary tasks and a main task at the same time. Therefore, the total target loss function of the deep learning network model is as follows:

[0217] Loss = Loss main + αLoss ra + βLoss sp

[0218] where α, α are two hyperparameters, which are respectively used to control the influence of the two auxiliary tasks on the multi-turn dialogue model.

[0219] As Figure 3 shown, this embodiment also provides a multi-turn dialogue system adopting the above method, including: a data collection module, a preprocessing module, an encoding module, a network training module, and a response selection module.

[0220] The data collection module is used to extract user conversations, responses involved in user conversations, and label the tags of relevant response discourses in the conversations to construct a training set.

[0221] The preprocessing module is used to preprocess the training samples in the training set, including word segmentation, stop word removal, and new word addition.

[0222] The encoding module is used to find the word vectors in the preprocessed dialogue context in the pre-trained word vector matrix to obtain the initial representation vectors of the context and the initial representation vectors of the response.

[0223] The network training module is used to input the initial representation vectors of the context and the initial representation vectors of the responses obtained by processing different data into a deep learning network. The pre-trained language models in the deep learning network share parameters to obtain enhanced responses after data augmentation, and multi-granularity representation vectors of the context and the responses, and then train the deep learning network with these. Using the probability that the representation vector belongs to a certain category and the annotations in the training set as the loss, the entire deep learning network is trained with the goal of minimizing the loss to obtain a deep learning network model that jointly optimizes response enhancement and span prediction.

[0224] The response selection module uses the deep learning network model that jointly optimizes response enhancement and span prediction, which has been trained, to analyze and process the input conversation context and candidate responses, and outputs the response that best matches the current context content among the candidate responses.

[0225] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.

[0226] The present application is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0227] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including an instruction device that implements the functions specified in Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0228] These computer program instructions can also be loaded onto a computer or other programmable data processing apparatus, so that a series of operation steps are executed on the computer or other programmable apparatus to produce a computer-implemented process, thereby the instructions executed on the computer or other programmable apparatus provide steps for implementing the functions specified in one process or a plurality of processes and / or boxes. Figure 1 One process or a plurality of processes and / or boxes Figure 1 steps for implementing the functions specified in one box or a plurality of boxes.

[0229] As described above, it is only the preferred embodiment of the present invention, and is not intended to limit the present invention to other forms. Any technicians familiar with the art may use the disclosed technical content to make changes or modifications into equivalent embodiments with equivalent changes. However, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the technical solution content of the present invention still belong to the protection scope of the technical solution of the present invention.

Claims

1. A multi-round dialogue method that jointly optimizes response enhancement and span prediction, characterized in that, It includes the following steps: Step A: Extract the user conversations, the responses involved in the user conversations, and label the tags of the relevant response words involved in the user conversations. The positive samples are the correct responses in the conversations, and the negative samples are the incorrect responses, to construct the training set UB; Step B: Use the training set UB to train the deep learning network model G that jointly optimizes response enhancement and span prediction, which is used to learn the local semantic relationships in the user conversations and the responses involved in the user conversations, and at the same time learn the content in the responses involved in the user conversations; Step C: Input the complete user conversation and the responses involved in the user conversation into the trained deep learning network model G to obtain the correct response for the complete user conversation; The specific steps of Step B include the following steps: Step B1: Encode each training sample in the training set UB to obtain the initial representation vector of the context The initial representation vector of the candidate response and the label y; Step B2: For each training positive sample in the training set UB, perform data augmentation on the word-level vectors in the response through two operations: randomly deleting or randomly shuffling, to obtain the initial representation vector of the response after data augmentation Use the initial representation vector of the context in Step B1 As the context corresponding to the augmented response, obtain The corresponding context representation vector Step B3: For each training sample in the training set UB, randomly intercept a dialogue segment from the context to obtain the initial representation vector of the dialogue segment Use the candidate response corresponding to the original context as the response of the obtained dialogue segment to obtain the initial representation vector of the candidate response corresponding to the dialogue segment Step B4: Input the characterization vectors and into the pre-trained language model to obtain the characterization vector of the optimized global semantic information Step B5: Input into the sigmoid layer, and calculate the gradients of the parameters in the deep learning network model G according to the target loss function loss main , and update the parameters using the stochastic gradient descent method; Step B6: The initial context representation vector and the initial representation vector of the reply after data augmentation are concatenated and then input into the pre-trained language model to obtain a label vector representing global semantic information Input into the sigmoid layer. According to the objective loss function loss ra , use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters; Step B7: Concatenate the initial representation vector of the dialogue segment and the response representation vector corresponding to the dialogue segment and input the concatenated result into the pre-trained language model to obtain Input into the LSTM network layer to learn temporal information and obtain Input into the sigmoid layer. According to the target loss function loss sp , use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters; Step B8: Based on the target loss functions loss main 、loss ra 、loss sp Obtain the total target loss function loss, and according to the total target loss function loss, use the backpropagation method to calculate the gradients of the parameters in the deep learning network model G, and use the stochastic gradient descent method to update the parameters; when the iterative change of the loss value generated by the deep learning network model G is less than the set threshold and no longer decreases or reaches the maximum number of iterations, terminate the training of the deep learning network model G; In Step B8, while fine-tuning, the pre-trained language model BERT uses a multi-task joint training framework, exports the parameters of the pre-trained language model BERT, and learns the semantic information contained in the response words and the correlation between the local context and the response words during the optimization process of the two auxiliary tasks; during the fine-tuning process, two auxiliary tasks and one main task are optimized simultaneously. Therefore, the total objective loss function of the deep learning network model is as follows: Loss=Loss main +αLoss ra +βLoss sp Where α and β are two hyperparameters, which are used to control the influence of the two auxiliary tasks on the multi-turn dialogue model respectively.

2. The multi-turn dialogue method for jointly optimizing response enhancement and span prediction according to claim 1, characterized in that The specific steps of Step B1 include the following steps: Step B11: Traverse the training set UB. Each training sample in UB is represented as D=(c, r, y). Tokenize c and r in the training sample D, remove stop words and add new words; where c is the context in the conversation, r is the response corresponding to this context, and y is the label of this response. The labels include {0, 1}, 0 means r is not the response corresponding to c, and 1 means r is the response corresponding to c; After the context c is tokenized and stop words are removed, add the new word [CLS] at the very front of the context, which is represented as: Among them, is the i-th word among the remaining words after the context c is segmented, stop words are removed, and new words are added, where i = 1, 2,..., n, and n is the number of remaining words after the context c is segmented, stop words are removed, and new words are added; After the response r is tokenized and stop words are removed, it is represented as: Among them, represents the i-th word among the remaining words in the reply r after word segmentation, stop word removal, and new word addition, where i = 1, 2,..., m, and m is the number of remaining words in the reply r after word segmentation, stop word removal, and new word addition; Step B12: Encode the context obtained in Step B11 after word segmentation, stop word removal, and new word addition to obtain the initial representation vector of context c Among them, It is expressed as: Among them, is the word vector corresponding to the i-th word obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the vocabulary V; Step B13: For the response obtained in Step B11 after word segmentation, stop word removal, and new word addition perform encoding to obtain the initial representation vector of the response r Among them, represents the word vector corresponding to the i-th word, which is obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

3. The multi-turn dialogue method for jointly optimizing response enhancement and span prediction according to claim 2, characterized in that, The specific steps of Step B2 include the following steps: Step B21: The response in Step B11 is enhanced by scrambling or deleting with a certain probability according to two strategies: Among them, represents the response obtained after random scrambling at the word granularity, represents the i-th word in the response r, where i = 1, 2,..., m, and m is the number of words in the response r after random scrambling; Among them, represents the response obtained after random deletion at the word granularity, represents the $i$-th word in the response $r$, where $i = 1, 2, \ldots, n$, $n < m$, and $n$ is the number of words in the response $r$ after random deletion; Obtain the response after data augmentation: Where || means that the response r is augmented by scrambling or deleting according to a certain probability; Step B22: Concatenate the response obtained in Step B21 to the context to obtain c1 = {u1, u2,..., u n , r1}, add the word [CLS] in front of u1 to obtain c1 = {[CLS], u1, u2,..., u n , r1}, and then encode c1 at the word granularity level to obtain the initial representation vectors of the context and the enhanced response Among them, denotes the word vector corresponding to the i-th word in where i = 1, 2,.., p, and p is the number of word vectors in obtained by looking up in the pre-trained word vector matrix where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

4. The multi-turn dialogue method for jointly optimizing response enhancement and span prediction according to claim 3, wherein The specific steps of Step B3 include the following steps: Step B31: Merge the context in Step B11 at the discourse level to obtain c2 = {u1, i2,..., i n}, where n represents the number of discourses in the context; randomly intercept a segment of dialogue in the context c2 as the local context for the span prediction task: c2 ′ = {u i , u i+1 ,..., u i+m-1} where u i represents the i-th sentence in context c2, i ∈ [1, n - m + 1], m represents the number of utterances in the dialogue segment, m ∈ (0, n); At c2 ′ = {u i , u i+1 ,..., u i+m-1} Add the new word [CLS] at the very front to get: Among them, represents the i-th word in c2 after adding new words, where i = 1, 2, …, l, and l is the number of words in c2 ′ ; ′ is the number of words in Step B32: Encode the context c2 of the dialogue segment intercepted in Step B31 ′ to obtain the initial representation vector of the dialogue segment Among them, denotes in the corresponding word vector, where \(i = 1, 2, \ldots, l\), and \(l\) is the number of word vectors in; obtained by looking up in the pre-trained word vector matrix where \(d\) represents the dimension of the word vector and \(|V|\) is the number of words in the dictionary \(V\); Use the response obtained in step B11 as dialogue segment c2 ′ and encode the response r2 to obtain the initial representation vector corresponding to the response of the dialogue segment in, Represents the i-th word The corresponding word vector, i = 1, 2, ..., m, m represents The number of word vectors in the pre-trained word vector matrix It can be found by searching in , where d represents the dimension of the word vector and |V| is the number of words in the dictionary V.

5. The multi-turn dialogue method for jointly optimizing response enhancement and span prediction according to claim 4, wherein The specific steps of Step B4 include the following steps: Step B41: Input the characterization vector and the characterization vector into the pre-trained language model BERT, and output h 1 : Among them, is the output of the i-th word vector of the pre-trained language model BERT, The calculation formula of is as follows: where x is the context representation vector and the response representation vector for each word vector; Step B42: Extract the representation vector of the global semantic information [CLS] optimized by two auxiliary tasks 6. The multi-turn dialogue method for jointly optimizing response enhancement and span prediction according to claim 5, characterized in that The specific steps of Step B5 include the following steps: Step B51: Input the result obtained in step B42 into a fully connected layer for dimensionality reduction, and calculate the probability distribution through the sigmoid activation function. The calculation process is as follows: Where W is a trainable parameter, b is a bias, and σ(·) is the sigmoid activation function; Step B52: Input the probability distribution g(c, r) obtained in Step B51 into the loss function to calculate the loss and perform iterative updates. Use cross-entropy as the loss function during the fine-tuning process of the response selection main task: Use the gradient accumulation strategy to update the gradient after accumulating 5 batches; the model uses the AdamW algorithm as the optimizer for gradient descent.

7. The multi-turn dialogue method for jointly optimizing response enhancement and span prediction according to claim 6, characterized in that The specific steps of Step B6 include the following steps: Step B61: Input the context obtained in step B22 and the reply after data augmentation into the pre-trained language model BERT for encoding: Step B62: Take out the representation vector of the [CLS] label in the vector after the BERT output in the vector Among them means taking out the vector of the first dimension in the vector; The obtained is input into the linear layer, and according to the target loss function loss ra , the gradients of the parameters in the deep learning network model G are calculated using the backpropagation method, and the parameters are updated using the stochastic gradient descent method; where W1 is a trainable parameter, b1 is a bias, and σ(·) is the sigmoid activation function; the cross-entropy is used as the loss function to calculate the loss value, the learning rate is updated through the gradient optimization algorithm Adam, and the model parameters are iteratively updated using backpropagation to train the model by minimizing the loss function; where the calculation formula for minimizing the loss function Loss is as follows: The step B7 specifically includes the following steps: Step B71: The initial representation vector of the dialogue segment obtained in step B32 and the reply representation vector corresponding to the obtained dialogue segment are concatenated to obtain where represents a cascading operation; Input into the pre-trained model to obtain Then is input into the LSTM network layer for fusion learning of temporal information, and the output after the final state hidden layer of LSTM is obtained. The updated state of the LSTM network layer at the current moment is as follows: c t = f t ⊙ c t-1 + i t ⊙ g t Then input it into the fully connected layer and then calculate the final score after passing through sigmoid; where W2 is a trainable parameter, b2 is a bias, and σ(·) is the sigmoid activation function; Step B72: Use the cross-entropy as the loss function to calculate the loss value, update the learning rate through the gradient optimization algorithm Adam, and iteratively update the model parameters using backpropagation to train the model by minimizing the loss function; where the calculation formula for minimizing the loss function Loss is as follows:

8. A multi-turn dialogue system using the method according to any one of claims 1-7, characterized in that, including: A data collection module for extracting user conversations, the responses involved in the user conversations, and labeling the tags of the relevant response words in the conversations to construct a training set; A preprocessing module for preprocessing the training samples in the training set, including word segmentation, stop word removal, and new word addition; An encoding module for looking up the word vectors in the preprocessed conversation context in the pre-trained word vector matrix to obtain the initial representation vectors of the context and the initial representation vector of the response; A network training module for inputting the initial representation vectors of the context and the response obtained by processing different data into a deep learning network. The pre-trained language model in the deep learning network shares parameters to obtain the enhanced response after data augmentation and the multi-granularity representation vectors of the context and the response, and trains the deep learning network with these. Using the probability that the representation vector belongs to a certain category and the annotation in the training set as the loss, the entire deep learning network is trained with the goal of minimizing the loss to obtain a deep learning network model that jointly optimizes response enhancement and span prediction; A response selection module that uses the trained deep learning network model that jointly optimizes response enhancement and span prediction to analyze and process the input conversation context and candidate responses, and outputs the response that best matches the current context content among the candidate responses.

Citation Information

Patent Citations

  • Local information perception dialogue method and system based on pre-training language model

    CN114443827A

  • Dialogue generation model training method and device, dialogue generation method and electronic equipment

    CN114492465A