A task-oriented dialogue method for contrastive learning enhanced dialogue state tracking
Through comparative learning, the training of the enhanced dialogue state tracking module is solved, and the problem of insufficient semantic information extraction capability of task-based dialogue systems on multi-domain and multiple round data sets is achieved, and more powerful context semantic information extraction capability is achieved.
Patent Information
- Application Number
- CN202210408339.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-19
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2042-04-19
AI Technical Summary
The existing task-based dialogue system lacks semantic information extraction capabilities on multi-domain and multi-round dialogue datasets, resulting in poor performance in context semantic information extraction.
The method of enhancing dialogue state tracking is adopted to enhance dialogue state tracking. Through the intention recognition module, the binary slot recognition module and the Span slot comparison learning enhancement module, the training of the dialogue state tracking module is optimized, and the model's ability to extract context semantic information is improved.
It effectively improves the semantic information extraction capability of the dialogue system in multiple fields and multiple rounds, and improves the system's understanding and processing ability of the entire input statement context semantic information.
Smart Images

Figure CN114722177B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer natural language processing, and in particular, to a task-based dialogue method for enhancing dialogue state tracking through contrastive learning. Background Art
[0002] In recent years, human-computer dialogue systems have received extensive attention in both the academic and industrial fields. In research, dialogue language understanding technology has gradually developed in the direction of deep learning. Dialogue management has experienced a development process from rules to supervised learning and then to reinforcement learning. Natural language generation has evolved from template generation and sentence planning to end-to-end deep learning models. In applications, products based on human-computer dialogue technology have emerged in an endless stream. Thanks to the development of deep learning technology, the robustness and accuracy of current dialogue systems have been greatly improved, and major companies have started to develop their own dialogue systems. Among them, task-based dialogue systems are the type of dialogue systems with the greatest demand and the most applications in practice. However, currently, most products only have relatively accurate semantic understanding capabilities for single-round dialogues, and their performance on multi-domain and multi-turn dialogue datasets is often unsatisfactory. Ultimately, it is due to the insufficient ability of the dialogue state tracking module in the system to extract context semantic information. Summary of the Invention
[0003] The present invention provides a task-based dialogue method for enhancing dialogue state tracking through contrastive learning, which can effectively solve the problem of semantic information extraction in multi-domain and multi-turn dialogue systems and provide new ideas for subsequent engineering applications.
[0004] To solve the above problems, the present invention includes the following steps:
[0005] Step 1: Collect Chinese multi-domain task-based dialogue data and perform data augmentation based on the publicly available Chinese multi-domain task-based dialogue CrossWoz.
[0006] Step 2: Through the dialogue state tracking module, the collected and sorted data is processed by the intent recognition module, the binary classification slot recognition module, and the Span slot contrastive learning reinforcement module to obtain the user's intent and the dialogue state in each round of dialogue.
[0007] Step 3: Screen out the data that meets the user's intent and dialogue state through the database query module.
[0008] Step 4: Perform decision-making learning on the dialogue state through the dialogue state decision-making module to generate more reasonable and diverse dialogue decisions.
[0009] Step 5: Generate accurate and appropriate responses through the dialogue response module.
[0010] Step 6: Friendly display the response content to the user through the web front-end module.
[0011] Furthermore, in step one, when collecting Chinese multi-domain task-oriented dialogue data, data augmentation is specifically performed based on the publicly available Chinese multi-domain task-oriented dialogue dataset CrossWoz: Define the dialogue goal, involved domains, and relevant dialogue templates in advance, and manually fill in the relevant status information into the templates to complete the dialogue content.
[0012] Furthermore, in step two, when the collected and sorted data is used by the dialogue state tracking module to obtain the user's intention and the dialogue state in each round of dialogue through the intention recognition module, binary classification slot recognition module, and Span slot contrast learning enhancement module, it is specifically as follows:
[0013] (1) The intention recognition module takes the user's dialogue information as the input Context, constructs "Are you asking about {slot}?" as the input Query, and performs an attention mechanism operation in the form of a question and answer. The operation result completes the task of intention recognition through a binary classification fully connected layer;
[0014] (2) The binary classification slot recognition module recognizes the binary classification slots that cannot be extracted in the dialogue. It takes the user's dialogue information as the input Context, constructs "Do you have a need for {slot}?" as the input Query, and performs an attention mechanism operation in the form of a question and answer. The operation result completes the task of binary classification slot recognition through a binary classification fully connected layer;
[0015] (3) The contrast learning enhancement module adopts a two-stage joint training method, taking the dialogue information of the user and the system and the previous round of dialogue state as the input.
[0016] Furthermore, in step (3), the two stages of the contrast learning enhancement module include:
[0017] 1) Adopt multi-classification contrast learning to enhance the classification effects of the dialogue states None, Do not care, and State value match;
[0018] 2) The State value match classification extracts the dialogue state corresponding to the Span slot from the input statement through a pointer network.
[0019] Furthermore, in step three, when the database query module filters out the data that meets the user's intention and dialogue state, it is specifically as follows: After obtaining the user state and decision information, use an SQL query statement to query and filter out the data that meets the conditions.
[0020] Further, the decision learning of the dialogue state by the dialogue state decision module in step 4 to generate more reasonable and diverse dialogue decisions is specifically as follows: offline decision training is carried out through templates, and the decision-making process is continuously optimized through the reinforcement learning Q-learn algorithm offline.
[0021] Further, the generation of accurate and appropriate responses by the dialogue response module in step 5 is specifically as follows: a large number of template projects are used to convert the dialogue decision information into corresponding response texts.
[0022] Further, the friendly display of the response content to the user by the web front-end module in step 6 is specifically as follows: the visualization of the dialogue interface is realized by using html+css technology, and the flask framework based on python is used at the back end to formulate the interface for the interaction between the front end and the back end.
[0023] The present invention incorporates contrastive learning into the dialogue state tracking training, optimizes the model's ability to extract Span slots, and thus improves the model's ability to extract the context semantic information of the entire input sentence. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] Figure 1 It is the overall flowchart of the dialogue system of the method of the present invention;
[0025] Figure 2 It is the overall architecture diagram of the dialogue state tracking module in the method of the present invention;
[0026] Figure 3 It is the model architecture diagram of the intention recognition module in the method of the present invention;
[0027] Figure 4 It is the model architecture diagram of the binary classification slot recognition module in the method of the present invention;
[0028] Figure 5 It is the model architecture diagram of the contrastive learning enhancement module in the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0029] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further describes the implementation embodiments of the present invention in detail with reference to the drawings.
[0030] This embodiment includes the following steps:
[0031] Step 1: Collect Chinese multi-domain task-oriented dialogue data. Based on the publicly available Chinese multi-domain task-oriented dialogue dataset CrossWoz, data augmentation is performed as follows: Define the dialogue goals, involved domains, and relevant dialogue templates in advance, and manually fill in the relevant status information into the templates to complete the dialogue content. The extended CrossWoz dataset includes five domains: scenic spots, hotels, restaurants, taxis, and trains, with a total of 6,122 times and 100,435 rounds of dialogue, including 72 slots. The average dialogue lasts about 17 rounds, belonging to multi-domain and multi-round dialogue data.
[0032] Step 2: Use the dialogue state tracking module to obtain the user's intention and the dialogue state in each round of dialogue from the collected and sorted data through the intention recognition module, binary classification slot recognition module, and Span slot contrast learning enhancement module. As Figure 2 shown, the dialogue state tracking module consists of three sub-modules: the intention recognition module, the binary classification slot recognition module, and the contrast learning enhancement module, where:
[0033] (1) As Figure 3 shown, the intention recognition module takes the user's dialogue information as the input Context, constructs "Are you asking about {slot}?" as the input Query, converts the text information into a 768-dimensional vector through the Word Embed Layer, and then performs the attention operation on its own sentence through the Context Attention Layer. The operation result is used as the input of the Context-Question Attention Layer to obtain the high-order mutual information between Context and Query. Finally, binary classification is performed through the Prediction Layer, and the result is a yes or no answer to the Query;
[0034] (2) As Figure 4 shown, the binary classification slot recognition module takes the user's dialogue information as the input Context, constructs "Do you have a need for {slot}?" as the input Query, converts the text information into a 768-dimensional vector through the Word Embed Layer, and then performs the attention operation on its own sentence through the Context Attention Layer. The operation result is used as the input of the Context-Question Attention Layer to obtain the high-order mutual information between Context and Query. Finally, binary classification is performed through the Prediction Layer, and the result is a yes or no answer to the Query.
[0035] (3) The contrastive learning enhancement module adopts a two-stage joint training method, taking the dialogue information of the user and the system and the previous dialogue state as inputs.
[0036] For the training in the first stage, as Figure 5 shown in the upper part, the input of the contrastive learning enhanced dialogue state tracking module is dialogue information and dialogue state, that is, the input Among them, D t-1 and D t are the dialogue information of the (t - 1)-th round and the t-th round, and B t-1 is the dialogue state of the (t - 1)-th round. is a connector, indicating that the content before and after is in a context relationship. [CLS] is a flag token, and its operation result is used as the vector representation of the entire input information; A t represents the system response content in the t-th round, and U t represents the user response content in the t-th round; [SEP] is a flag token, and its operation result is used as the vector representation for separating the user content and the system content; As the input form of the dialogue state information, where [SLOT] is a flag token, and its operation result is used as the vector representation of a specific dialogue state, and S j is the selected slot, and its value is one of the 72 slots in the training set, such as Figure 5 the hotel name in it; V is the dialogue state corresponding to the selected slot, such as Figure 5 none in it. Immediately afterwards, the Transformer pre-trained model used is bert-base-uncased, and the hidden vector of the entire input statement after the self-attention operation is obtained through 8 Heads and 12 Layers corresponding to the hidden vectors of [CLS] and [SLOT] j That is:
[0037]
[0038] Furthermore, through a fully connected layer, the probabilities of None ( Figure 5 the classification number 0 in the first stage in it), Do not care (classification number 1, due to space reasons, Figure 5 this classification situation is not drawn in it), and Slot value match ( Figure 5 the classification number 2 in the first stage in it) are obtained, that is:
[0039]
[0040] Among them, W opr is the parameter illustration of the fully connected layer, and softmax is the activation function.
[0041] Furthermore, in the face of complex and diverse dialogue data, the F1 of classification in this stage is often not high, especially for the "Don't care" class with scarce data. In order to enhance its classification accuracy, multi-classification comparative learning is added to the training process in this stage. The classification accuracy is improved by bringing the same class closer and alienating different classes. The loss function used is the multi-classification InfoNCE. The positive sample is the data of the same class in the batch, and the negative sample is the data of different classes in the batch, that is:
[0042]
[0043] Among them, n is the number of categories, E represents the expectation, q is the training sample, k is the i is the comparison sample i, k corresponding to q + is a positive sample of q, τ is a hyperparameter, and K represents all training samples in the same training round during training.
[0044] For the second stage of training, Figure 5 As shown in the lower part, the second stage of training is performed on the dialogue state whose classification result is Slot value match, and its vector representation after the Transformer layer is extracted [SLOT] k , perform attention operation with the hidden vector of the entire input sequence, that is:
[0045]
[0046] Where T stands for transpose.
[0047] Furthermore, H′ t The probability of the start and end positions of the Span slot is obtained through the pointer network composed of the START and END fully connected layers, namely:
[0048]
[0049] Among them, W start is the parameter matrix of the START fully connected layer, W end is the parameter matrix of the END fully connected layer.
[0050] Finally, according to the probability vector and Find the element subscripts with the highest probability as the starting and ending points of the Span slot, and extract the dialogue state corresponding to the Span slot from the input sentence.
[0051] Step 3: The database query module converts the dialogue intention and dialogue status obtained in Step 2 above into a database query statement. For example, "Intent: Restaurant - Name, greet, Bi - slot: None, Span - slot: Restaurant - per capita consumption: 100 - 150 yuan, restaurant rating: above 4.5 points" is converted into the database query statement "SELECT Name from Restaurant WHERE per capita consumption BETWEEN 100 AND 150 AND rating > 4.5" to filter out the data that meets the conditions from the database.
[0052] Step 4: After obtaining the dialogue intention, dialogue status in Step 2 above and the data query result in Step 3, the dialogue decision - making module uses the Q - learn reinforcement learning algorithm to sample it and filters out the data that best meets the user's expectations.
[0053] Step 4: The dialogue reply module fills the result of Step 4 above into the module in the form of rules to generate the user reply message.
[0054] Step 5: The web front - end module aims to give the user a beautiful and clear display result of the dialogue information. In order to make the system more convenient and intuitive for the user to use, the present invention implements the visualization of the dialogue interface using html + css technology. On the back - end, the flask framework based on python is used to formulate the interface for the interaction between the front - end and the back - end, so as to connect the interaction between the interface and the underlying algorithm model. Flask is a lightweight web framework written in the python language. Developers can use this framework to quickly implement a high - performance and customizable web service, which is very suitable for the platform business of quickly building the entire dialogue system service. The main interface realizes the function of mutual transmission of text messages when the user and the robot are having a dialogue. At the same time, the underlying algorithm model will automatically extract different types of recognition results according to the semantics of the sentence and return them to the front - end for display.
[0055] Experimental data table
[0056]
[0057] Through experiments, on the CrossWOZ test set, the intention recognition module achieved an F1 of 93.28%, the binary - classification Slot recognition module achieved an F1 of 90.09%, and the dialogue Span slot extraction module using NER technology achieved an F1 of 98.56%. This shows that the present invention can effectively solve the problem of semantic information extraction in the dialogue system under multi - domain and multi - turn scenarios, providing new ideas for subsequent engineering applications.
Claims
1. A task-oriented dialogue method for contrastive learning enhanced dialogue state tracking, characterized in that the method comprises the following steps: Step 1, collect Chinese multi-domain task-oriented dialogue data and perform data enhancement on the basis of the dialogue data; Step 2, use the dialogue state tracking module to obtain the user's intention and the dialogue state in each round of dialogue through the intention recognition module, the binary classification slot recognition module, and the Span slot contrastive learning reinforcement module for the collected and sorted data; Step 3, screen out the data that meets the user's intention and dialogue state through the database query module; Step 4, perform decision-making learning on the dialogue state through the dialogue state decision-making module to generate more reasonable and diverse dialogue decisions; Step 5, generate accurate and appropriate responses through the dialogue response module; Step 6, friendly display the response content to the user through the web front-end module; The contrastive learning reinforcement module adopts a two-stage joint training method, taking the dialogue information of the user and the system and the previous round of dialogue state as inputs; The first stage: adopt multi-classification contrastive learning to enhance the classification effects of dialogue state None, Do not care, and State value match; The second stage: the State value match classification extracts the corresponding dialogue state of the Span slot from the input statement through a pointer network; In Step 2, the intention recognition module takes the user's dialogue information as the input Context, constructs "Are you asking about {slot}?" as the input Query, and performs attention mechanism operations in the form of a question and answer. The operation result completes the task of intention recognition through a binary classification fully connected layer; In Step 2, the binary classification slot recognition module recognizes the binary classification slots that cannot be extracted in the dialogue. It takes the user's dialogue information as the input Context, constructs "Do you have a need for {slot}?" as the input Query, and performs attention mechanism operations in the form of a question and answer. The operation result completes the task of binary classification slot recognition through a binary classification fully connected layer.
2. A task-oriented dialogue method for contrastive learning enhanced dialogue state tracking according to claim 1, characterized in that: The dialogue data described in Step 1 adopts the CrossWoz dataset.
3. A task-oriented dialogue method for contrastive learning enhanced dialogue state tracking according to claim 1, characterized in that: Step 3 is specifically: after obtaining the user state and decision-making information, query and screen out the qualified data through an SQL query statement.
4. A task-oriented dialogue method for contrastive learning enhanced dialogue state tracking according to claim 1, characterized in that: Step 4 is specifically: perform decision-making training offline through templates, and continuously optimize the decision-making process through the reinforcement learning Q-learn algorithm offline.
5. A task-oriented dialogue method for contrastive learning enhanced dialogue state tracking according to claim 1, characterized in that: Step 5 is specifically: adopt a large number of template projects to convert the dialogue decision-making information into corresponding response texts.
6. The task-based dialogue method for contrastive learning enhanced dialogue state tracking according to claim 1, characterized in that: Step six is specifically: The visualization of the dialogue interface is realized by using html+css technology, and an interface for front-end and back-end interaction is formulated using the flask framework based on python at the back-end.
Citation Information
Patent Citations
Dialogue intention recognition method and recognition system based on multi-task learning
CN112417894A
Dialogue intention recognition method and system based on feature matching and domain self-adaption
CN113672718A