Conversation state tracking method and system based on state correction
By using the deep learning network model in dialogue state tracking, the dialogue state is corrected and integrated, and the problem of insufficient accuracy of dialogue state tracking in the prior art is solved, and more efficient and personalized dialogue state tracking is achieved.
Patent Information
- Application Number
- CN202510126293.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-27
- Publication Date
- 2025-05-13
AI Technical Summary
The existing dialogue state tracking methods are insufficiently accurate when dealing with complex dialogues and the dialogue information is not expressed clearly enough, resulting in information redundancy and inconsistency.
The dialogue state tracking model based on the deep learning network is adopted, and the dialogue state in the input data is corrected, the groove information is extracted and fusion is performed, the context information is deeply perceived, and the dialogue state is dynamically adjusted to improve accuracy.
It significantly improves the accuracy of dialogue status tracking, can flexibly respond to changes in user intentions, provide more accurate and personalized services, and improves user experience.
Smart Images

Figure CN119988561A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer dialogue models, and in particular to a dialogue state tracking method and system based on state correction. Background Art
[0002] AI chatbots are software applications that use artificial intelligence (AI) and natural language processing (NLP) to simulate human conversations with customers. They can answer common questions, provide information, and perform simple tasks, such as booking appointments, processing payments, or updating account details. AI chatbots can be integrated with various platforms, such as websites, mobile apps, social media, or messaging apps to provide a vast array of customer service without the need for human agents.
[0003] The core technologies of AI chatbots can be divided into two major categories: open domain dialogue systems and task-oriented dialogue systems. However, since human dialogue is inherently complex and ambiguous, creating an open domain dialogue system that can perform arbitrary tasks is an unsolved problem. The current practice focuses on building task-oriented dialogue (TOD) systems that are limited to specific domains (such as flight bookings). For example, consider a hotel reservation dialogue system. The user may make a request similar to: "I want to book a hotel in the city center tomorrow night, single room, for two nights." The task-oriented dialogue system needs to understand the user's requirements, which may involve entity recognition (date, location, room type, etc.), and interact with the hotel reservation system to complete the reservation.
[0004] Dialogue State Track (DST) is an important part of the task-based dialogue task flow in natural language processing. Its basic goal is to obtain the current dialogue state based on the dialogue context. The dialogue state is a summary of the user's goals from the beginning of the dialogue to the current dialogue, usually in the form of a combination of multiple slot-value pairs, and sometimes also includes information such as the domain to which the dialogue belongs and the user's intention. Dialogue state tracking refers to the process of inferring and updating the current dialogue state by combining information such as the dialogue history, the current dialogue, and the previous round of dialogue state. The goal of a task-oriented dialogue system is to give appropriate answers to questions raised by users in a specified domain, or to give candidate information to users' requests, communicate with users naturally, and complete the requirements raised by users, so that users themselves feel that they are talking to "people". It is of great help in accelerating social development and progress, increasing user efficiency, and improving users' life experience. In addition, it has a wide range of commercial uses and has become a field that major companies are competing to research, with great research value and significance.
[0005] In recent years, deep learning methods have been widely used in many fields of natural language processing. Deep learning is used for dialogue state tracking, which can automatically extract semantic feature information from the dialogue context without the need for heavy manual rule design. Some existing methods consider that slots are related, and they believe that slots are not conditionally independent, such as hotel star rating and price range. Modeling slot relationships can be divided into explicit and implicit. One is to consider these relationships explicitly using self-attention. Similarly, Manotumruksa adopted a hybrid architecture to achieve sequence value prediction based on the GPT-2 model while using the graph attention network (GAT) to model the relationship between slots and values. Another type of work uses knowledge obtained from domain ontology, such as using its hierarchical structure to implicitly model the relationship between slots. These relationships can also be represented by a graph with slots and domains as nodes. Wang first built a pattern graph based on the ontology, and then used GAT to fuse the information in the dialogue history and the pattern graph. Chen extended this method by dynamically updating the slot relationships in the pattern graph based on the dialogue context. Guo used both dialogue turns and slot-value pairs as nodes, only considering relevant turns and solving coreference relationships. Other approaches consider adapting models to new domains. Dingliwal uses meta-learning to leverage meta-learned model parameters and initialize fine-tuning on the target domain. Work around pattern-based datasets uses slot descriptions to handle unseen domains and slots. One drawback of these approaches is their reliance on similarities between the unseen domain and the initial fine-tuning domain. Another group of approaches attempts to mine external knowledge from other tasks with richer resources. Hudeček uses frame network semantic analysis as weak supervision to identify potential slots. Li proposed different methods to pre-train on reading comprehension data before applying the model to DST. Similarly, Shin redefined DST as a dialogue summarization task based on templates and leveraging external annotated data. It is worth noting that the above-mentioned PLM adaptation methods allow for more effective learning when less data is available, and are also potential solutions to the problem of data scarcity. Along this line of thought, Mi proposed a small-sample DST self-learning method that is complementary to TOD-BERT.
[0006] However, the above solutions only model the relationship between slots and dialogue contexts and the correlation between slots. The representation information of the dialogue text is not clear enough, resulting in insufficient expression of dialogue information, using the entire dialogue history in the training phase, introducing too many dialogues unrelated to the slots, resulting in dialogue information redundancy, and using the real dialogue state for training during training and using the previous round of dialogue states that may be wrong for prediction in the prediction phase, which leads to inconsistency of dialogue information. Summary of the invention
[0007] The object of the present invention is to provide a method and system for tracking a dialog state based on state correction, which method and system are beneficial to improving the accuracy of dialog state tracking.
[0008] In order to achieve the above object, the technical solution adopted by the present invention is: a method for tracking a dialog state based on state correction, comprising the following steps: Step A: Collect conversation data and establish training set DS; Step B: constructing a dialogue state tracking model based on a deep learning network, wherein the dialogue state tracking model corrects the dialogue state in the input data, extracts and fuses slot information, deeply perceives context information, and thereby more accurately obtains user intent, and corrects dialogue state errors during the prediction process, thereby improving dialogue state tracking accuracy; using the training set DS, the dialogue state tracking model based on the deep learning network is trained to obtain a trained dialogue state tracking model; Step C: Input the dialogue to be tracked into the trained dialogue state tracking model and output the corresponding dialogue state.
[0009] Furthermore, the specific implementation method of step A is: Step A1: Construct a dialogue training dataset and collect the following: (1) User input: Questions raised by users in a dialogue, in the form of text input; (2) System response: The dialogue system’s response to user input, in the form of text reply; (3) Dialogue history: Record each round of interaction in a dialogue, including user input, system reply, and any contextual information involved; this helps to understand the coherence and contextual relationship of the dialogue; (4) Slot candidate values: Possible values of a slot, which are pre-defined values; slots are categories of key entities in a dialogue; (5) Dialogue state: Tracks state information in a dialogue, including the current state of the dialogue and the user’s contextual information; exists in the form of slot-value pairs consisting of slots and values, where values are key entities in a dialogue; Step A2: The dialogue training set is represented as , where N represents the number of training samples, that is, the number of dialogue samples; V represents all slot candidate values; n represents the number of rounds of each dialogue, represents a training sample of a round in the dialogue training set; Represents the current round of dialogue, where R t is the system discourse, U t is the user's words; S is the slot, , j∈J, J is the total number of slots; Represents the current conversation state, which is composed of slots and corresponding values. j represents the slot pair S in round t-1 j The value of V j ∈V, , and the initial value of each slot is none, H s Indicates the slot that needs to be tracked.
[0010] Furthermore, the step B specifically includes the following steps: Step B1: Modify the dialogue state of each sample in the training set, specifically by deleting or adding the slot value of each slot in the dialogue state with a certain probability to obtain a new dialogue state , using the new dialogue state to replace the dialogue state in the original training set for subsequent training; Step B2: Initially encode each round of dialogue samples in the training set DS. Each round of dialogue samples includes a round of dialogue between the user and the system and the current round of dialogue status, so as to obtain the initial features H of each round of dialogue, dialogue status, slots to be tracked, and corresponding slot candidate values. t , , , ; Step B3: The initial dialogue feature H obtained in step B2 t The slot H that needs to be tracked s Input to the specific slot information extraction module to obtain the dialog semantic representation vector of the specific slot after information enhancement ; The initial features of the dialogue state obtained in step B1 With slot H s Input to the specific slot information extraction module to obtain the paragraph-level dialogue semantic representation vector P of the specific slot after information enhancement t ; Step B4: The conversation semantic representation vector obtained in step B3 And the paragraph-level dialogue semantic representation vector P t Input into the dynamic information fusion module to obtain the final slot representation vector ; Step B5: The slot characterization vector obtained in step B4 The linear layer determines whether the slot is over-predicted or under-predicted. Over-predicted slot values are deleted, and under-predicted slot values are re-predicted. Step B6: The slot characterization vector obtained in step B4 The candidate slot value of the corresponding slot is represented by the encoded vector Perform similarity matching and select the most similar candidate slot value as the predicted value; compare the predicted value with the original value corresponding to the slot to calculate the loss, use the back propagation algorithm to calculate the gradient of each parameter in the deep learning network, and use the stochastic gradient descent algorithm to update the parameters; Step B7: When the loss value generated by the dialog state tracking model is less than a set threshold or reaches the maximum number of training times, the training of the dialog state tracking model is terminated.
[0011] Furthermore, in step B1, the slot in the dialog state is judged. If the slot value of the slot is not none, it is deleted with a set probability, and the slot whose slot value is deleted is recorded with a one-hot vector, which is expressed as ,in The value is 0 or 1, 0 means no deletion, 1 means deletion, t-1 is the t-1th round of dialogue, and J is the Jth slot among all slots; if the slot value of the slot is none, a value is randomly selected from the candidate values of the slot with probability to fill it, which is expressed as ,in The value is 0 or 1, 0 means no filling, 1 means filling; the modified dialog state is represented as .
[0012] Furthermore, in step B2, the current round of dialogue and the dialogue state are concatenated and input into the BERT encoding, and the output is , where [CLS] and [SEP] are special tokens used to separate the various parts of the input. The [CLS] tag is used to obtain the vector representation of the input model. [SEP] tags the end. BERT is a pre-trained model. ; where L h is the length of the current conversation, L b Indicates the length of the current dialogue state, d is the dimension of the token representation vector, , , , .
[0013] Furthermore, the step B3 specifically includes the following steps: Step B31: The current conversation representation H outputted in step B22 is t With slot H s After modeling by the specific slot information extraction module, the conversation semantic representation vector of the specific slot after information enhancement is obtained : Input Dialogue Representation , where T is the number of time steps and d t is the dialogue representation dimension; slot representation , where S is the number of slots, d s Represents the slot dimension; H t and H s Mapped to query Q, key K and value V respectively through linear transformation: in, , , is the learnable weight matrix, d k and d v are the dimensions of key and value respectively; calculate the similarity between query Q and key K to obtain the score matrix of specific slot information: Among them, QK T yes A matrix of dimension , representing the specific slot weights between the dialogue representation and the slot value representation; Then, the conversation representation H is calculated through a linear layer t and slot value characterization H s The weighted representation between: Among them, W s is a learnable parameter; after the specific slot information extraction module, the dialogue representation H is obtained t and slot H s A weighted representation between , which captures the importance and relevance of different slot values in the conversation context; Step B32: Represent the dialog state obtained in step B21 Same as slot H s Input into the specific slot information extraction module for modeling, and obtain the paragraph-level dialogue semantic representation vector P t .
[0014] Furthermore, the step B4 specifically includes the following steps: Step B41: In order to enrich the slot semantics, the conversation semantic representation vector obtained in step B31 is transformed into and the paragraph-level conversation semantic representation vector P obtained in step B32 t Input to the dynamic information fusion module for cross attention fusion and calculate P t The cross attention matrix: Calculate P t right The cross attention matrix: Where d is the hidden layer dimension, , All correspond The learnable parameter matrix, , All correspond to P t The learnable parameter matrix, and They are P t The query vector and P t right The query vector, Q X , Q Y are conversation semantic query vector and paragraph-level conversation semantic query vector respectively, T represents matrix transpose, A X , A Y They are and P t The cross attention matrix; Step B42: The cross attention matrix A obtained from step B41 X , A Y , compute the cross-context representation: in, express With P t The interactive context representation of Indicates P t and Interaction context representation; Step B43: Calculate the two cross-context representations obtained in step B42 , The fusion weight of , and the two are fused according to the fusion weight: in, Will and The smaller dimension is aligned to the larger one, and the smaller dimension is filled with 0; , All are learnable parameter matrices; is the activation function, Represents the matrix dot product, and finally obtains the fused slot representation vector .
[0015] Furthermore, the step B5 specifically includes the following steps: Step B51: In order to improve the robustness of the model, the slot representation vector obtained in step B41 is Make judgments to identify over-predicted and under-predicted slots; The probability distribution of whether it is an over-prediction slot is: The probability distribution of whether there is a small prediction slot is: Among them, W over , W lack is a learnable parameter, Sigmoid is a logistic regression function; for Delete the slot value of Re-predict the slot value; Step B52: Calculate the excess loss using the two probabilities obtained in step B51 and lack of loss : .
[0016] Furthermore, the step B6 specifically includes the following steps: Step B61: For each slot, first encode the slot candidate value through BERT to obtain the slot value vector representation: in, Indicates the i-th candidate value of the j-th slot, and finally takes The [cls] bit represents the final value ; Encode each candidate value to get the candidate value set ,Since the number of candidate values for each slot is different, the value range of i is different; Step 62: Combine all candidate value representations obtained in B61 with the slot representation vector obtained in B51 Calculate the semantic distance and then select the slot value with the minimum distance as slot S j The final prediction result; use L2 norm as the distance metric; in the training phase, calculate the final slot S j The true value of Probability for: The value with the maximum probability is taken as the predicted value; represents the exponential function, represents the L2 norm; Step B63: The model is trained to maximize the joint probability of all slots, i.e. ; The loss function for each round t is defined as the accumulation of negative log-likelihood: Therefore, the total loss is: Step B64: The total loss calculated in step B63 is updated with the learning rate through the stochastic gradient descent algorithm Adam, and the model parameters are updated iteratively using back propagation to minimize the total loss to train the model.
[0017] The present invention also provides a dialogue state tracking system based on state correction, including a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the above method can be implemented.
[0018] Compared with the prior art, the present invention has the following beneficial effects: the present invention provides a method and system for tracking a dialogue state based on state correction, so as to solve the technical problems of the prior art such as inaccuracy in dialogue state tracking and insufficient processing of complex dialogues. The method adopts a dynamic dialogue state correction mechanism to dynamically adjust the dialogue state by analyzing the user's input and dialogue context in real time. This method can flexibly respond to changes in user intentions and improve the accuracy of dialogue state tracking; the present invention is based on a deep learning network and can capture complex patterns and relationships in dialogues, so as to better understand the nuances of language and improve the accuracy of state tracking; the present invention attaches importance to the contextual information of the dialogue and understands the current user intention by analyzing the dialogue history. This context-aware capability makes state tracking more accurate and can effectively handle multi-round dialogue situations. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] Figure 1 is a flow chart of a method implementation of an embodiment of the present invention; Figure 2 It is an architecture diagram of a dialogue state tracking model based on a deep learning network in an embodiment of the present invention. DETAILED DESCRIPTION
[0020] The present invention will be further described below in conjunction with the accompanying drawings and embodiments.
[0021] It should be noted that the following detailed descriptions are exemplary and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0022] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0023] like Figure 1 As shown, this embodiment provides a method for tracking a dialog state based on state correction, comprising the following steps: Step A: Collect conversation data and establish training set DS; Step B: constructing a dialogue state tracking model based on a deep learning network, wherein the dialogue state tracking model corrects the dialogue state in the input data, extracts and fuses slot information, deeply perceives context information, and thereby more accurately obtains user intent, and corrects dialogue state errors during the prediction process, thereby improving dialogue state tracking accuracy; using the training set DS, the dialogue state tracking model based on the deep learning network is trained to obtain a trained dialogue state tracking model; Step C: Input the dialogue to be tracked into the trained dialogue state tracking model and output the corresponding dialogue state.
[0024] In this embodiment, the specific implementation method of step A is: Step A1: Construct a dialogue training dataset and collect the following content: (1) User input: Questions raised by users in a dialogue, in the form of text input; (2) System response: The dialogue system’s response to user input, in the form of text reply; (3) Dialogue history: Record each round of interaction in a dialogue, including user input, system reply, and any contextual information involved; this helps to understand the coherence and contextual relationship of the dialogue; (4) Slot candidate values: Possible values of a slot, which are pre-defined values; slots are categories of key entities in a dialogue; (5) Dialogue state: Tracks state information in a dialogue, including the current state of the dialogue and the user’s contextual information; exists in the form of slot-value pairs consisting of slots and values, with values being key entities in a dialogue.
[0025] Step A2: The dialogue training set is represented as , where N represents the number of training samples, that is, the number of dialogue samples; V represents all slot candidate values; n represents the number of rounds of each dialogue, represents a training sample of a round in the dialogue training set; Represents the current round of dialogue, where R t is the system discourse, U t is the user's words; S is the slot, , j∈J, J is the total number of slots; Represents the current conversation state, which is composed of slots and corresponding values. j represents the slot pair S in round t-1 j The value of V j ∈V, , and the initial value of each slot is none, H s Indicates the slot that needs to be tracked.
[0026] Figure 2 : is an architecture diagram of a dialogue state tracking model based on a deep learning network in this embodiment. In this embodiment, the specific implementation method of step B is as follows.
[0027] Step B1: Modify the dialogue state of each sample in the training set, specifically by deleting or adding the slot value of each slot in the dialogue state with a certain probability to obtain a new dialogue state , use the new dialogue state to replace the dialogue state in the original training set for subsequent training.
[0028] Specifically, the slot in the conversation state is judged. If the slot value of the slot is not none, it is deleted with a set probability, and the slot whose slot value is deleted is recorded as a one-hot vector, which is expressed as ,in The value is 0 or 1, 0 means no deletion, 1 means deletion, t-1 is the t-1th round of dialogue, and J is the Jth slot among all slots; if the slot value of the slot is none, a value is randomly selected from the candidate values of the slot with probability to fill it, which is expressed as ,in The value is 0 or 1, 0 means no filling, 1 means filling; the modified dialog state is represented as .
[0029] Step B2: Initially encode each round of dialogue samples in the training set DS. Each round of dialogue samples includes a round of dialogue between the user and the system and the current round of dialogue status, so as to obtain the initial features H of each round of dialogue, dialogue status, slots to be tracked, and corresponding slot candidate values. t , , , .
[0030] Specifically, the current round of dialogue and the dialogue state are concatenated and input into the BERT encoder, and the output is , where [CLS] and [SEP] are special tokens used to separate the various parts of the input. The [CLS] tag is used to obtain the vector representation of the input model. [SEP] tags the end. BERT is a pre-trained model. ; where L h is the length of the current conversation, L b Indicates the length of the current dialogue state, d is the dimension of the token representation vector, , , , .
[0031] Step B3: The initial dialogue feature H obtained in step B2 t The slot H that needs to be tracked s Input to the specific slot information extraction module to obtain the dialog semantic representation vector of the specific slot after information enhancement ; The initial features of the dialogue state obtained in step B1 With slot H s Input to the specific slot information extraction module to obtain the paragraph-level dialogue semantic representation vector P of the specific slot after information enhancement t .
[0032] In this embodiment, step B3 specifically includes the following steps: Step B31: The current conversation representation H outputted in step B22 is t With slot H s After modeling by the specific slot information extraction module, the conversation semantic representation vector of the specific slot after information enhancement is obtained : Input Dialogue Representation , where T is the number of time steps and d t is the dialogue representation dimension; slot representation , where S is the number of slots, d s Represents the slot dimension; H t and H s Mapped to query Q, key K and value V respectively through linear transformation: in, , , is the learnable weight matrix, d k and d v are the dimensions of key and value respectively; calculate the similarity between query Q and key K to obtain the score matrix of specific slot information: Among them, QK T yes A matrix of dimension denoting the specific slot weights between the dialogue representation and the slot value representation.
[0033] Then, the conversation representation H is calculated through a linear layer t and slot value characterization H sThe weighted representation between: Among them, W s is a learnable parameter; after the specific slot information extraction module, the dialogue representation H is obtained t and slot H s The weighted representation between , which captures the importance and relevance of different slot values in the conversation context.
[0034] Step B32: Represent the dialog state obtained in step B21 Same as slot H s Input into the specific slot information extraction module for modeling, and obtain the paragraph-level dialogue semantic representation vector P t .
[0035] Step B4: The conversation semantic representation vector obtained in step B3 And the paragraph-level dialogue semantic representation vector P t Input into the dynamic information fusion module to obtain the final slot representation vector .
[0036] In this embodiment, step B4 specifically includes the following steps: Step B41: In order to enrich the slot semantics, the conversation semantic representation vector obtained in step B31 is transformed into and the paragraph-level conversation semantic representation vector P obtained in step B32 t Input to the dynamic information fusion module for cross attention fusion and calculate P t The cross attention matrix: Calculate P t right The cross attention matrix: Where d is the hidden layer dimension, , All correspond The learnable parameter matrix, , All correspond to P t The learnable parameter matrix, and They are P tThe query vector and P t right The query vector, Q X , Q Y are conversation semantic query vector and paragraph-level conversation semantic query vector respectively, T represents matrix transpose, A X , A Y They are and P t The cross attention matrix.
[0037] Step B42: The cross attention matrix A obtained from step B41 X , A Y , compute the cross-context representation: in, express With P t The interactive context representation of Indicates P t and Interaction context representation.
[0038] Step B43: Calculate the two cross-context representations obtained in step B42 , The fusion weight of , and the two are fused according to the fusion weight: in, Will and The smaller dimension is aligned to the larger one, and the smaller dimension is filled with 0; , All are learnable parameter matrices; is the activation function, Represents the matrix dot product, and finally obtains the fused slot representation vector .
[0039] Step B5: The slot characterization vector obtained in step B4 The linear layer is used to determine whether the slot is over-predicted or under-predicted. The over-predicted slot values are deleted and the under-predicted slot values are re-predicted.
[0040] In this embodiment, step B5 specifically includes the following steps: Step B51: In order to improve the robustness of the model, the slot representation vector obtained in step B41 is A judgment is made to identify over-predicted and under-predicted slots.
[0041] The probability distribution of whether it is an over-prediction slot is: The probability distribution of whether there is a small prediction slot is: Among them, W over , W lack is a learnable parameter, Sigmoid is a logistic regression function; for Delete the slot value of The slot value is re-predicted.
[0042] Step B52: Calculate the excess loss using the two probabilities obtained in step B51 and lack of loss : .
[0043] Step B6: The slot characterization vector obtained in step B4 The candidate slot value of the corresponding slot is represented by the encoded vector Perform similarity matching and select the most similar candidate slot value as the predicted value; compare the predicted value with the original value corresponding to the slot to calculate the loss, use the back propagation algorithm to calculate the gradient of each parameter in the deep learning network, and use the stochastic gradient descent algorithm to update the parameters.
[0044] In this embodiment, step B6 specifically includes the following steps: Step B61: For each slot, first encode the slot candidate value through BERT to obtain the slot value vector representation: in, Indicates the i-th candidate value of the j-th slot, and finally takes The [cls] bit represents the final value ; Encode each candidate value to get the candidate value set ,Since the number of candidate values for each slot is different, the value range of i is different.
[0045] Step 62: Combine all candidate value representations obtained in B61 with the slot representation vector obtained in B51 Calculate the semantic distance and then select the slot value with the minimum distance as slot S j The final prediction result; use L2 norm as the distance metric; in the training phase, calculate the final slot S j The true value of Probability for: The value with the maximum probability is taken as the predicted value; represents the exponential function, represents the L2 norm.
[0046] Step B63: The model is trained to maximize the joint probability of all slots, i.e. ; The loss function for each round t is defined as the accumulation of negative log-likelihood: Therefore, the total loss is: Step B64: The total loss calculated in step B63 is updated with the learning rate through the stochastic gradient descent algorithm Adam, and the model parameters are updated iteratively using back propagation to minimize the total loss to train the model.
[0047] Step B7: When the loss value generated by the dialog state tracking model is less than a set threshold or reaches the maximum number of training times, the training of the dialog state tracking model is terminated.
[0048] Based on the above technical solutions provided in this embodiment, this method has achieved the following outstanding technical effects: (1) Improved accuracy: Through dynamic state correction and deep learning model integration, the present invention significantly improves the accuracy of dialogue state tracking and can accurately capture the user's intentions and needs.
[0049] (2) Enhanced adaptability: Context-aware analysis and user feedback loops enhance the system’s adaptability to complex conversational situations, enabling the system to flexibly respond to changes in user intent and provide more personalized services.
[0050] (3) Improving user experience: Accurate status tracking and highly personalized services significantly improve the user experience, reduce the user's operational burden, and make the interaction more smooth and natural.
[0051] This embodiment also provides a dialogue state tracking system based on state correction, including a memory, a processor, and computer program instructions stored in the memory and executable by the processor. When the processor executes the computer program instructions, the above method steps can be implemented.
[0052] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0053] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0054] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0055] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 A step that specifies a function in one or more boxes.
[0056] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.
Claims
1. A method for tracking a dialog state based on state correction, characterized in that: The following steps are involved: Step A: Collect conversation data and establish training set DS; Step B: Construct a dialogue state tracking model based on a deep learning network. The dialogue state tracking model corrects the dialogue state in the input data, extracts and fuses slot information, deeply perceives context information, and thus more accurately obtains user intent. It also corrects dialogue state errors during the prediction process, thereby improving dialogue state tracking accuracy. Use the training set DS to train a dialogue state tracking model based on a deep learning network to obtain a trained dialogue state tracking model; Step C: Input the dialogue to be tracked into the trained dialogue state tracking model and output the corresponding dialogue state.
2. The method for tracking a dialog state based on state correction according to claim 1, characterized in that: The specific implementation method of step A is: Step A1: Construct a dialogue training dataset and collect the following: (1) User input: Questions raised by users in a dialogue, in the form of text input; (2) System response: The dialogue system’s response to user input, in the form of text reply; (3) Dialogue history: Record each round of interaction in a dialogue, including user input, system reply, and any contextual information involved; this helps to understand the coherence and contextual relationship of the dialogue; (4) Slot candidate values: Possible values of a slot, which are pre-defined values; slots are categories of key entities in a dialogue; (5) Dialogue state: Tracks state information in a dialogue, including the current state of the dialogue and the user’s contextual information; exists in the form of slot-value pairs consisting of slots and values, where values are key entities in a dialogue; Step A2: The dialogue training set is represented as , where N represents the number of training samples, that is, the number of dialogue samples; V represents all slot candidate values; n represents the number of rounds of each dialogue, represents a training sample of a round in the dialogue training set; Represents the current round of dialogue, where R t is the system discourse, U t is the user's words; S is the slot, , j∈J, J is the total number of slots; Represents the current conversation state, which is composed of slots and corresponding values. j represents the slot pair S in round t-1 j The value of V j ∈V, , and the initial value of each slot is none, H s Indicates the slot that needs to be tracked.
3. The method for tracking a dialog state based on state correction according to claim 2, characterized in that: The step B specifically comprises the following steps: Step B1: Modify the dialogue state of each sample in the training set, specifically by deleting or adding the slot value of each slot in the dialogue state with a certain probability to obtain a new dialogue state , using the new dialogue state to replace the dialogue state in the original training set for subsequent training; Step B2: Initially encode each round of dialogue samples in the training set DS. Each round of dialogue samples includes a round of dialogue between the user and the system and the current round of dialogue status, so as to obtain the initial features H of each round of dialogue, dialogue status, slots to be tracked, and corresponding slot candidate values. t , , , ; Step B3: The initial dialogue feature H obtained in step B2 t The slot H that needs to be tracked s Input to the specific slot information extraction module to obtain the dialog semantic representation vector of the specific slot after information enhancement ; The initial features of the dialogue state obtained in step B1 With slot H s Input to the specific slot information extraction module to obtain the paragraph-level conversation semantic representation vector P of the specific slot after information enhancement t ; Step B4: The conversation semantic representation vector obtained in step B3 And the paragraph-level dialogue semantic representation vector P t Input into the dynamic information fusion module to obtain the final slot representation vector ; Step B5: The slot characterization vector obtained in step B4 The linear layer determines whether the slot is over-predicted or under-predicted. Over-predicted slot values are deleted, and under-predicted slot values are re-predicted. Step B6: The slot characterization vector obtained in step B4 The candidate slot value of the corresponding slot is represented by the encoded vector Perform similarity matching and select the most similar candidate slot value as the predicted value; compare the predicted value with the original value corresponding to the slot to calculate the loss, use the back propagation algorithm to calculate the gradient of each parameter in the deep learning network, and use the stochastic gradient descent algorithm to update the parameters; Step B7: When the loss value generated by the dialog state tracking model is less than a set threshold or reaches the maximum number of training times, the training of the dialog state tracking model is terminated.
4. The method for tracking a dialog state based on state correction according to claim 3, characterized in that: In step B1, the slot in the dialog state is judged. If the slot value of the slot is not none, it is deleted with a set probability, and the slot whose slot value is deleted is recorded as a one-hot vector, which is expressed as ,in The value is 0 or 1, 0 means no deletion, 1 means deletion, t-1 is the t-1th round of dialogue, and J is the Jth slot among all slots; if the slot value of the slot is none, a value is randomly selected from the candidate values of the slot with probability to fill it, which is expressed as ,in The value is 0 or 1, 0 means no filling, 1 means filling; the modified dialog state is represented as .
5. The method for tracking a dialog state based on state correction according to claim 4, characterized in that: In step B2, the current round of dialogue and the dialogue state are concatenated and input into the BERT encoder, and the output is , where [CLS] and [SEP] are special tokens used to separate the various parts of the input. The [CLS] tag is used to obtain the vector representation of the input model. [SEP] tags the end. BERT is a pre-trained model. ; where L h is the length of the current conversation, L b Indicates the length of the current dialogue state, d is the dimension of the token representation vector, , , , .
6. The method for tracking a dialog state based on state correction according to claim 5, characterized in that: The step B3 specifically comprises the following steps: Step B31: The current conversation representation H outputted in step B22 is t With slot H s After modeling by the specific slot information extraction module, the conversation semantic representation vector of the specific slot after information enhancement is obtained : Input Dialogue Representation , where T is the number of time steps and d t is the dialogue representation dimension; slot representation , where S is the number of slots, d s Represents the slot dimension; H t and H s Mapped to query Q, key K and value V respectively through linear transformation: in, , , is the learnable weight matrix, d k and d v are the dimensions of key and value respectively; calculate the similarity between query Q and key K to obtain the score matrix of specific slot information: Among them, QK T yes A matrix of dimension , representing the specific slot weights between the dialogue representation and the slot value representation; Then, the conversation representation H is calculated through a linear layer t and slot value characterization H s The weighted representation between: Among them, W s is a learnable parameter; after the specific slot information extraction module, the dialogue representation H is obtained t and slot H s A weighted representation between , which captures the importance and relevance of different slot values in the conversation context; Step B32: Represent the dialog state obtained in step B21 Same as slot H s Input into the specific slot information extraction module for modeling, and obtain the paragraph-level dialogue semantic representation vector P t .
7. The method for tracking a dialog state based on state correction according to claim 6, characterized in that: The step B4 specifically comprises the following steps: Step B41: In order to enrich the slot semantics, the conversation semantic representation vector obtained in step B31 is transformed into and the paragraph-level conversation semantic representation vector P obtained in step B32 t Input to the dynamic information fusion module for cross attention fusion and calculate P t The cross attention matrix: Calculate P t right The cross attention matrix: Where d is the hidden layer dimension, , All correspond The learnable parameter matrix, , All correspond to P t The learnable parameter matrix, and They are P t The query vector and P t right The query vector, Q X , Q Y are conversation semantic query vector and paragraph-level conversation semantic query vector respectively, T represents matrix transpose, A X , A Y They are and P t The cross attention matrix; Step B42: The cross attention matrix A obtained from step B41 X , A Y , compute the cross-context representation: in, express With P t The interactive context representation of Indicates P t and Interaction context representation; Step B43: Calculate the two cross-context representations obtained in step B42 , The fusion weight of , and the two are fused according to the fusion weight: in, Will and The smaller dimension is aligned to the larger one, and the smaller dimension is filled with 0; , All are learnable parameter matrices; is the activation function, Represents the matrix dot product, and finally obtains the fused slot representation vector .
8. The method for tracking a dialog state based on state correction according to claim 7, characterized in that: The step B5 specifically comprises the following steps: Step B51: In order to improve the robustness of the model, the slot representation vector obtained in step B41 is Make judgments to identify over-predicted and under-predicted slots; The probability distribution of whether it is an over-prediction slot is: The probability distribution of whether there is a small prediction slot is: Among them, W over , W lack is a learnable parameter, Sigmoid is a logistic regression function; for Delete the slot value of Re-predict the slot value; Step B52: Calculate the excess loss using the two probabilities obtained in step B51 and lack of loss : 。 9. The method for tracking a dialog state based on state correction according to claim 8, characterized in that: The step B6 specifically comprises the following steps: Step B61: For each slot, first encode the slot candidate value through BERT to obtain the slot value vector representation: in, Indicates the i-th candidate value of the j-th slot, and finally takes The [cls] bit represents the final value ; Encode each candidate value to get the candidate value set ,Since the number of candidate values for each slot is different, the value range of i is different; Step 62: Combine all candidate value representations obtained in B61 with the slot representation vector obtained in B51 Calculate the semantic distance and then select the slot value with the minimum distance as slot S j The final prediction result; use L2 norm as the distance metric; in the training phase, calculate the final slot S j The true value of Probability for: The value with the maximum probability is taken as the predicted value; represents the exponential function, represents the L2 norm; Step B63: The model is trained to maximize the joint probability of all slots, i.e. ; The loss function for each round t is defined as the accumulation of negative log-likelihood: Therefore, the total loss is: Step B64: The total loss calculated in step B63 is updated by the stochastic gradient descent algorithm Adam, and the model parameters are updated iteratively using back propagation to minimize the total loss to train the model.
10. A dialogue state tracking system based on state correction, characterized in that: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 9 can be implemented.