A zero-shot dialogue state tracking method

By adopting a zero-sample dialogue state tracking method based on pre-trained models and mixed semantic experts in multi-domain dialogue state tracking, the problem of insufficient zero-sample generalization ability in the existing technology is solved, and more accurate dialogue state tracking and performance improvement is achieved.

CN118395994BActive Publication Date: 2025-05-13INSTITUTE OF INFORMATION ENGINEERING CHINESE ACADEMY OF SCIENCES
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410327987.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-03-21
Publication Date
2025-05-13
Estimated Expiration
2044-03-21

AI Technical Summary

Technical Problem

The prior art lacks effective zero-sample generalization capabilities in multi-domain dialogue state tracking, especially in semantic decoupling, and it is difficult to accurately map unseen samples to areas where the samples have been seen.

Method used

Using the zero-sample dialogue state tracking method based on pre-trained models and mixed semantic experts, the non-search samples are mapped to the corresponding semantic experts through the steps of division, resolution and merging, improving the performance performance of zero-sample.

Benefits of technology

Through this method, the performance of multi-field zero-sample migration tasks is significantly improved, more accurate dialogue state tracking is achieved, and the generalization ability of the model is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118395994B_ABST
    Figure CN118395994B_ABST
Patent Text Reader

Abstract

The present invention discloses a zero-shot dialogue state tracking method, and its steps include: 1) Division stage: For each piece of dialogue text C with dialogue state annotation t , the pre-trained language model f is used to convert the dialogue text C t into a dialogue text vector e t . Then, clustering technology is used to classify each dialogue text vector into its different subsets, obtaining K subsets; 2) Solution stage: Each text vector in the subset is used as a sample, and each obtained subset is used to train a semantically independent state tracking model, obtaining a total of K trained state tracking models; 3) Merging stage: First, relationship mining is performed, and a given dialogue text C' t is converted into a semantic vector e' t . The relationship δ between the semantic space of each subset and the dialogue text C' t is calculated; then, aggregation inference is performed, and the dialogue state corresponding to the dialogue text C' t is predicted according to each trained state tracking model and its corresponding relationship δ.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of text generation, and relates to several technical methods for extracting dialogue states (user goals, user needs) from task-based dialogue data, specifically a dialogue state tracking method based on a pre-trained model and hybrid semantic experts. Background Art

[0002] Dialogue State Tracking (DST) plays a vital role in task-driven dialogue systems. The goal of this task is to understand the user's needs and goals and extract the dialogue state. The dialogue state refers to a set of multiple slot-value pairs. Accurate dialogue state tracking performance can help downstream modules, such as the dialogue manager.

[0003] However, it is difficult and expensive to collect and annotate dialogue states. With the emergence of single-domain to multi-domain dialogue, this problem becomes more acute. In order to train a multi-domain to dialogue state tracking model, dialogue annotators must annotate slot-value pairs for each turn and each domain of the dialogue. In order to make the dialogue state model more robust, various models improve the zero-shot performance from the data level or model level. The first type of method is to synthesize new dialogue samples or introduce other large-scale annotated datasets. Some of them use existing ontology information and annotated datasets to synthesize new dialogues. Another part of the work uses a variety of annotated datasets, such as Question Answering (QA) and Dialogue Summarization, to overcome the data sparsity problem. This type of method using other task datasets is also called zero-shot cross-task transfer method. The second type of method is to develop efficient models or frameworks to improve the scalability of dialogue state tracking models, such as span-based approaches, copy-augmented decoders, or large-scale pre-trained language models. In the dialogue state tracking task, some people have used the description information of the slot as a prompt to generate the slot value result; others have further improved the performance of zero-shot by modeling three dependencies on the slot based on the prompt. Although the existing models have achieved certain results, we believe that these data or model-level enhancement methods have not explored the essence of zero-shot generalization, mainly because they lack the ability of semantic decoupling to map the area of ​​unseen samples to seen samples.

[0004] To intuitively explain how visible samples in the semantic space can help unseen samples, we give an example. Assuming that there is an unseen sample from the "train" domain, the sample is "I plan to travel from Beijing to Hong Kong by train on Sunday", the dialogue semantic space of "booking a room" can help predict the unseen slot "train booking time"; the dialogue semantic space of "booking a taxi" can help predict the "train departure point" slot and the "train arrival point" slot. In other words, a new unseen sample may be difficult to reason directly due to its complex combination, but it can be easily predicted by mapping related semantics. However, decoupling at the semantic level has always been challenging and unstable, especially for scenarios that require accurate semantic division. Summary of the invention

[0005] In view of the problems existing in the prior art, the purpose of the present invention is to provide a zero-sample dialogue state tracking method based on a pre-trained model and a hybrid semantic expert. The present invention proposes a simple but efficient solution - division, resolution and merging to map unseen samples to corresponding accurate semantic experts. This idea provides the flexibility of mapping unseen samples to corresponding semantic experts by explicitly dividing visible samples into different semantic spaces and training corresponding experts. Such data-level decoupling provides flexibility in mapping unseen samples to corresponding semantic experts. Ultimately, the prediction results from multiple hybrid experts can improve the performance of zero samples. In the experiment, we implemented the solution in the current pre-trained model T5 and adapter (Adapter), and verified the effectiveness and versatility of the method.

[0006] The technical solution of the present invention is:

[0007] A zero-sample dialogue state tracking method, the steps of which include:

[0008] 1) Division stage: For each dialogue state labeled dialogue text C t , using the pre-trained language model f to transform the conversation text C t Converted into dialogue text vector e t , and then use clustering technology to classify each dialogue text vector into its different subsets to obtain K subsets;

[0009] 2) Solution stage: Take each text vector in the subset as a sample, and use each obtained subset to train a semantically independent state tracking model, and obtain a total of K trained state tracking models;

[0010] 3) Merging stage: First, perform relationship mining to merge a given conversation text C′ t Converted to semantic vector e′ t , calculate the semantic space and dialogue text C′ of each subset tThen perform aggregate reasoning and predict the dialogue text C′ based on each trained state tracking model and its corresponding relationship δ. t The corresponding dialog status.

[0011] Furthermore, the semantic space μ corresponding to the kth subset k and the dialogue text C′ t The relationship between Where d represents the distance function and τ represents the temperature coefficient.

[0012] Furthermore, the corresponding semantic space μ is obtained by averaging all text vectors in the kth subset. k .

[0013] Furthermore, parameter-level aggregate reasoning is used to predict the conversation text C′ t The corresponding dialogue state is as follows: using the semantic space μ corresponding to the kth subset k and the dialogue text C′ t The relationship between δ(C′ t ,μ k ) and the state tracking model parameters after training with the kth subset Initialize a new dialogue state tracking model, and then sum up the K new dialogue state tracking models to get the model Then use the model φ' to predict the dialogue text C' t The corresponding dialog status.

[0014] Furthermore, character-level aggregate reasoning is used to predict the conversation text C′ t The corresponding dialogue state is as follows: first, the state tracking model trained on the kth subset is used to predict the dialogue text C′ t The corresponding dialogue state π k , and then use the semantic space μ corresponding to the kth subset k and the dialogue text C′ t The relationship between δ(C′ t ,μ k ) and the corresponding dialogue state π k Predict the dialogue text C′ t The prediction results for each character in the final dialogue state y m The dialogue text C′ t The prediction result of the mth character in .

[0015] Furthermore, the pre-trained language model T5 is used as the state tracking model.

[0016] Furthermore, the pre-trained language model f is a pre-trained language model T5 or BERT.

[0017] A server, characterized in that it comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing each step in the above method.

[0018] A computer-readable storage medium stores a computer program, wherein the computer program implements the steps of the above method when executed by a processor.

[0019] The advantages of the present invention are as follows:

[0020] The performance of the model on MultiWOZ is compared with several important benchmarks. The experimental results show that the model of the present invention achieves good results on zero-shot transfer tasks in multiple fields. The specific results are shown in Table 1 below.

[0021] Table 1 Experimental results of different models on zero samples

[0022] Model Attraction Hotel Restaurant Taxi Train Average TRADE 19.87 13.70 11.52 60.58 22.37 25.67 MA-DST 22.46 16.28 13.56 59.27 22.76 26.87 T5-DST 35.51 22.48 25.04 65.93 37.82 37.36 TransferQA 31.25 22.72 26.28 61.87 36.72 35.77 Ours(Param) 41.28 26.15 31.05 66.64 38.72 40.76 Ours(Token) 41.35 22.72 33.76 66.90 43.81 42.71

[0023] In order to understand the role of each stage, the present invention also conducted an ablation experiment to prove the effectiveness of each stage. First, the effects of different clustering algorithms, including Kmeans, Birch, Agglomerative and GMM in the hotel field, were studied. According to the implementation results, it was found that all clustering effects were better than directly using the T5 model without clustering, which also proved the effectiveness and stability of the framework of the present invention; in addition, by comparison, it was found that the GMM algorithm achieved the best results, but the results of the Kmeans algorithm at the character level were also very good. Therefore, the present invention believes that a more efficient clustering algorithm can bring better performance improvement.

[0024] The present invention observes the influence of the number of partitions through additional experiments. The model with different K values ​​is tested on the hotel domain. It is found that the accuracy performance of the joint objective first increases with the value of K, and then decreases on the basis of T5. The results show that for T5 small, the optimal number of subsets is 2 for the T5 basic model. Note that the model of the present invention depends on the data distribution and data partitioning, which means that the zero-shot performance may not increase linearly with the increase of K.

[0025] The present invention also experimentally observes different mapping methods. Mapping unseen samples to existing subsets and obtaining mapping weights are the core of the merging process. In addition to adopting weights by inferring from the trained clustering model, the present invention also tried two other weights: 1) argmax: assigning 1 to the subset with the maximum mapping probability and 0 to other subsets; 2) average: assigning uniform probability to all subsets. The experimental results show that directly using the inference weights shows the best performance of parameter-level and token-level integrated reasoning. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] Figure 1 It is a flow chart of the method of the present invention. DETAILED DESCRIPTION

[0027] The present invention is further described in detail below in conjunction with the accompanying drawings. The examples given are only used to explain the present invention but not to limit the scope of the present invention.

[0028] In order to overcome the problem that existing zero-shot transfer methods cannot map and use existing semantic spaces for effective feature transfer, this paper proposes a divide-solve-merge zero-shot transfer framework for dialogue state tracking tasks. Figure 1 As shown in FIG. 1 , the framework follows three steps, namely, division, resolution, and merging. The division process uses a text encoder to encode visible dialogue samples into a semantic space, and these samples are then divided into several subsets by clustering; in the resolution phase, the present invention trains a semantically independent dialogue state tracking expert by using the labeled samples of each subset; finally, in the merging phase, the present invention first estimates the relationship between visible samples and invisible samples, and then realizes relationship-based reasoning by mixing multiple semantically independent experts.

[0029] The technical solution adopted by the present invention is specifically:

[0030] 1) Partitioning stage: The goal of data partitioning is to obtain (ideally) semantically independent spaces for visible data. Previous work has shown that semantically decoupled representations can effectively improve zero-shot generalization in Computer Vision (CV) and Natural Language Processing (NLP), but has not yet been explored in the field of dialogue. This paper believes that explicit partitioning at the data level is simpler and more interpretable than implicit partitioning at the semantic level.

[0031] Given a conversation text, the segmentation algorithm should consider multiple features, including domain, speaker intent, sentence keywords, etc., but these are difficult to obtain in real scenarios. This patent proposes using a simple clustering algorithm, such as K-means, to achieve subset segmentation.

[0032] Specifically, given a dialogue text C t , the pre-trained language model f is used to transform the conversation text C t Converted into dialogue text vector e t =Agg[f(C t )], where Agg represents an aggregation operation, such as mean pooling. The present invention uses the encoder of the pre-trained language model T5 as the model f. Here, the parameters in the model f are fixed.

[0033] Then the present invention uses clustering technology to classify a large number of dialogue text vectors into their different subsets. t After clustering, the subsets classified into the kth subset can be formulated as:

[0034] D k =clustering(e t ),k∈{1,2,…,K}

[0035] Among them, D k Represents the kth subset of the cluster, where K represents the total number of subsets.

[0036] 2) Solution phase: In the solution phase, the data subsets obtained in the first step will be used to train semantically independent state tracking models. Specifically, each subset contains several sample pairs of (dialogue text, dialogue state). Each dialogue state contains multiple slot-slot value pairs. A dialogue state tracking model is trained using the labeled data in a data subset. The input of the dialogue state tracking model is the dialogue text and the description information of a slot. This information is provided by the dataset itself. The output of the dialogue state tracking model is the slot value corresponding to this slot. That is, each dialogue state tracking model can only predict the slot value result of one slot at a time. Then, K subsets can obtain corresponding K dialogue state tracking models. Among them, in subset D k The loss function of the dialogue state tracking model trained on can be expressed as:

[0037]

[0038] in, represents the model parameters in the kth dialogue state tracking model (i.e., T5 model), s j Represents the description information of the jth slot in the conversation state; N k Represents the subset D k The number of all labeled samples in the conversation state, J represents the number of slot values ​​in the conversation state; v j Represents the slot value prediction result of the jth slot, represents the slot value label in the annotated data, C t Represents the dialogue text at time t, and T represents the number of all dialogue rounds. In order to obtain knowledge from the pre-trained model and avoid overfitting, the present invention uses the pre-trained language model T5 as the dialogue state tracking model. The complete T5 model, i.e., the encoder-decoder structure, is used here. The loss function L k Used to update and learn the parameters in the T5 model

[0039] 3) Merging stage: There are two operations in the merging stage, namely, relation mining and aggregate reasoning. In the relation mining stage, given an unseen sample, the present invention maps its conversation text C′ t to the corresponding semantic vector e′ t Then, the semantic space μ corresponding to the kth subset k and unseen samples C′ t The relationship between δ(C′ t ,μ k ) can be calculated as:

[0040]

[0041] Where d represents the distance function; τ represents the temperature coefficient. k Represents the prototype of the semantic space, and the corresponding semantic space μ is obtained by averaging all semantic vectors in the kth subset k In the aggregation reasoning stage, the present invention considers two aggregation strategies to realize multi-expert reasoning based on relationships. These two strategies are widely used in the field of artificial intelligence. These two strategies are parameter-level aggregation and character-level aggregation. The following introduces these two aggregation methods respectively.

[0042] 1) In the aggregation at the parameter level, the present invention uses the previously trained dialogue state tracking model parameters to initialize a new dialogue state tracking model. The specific process is as follows:

[0043]

[0044] y m = argmaxP(v j |C′ t ,s j ; φ')

[0045] Among them, C′ t represents the conversation text vector of the unknown sample, φ' represents the model parameters after the K conversation state tracking models are merged, and y m Represents the predicted m-th character result.

[0046] 2) In character-level aggregation, the present invention calculates the prediction results of unknown samples by combining the prediction results of the existing dialogue state tracking model. Specifically, the prediction process of the unknown sample in the jth slot is as follows:

[0047]

[0048]

[0049] Among them, y (<m) Represents the characters predicted before the mth character, π k is the word prediction probability distribution of the mth character, y m That is, the prediction result of the mth character.

[0050] Finally, the predicted slot value of the unknown sample in the jth slot is the word composed of characters, represented by {y 1 ,y 2 ,...,y M}, where M represents the number of all characters. All slot-slot value pairs eventually constitute the dialogue state of the unknown sample.

[0051] Dataset selection: The present invention evaluates the method proposed in this patent on the widely used MultiWOZ and SGD public datasets. The MultiWOZ dataset contains 100,000 dialogue data and 7 dialogue domains. Each dialogue data contains one or more domains. This patent follows the existing public data preprocessing procedures and evaluation settings. Among them, restaurants, attractions, hotels and taxis are widely used in zero-sample cross-domain experiments. The SGD dataset contains 160,000 multi-domain dialogue datasets and 16 dialogue domains. The test dataset contains datasets that are not seen in the training samples and is used for zero-sample cross-domain experiments.

[0052] Evaluation indicators: This patent follows the existing work and uses two indicators, namely slot accuracy and joint accuracy. Slot accuracy is to calculate the correct proportion of each independent slot prediction, and joint accuracy measures the correct proportion of all the slot sets in the dialogue round, that is, only when all the slots in the same round are predicted correctly, the current slot set prediction is considered correct. In the dialogue state tracking experiment of zero-shot transfer, the model contains all data except the new domain during the training phase, and only the data in the new domain is used during the evaluation.

[0053] Experimental setup: This patent uses Pytorch as the implementation framework. 1) In the data partitioning stage, the present invention uses T5-base as the text encoder and applies average pooling to obtain the vector representation of the conversation text. The present invention selects Kmeans as the clustering method and sets the number of clusters to 3; 2) In the solution stage, T5 is used as an expert in conversation state tracking. And only about 0.8M parameters are added to T5-small and 3.6M parameters are added to T5-base. The present invention freezes the parameters in the transformer, uses a learning rate of 1e-4, and only adjusts the parameters of the adapter. For all experiments, the number of batch data is 16, and AdamW is used as the optimizer. 3) In the merging stage, the temperature at the parameter level is set to 0.2, and the temperature at the character level is set to 2.

[0054] Although the specific embodiments of the present invention are disclosed for the purpose of illustration, the purpose is to help understand the content of the present invention and implement it accordingly, those skilled in the art will understand that various substitutions, changes and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the best embodiment, and the scope of the present invention is subject to the scope defined in the claims.

Claims

1. A zero-sample dialogue state tracking method, the steps comprising: 1) Division stage: For each dialogue state labeled dialogue text C t , using the pre-trained language model f to transform the conversation text C t Converted into dialogue text vector e t , and then use clustering technology to classify each dialogue text vector into its different subsets to obtain K subsets; 2) Solution stage: Take each text vector in the subset as a sample, and use each obtained subset to train a semantically independent state tracking model, and obtain a total of K trained state tracking models; 3) Merging stage: First, perform relationship mining to merge a given conversation text C′ t Converted to semantic vector e′ t , calculate the semantic space and dialogue text C′ of each subset t Then perform aggregate reasoning and predict the dialogue text C′ based on each trained state tracking model and its corresponding relationship δ. t The corresponding dialogue state; among them, the semantic space μ corresponding to the kth subset k and the dialogue text C′ t The relationship between Where d represents the distance function and τ represents the temperature coefficient.

2. The method according to claim 1, characterized in that The corresponding semantic space μ is obtained by averaging all text vectors in the kth subset k .

3. The method according to claim 1 or 2, characterized in that: Use parameter-level aggregate reasoning to predict the conversation text C′ t The corresponding dialogue state is as follows: using the semantic space μ corresponding to the kth subset k and the dialogue text C′ t The relationship between δ(C ′ t ,μ k ) and the state tracking model parameters after training with the kth subset Initialize a new dialogue state tracking model, and then sum up the K new dialogue state tracking models to get the model Then use the model φ' to predict the dialogue text C' t The corresponding dialog status.

4. The method according to claim 1 or 2, characterized in that: Use character-level aggregate reasoning to predict the conversation text C′ t The corresponding dialogue state is as follows: first, the state tracking model trained on the kth subset is used to predict the dialogue text C′ t The corresponding dialogue state π k , and then use the semantic space μ corresponding to the kth subset k and the dialogue text C′ t The relationship between δ(C ′ t ,μ k ) and the corresponding dialogue state π k Predict the dialogue text C′ t The prediction results for each character in the final dialogue state y m The dialogue text C′ t The prediction result of the mth character in .

5. The method according to claim 1 or 2, characterized in that: Use the pre-trained language model T5 as the state tracking model.

6. The method according to claim 1 or 2, characterized in that: The pre-trained language model f is a pre-trained language model T5 or BERT.

7. A server, characterized in that: The method comprises a memory and a processor, wherein the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program comprises instructions for executing each step in the method according to any one of claims 1 to 6.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Implementation method and device of generative dialogue state tracking model

    CN114841069A

  • Text recognition model training method, model training device and electronic equipment

    CN114841148A