Dialogue state tracking method and system based on dynamic comparison and context sorting

Through dynamic comparison and context sorting methods, the slot state prediction of the pre-trained language model is optimized, which solves the accuracy and stability problems of the dialogue state tracking method in complex dialogues and achieves more efficient slot value prediction and dialogue state tracking.

CN120806150APending Publication Date: 2025-10-17FUZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510911845.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-10-17

AI Technical Summary

Technical Problem

Existing dialogue state tracking methods lack accuracy and stability when faced with complex and ambiguous dialogues, making it difficult to effectively predict slot values ​​in open domains.

Method used

A method based on dynamic comparison and context sorting is adopted. By constructing an instruction fine-tuning dataset and quantifying slot contributions, combined with a dynamic comparison mechanism, the slot state prediction of the pre-trained language model is optimized, the generation of incorrect slot values ​​is reduced, and the sensitivity of the dialogue context is improved.

Benefits of technology

It improves the accuracy and stability of dialogue state tracking, enhances the model's ability to distinguish between similar slot values, reduces redundant information interference, ensures the transmission of key information, and improves the generalization ability of tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120806150A_ABST
    Figure CN120806150A_ABST
Patent Text Reader

Abstract

The invention relates to a dialogue state tracking method and system based on dynamic comparison and context sequencing, and the method comprises the steps: A, collecting dialogue data of a user and a system in different scenes, and constructing a data set; b, in a training data preprocessing stage, a data set is finely adjusted based on a data set construction instruction, and a data set for quantifying slot contribution degree is constructed according to a migration mode of a slot state in a dialogue round; in the training stage, model training is carried out based on an instruction fine tuning task and an auxiliary sorting task: the instruction fine tuning task is carried out on the basis of a pre-training language model, so that the instruction fine tuning task adapts to an input format and target output of a dialogue state tracking task; meanwhile, an auxiliary sorting task is carried out, and key clues are captured by using dialogue context information; and step C, in a reasoning stage, inputting dialogue data of a user and a system into the pre-trained language model after fine tuning, and performing state prediction in combination with a dynamic comparison mechanism. The method and the system are beneficial to improving the accuracy of dialogue state tracking.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the field of natural language processing, and particularly relates to a dialogue state tracking method and system based on dynamic comparison and context ranking. BACKGROUND

[0002] Dialogue state tracking (DST) is a crucial module in task-oriented dialogue systems. Its goal is to monitor the user goals hidden in the dialogue history and represent them as dialogue states consisting of a series of (domain, slot, slot value) triplets. The continuous progress of information technology enables people to access information, applications and services almost instantaneously in a wireless connected manner at any time and place. However, while various Internet applications and personal applications make daily life and travel more and more convenient, more and more applications increase the time cost and use difficulty of users, thus increasing the demand for virtual personal intelligent assistants. Dialogue state tracking is a core technology for realizing virtual personal intelligent assistants, but human dialogue is essentially complex and ambiguous, and there is still a long way to go to create an open-domain dialogue artificial intelligence that can face any scenario.

[0003] In order to solve the disadvantage of manually creating features to predict slot values, Zilu et al. proposed a model called neural belief tracker (NBT). NBT uses a convolutional neural network (CNN) instead of manually created features to predict slot values in the word embedding processing part, and the performance has been greatly improved compared with previous DST models. Inspired by this pioneering work, many neural DST methods based on long short-term memory (LSTM) networks and bidirectional gated recurrent unit (BiGRU) networks have been proposed to further improve the performance of neural DST models.

[0004] Researchers usually regard DST as a multi-classification or multi-hop classification task, and propose ontology-based dialogue state tracking methods. Henderson et al. were the first to use a deep learning model in the DST task. They integrated many feature functions (such as SLU score, Rank score, Affirm score, etc.) as inputs to the neural network, and then predicted the probability of each slot value pair.

[0005] To overcome the limitations of fixed ontology, researchers proposed DST methods based on open vocabulary, which not only reduced the model complexity and time complexity of the DST task, but also facilitated the end-to-end training of task-oriented dialogue systems. Wu et al. proposed the TRADE model, which also applied the copying mechanism and used a soft-gated pointer generator to generate slot values based on domain-slot pairs and dialogue context encoding. Cheng et al. proposed the tree encoder-decoder TED architecture, which uses a hierarchical tree structure to represent dialogue states and system behavior. The TED architecture generates the dialogue state of the current round in tree structure according to the dialogue history, dialogue behavior and dialogue state of the last round. Chen et al. established an interactive encoder to model the dependencies between rounds. In addition, they used attention mechanisms to build slot-level context environments for users and systems, respectively, and the state generator copied slot values from the dialogue context based on this.

[0006] Recently, pre-trained language models have received extensive attention from industry and academia, and various pre-trained language models have been released, such as BERT and GPT-2. Since the models are trained on large-scale corpora, they have shown strong performance in downstream tasks. Lee et al. first introduced the BERT model into the DST task and proposed the SUMBT framework, which uses BERT to learn the relationship between slots and dialogue context through slot-word attention mechanisms. Inspired by the SUMBT model, Shan et al. proposed the CHAN architecture, which applies BERT for multi-task learning and generates dialogue states. They first encode the word-level and round-level dialogue content, then retrieve relevant information for each slot from the dialogue content by applying word-level and round-level attention, and predict the corresponding slot values based on the retrieved information. Heck et al. proposed a new DST method TripPy, which does not require maintaining a list of candidate values, and uses three copying mechanisms to extract slot values: directly extracting slot values from user input, copying slot values from system notifications memory that informs the system's operations, and copying slot values from different slots that have already been included in the dialogue state. Hosseini-Asl et al. adopted GPT-2 as a dialogue context encoder and transformed DST into a language generation task, which also achieved good results. SUMMARY

[0007] The purpose of the present application is to provide a dialogue state tracking method and system based on dynamic comparison and context ranking, which can improve the accuracy of dialogue state tracking.

[0008] To achieve the above purpose, the technical scheme adopted by the present application is: a dialogue state tracking method based on dynamic comparison and context ranking, characterized by comprising the following steps:

[0009] Step A: Collect the dialogue data between the user and the system in different scenarios, and construct a dataset TS;

[0010] Step B: In the training data preprocessing stage, construct the instruction fine-tuning dataset based on the dataset TS, and construct the dataset quantifying the slot contribution degree according to the migration mode of the slot state in the dialogue round. In the training stage, the model is trained based on the instruction fine-tuning task and the auxiliary sorting task: the instruction fine-tuning task is performed based on the pre-trained language model to adapt it to the input format and target output of the dialogue state tracking task; at the same time, the auxiliary sorting task is performed to capture key clues using dialogue context information and improve the sensitivity of the pre-trained language model to slot state changes;

[0011] Step C: In the inference stage, input the dialogue data between the user and the system into the fine-tuned pre-trained language model, and perform state prediction combined with the dynamic comparison mechanism.

[0012] Further, the step A specifically comprises the following steps:

[0013] Step A1: Collect the dialogue data between the user and the system in different scenarios; the dialogue data in one scenario includes multiple round dialogues, and each round dialogue includes user utterance and system response, wherein the user utterance represents the input text of the user in one round dialogue, and the system response represents the response of the system to the user utterance; each round dialogue in each scenario contains a set of dialogue state data, which includes:

[0014] Slot: represents a pre-defined information category in the dialogue, used to identify the specific component of the user's intent;

[0015] Slot value: the actual information filled into the slot, representing the specific data corresponding to the slot extracted from the user's input;

[0016] Dialogue state: represented as a set of key-value pairs, where the key represents the slot name and the value represents the slot value;

[0017] Dialogue history: contains the sequence of all user utterances and system responses before the current round;

[0018] Step A2: Take the dialogue data between the user and the system in one scenario as one dialogue sample, and multiple dialogue samples together constitute the dataset TS; one dialogue sample in the dataset TS includes:

[0019] The current dialogue of T round dialogues is represented as D t = (R t , U t ), where R t represents the system response, and U trepresenting a user utterance;

[0020] dialog history of T-round dialogue: the dialog history of the tth round is represented as H t = D1 ⊕ D2 ⊕... ⊕ Di t-1 ;

[0021] dialog state of T-round dialogue: the dialog state of the tth round is represented as a set of key-value pairs where S i is the name of the ith slot, is the slot value corresponding to the ith slot in the tth round, and I is the number of all slots in different domains;

[0022] Then, by collecting the slot values of all slots in the data set TS, a slot value set V s all ;

[0023] Step A3: dividing the data set TS into a training set D train and a validation set D valid , wherein the training set is used for model training and parameter optimization, and the validation set is used to evaluate the performance of the model to prevent overfitting.

[0024] Further, the step B specifically comprises the following steps:

[0025] Step B1: constructing an instruction fine-tuning data set based on the data set TS to improve the understanding and generalization ability of the pre-trained language model for task instructions in the training phase; the instruction fine-tuning data set includes a positive instruction sample set and a reverse instruction sample set, and the samples in the positive instruction sample set and the reverse instruction sample set are composed of three parts: instruction, dialog data and output slot value; the instruction is a pre-designed instruction, and the dialog data is given by the dialog sample; for the output slot value, the output slot value of the positive instruction sample in the positive instruction sample set is the true value, corresponding to the slot value in the dialog sample; the output slot value of the reverse instruction sample in the reverse instruction sample set is constructed as follows: for each non-empty slot, an interference value is constructed: if the slot has a value in the historical dialogue, the historical value is selected as the interference value; if there is no historical value, select the slot value closest to the true value in semantics from all slot values as the interference value; if there is neither historical value nor semantic closest value, randomly select a slot value as the interference value;

[0026] Then divide the instruction fine-tuning data set into an instruction fine-tuning training set and an instruction fine-tuning validation set

[0027] Step B2: constructing a data set D that quantifies the contribution degree of the slot according to the migration mode of the slot state in the dialogue rounds To enable the model to learn the importance of different slots in the dialogue process, the following strategy is introduced when calculating the contribution: if the slot value changes in the current round, a higher contribution is given; if the slot value remains unchanged in the current round, a lower contribution is given to reflect the influence of its stability; in addition, the influence of the contribution is controlled through a time decay mechanism, and the contribution of recent changes is greater than that of long-term changes; then, the contribution of all rounds is normalized to ensure that the total contribution is distributed within a reasonable range;

[0028] Step B3: Construct the instruction fine-tuning task and use the instruction fine-tuning training set obtained in step B1 Perform the instruction fine-tuning task to obtain the loss of the instruction fine-tuning task;

[0029] Step B4: Construct the auxiliary sorting task and use the dataset D constructed in step B2 s Perform the auxiliary sorting task to obtain the loss of the auxiliary sorting task;

[0030] Step B5: Jointly train the instruction fine-tuning task and the auxiliary sorting task for multi-task learning, that is, weight the losses of the two tasks to obtain the total loss, and perform model parameter fine-tuning training; Specifically, the total loss is optimized by LoRA to fine-tune the parameters of the pre-trained language model; During the training optimization process, only the parameters of the low-rank matrix W LoRA are adjusted, and the parameters other than the low-rank matrix W LoRA are frozen, and the instruction fine-tuning verification set is used to verify the training process to prevent overfitting of the instruction fine-tuning task.

[0031] Further, the step B1 specifically includes the following steps:

[0032] Step B11: For the true value of each slot s The following three strategies are used to construct the interference value:

[0033] 1) History value priority strategy;

[0034] If the slot s has appeared in the dialogue history, the historical value is selected as the interference value; Specifically: define the set of historical values of the slot s appearing in the dialogue history as If , the most recent historical value is selected from as the interference value:

[0035]

[0036] Wherein, represents the interference value, v represents the historical value in the set , and f recency(v) represents the time of the historical value v appearing in the dialogue history;

[0037] If , the historical value cannot be selected;

[0038] 2) Semantic similarity priority strategy;

[0039] If there is no historical value, calculate the semantic similarity of each slot value v in the slot value set V s all to the real value , and select the slot value closest in semantics to the real value as the interference value:

[0040]

[0041] where F() is an embedding function, cos() is a cosine similarity, represents the set after removing the real value s all from the slot value set V , to avoid selecting the real value as the slot value;

[0042] Set a similarity threshold τ, if the semantic similarity between the selected slot value and the real value is less than the similarity threshold, i.e. , the semantic closest value cannot be selected;

[0043] 3) Random strategy;

[0044] If the above two strategies cannot select the interference value, randomly select a slot value from the slot value set V s all as the interference value:

[0045]

[0046] where Uniform() is a value randomly selected from the set with equal probability;

[0047] Step B12: After obtaining the interference value , construct the forward instruction sample and the reverse instruction sample in the structure of instruction + dialogue data + output slot value; when constructing the complete data sample, first standardize the dialogue data, and its format is:

[0048] Dialog = [USER] U1 [SYSTEM] S1 [USER] U2 [SYSTEM] S2…

[0049] Wherein, Dialog represents the standardized representation of the dialogue data, [USER] represents the prefix prompt of the user's utterance, U1 is the user's input in the first round, [SYSTEM] represents the prefix prompt of the system's response, and S1 is the system's reply in the first round;

[0050] For the forward instruction sample, the real value is used as the output slot value, which is constructed as follows:

[0051]

[0052] Among them, Instruction_P is the forward instruction;

[0053] For the reverse instruction sample, the interference value is used as the output slot value, which is constructed as follows:

[0054]

[0055] Among them, Instruction_N is the reverse instruction;

[0056] Then construct the forward instruction sample set D from all forward instruction samples pos , construct the reverse instruction sample set D from all reverse instruction samples neg , and then the forward instruction sample set D pos and reverse instruction sample set D neg Construct instruction fine-tuning dataset D i :

[0057] D i =D pos +D neg

[0058] Step B13: Fine-tune the instructions to the dataset D i Divided into instruction fine-tuning training set and instruction fine-tuning validation set

[0059] Furthermore, the step B2 specifically includes the following steps:

[0060] Step B21: Define slot contribution To measure the influence of the slot in the current round; the strategy for calculating contribution is:

[0061] If the slot value changes in the current round, a higher contribution is given;

[0062] If the slot value remains unchanged in the current round, a lower contribution is given to reflect its stability impact;

[0063] The time decay factor is used to control the impact of contribution, and the contribution of recent changes is greater than that of long-term changes;

[0064] Let the slot value of the dialogue in the current turn t be The slot value of the previous turn is The slot contribution is The calculation is as follows:

[0065]

[0066] where Indicator() is an indicator function for determining whether the slot value changes, a is the contribution weight when the slot value changes, and β is the contribution weight when the slot value does not change;

[0067] In order to ensure reasonable modeling of historical dialogues, a time decay factor is further introduced to dynamically adjust the time influence of the contribution:

[0068]

[0069] where the time decay factor γ t is defined as:

[0070] γ t = e -λ(T-t)

[0071] where T represents the total number of turns of the current dialogue, and λ is a parameter for controlling the decay rate;

[0072] Step B22: Normalize the contribution of all turns; in addition, adjust the normalized contribution by a regularization term to control the smoothness of the distribution; let the slot set of the current dialogue be S, and the turn set of the dialogue be T, then the normalized contribution of all slots is :

[0073]

[0074] where ∈ is a small positive number set;

[0075] Step B23: Construct a data set D s quantifying the slot contribution as follows:

[0076] Let the dialogue set be D = {d1, d2, …, d i ,…, d N}, each data d i contains at most T i turns of interaction, T i is the number of turns of the complete dialogue sample to which d i belongs, and the slot set of each turn is S; for the first n turns of dialogue of each data d i , n ≤ T i , the contribution of all slots is calculated, and the data set D s is constructed:

[0077]

[0078] wherein t represents a dialogue turn, s is a slot, is a normalized contribution of slot s in turn t; dataset D s Records the contribution information of each slot in different turns of all dialogues, which is used for subsequent auxiliary ranking tasks.

[0079] Further, the specific implementation method of step B3 is:

[0080] Instruction fine-tuning training set constructed based on step B1 Define the input and output formats of the instruction fine-tuning task; let the input dialogue sequence be X=(x1, x2, …, x k ), k is the length of the input sequence, the target output sequence is Y=(y1, y2, …, y o ), o is the length of the output sequence, and only the part of the sequence containing the slot value position in the target output sequence participates in loss calculation; the fine-tuning target is to maximize the generation probability of the pre-trained language model on the slot filling part, that is, to minimize the loss:

[0081]

[0082] wherein L DST is the instruction fine-tuning task loss, i.e., the main task loss, θ represents the model parameters, P(y i |X, y <i ; θ) is the probability predicted by y <i at the current time step based on the generated sequence y i , and only the part of the sequence containing the slot value position is calculated for loss to avoid unnecessary gradient update.

[0083] Further, step B4 specifically includes the following steps:

[0084] Step B41: For the auxiliary ranking task, let the input dialogue sequence be X=(x1, x2, …, x k ), and insert a special mark [rank] at the end of each round of dialogue, i.e., construct an extended sequence X ′ :

[0085]

[0086] wherein T is the number of turns of the input dialogue sequence, [rank] is inserted after each turn, L1, L2, L T represent the first turn, the second turn, and the Tth turn of dialogue, respectively.

[0087] The extended sequence is input into the pre-trained language model to obtain the hidden layer output:

[0088] H = f θ (X ′ )

[0089] where H is the hidden state of the last layer of the pre-trained language model; f θ () is the pre-trained language model;

[0090] From the hidden layer output H, extract the hidden state representation of all [rank] positions, and let its index set be R = {L1, L2, …, L T}, then get the hidden state representation set H R :

[0091] H R = (h1, h2, …, h t ,…, h T )

[0092] where h t represents the hidden state representation of the tth round of dialogue;

[0093] Step B42: Calculate the predicted ranking scores of all rounds of dialogue and form a predicted ranking score sequence:

[0094] r ′ = H R W rank = (r ′ 1, r ′ 2, …, r t ′ ,…, r ′ T )

[0095] where W rank ∈ R d×1 is a ranking weight matrix, which is updated with the training of the model, r t ′ is the ranking score of the tth round of dialogue in the input dialogue predicted by the pre-trained language model, and r ′ is the ranking score sequence of all rounds of dialogue in the input dialogue predicted by the pre-trained language model;

[0096] Step B43: In the contribution data set D s constructed in step B2, the normalized contribution of each slot is recorded for each round. According to the input dialogue sequence X, the total number of rounds T, and the predicted slot s, the contribution of all rounds before round T is obtained from the data set D s , which is directly used as the target ranking score sequence:

[0097]

[0098] wherein, is the target ranking score of the t-th round of dialogue, * is a sequence of target ranking scores;

[0099] Step B44: Calculate the difference between the predicted ranking score and the target ranking score using mean square error:

[0100]

[0101] wherein, L rank is the auxiliary ranking task loss.

[0102] Further, the step B5 specifically comprises the following steps:

[0103] Step B51: Calculate the total loss L DST by combining the instruction fine-tuning task loss L ranK and the auxiliary ranking task loss L total :

[0104] L total = L DST + λL rank

[0105] wherein, λ is a hyperparameter used to adjust the relative importance of the auxiliary ranking task and the instruction fine-tuning task;

[0106] Use the total loss L total for model training, and optimize the model parameters through backpropagation;

[0107] Step B52: To optimize the total loss L toTaL , introduce LoRA parameters to efficiently fine-tune the pre-trained language model parameters; let the weight matrix of the pre-trained language model be W, and LoRA only performs low-rank decomposition on the weights of the specified layer, that is, introduce trainable parameters A and B for adjustment, so that the weight change ΔW updated by the model during fine-tuning is:

[0108] ΔW = AB

[0109] During forward propagation, the effective weight of the model becomes:

[0110] W' = W + ΔW = W + AB

[0111] wherein, W remains frozen, and only A and B are adjusted to adapt to the task requirements, ensuring that fine-tuning only targets task-related features; in addition, the instruction fine-tuning validation set is used to verify the training process to prevent overfitting of the instruction fine-tuning task.

[0112] Further, the step C specifically comprises the following steps:

[0113] Step C1: In the inference process of each data, first construct the dialogue context containing two different instructions as the model input; mark the input containing the forward instruction as x pos ; mark the input containing the reverse instruction as x neg ; To ensure that the model generates more stable states in the inference stage, the model compares the logits generated by the forward instruction with the logits generated by the reverse instruction, and calculates the final logits through a dynamic comparison mechanism:

[0114] L t = logits(x t |x p x <t ; θ) - β * logits(x t |x r x <t ; θ)

[0115] Where logits is the unnormalized probability distribution calculated by the model for the input, and β is the dynamic suppression factor, which controls the influence of the logits generated by the reverse prompt on the final logits.

[0116] The final probability distribution is normalized by softmax:

[0117] p(x t |x p x <t ; θ) := softmax(L t )

[0118] Step C2: At each time step, calculate the softmax probability distribution P and Q according to the logits of the next token generated by the forward instruction and the reverse instruction, respectively, and further calculate the JS divergence:

[0119]

[0120] Where KL() is used to calculate the KL divergence of two probability distributions.

[0121] Normalize the JS divergence to the interval [0, 1] to get:

[0122]

[0123] Then, set the dynamic adjustment coefficient as:

[0124] β = 1 - JS norm

[0125] The closer beta is to 1, the more similar the distribution generated by the positive instruction and the reverse instruction, thereby enhancing the inhibition of the reverse prompt; on the contrary, the smaller beta is, the greater the difference between the distribution generated by the positive instruction and the reverse instruction, thereby reducing the influence of the reverse prompt.

[0126] The application further provides a dialogue state tracking system based on dynamic comparison and context ordering, comprising a memory, a processor and computer program instructions stored in the memory and capable of being executed by the processor, when the processor executes the computer program instructions, the above-mentioned method steps can be realized.

[0127] Compared with the prior art, the application has the following beneficial effects:

[0128] 1) The application adopts a dynamic comparison mechanism, generates correct and incorrect slot values by constructing different instruction guidance models, and dynamically compares the probability distribution, thereby effectively suppressing the generation of incorrect slot values. This method enhances the discrimination ability of the generative dialogue state tracking (DST) model between similar slot values, enables the model to more accurately select the slot value that best matches the current dialogue context, reduces incorrect predictions, and improves the accuracy of dialogue state tracking.

[0129] 2) The application optimizes the way of using historical dialogue information through an auxiliary ordering task, so that the model can pay more attention to the historical dialogue turns related to the current slot while retaining the complete context. This method reduces the interference of redundant information while ensuring the transmission of key information, improves the stability and generalization ability of the dialogue state tracking task. BRIEF DESCRIPTION OF DRAWINGS

[0130] Figure 1 is the implementation flowchart of the dialogue state tracking method based on dynamic comparison and context ordering provided by the embodiment of the application;

[0131] Figure 2 is the implementation principle diagram of the dialogue state tracking method based on dynamic comparison and context ordering provided by the embodiment of the application. DETAILED DESCRIPTION

[0132] The application will be further described below in conjunction with the drawings and embodiments.

[0133] It should be noted that the following detailed description is exemplary and is intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as generally understood by those skilled in the art to which the present application belongs.

[0134] It is to be noted that the terms used herein are merely for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise, and it will be further understood that the terms "comprise" and / or "include" when used in this specification, specify the presence of stated features, steps, operations, devices, components and / or combinations thereof.

[0135] As shown in Figure 1 The embodiment provides a dialogue state tracking method based on dynamic comparison and context ranking, comprising the following steps:

[0136] Step A: Collect dialogue data of users and systems in different scenarios to construct a data set TS;

[0137] Step B: In the training data preprocessing stage, construct an instruction fine-tuning data set based on the data set TS, and construct a data set quantifying the contribution of slots according to the migration mode of slot states in the dialogue round. In the training stage, the model is trained based on the instruction fine-tuning task and the auxiliary ranking task: perform the instruction fine-tuning task based on the pre-trained language model to adapt it to the input format and target output of the dialogue state tracking task, and improve the instruction following ability; at the same time, perform the auxiliary ranking task with context awareness, capture key clues using dialogue context information, and improve the sensitivity of the pre-trained language model to slot state changes, thereby optimizing the prediction accuracy;

[0138] Step C: In the inference stage, input the dialogue data of the user and the system into the fine-tuned pre-trained language model, and perform state prediction combined with the dynamic comparison mechanism; the dynamic comparison mechanism effectively reduces the generation of incorrect states, so that the model can more accurately output the current dialogue state, and improve the reliability and stability of the dialogue state tracking task.

[0139] In the embodiment, the step A specifically comprises the following steps:

[0140] Step A1: Collect dialogue data of users and systems in different scenarios; the dialogue data in one scenario includes multiple round dialogues, and each round dialogue includes user utterance and system response, wherein the user utterance represents the input text of the user in one round dialogue, and the system response represents the response of the system to the user utterance; each round dialogue in each scenario contains a set of dialogue state data, which includes:

[0141] Slot: represents a pre-defined information category in the dialogue, used to identify specific components of the user's intent;

[0142] Slot value: actual information filled into the slot, representing specific data corresponding to the slot extracted from the user's input;

[0143] Dialog state: represented as a set of key-value pairs, where the key represents the slot name and the value represents the slot value;

[0144] Dialog history: contains the sequence of all user utterances and system responses before the current turn.

[0145] Step A2: Take the dialog data between the user and the system in a scene as a dialog sample, and multiple dialog samples collectively constitute a dataset TS; a dialog sample in the dataset TS includes:

[0146] Current dialog of T-turn dialog: the current dialog of the t-turn is represented as D t = (R t , U t ), where R t represents the system response, and U t represents the user utterance;

[0147] Dialog history of T-turn dialog: the dialog history of the t-turn is represented as H t = D1 t-1 ;

[0148] Dialog state of T-turn dialog: the dialog state of the t-turn is represented as a set of key-value pairs where S i is the name of the i-th slot, is the slot value corresponding to the i-th slot in the t-turn, and I is the number of all slots in different domains.

[0149] Then, by collecting the slot values of all slots in the dataset TS, a slot value set V s all is formed.

[0150] Step A3: Divide 80% of the data in the dataset TS into a training set D train and 20% of the data into a validation set D valid , where the training set is used for model training and parameter optimization, and the validation set is used to evaluate the performance of the model to prevent overfitting.

[0151] Figure 2 is the implementation schematic diagram of the dialog state tracking based on dynamic comparison and context ranking provided in the embodiment. As shown in Figure 2 , the specific implementation steps of step B are as follows.

[0152] Step B1: constructing an instruction fine-tuning dataset based on the dataset TS to improve the understanding and generalization ability of the pre-trained language model for task instructions in the training stage; the instruction fine-tuning dataset includes a positive instruction sample set and a reverse instruction sample set, and the samples in the positive instruction sample set and the reverse instruction sample set are each composed of three parts: an instruction, dialogue data, and an output slot value; the instruction is a pre-designed instruction, and the dialogue data is given by a dialogue sample; for the output slot value, the output slot value of the positive instruction sample in the positive instruction sample set is the true value, corresponding to the slot value in the dialogue sample; the output slot value of the reverse instruction sample in the reverse instruction sample set is constructed as follows: for each non-empty slot, an interference value is constructed: if the slot has appeared in the historical dialogue, the historical value is selected as the interference value; if there is no historical value, the slot value closest in semantics to the true value is selected from all slot values as the interference value; if there is neither a historical value nor a value closest in semantics, a slot value is randomly selected as the interference value.

[0153] The instruction fine-tuning dataset is then divided into an instruction fine-tuning training set and an instruction fine-tuning validation set

[0154] In this embodiment, the step B1 specifically includes the following steps:

[0155] Step B11: for each slot s, the true value The following three strategies are used to construct the interference value:

[0156] 1) Historical value priority strategy

[0157] If the slot s has appeared in the dialogue history, the historical value is selected as the interference value; specifically: define the set of historical values of the slot s appearing in the dialogue history as If , the most recent historical value is selected from as the interference value:

[0158]

[0159] wherein, indicates the interference value, v indicates the historical value in the set , and f recency (v) indicates the calculation of the occurrence time of the historical value v in the dialogue history.

[0160] If , the historical value cannot be selected.

[0161] 2) Semantic similarity priority strategy

[0162] If there is no historical value, the slot value set V s alleach slot value v and the true value The semantic similarity of each slot value v and the true value

[0163]

[0164] where F() is an embedding function using an embedding model (such as Sentence-BERT), and cos() is the cosine similarity, The set of slot values V s all from which the true value is removed to avoid selecting the true value as the interference value.

[0165] A similarity threshold τ is set. If the semantic similarity of the selected slot value and the true value is less than the similarity threshold, i.e. the semantic closest value cannot be selected.

[0166] 3) Random strategy

[0167] If the above two strategies cannot select the interference value, a slot value is randomly selected from the set of slot values V s all as the interference value:

[0168]

[0169] where Uniform() is a function that randomly selects a value from a set with equal probability.

[0170] After obtaining the interference value , forward instruction samples and reverse instruction samples are constructed in the structure of instruction + dialogue data + output slot value. When constructing complete data samples, the dialogue data is first standardized and represented in the format:

[0171] Dialog = [USER] U1 [SYSTEM] S1 [USER] U2 [SYSTEM] S2…

[0172] where Dialog represents the standardized representation of the dialogue data, [USER] represents the prefix prompt of the user's speech, U1 is the input of the user in the first round, [SYSTEM] represents the prefix prompt of the system response, and S1 is the reply of the system in the first round.

[0173] For forward instruction samples, the true value is used as the output slot value, which is constructed as follows:

[0174]

[0175] Among them, Instruction_P is a forward instruction, specifically: "Please track the slots in the conversation and the corresponding slot values."

[0176] For the reverse instruction sample, the interference value is used as the output slot value, which is constructed as follows:

[0177]

[0178] Among them, Instruction_N is the reverse instruction, specifically: "Please track the slots in the conversation and the corresponding slot values, but deliberately return the wrong slot values."

[0179] Then construct the forward instruction sample set D from all forward instruction samples pos , construct the reverse instruction sample set D from all reverse instruction samples neg , and then the forward instruction sample set D pos and reverse instruction sample set D neg Construct instruction fine-tuning dataset D i :

[0180] D i =D pos +D neg

[0181] Step B13: Fine-tune the instructions to dataset D i Divided into instruction fine-tuning training set and instruction fine-tuning validation set

[0182] Step B2: Based on the transition pattern of slot states in the dialogue rounds, construct a dataset D to quantify slot contribution s , so that the model can learn the importance of different slots in the dialogue process; when calculating the contribution, the following strategy is introduced: if the slot value changes in the current round, a higher contribution is given; if the slot value remains unchanged in the current round, a lower contribution is given to reflect the influence of its stability; in addition, the influence of the contribution is controlled by the time decay mechanism, and the contribution of recent changes is greater than that of long-term changes; then, the contribution of all rounds is normalized to ensure that the total contribution is distributed within a reasonable range, avoiding the excessively high or low contribution values ​​of some slots affecting the overall weight of the model.

[0183] In this embodiment, step B2 specifically includes the following steps:

[0184] Step B21: In the dialogue state tracking task, the state changes of each slot s in different rounds t are crucial to understanding the dialogue; therefore, the slot contribution is defined as To measure the influence of the slot in the current round; the strategy for calculating contribution is:

[0185] If the slot value changes in the current turn, a higher contribution is given;

[0186] If the slot value remains unchanged in the current turn, a lower contribution is given to reflect the influence of its stability;

[0187] The influence of the contribution is controlled by a time decay factor, and the contribution of recent changes is greater than that of long-term changes.

[0188] Let the slot value of the dialogue in the current turn t be The slot value of the previous turn is The slot contribution is The calculation is as follows:

[0189]

[0190] Where Indicator() is an indicator function for determining whether the slot value changes, a is the contribution weight when the slot value changes, and β is the contribution weight when the slot value does not change.

[0191] In order to ensure reasonable modeling of historical dialogues, a time decay factor is further introduced to dynamically adjust the time influence of the contribution:

[0192]

[0193] Where the time decay factor γ t is defined as:

[0194] γ t = e -λ(T-t)

[0195] Where T represents the total number of turns of the dialogue, and λ is a parameter for controlling the decay rate. When λ takes a small value, the contribution of earlier turns still has a high influence; otherwise, the contribution of earlier turns decays rapidly, retaining only recent information.

[0196] Step B22: Since the contributions of different slots can have different numerical ranges, in order to avoid some slots having too large or too small contributions to the whole, the contributions of all turns are normalized; in addition, the normalized contributions are adjusted by a regularization term to control the smoothness of the distribution; let the slot set of the current dialogue be S, then the normalized contribution of all slots is :

[0197]

[0198] This normalization ensures that the sum of the contributions of all slots is 1, making the contributions of different slots comparable, while avoiding the influence of individual slots on the modeling of the dialogue state being too large; in addition, ε is a small positive number set (such as 10 -6), for preventing zero division error and ensuring the numerical stability of the normalized contribution degree.

[0199] Step B23: Constructing the data set D of quantifying the contribution degree of the slot s As follows:

[0200] Let the dialogue set be D = {d1, d2, …, d i ,…,d N}, and the data d i in the dialogue set D represent the ith dialogue sample, and the data d i contains T i rounds of dialogue, T i is the number of rounds of the ith dialogue sample, the slot set S of each round of dialogue, and N is the total number of dialogue samples; for the first t rounds of dialogue of the data d i , the contribution degrees of all slots are calculated, and the data set D s is constructed:

[0201]

[0202] wherein t represents the round of dialogue, s represents the slot, , and represents the normalized contribution degree of the slot s in the round t; the data set D s records the contribution degree information of each slot in different rounds of all dialogues, which is used for subsequent auxiliary sorting tasks.

[0203] Step B3: Constructing the instruction fine-tuning task and using the instruction fine-tuning training set obtained in step B1 to perform the instruction fine-tuning task, and obtaining the loss of the instruction fine-tuning task.

[0204] In the embodiment, the specific implementation method of the step B3 is as follows:

[0205] The instruction fine-tuning training set constructed based on the step B1 defines the input and output formats of the instruction fine-tuning task; let the input dialogue sequence be X = (x1, x2, …, x k ), k be the length of the input sequence, the target output sequence be Y = (y1, y2, …, y o ), o be the length of the output sequence, and only the part of the sequence containing the slot value position in the target output sequence participate in the loss calculation; the fine-tuning target is to maximize the generation probability of the pre-trained language model on the slot filling part, that is, to minimize the loss:

[0206]

[0207] wherein L DST is the loss of the instruction fine-tuning task, that is, the main task loss, θ represents the model parameter, and P(y i |X,y<i ; θ) is the probability of the generated sequence y <i at the current time step, only the partial sequence containing the slot value position is calculated to avoid unnecessary gradient updates. i

[0208] Step B4: build an auxiliary ranking task and use the dataset D s built in step B2 to perform the auxiliary ranking task to obtain the loss of the auxiliary ranking task.

[0209] In this embodiment, step B4 specifically comprises the following steps:

[0210] Step B41: for the auxiliary ranking task, set the input dialogue sequence as X = (x1, x2, …, x k ), insert a special mark [rank] at the end position of each round of dialogue, that is, construct an extended sequence X ′ :

[0211]

[0212] wherein T is the number of rounds of the input dialogue sequence, [rank] is inserted after each round, L1, L2, L T represent the first round, the second round, and the Tth round of dialogue, respectively.

[0213] The extended sequence is input into the pre-trained language model to obtain the hidden layer output:

[0214] H = f θ (X ′ )

[0215] wherein H ∈ R (k+T)×d is the hidden state of the last layer of the pre-trained language model, d is the dimension of the hidden layer; f θ () is the pre-trained language model.

[0216] From the hidden layer output H, extract the hidden state representation of all [rank] positions, set its index set as R = {L1, L2, …, L T}, and obtain the hidden state representation set H R :

[0217] H R = (h1, h2, …, h t , …, h T )

[0218] wherein H R ∈ R T×d , h t represents the hidden state representation of the tth round of dialogue.

[0219] ​Step B42: Calculate the predicted ranking scores of all dialogue turns and form a predicted ranking score sequence:

[0220] r ′ =H R W rank =(r ′ 1,r ′ 2,…,r t ′ ,…,r ′ T )

[0221] Among them, W rank ∈R d×1 is the sorting weight matrix, which is updated as the model is trained, r t ′ is the ranking score of the t-th round of the input dialogue predicted by the pre-trained language model, r ′ Sequence of ranked scores for all dialogue turns in the input dialogue predicted by the pre-trained language model.

[0222] Step B43: Contribution dataset D constructed in step B2 s In each round of each data, the normalized contribution of each slot is recorded. According to the input dialogue sequence X, the total number of rounds T and the predicted slot s, the data set D s The contribution of all rounds before round T is obtained and directly used as the target ranking score sequence:

[0223]

[0224] in, is the target ranking score for the t-th round of dialogue, r * Sort the score sequence for the target;

[0225] Step B44: Calculate the difference between the predicted ranking score and the target ranking score using the mean square error (MSE):

[0226]

[0227] Among them, L rank This loss directly reflects the model's fitting effect on the target ranking score calculated by slot-normalized contribution in step B2, thereby guiding the model to optimize its ability to rank the importance of each round of dialogue.

[0228] Step B5: Jointly train the instruction fine-tuning task and the auxiliary ranking task for multi-task learning, i.e., weight the losses of the two tasks to obtain a total loss, and perform model parameter fine-tuning training; specifically, optimize the total loss through LoRA (Low-Rank Adaptation) to efficiently fine-tune the parameters of the pre-trained language model; during the training optimization process, only the parameters of the low-rank matrix W LoRA are adjusted, and the parameters other than the low-rank matrix W LoRA are frozen to reduce computational overhead and improve training efficiency, while the instruction fine-tuning validation set is used to verify the training process to prevent overfitting of the instruction fine-tuning task; the goal is to learn task-related features through LoRA to optimize the model's performance in the dialogue state tracking task; through multi-task learning, it is ensured that the model can not only accurately predict slot values but also effectively model dialogue context information, thereby optimizing the overall dialogue state tracking performance.

[0229] In the present embodiment, step B5 specifically comprises the following steps:

[0230] Step B51: Calculate the total loss L DST by the instruction fine-tuning task loss L rank and the auxiliary ranking task loss L total :

[0231] L total =L DST +λL rank

[0232] where λ is a hyperparameter used to adjust the relative importance of the auxiliary ranking task and the instruction fine-tuning task, which is usually adjusted through cross-validation or experience.

[0233] Use the total loss L total for model training to optimize the model parameters through backpropagation.

[0234] Step B52: To optimize the total loss L total , introduce LoRA parameters to efficiently fine-tune the pre-trained language model parameters; let the weight matrix of the pre-trained language model be W, and LoRA only performs low-rank decomposition on the weights of the specified layer, i.e., introduce trainable parameters A and B for adjustment, so that the weight change ΔW updated by the model during the fine-tuning process is:

[0235] ΔW=AB

[0236] where A∈R d×r and B∈R r×d , d is the specified hidden dimension of the original pre-trained language model, and r is the rank specified by LoRA, r << d to ensure that the computational load of parameter updating is small.

[0237] In the forward propagation process, the effective weight of the model becomes:

[0238] W' = W + AW = W + AB

[0239] where W remains frozen, and only A and B are adjusted to adapt to the task requirements, ensuring that fine-tuning is only for task-related features; in addition, the instruction fine-tuning validation set The training process is verified to prevent overfitting of the instruction fine-tuning task.

[0240] As Figure 2 shown, the specific implementation steps of step C are as follows.

[0241] Step C1: In the inference process of each piece of data, first construct a dialogue context containing two different instructions as the model input; for the input containing the forward instruction, denoted as x pos ; for the input containing the reverse instruction, denoted as x neg ; to ensure that the model generates more stable states in the inference stage, the model compares the logits generated by the forward instruction with the logits generated by the reverse instruction, and calculates the final logits through a dynamic comparison mechanism:

[0242] L t = logits(x t |x p x <t ; θ) - β * logits(x t |x r x <t ; θ)

[0243] Where logits is the unnormalized probability distribution calculated by the model for the input, and β is the dynamic suppression factor, which controls the influence of the logits generated by the reverse prompt on the final logits.

[0244] The final probability distribution is normalized by softmax:

[0245] p(x t |x p x <t ; θ) := softmax(L t )

[0246] Step C2: At each time step, the softmax probability distributions P and Q are calculated according to the logits of the next token generated by the forward instruction and the reverse instruction, respectively, and the JS divergence is further calculated:

[0247]

[0248] wherein KL() is used to calculate the KL divergence of two probability distributions.

[0249] Since the maximum value of JS divergence under the natural logarithm is ln2, the JS divergence is normalized to the interval [0, 1] to obtain:

[0250]

[0251] Then, the dynamic adjustment coefficient is set as:

[0252] β = 1 - JS norm

[0253] In this way, when the distribution generated by the positive instruction and the negative instruction is more similar (the JS divergence is smaller), β is closer to 1, thereby enhancing the inhibition of the negative prompt; on the contrary, when the distribution generated by the positive instruction and the negative instruction is more different, β is smaller, thereby reducing the influence of the negative prompt; this mechanism realizes the dynamic balance of the positive and negative prompts, and further improves the accuracy of the finally generated token.

[0254] In this embodiment, the performance of the method and other baseline methods on the MultiWOZ2.1 dataset and the like is compared, and the results are shown in Table 1. As can be seen from Table 1, the performance of the method is better than that of other methods.

[0255] Table 1 Comparison of performance of the method and baseline methods on MultiWOZ2.1 and the like

[0256]

[0257] The embodiment also provides a dialogue state tracking system based on dynamic comparison and context ranking, which includes a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor, and when the processor executes the computer program instructions, the method steps described above can be realized.

[0258] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can be in the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can be in the form of a computer program product implemented on one or more computer usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer usable program code.

[0259] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 one or more flow or blocks

[0260] These computer program instructions can also be stored in a computer readable memory that can direct a computer or other programmable data processing apparatus to function in a particular manner, such that the instructions stored in the computer readable memory produce an article of manufacture including instructions which implement the function specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 one or more flow or blocks

[0261] The computer program instructions can also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process such that the instructions which execute on the computer or other programmable apparatus provide steps for implementing the functions specified in the flowchart block or blocks. Figure 1 one or more flow or blocks Figure 1 one or more flow or blocks

[0262] The above description is only preferred embodiments of the present application, and is not intended to limit the present application to other forms described above. Any person skilled in the art can make modifications or improvements to the above-mentioned technical contents disclosed in the present application, or make equivalent embodiments with equivalent changes. However, any simple modification, equivalent change and modification made according to the technical essence of the present application to the above-mentioned embodiments, without departing from the technical solution of the present application, still falls within the protection scope of the present application.

Claims

1. A method for tracking conversation states based on dynamic comparison and context sorting, characterized in that: The following steps are involved: Step A: Collect conversation data between users and the system in different scenarios and build a dataset TS; Step B: During the training data preprocessing phase, a command fine-tuning dataset is constructed based on the TS dataset. Furthermore, a dataset for quantifying slot contributions is constructed based on the slot state transition patterns during conversational turns. During the training phase, model training is performed based on instruction fine-tuning and auxiliary sorting tasks. The instruction fine-tuning task is performed on the pre-trained language model to adapt it to the input format and target output of the dialogue state tracking task. Simultaneously, the auxiliary sorting task utilizes dialogue context information to capture key clues and improve the pre-trained language model's sensitivity to slot state changes. Step C: In the inference phase, the conversation data between the user and the system is input into the fine-tuned pre-trained language model, and the state prediction is performed in combination with the dynamic comparison mechanism.

2. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 1, characterized in that: The step A specifically comprises the following steps: Step A1: Collect conversation data between the user and the system in different scenarios. The conversation data in one scenario includes multiple conversation rounds. Each conversation round includes user utterances and system responses. User utterances represent the user's input text in a conversation round, and system responses represent the system's response to the user utterances. Each conversation round in each scenario contains a set of conversation state data, which includes: Slot: represents a predefined information category in a conversation and is used to identify a specific component of user intent; Slot value: The actual information filled in the slot, representing the specific data corresponding to the slot extracted from the user input; Dialog state: represented as a set of key-value pairs, where the key represents the slot name and the value represents the slot value; Dialogue history: Contains the sequence of all user utterances and system responses before the current turn; Step A2: The conversation data between the user and the system in a scenario is taken as a conversation sample. Multiple conversation samples together constitute a data set TS. A conversation sample in the data set TS includes: The current dialogue of T rounds of dialogue: The current dialogue of round t is represented by D t =(R t ,U t ), where R t Indicates the system response, U t Represents user utterances; The dialogue history of T rounds of dialogue: The dialogue history of the tth round is expressed as The dialogue state of T rounds of dialogue: The dialogue state of the tth round is represented as a set of key-value pairs Among them S i is the name of the i-th slot, is the slot value corresponding to the i-th slot in round t, and I is the number of all slots in different fields; Then, by collecting the slot values ​​of all slots in the data set TS, a slot value set V is formed s all ; Step A3: Divide the dataset TS into training set D by random sampling train and validation set D valid The training set is used for model training and parameter optimization, and the validation set is used to evaluate the performance of the model to prevent overfitting.

3. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 1, characterized in that: The step B specifically comprises the following steps: Step B1: Construct an instruction fine-tuning dataset based on the dataset TS to improve the pre-trained language model's understanding and generalization ability of task instructions during the training phase; the instruction fine-tuning dataset includes a forward instruction sample set and a reverse instruction sample set, and the samples in the forward instruction sample set and the reverse instruction sample set are composed of three parts: instructions, dialogue data, and output slot values; the instructions are pre-designed instructions, and the dialogue data are given by dialogue samples; for the output slot values, the output slot values ​​of the forward instruction samples in the forward instruction sample set are true values, corresponding to the slot values ​​in the dialogue samples; the output slot values ​​of the reverse instruction samples in the reverse instruction sample set are constructed as follows: for each non-empty slot, construct an interference value: if the slot has a value in the historical dialogue, select the historical value as the interference value; if there is no historical value, select the slot value that is semantically closest to the true value from all slot values ​​as the interference value; if there is neither a historical value nor a semantically closest value, randomly select a slot value as the interference value; Then the instruction fine-tuning dataset is divided into the instruction fine-tuning training set and instruction fine-tuning validation set Step B2: Based on the transition pattern of slot states in the dialogue rounds, construct a dataset D to quantify slot contribution s , so that the model can learn the importance of different slots in the dialogue process; when calculating the contribution, the following strategy is introduced: if the slot value changes in the current round, a higher contribution is given; if the slot value remains unchanged in the current round, a lower contribution is given to reflect its stability; in addition, the influence of the contribution is controlled by the time decay mechanism, so that the contribution of recent changes is greater than that of long-term changes; then, the contribution of all rounds is normalized to ensure that the total contribution is distributed within a reasonable range; Step B3: Construct the instruction fine-tuning task and use the instruction fine-tuning training set obtained in step B1 Perform the instruction fine-tuning task and obtain the loss of the instruction fine-tuning task; Step B4: Construct an auxiliary sorting task and use the dataset D constructed in step B2 s Perform auxiliary sorting tasks and obtain the loss of the auxiliary sorting tasks; Step B5: Combine the training instruction fine-tuning task and the auxiliary sorting task for multi-task learning, that is, weight the losses of the two tasks to obtain the total loss, and perform model parameter fine-tuning training; specifically, fine-tune the parameters of the pre-trained language model by optimizing the total loss through LoRA; during the training optimization process, only adjust the low-rank matrix W LoRA The parameters of the low-rank matrix W are frozen. LoRA Parameters other than , and use instructions to fine-tune the validation set The training process is validated to prevent overfitting on the instruction fine-tuning task.

4. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 3, characterized in that: The step B1 specifically includes the following steps: Step B11: For each slot s, the true value The following three strategies are used to construct interference values: 1) Historical value priority strategy; If the slot s has a value in the conversation history, the historical value is selected as the interference value first; specifically, the set of historical values ​​that the slot s has appeared in the conversation history is defined as like Then from Select the most recent historical value as the interference value: in, represents the interference value, v represents the set The historical value, f recency (v) represents the time when the calculated history value v appears in the conversation history; like Then the historical value cannot be selected; 2) semantic similarity priority strategy; If there is no historical value, calculate the slot value set V s all Each slot value v and the true value The semantic similarity of , and select the slot value closest to the true value as the interference value: Among them, F() is the embedding function, cos() is the cosine similarity, Represents the value from the slot set V s all Remove the true value The set after that is used to avoid the selected slot value being the real value; Set the similarity threshold τ. If the semantic similarity between the selected slot value and the true value is less than the similarity threshold, that is, Then the semantically closest value cannot be selected; 3) Random strategy; If the above two strategies cannot select the interference value, then randomly select the value from the slot value set. Select a slot value as the interference value: Among them, Uniform() is to randomly select a value from the set with equal probability; Step B12: Obtaining the interference value After that, construct the forward instruction sample and the reverse instruction sample with the structure of instruction + dialogue data + output slot value. When constructing the complete data sample, first standardize the dialogue data, and its format is: Dialog=[USER]U1[SYSTEM]S1[USER]U2[SYSTEM]S2…[USER]U T [SYSTEM]S T Among them, Dialog represents the standardized dialog data, [USER] represents the prefix prompt of the user's speech, [SYSTEM] represents the prefix prompt of the system response, and U t is the user’s input in round T, S T is the system’s response in round T; For the forward instruction sample, the real value is used as the output slot value, which is constructed as follows: Among them, Instruction_P is the forward instruction; For the reverse instruction sample, the interference value is used as the output slot value, which is constructed as follows: Among them, Instruction_N is the reverse instruction; Then construct the forward instruction sample set D from all forward instruction samples pos , construct the reverse instruction sample set D from all reverse instruction samples neg , and then the forward instruction sample set D pos and reverse instruction sample set D neg Construct instruction fine-tuning dataset D i : D i =D pos +D neg Step B13: Fine-tune the instructions to dataset D i Divided into instruction fine-tuning training set and instruction fine-tuning validation set 5. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 4, characterized in that: The step B2 specifically includes the following steps: Step B21: Define slot contribution To measure the influence of the slot in the current round; the strategy for calculating contribution is: If the slot value changes in the current round, a higher contribution is given; If the slot value remains unchanged in the current round, a lower contribution is given to reflect its stability impact; The time decay factor is used to control the impact of contribution, and the contribution of recent changes is greater than that of long-term changes; Let the slot value of the conversation in the current round t be The slot value of the previous round is Slot contribution The calculation is as follows: Among them, Indicator() is an indicator function used to determine whether the slot value changes, α is the contribution weight when the slot value changes, and β is the contribution weight when the slot value remains unchanged; In order to ensure the reasonable modeling of historical conversations, a time decay factor is further introduced to dynamically adjust the time impact of contribution: Among them, the time decay factor γ t Defined as: c t =e -λ(T-t) Where T represents the total number of rounds in the current dialogue, and λ is the parameter that controls the decay rate; Step B22: Normalize the contribution of all rounds; in addition, adjust the normalized contribution through the regularization term to control the smoothness of the distribution; let the slot set of the current dialogue be S and the dialogue round set be T, then the normalized contribution of all slots is for: Among them, ∈ is a set minimum positive number; Step B23: Construct a dataset D to quantify slot contribution s as follows: Let the conversation set be D={d1,d2,…,d i ,…,d N }, each data d i Contains up to T i Wheel interaction, T i d i The number of rounds of the complete conversation sample, the slot set of each round is S; for each data d i The first n rounds of dialogue, n≤T i , calculate the contribution of all slots and construct the data set D s : Among them, t represents the conversation turn, s is the slot, is the normalized contribution of slot s in round t; dataset D s Record the contribution information of each slot in different rounds of all conversations for subsequent auxiliary sorting tasks.

6. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 5, characterized in that: The specific implementation method of step B3 is: Based on the instruction fine-tuning training set constructed in step B1 Define the input and output format of the instruction fine-tuning task; suppose the input dialogue sequence is X=(x1,x2,…,x k ), k is the length of the input sequence, and the target output sequence is Y=(y1,y2,…,y o ), o is the output sequence length, and only the part of the target output sequence containing the slot value position participates in the loss calculation; the fine-tuning goal is to maximize the generation probability of the pre-trained language model for the slot filling part, that is, to minimize the loss: Among them, L DST is the instruction fine-tuning task loss, that is, the main task loss, θ represents the model parameters, P(y i |X,y <i ; θ) is the generated sequence y at the current time step <i Prediction y i The probability of , is used to calculate the loss only for the partial sequence containing the slot value position to avoid unnecessary gradient updates.

7. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 6, characterized in that: The step B4 specifically includes the following steps: Step B41: For the auxiliary sorting task, let the input dialogue sequence be X = (x1, x2, ..., x k ), insert a special marker [rank] at the end of each round of dialogue, that is, construct an extended sequence X ′ : Where T is the number of rounds of the input dialogue sequence, and [rank] is inserted after each round, L1, L2, L T Respectively represent the 1st, 2nd, and Tth rounds of dialogue; Feed the expanded sequence as input to the pre-trained language model and get the hidden layer output: H=f θ (X ′ ) Among them, H is the hidden state of the last layer of the pre-trained language model; f θ () is the pre-trained language model; From the hidden layer output H, extract the hidden state representation of all [rank] positions, and set its index set to R = {L1, L2, ..., L T }, then we get the hidden state representation set H R : H R =(h1,h2,…,h t ,…,h T ) Among them, h t Represents the hidden state representation of the t-th round of dialogue; Step B42: Calculate the predicted ranking scores of all dialogue turns and form a predicted ranking score sequence: r′=H R W rank =(r′1,r′2,…,r t ′,…,r′ T ) Among them, W rank ∈R d×1 is the sorting weight matrix, which is updated as the model is trained, r′ t is the ranking score of the tth round of dialogue in the input dialogue predicted by the pre-trained language model, and r′ is the ranking score sequence of all rounds of dialogue in the input dialogue predicted by the pre-trained language model; Step B43: Contribution dataset D constructed in step B2 s In each round of each data, the normalized contribution of each slot is recorded. According to the input dialogue sequence X, the total number of rounds T and the predicted slot s, the data set D s The contribution of all rounds before round T is obtained and directly used as the target ranking score sequence: in, is the target ranking score for the t-th round of dialogue, r * Sort the score sequence for the target; Step B44: Calculate the difference between the predicted ranking score and the target ranking score using the mean squared error: Among them, L rank is the auxiliary ranking task loss.

8. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 7, characterized in that: The step B5 specifically includes the following steps: Step B51: Fine-tune the task loss L through instructions DST and auxiliary ranking task loss L rank Calculate the total loss L total : THE total =L DST +λL rank Among them, λ is a hyperparameter used to adjust the relative importance of the auxiliary sorting task and the instruction fine-tuning task; The total loss L total Used for model training and optimizing model parameters through back propagation; Step B52: To optimize the total loss L total , LoRA parameters are introduced to efficiently fine-tune the parameters of the pre-trained language model; let the weight matrix of the pre-trained language model be W, LoRA only performs low-rank decomposition on the weights of the set layer, that is, introduces trainable parameters A and B for adjustment, so that the weight change ΔW updated by the model during fine-tuning is: ΔW=AB During the forward propagation process, the effective weights of the model become: W′=W+ΔW=W+AB Among them, W remains frozen, and only A and B are adjusted to adapt to the task requirements, ensuring that fine-tuning only targets task-related features; in addition, the instruction fine-tuning validation set The training process is validated to prevent overfitting on the instruction fine-tuning task.

9. The method for tracking conversation state based on dynamic comparison and context sorting according to claim 1, characterized in that: The step C specifically comprises the following steps: Step C1: In the reasoning process of each data, first construct a dialogue context containing two different instructions as the model input; the input containing the forward instruction is denoted as x pos ; For input containing reverse instructions, record it as x neg To ensure that the model generates a more stable state during the inference phase, the model compares the logits generated by the forward instructions with the logits generated by the reverse instructions, and calculates the final logits through a dynamic comparison mechanism: L t =logits(x t |x p x <t ;θ)-β*logits(x t |x r x <t ;θ) Among them, logits is the unnormalized probability distribution obtained after the model calculates the input, and β is the dynamic inhibition factor, which controls the impact of the reverse prompt generated logits on the final logits; The final probability distribution is normalized by softmax: p(x t |x p x <t ;θ):=softmax(L t ) Step C2: At each time step, the softmax probability distribution P and Q are calculated based on the logits of the next token generated by the forward instruction and the reverse instruction, and the JS divergence is further calculated: Among them, KL() is used to calculate the KL divergence of two probability distributions; Normalizing the JS divergence to the [0,1] interval, we get: Then, set the dynamic adjustment coefficient to: β=1-JS norm When the distributions generated by the forward and reverse instructions are more similar, β is closer to 1, thereby enhancing the inhibition of the reverse prompt; conversely, when the distributions generated by the forward and reverse instructions are more different, β is smaller, thereby reducing the influence of the reverse prompt.

10. A conversation state tracking system based on dynamic comparison and context sorting, characterized by: The method comprises a memory, a processor, and computer program instructions stored in the memory and capable of being executed by the processor. When the processor executes the computer program instructions, the method according to any one of claims 1 to 9 can be implemented.