A hallucination correction action compression method for task-oriented dialogue policy optimization
Patent Information
- Application Number
- CN202611012353.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-08
- Publication Date
- 2026-09-29
AI Technical Summary
[0010]本发明属于任务型对话系统的强化学习与策略优化技术领域,针对现有任务型对话策略学习过程中存在的动作空间规模大、无效动作探索比例高、大语言模型候选动作易产生幻觉、动作价值估计不稳定以及策略收敛速度慢等问题,提出了一种用于任务型对话策略优化的幻觉校正动作压缩方法及系统
[0033]与现有直接在完整动作空间中进行强化学习探索,或直接采用原始大语言模型生成候选动作的方法相比,本发明具有以下实质性特点与显著进步:
Smart Images

Figure CN122840069A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of reinforcement learning and policy optimization technology for task-oriented dialogue systems. It relates to a method and system for illusion correction action compression, which is a reinforcement learning policy optimization framework that combines large language model action prior correction, executable action space compression and distributed Q-learning. It is used to reduce the interference of invalid actions, reduce the action search scale, improve the accuracy of action value estimation, accelerate policy convergence and enhance the reliability and generalization ability of the system in task-oriented dialogue systems. Background Technology
[0002] Task-Oriented Dialogue (TOD) systems, as an important research direction in the field of human-computer interaction, aim to help users complete specific tasks through multi-turn natural language communication, such as booking tickets, ordering food, booking travel, searching for hotels, providing intelligent customer service, processing work orders, or handling business transactions. According to traditional system architecture, task-oriented dialogue systems typically consist of modules such as natural language understanding, dialogue state tracking, dialogue strategy learning, and natural language generation. Among these, the dialogue strategy learning module is responsible for selecting system actions based on the current dialogue state and is a key link affecting task completion rate, number of interaction rounds, and long-term benefits.
[0003] In reinforcement learning frameworks, dialogue policy learning is typically modeled as a Markov decision process. The agent observes the dialogue state in the current turn. Select dialogue action Rewards for environmental or user feedback and transition to the next state. Ultimately, through optimization strategies Maximizing long-term cumulative rewards. This modeling approach enables dialogue systems to continuously learn better action selection strategies through multi-turn interactions, and is therefore widely used in policy optimization for complex task-oriented dialogue systems.
[0004] However, with the expansion of business domains, the increase in user intents, the refinement of slot types, and the complexity of business rules, the state space and action space of task-oriented dialogue systems are rapidly expanding. In scenarios such as ticket booking, travel, hotel reservations, and intelligent customer service, a system action often involves not only action type but also multi-dimensional information such as domain, intent, slot, slot value, constraints, and execution rules. Different actions may be similar in linguistic expression, but they differ significantly in task executability and constraint satisfaction. If reinforcement learning methods directly explore the complete action space, they are prone to generating a large number of invalid actions and low-quality interaction samples, leading to problems such as high action search complexity, low training efficiency, and slow policy convergence.
[0005] To alleviate the aforementioned problems, recent methods have attempted to incorporate large language models into the task-oriented dialogue policy optimization process. Large language models possess strong semantic understanding, contextual reasoning, and natural language generation capabilities, enabling them to generate candidate dialogue actions based on the current dialogue state, thus providing state-related action priors for reinforcement learning. Compared to random exploration or full action space search, candidate action generation based on large language models can, to some extent, narrow the action search scope, allowing reinforcement learning to more focusedly evaluate semantically relevant actions and improving exploration efficiency in the early stages of training.
[0006] However, directly using large language models to generate candidate actions still has significant drawbacks. The training objective of large language models is usually to predict the next lexical unit based on context, rather than directly addressing constraint satisfaction, action feasibility, and long-term reward maximization in task-oriented dialogue. Therefore, when generating action candidates, large language models may tend to output actions that appear linguistically reasonable but are inconsistent with the current task state, such as omitting key user constraints, fabricating unusable actions, or generating actions irrelevant to the current domain. While such actions may have some rationality at the natural language level, they are unexecutable or low-value actions at the dialogue system execution level, typically manifesting as the illusion problem in the action generation process of large language models.
[0007] Most existing methods for action space compression using large language models directly employ the action priors output by the original large language model. The generation of candidate action sets lacks a mechanism to correct the reliability of candidate actions before action compression. If the original action priors contain hallucinatory actions, subsequent reinforcement learning will estimate action value within a contaminated compressed candidate action space, which can easily lead to... Value estimation bias, increased invalid interaction samples, and policy training instability. Meanwhile, traditional scalar... Value estimation methods are unable to fully characterize the uncertainty of action rewards in complex dialogue environments, which further limits the stability and generalization ability of policy optimization.
[0008] In summary, existing technologies generally suffer from the following shortcomings: First, traditional reinforcement learning methods explore the entire action space, making it difficult to effectively reduce the complexity of action search in large-scale task-oriented dialogues; second, existing large language model-assisted action compression methods rely on uncorrected original action priors, making it difficult to suppress illusory actions and non-executable actions; third, once candidate actions are contaminated by illusions, they will interfere with subsequent action value estimation and policy updates in reinforcement learning; fourth, traditional scalar value estimation is difficult to express the distribution and uncertainty of action rewards, resulting in insufficient training stability in complex dialogue environments.
[0009] Therefore, the industry urgently needs a technical approach that can perform reliability correction on action priors of large language models before action space compression, and perform stable policy optimization within the compressed executable candidate action space. This invention is proposed against this backdrop, aiming to achieve reliable action prior correction, executable action space compression, and distributed... By combining learning with practical application, we can reduce illusory action interference, lower the action search burden, improve the stability of action value estimation, accelerate policy convergence, and enhance the generalization ability of multi-domain complex task-oriented dialogue systems. Summary of the Invention
[0010] This invention belongs to the field of reinforcement learning and policy optimization technology for task-oriented dialogue systems. Addressing the problems existing in current task-oriented dialogue policy learning processes, such as large action space size, high proportion of invalid action exploration, susceptibility to illusions arising from large language model candidate actions, unstable action value estimation, and slow policy convergence speed, this invention proposes an illusion correction action compression method and system for task-oriented dialogue policy optimization. This method introduces a large language model action prior correction mechanism, an executable action space compression mechanism, and a distributed... The learning mechanism enables reinforcement learning policy optimization to be performed on a more reliable and compact set of candidate actions, thereby improving action selection quality, reducing training overhead, and enhancing policy convergence stability.
[0011] This method, based on the traditional reinforcement learning dialogue strategy optimization framework, introduces a closed-loop structure of "hallucination-induced correction—action mapping compression—distributed value assessment—classification cross-entropy optimization". Specifically, the system first generates action priors using the original large language model; then constructs a biased large language model with enhanced hallucination tendencies; next, by comparing the prediction distributions of the original model and the biased model, it suppresses lexical and action components related to hallucination tendencies to obtain corrected action priors; further, it maps free-form action proposals to the structured action space executable by the task-oriented dialogue system, forming a compressed candidate action set; finally, it utilizes distributed cross-entropy optimization on this candidate set. The network estimates the action reward distribution and updates the policy network based on the classification cross-entropy loss between the target reward distribution and the predicted reward distribution. Its core process includes the following steps:
[0012] S1, Dialogue state acquisition, in the current round of a multi-turn task-oriented dialogue. Get the current dialogue state The dialogue state includes user input, historical dialogue context, domain information, intent information, slot filling information, slot position confidence, constraints, and historical system actions, which are used to characterize the current task progress and the basis for action decisions.
[0013] S2, Pre-generation of original actions, including the current dialogue state. The original large language model is input, and the original large language model generates action priors for candidate dialogue actions based on the state context and action cue templates. The action prior is used to represent the distribution of actions that the original large language model believes are suitable for execution in the current state.
[0014] S3, Illusion-Induced Offset Model Construction: Based on the illusion-induced dataset, the original large language model is offset-induced to obtain an offset large language model. This offset large language model is used to enhance and explicitly represent the generation tendencies in the original model related to task constraint inconsistencies, slot errors, unexecutable actions, or unreliable semantics, rather than being directly used for the final action decision.
[0015] S4, Hallucination Correction Action Prior Generation. The prediction distributions of the original large language model and the offset large language model are calculated separately under the same input prefix. The hallucination tendency component in the offset model is weakened through distribution comparison to obtain the hallucination-corrected lexical distribution and action prior. This step allows action compression to no longer rely directly on the unprocessed output of the original large language model, but instead generates candidate actions based on corrected priors that better fit task constraints and execution rules.
[0016] S5, Action Mapping and Spatial Compression. Free-form action proposals are generated based on illusion-corrected action priors, and the action mapping function converts the free-form outputs into structured actions executable by the task-oriented dialogue system. Action proposals that fail to meet slot integrity, business rules, or executor interface requirements are filtered, downweighted, or invalidated. Subsequently, actions are selected or sampled from the executable action distribution. Given several candidate actions, construct a compressed candidate action set for the current state. .
[0017] S6, Distributed Action evaluation and strategy optimization. This involves compressing the candidate action set. Above, utilizing distributed systems The network estimates the reward distribution for each candidate action and calculates the action value from the expected reward distribution. The system constructs a policy distribution based on the candidate action values, selects actions to interact with the environment, obtains rewards and the next state, writes state transition samples into the experience replay pool, and updates the principal function based on the classification cross-entropy loss between the target reward distribution and the predicted reward distribution. Network and Target The network continues until the strategy converges.
[0018] In a preferred embodiment, task-oriented dialogue policy optimization is modeled as a Markov decision process:
[0019]
[0020] in, Represents the dialogue state space. Represents the action space, Represents the state transition function. Represents the reward function, This represents the discount factor. In the current round... The agent observes the state. Execute actions Receive rewards and transition to the next state. This invention does not require replacing existing natural language understanding, dialogue state tracking, or action execution modules. Instead, it adds an illusion correction action compression mechanism during policy learning, enabling the policy network to complete value estimation and action selection from a smaller and more reliable set of actions.
[0021] In a preferred embodiment, the hallucination-induced shift model is constructed based on a hallucination-induced dataset. Let the hallucination-induced dataset be:
[0022]
[0023] in, Indicates the dialogue status. This represents user input. This indicates unreliable output or hallucination-induced output. (This refers to the parameters of the original large language model.) Introducing parameter offset Obtain the large offset language model parameters And optimize through the following objectives:
[0024]
[0025] The resulting large offset language model can more effectively represent the generation direction related to unreliable outputs, providing a reference for subsequent distribution comparison and illusion suppression.
[0026] In a preferred embodiment, to improve the stability and feasibility of the hallucination correction action compression process, the present invention may further include the following optional improvements:
[0027] (1) Lightweight construction of the migration model: The hallucination-induced migration model can be constructed using low-rank adaptation, prefix parameters, or adapter parameters, etc., and only learns the parameter migration used to characterize the hallucination tendency. This reduces model training costs while retaining the general capabilities of the original large language model.
[0028] (2) Stabilization of corrected scores: In calculating corrected scores When this is done, a smoothing term can be added to the original model probability and the offset model probability, and combined with temperature scaling or normalization processing, to avoid unstable generation distribution caused by extremely small probability or overly strong correction.
[0029] (3) Constrained implementation of action mapping: Action mapping function Based on preset action modes, slot types, slot value ranges, and execution conditions, free-form action proposals can be converted into standard executable actions; outputs that cannot match or lack necessary slots are not included in the valid candidate set.
[0030] (4) Distributed Support range constraint: When constructing the target reward distribution, the set of supporting atoms can be set according to the task reward range. The upper and lower bounds, and the target value Truncation followed by projection improves the numerical stability of cross-entropy training for classification.
[0031] (5) Consistency of the next state candidate set: When calculating the target value, the next state Candidate action set The same illusion correction, action mapping, and candidate compression process as the current state is used to maintain consistency between the training and decision-making phases.
[0032] In another embodiment, the method of the present invention can also be implemented as an illusion correction action compression device for task-oriented dialogue strategy optimization. The device includes a processor and a memory, the memory storing instructions executable on the processor. When executed, these instructions perform the aforementioned dialogue state acquisition, original action prior generation, offset model construction, illusion correction action prior generation, action mapping compression, and distributed... The process includes evaluation and optimization of the cross-entropy strategy for classification. In the form of a computer-readable storage medium or computer program product, the instructions stored in the computer-readable medium, after being executed by a processor, similarly complete the entire process of illusion correction action compression and reinforcement learning strategy optimization. This device, medium, and program product implementation is merely an engineering embodiment of the method of this invention and does not constitute a limitation on the core technical solution.
[0033] Compared with existing methods that directly explore reinforcement learning in the complete action space or directly use the original large language model to generate candidate actions, this invention has the following substantial features and significant progress:
[0034] (1) The a priori action is corrected by hallucination:
[0035] By comparing the distribution of the offset large language model with that of the original large language model, the generated components related to hallucination tendencies are suppressed, making the candidate actions more consistent with the current dialogue state, user constraints, and task execution conditions.
[0036] (2) Action compression targets the executable space:
[0037] By using an action mapping function, free-form actions are projected into the executable action space, and actions that cannot be mapped or do not meet the constraints are filtered out, thereby improving the effectiveness of compressing the candidate action set.
[0038] (3) The candidate action set is more compact:
[0039] By using probabilistic aggregation, structured deduplication, and candidate truncation, multiple semantically similar free-form outputs are grouped into a finite number of executable actions, reducing the overhead of reinforcement learning action evaluation.
[0040] (4) The motion value estimation is more stable:
[0041] Through distributed The network estimates the distribution of candidate action rewards, rather than just estimating a single scalar. This improves the stability of value estimation in complex dialogue environments.
[0042] (5) Strategy updates form a closed loop:
[0043] The system employs a consistent action correction and compression process in both the current and next states, and updates the policy network based on classification cross-entropy loss, thus forming a closed loop of action compression, value assessment, and policy optimization.
[0044] In summary, this invention systematically solves the problems of unreliable candidate actions, value estimation due to hallucination interference in large language models, and policy training oscillations in existing task-oriented dialogue policy optimization through a closed-loop path of "hallucination induction—prior correction—action compression—distribution evaluation—policy optimization". By introducing a prior action generation mechanism for hallucination correction and an executable action space compression mechanism, this invention can stably achieve higher policy returns and faster convergence speed in multi-task and multi-scenario environments, significantly improving the task success rate of dialogue systems, and possessing clear technical effects and industrialization potential. Attached Figure Description
[0045] Figure 1 This is a flowchart of the method steps of the present invention;
[0046] Figure 2 This is a diagram of the algorithm architecture of the present invention; Detailed Implementation
[0047] The following specific examples illustrate the implementation of this disclosure. Those skilled in the art can easily understand other advantages and effects of this disclosure from the content disclosed in this specification. Obviously, the described embodiments are only a part of the embodiments of this disclosure, and not all of them. This disclosure can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this disclosure. It should be noted that, in the absence of conflict, the following embodiments and features in the embodiments can be combined with each other. Based on the embodiments in this disclosure, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this disclosure.
[0048] It should be noted that various aspects of embodiments within the scope of the appended claims are described below. It will be apparent that the aspects described herein can be embodied in a wide variety of forms, and any particular structure and / or function described herein is merely illustrative. Based on this disclosure, those skilled in the art will understand that one aspect described herein can be implemented independently of any other aspect, and two or more of these aspects can be combined in various ways. For example, any one or more aspects described herein can be used to implement the device and / or method; the invention is not limited thereto, and the same technical objective can be achieved through other structures or functionalities.
[0049] As shown in Figure 1, the present invention provides an illusion correction action compression method for task-oriented dialogue strategy optimization. In the process of multi-round task-oriented dialogue interaction and reinforcement learning strategy training, the method uses interaction rounds... As a time-series node, each round sequentially completes the following steps: dialogue state acquisition, generation of action priors from the original large language model, comparison with the hallucination-induced offset model, construction of corrected action priors, generation of free-form action proposals, mapping of executable actions, compression of candidate action sets, and distributed processing. Network action evaluation and policy optimization feedback. This method introduces an illusion correction mechanism before action space compression, so that reinforcement learning no longer directly relies on the candidate actions output by the original large language model. Instead, it performs value estimation and action selection from a set of corrected and verified executable candidate actions, thereby reducing the interference of illusion actions, faulty slot actions, and unexecutable actions on the policy learning process.
[0050] As the dialogue progressed... At this time, the system first obtains the current dialogue state. The current dialogue state This can include information such as current user input, historical dialogue context, user intent, filled slots, and historical system actions. For example, in a travel booking task, the dialogue status can include fields such as departure point, destination, travel date, travel time, whether it's a direct route, price cap, number of passengers, and order status. This status information is encoded and input into the strategy learning module to generate a set of candidate actions for the current round.
[0051] The system then displays the current dialogue status. Input the original large language model, and the original large language model generates action priors based on the preset action cue template. This action prior reflects the distribution of actions that the original large language model deems appropriate to perform in the current dialogue context. This differs from traditional reinforcement learning, which directly operates within the complete action space. Unlike traditional search methods, this invention first utilizes the contextual understanding capabilities of a large language model to obtain state-related action suggestions, thereby reducing the search scope for irrelevant actions. However, since the original large language model may generate actions inconsistent with the current task constraints, such as omitting key slots, generating incorrect slot values, calling non-existent action interfaces, or outputting actions irrelevant to the current domain, this invention does not directly use the original action priors as the final candidate action set.
[0052] To identify unreliable components in the action priors of the original large language model, the system further constructs a hallucination-induced offset model. Specifically, this can be based on a hallucination-induced dataset. The original large language model is trained by offset, where Indicates the status of historical dialogue. This represents user input. This represents unreliable output samples. Unreliable output samples may include actions with incorrect slot values, actions with missing constraints, actions with domain mismatches, and actions inconsistent with actual user needs. This is achieved by analyzing the original model parameters... Introducing parameter offset , thus obtaining the offset model This allows the offset model to more centrally represent the generative directions related to hallucination tendencies in the original large language model. The offset model is not used for final action decisions, but rather serves as a reference model for hallucination directions, used for subsequent correction of action priors.
[0053] In the autoregressive generation process, the original large language model uses the same input prefix. Next output the prediction distribution of the next word:
[0054]
[0055] in, Indicates candidate word elements, This represents the unnormalized score output by the original large language model. Correspondingly, the offset model outputs the predicted distribution. The system compares the distributions of the two to obtain a correction score:
[0056]
[0057] in, These are the enhancement coefficients for the reliable components of the original model. If a certain word has a high probability in the offset model, it indicates that the word may be related to the direction of hallucination generation. By subtracting the log probability of the offset model from the correction score, the weight of that word in the final generation can be reduced. Furthermore, the system calculates the correction word distribution based on the correction score:
[0058]
[0059] in, The system represents the vocabulary. Through the above methods, while preserving the semantic understanding capabilities of the original large language model, it suppresses potential illusionary lexical units, erroneous slot lexical units, and task-irrelevant lexical units, thereby obtaining a more reliable action generation distribution.
[0060] Based on the corrected lexical distribution, the system generates one or more free-form action proposals. The free-form action proposals can be in natural language or semi-structured text. For example, when a user suggests, "I want to book a direct flight from Beijing to Shanghai after 7 PM tomorrow, with a price under 300," the original large language model might generate action proposals such as "search for flights from Beijing to Shanghai," "recommend hotels in Shanghai," or "search for flights tomorrow night." After illusion correction, action proposals more consistent with the current travel booking task, time constraints, direct flight constraints, and price constraints will have a higher probability, while action proposals lacking key constraints or deviating from the current domain will be weakened.
[0061] Subsequently, the system uses the action mapping function This invention maps free-form action proposals to structured actions executable by a task-oriented dialogue system. Action proposals that cannot match the action name, lack necessary slots, have incorrect slot types, violate business processes, or cannot be invoked by the backend executor are filtered or assigned low weight. Thus, this invention transforms free-text actions generated by a large language model into structured actions that a task-oriented dialogue system can directly evaluate and execute, avoiding the waste of training resources by the reinforcement learning module on unexecutable actions.
[0062] In the executable action space, the system performs probabilistic aggregation on the mapping results. If multiple free-form action proposals are mapped to the same executable action... Then, by summing up their correction probabilities, we get:
[0063]
[0064] in, Indicates a free-form action proposal Mapped to executable actions This aggregation method allows semantically similar but differently expressed action proposals to be grouped into a single structured action, reducing duplicate candidate actions. The system then ranks the candidate actions based on their prior probability of correction, action validity, slot matching degree, and the degree to which current state constraints are satisfied, and selects the top-ranked actions. These actions constitute the compressed candidate action set:
[0065]
[0066] in, This represents the set of executable actions in the current state. If the number of valid candidate actions is less than [a certain value], then [the set is missing]. The system can supplement the current set of legal actions with safety actions such as requesting missing slots, confirming user constraints, or informing that there are no matching results, so as to ensure that there is a set of evaluable candidate actions in each round.
[0067] After obtaining the compressed candidate action set After that, the system enters a distributed state. Action evaluation phase. Let the discrete support atom set be:
[0068]
[0069] For candidate actions ,distributed The network outputs its reward distribution:
[0070]
[0071] in, Indicates the main Network parameters, Indicates action The reward falls on the supporting atoms The predicted probability. Main The network outputs the unnormalized score for each atom. The predicted probability is obtained by using the softmax function:
[0072]
[0073] The scalar value of a candidate action is obtained from the expectation of the reward distribution:
[0074]
[0075] The system in the candidate action set Compare the value of each action and construct the strategy distribution:
[0076]
[0077] in, This is a temperature parameter. The system can distribute sampled actions according to this strategy, or select the action with the highest value as the system action for the current round. The selected action is sent to the dialogue manager or backend task executor for execution and receives environmental feedback rewards. and the next state .
[0078] After execution, the system will provide a state transition sample. Write to the experience replay pool During the training phase, the system samples mini-batch samples from the experience replay pool and prepares for the next state. Re-execute the same illusion correction action compression process as the current state to obtain the next state candidate action set. Based on the goal Network parameters Calculate the target value:
[0079]
[0080] in, This is a discount factor. To avoid the loss of return information caused by simply regressing the scalar objective, this invention uses the objective value... Projected onto discrete supporting atom set This forms the target return distribution. Let:
[0081]
[0082] The target is distributed in atoms The probability of getting it is:
[0083]
[0084] in, Indicates the Gaussian width. Indicates the supporting atomic spacing, This represents the standard normal distribution function. The network updates by minimizing the classification cross-entropy loss between the target distribution and the predicted distribution:
[0085]
[0086] Target Network parameters are updated using a soft update method:
[0087] in, The coefficients are soft update coefficients. Through the above training process, the current state and the next state both adopt consistent action correction and candidate compression rules, ensuring that the action space distribution is consistent between the policy decision stage and the network training stage, and reducing the value estimation bias caused by inconsistent candidate sets.
[0088] In engineering implementation, the method of this invention can be embedded before the policy learning module of an existing task-oriented dialogue system without requiring replacement of the natural language understanding module, dialogue state tracking module, or action execution module. The natural language understanding module is responsible for parsing user input, and the dialogue state tracking module is responsible for maintaining the state. The hallucination correction action compression module of the present invention is responsible for compressing actions according to... Construct a set of candidate actions for compression ,distributed The network is responsible for performing value assessment and action selection within this set. Therefore, this invention can be incrementally integrated into existing task-oriented dialogue systems.
[0089] Through the above implementation process, this invention achieves "original action prior generation—illusion offset direction characterization—distribution comparison correction—executable action mapping—candidate action compression—distributed..." The closed-loop path of "evaluation-classification cross-entropy policy update" can suppress illusory actions from entering the candidate set during multi-round task-oriented dialogue policy training, reduce the blind search of the entire action space by reinforcement learning, improve the stability of action reward estimation, and accelerate policy convergence.
[0090] The above embodiments are merely illustrative of several implementation methods of this application, and their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the protection scope of this application. Therefore, the protection scope of this application should be determined by the appended claims.
Claims
1. A method for hallucination correction action compression for task-oriented dialogue strategy optimization, characterized in that, Includes the following steps: S1, Dialogue State Acquisition: In the current round of a multi-turn task-oriented dialogue. Get the current dialogue state The current dialogue state includes user input, historical dialogue context, domain information, intent information, slot filling information, and constraint information. S2, Original Action Prior Generation: The current dialogue state is generated... Input the original large model and obtain the action prior of the original large model regarding candidate dialogue actions. ,in Indicates a dialogue action. S3, Construction of Hallucination-Induced Misalignment Model: Misalignment is induced on the original large model based on the hallucination-induced dataset to obtain a large misalignment model with enhanced hallucination tendency. S4, Hallucination-Corrected Action Prior Generation: The prediction distributions of the original large model and the offset large model under the same input prefix are compared. By enhancing the reliable predictions in the original model and suppressing the hallucination-prone predictions in the offset model, the hallucination-corrected action prior is obtained. . S5, Action Space Compression: Based on the hallucination-corrected action priors, candidate action proposals are generated, and the candidate action proposals are mapped to the executable action space through an action mapping function to construct a compressed candidate action set for the current state. . S6, Distributed Q-action evaluation: In the compressed candidate action set The reward distribution for each candidate action is estimated using a distributed Q-network. S7, Action Selection and Environment Interaction: Constructing a strategy distribution based on the action value of candidate actions. And select or sample actions from the policy distribution. Interact with the task-oriented dialogue environment to earn rewards. and the next state . S8, Policy Network Update: Transitioning State Samples Store the data in the experience replay pool, update the parameters of the distributed Q-network using classification cross-entropy loss, until convergence and the optimal policy is output. .
2. The method according to claim 1, characterized in that, The task-oriented dialogue strategy optimization is modeled as a Markov decision process: in, Represents the dialogue state space. Represents the action space. Represents the state transition function. Represents the reward function, This represents the discount factor.
3. The method according to claim 2, characterized in that, The action prior of the original large model is: And construct an original set of candidate actions based on the action priors: in, This indicates the number of candidate actions.
4. The method according to claim 1, characterized in that, The hallucination-induced offset model is constructed based on a hallucination-induced dataset. For the parameters of the original large model Introducing parameter offset Obtain the parameters of the large offset model. And optimize the parameter offset by the following objectives: in, Indicates the dialogue status. This represents user input. Indicates hallucination-induced output. This indicates the number of samples induced by hallucination.
5. The method according to claim 4, characterized in that, The prior generation of the hallucination correction action includes: the next lexical prediction distribution of the original large language model in the autoregressive generation process is: The correction score is obtained by comparing the prediction distributions of the original large language model and the offset large language model: The corrected word distribution is obtained based on the corrected score: in, For comparison coefficients, For vocabulary list.
6. The method according to claim 5, characterized in that, The action space compression includes: based on the current dialogue state And action cue templates, generating free-form action proposals from hallucination-corrected action priors. ; through projection function Map the free-form action proposals to executable actions. And define the corrective action prior in the executable action space:
7. The method according to claim 6, characterized in that, The set of compression candidate actions is defined as follows: in, This represents the space of executable actions that satisfy business rules and execution constraints in the current state.
8. The method according to any one of claims 7, characterized in that, The distributed Q-action evaluation includes: modeling the candidate action reward as defined in... Discrete supporting atoms Classification distribution on: in, Indicates the main Q network parameters. Indicates concentration at the atom Point mass distribution, Indicates candidate actions In atoms The predicted probability.
9. The method according to claim 8, characterized in that, The predicted probabilities and the scalar action values of the candidate actions are respectively:
10. The method according to claim 9, characterized in that, The strategy is distributed across a compressed candidate action set. The above definition is: in, For temperature parameters, action according to Select or sample.
11. The method according to claim 10, characterized in that, The policy network update includes: updating from the experience replay pool. Mid-sample state transition sample; based on the next state Construct a set of candidate actions for compression Based on the target Q network parameters Calculate the target value:
12. The method according to claim 11, characterized in that, The target value Projected onto the class supporting atom set Above, construct the target return distribution: in, Indicates the atomic spacing width. This represents the cumulative distribution function of the standard normal distribution.
13. The method according to claim 12, characterized in that, The classification cross-entropy loss function is: The main Q network parameters are updated by minimizing the classification cross-entropy loss function. And update the target Q network parameters via soft update: in This is the soft update coefficient.