Large model processing method for task-oriented dialogue, electronic equipment and storage medium

By introducing a two-stage processing process of explicit confidence labels and slow thinking models in task-based dialogue systems, the shortcomings of existing systems in complex task processing and output reliability evaluation are solved, and higher reliability and efficiency are achieved.

CN120067263APending Publication Date: 2025-05-30AISPEECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510168227.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing task-based dialogue systems have shortcomings in handling complex tasks and evaluating the reliability of model outputs, including poor flexibility, stiff dialogue, risk of overfitting, lack of unified confidence evaluation criteria, and high error rates.

Method used

An explicit three-categorical confidence label (determined accepted, uncertain, explicit rejected) and slow thinking models are introduced, and user input is automatically analyzed through a two-stage processing process, assigned processing paths, and further in-depth thinking is carried out in uncertain situations to determine the final result.

Benefits of technology

It significantly improves the reliability and success rate of task execution, reduces error rate and resource waste, improves system efficiency, enhances user trust and experience, and promotes the popularization of AI and automation technologies in various industries.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067263A_ABST
    Figure CN120067263A_ABST
Patent Text Reader

Abstract

The embodiment of the invention discloses a task-oriented dialogue-oriented large model processing method, electronic equipment and a storage medium, and the method comprises the steps: automatically analyzing a result which should be extracted or predicted and a corresponding confidence label according to a dialogue history input by a user, and enabling the confidence label to comprise a result which is determined to be accepted, a result which is uncertain and a result which is definitely rejected; according to the confidence label corresponding to each prediction result, the prediction results are distributed to different processing paths to be processed, the received results are directly output as final results, the user is informed that the results are unreliable after the direct rejection is explicitly rejected, and the user is informed that the results are not reliable. If not, submitting to the slow-thinking model of the next stage for further decision processing; for the uncertain part of the previous stage, performing deeper and detailed slow thinking by using the slow thinking model so as to determine a final processing result; after the task is completed, a final result is returned to the user, and meanwhile, the confidence coefficient is labeled according to the feedback of the task.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of task-oriented dialogue, and particularly relates to a large model processing method, an electronic device, and a storage medium for task-oriented dialogue. Background Art

[0002] In related technologies, a task-oriented dialogue system refers to a dialogue system specifically designed to interact with users to complete specific tasks or provide specific services. Such systems usually understand the user's intention through natural language processing (NLP, Natural Language Processing) technology and generate appropriate responses according to predefined logic or rules. Among them, NLP is a branch of artificial intelligence, which is a technology and method that enables a computer to understand, generate, and interact with human natural language. The most similar existing technologies on the market currently include: rule-based task-oriented dialogue systems, task-oriented dialogue systems based on pre-trained language models, and task-oriented dialogue systems based on general large language models.

[0003] Model output confidence evaluation technology refers to a technology in machine learning and deep learning models that evaluates the confidence level of the model in its generated predictions or output results. It is usually used to judge the decision reliability of the model, help users understand the output quality of the model, and conduct risk management in specific application scenarios. In the field of deep learning, many methods calculate the output confidence value through the internal representation of the model or the results of multiple samplings, and set thresholds to measure the certainty degree of the model in its output.

[0004] Among them, rule-based task-oriented dialogue systems: These systems complete specific tasks, such as booking tickets, querying the weather, ordering meals, etc., by means of rules or templates through dialogue with users. Task-oriented dialogue systems based on pre-trained language models: These systems use advanced deep learning technology to process user input and generate responses. Such systems are usually pre-trained on large-scale text data to learn rich language representations, and then fine-tuned on specific tasks to achieve high performance. Task-oriented dialogue systems based on general large language models: General large language models (LLMs, Large Language Models) themselves have strong transfer and generalization capabilities, can be applied to different tasks without special training, and can achieve high performance. Among them, LLM is an artificial intelligence model that can process and generate natural language through massive data training. Summary of the Invention

[0005] Embodiments of the present invention provide a large model processing method, an electronic device, and a storage medium for task-oriented dialogue, which are used to solve at least one of the above technical problems.

[0006] In a first aspect, an embodiment of the present invention provides a large model processing method for task-oriented dialogue, including: automatically analyzing the results to be extracted or predicted and the corresponding confidence labels according to the dialogue history input by the user, where the confidence labels include definitely accepted, uncertain, and definitely rejected; according to the confidence labels corresponding to each predicted result, allocating the predicted results to different processing paths for processing, where the definitely accepted ones are directly output as the final results, the definitely rejected ones are directly rejected and the user is informed that the results are unreliable, and the uncertain ones are submitted to the slow thinking model in the next stage for further decision-making processing; for the uncertain part in the previous stage, using the slow thinking model to conduct more in-depth and detailed slow thinking to determine the final processing result; after the task is completed, returning the final result to the user, and at the same time adjusting the confidence label ratio according to the feedback of the task.

[0007] In a second aspect, an embodiment of the present invention further provides a computer program product, where the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, and the computer program includes program instructions. When the program instructions are executed by a computer, the computer is enabled to execute the steps of the large model processing method for task-oriented dialogue according to any embodiment of the present invention.

[0008] In a third aspect, an embodiment of the present invention further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, where the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the steps of the method in the first aspect.

[0009] In a fourth aspect, an embodiment of the present invention further provides a storage medium, on which a computer program is stored, and it is characterized in that when the computer program is executed by a processor, the steps of the method in the first aspect are implemented.

[0010] In the method of the embodiments of the present application, by optimizing the decision-making mechanism in the task-oriented dialogue system, and by introducing explicit three-category confidence labels and further slow thinking for uncertain results, the reliability and success rate of task execution are significantly improved, the error rate and resource waste are reduced, and the system efficiency is enhanced. The chain reaction includes enhancing user trust and experience, reducing manual intervention, lowering costs, facilitating the processing of more complex tasks, promoting the popularization of AI and automation technologies in various industries, enhancing safety and risk control capabilities, and at the same time accelerating cross-industry cooperation and innovation, providing support for the optimization and application expansion of future intelligent systems. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0012] Figure 1 The processing flow provided by an embodiment of the present invention; Figure 2 The flowchart of an embodiment of the large model processing method for task-oriented dialogue provided by an embodiment of the present invention; Figure 3 The performance of the model provided by an embodiment of the present invention on the dialogue state tracking task dataset MultiWOZ2.2 test set; Figure 4 The explicit knowledge boundary model (EKBM, Explicit Knowledge Boundary Modeling) provided by an embodiment of the present invention; Figure 5 The reliability adaptive data construction process provided by an embodiment of the present invention; Figure 6 The reliability preference data construction process provided by an embodiment of the present invention; Figure 7 The reliability of the MultiWOZ 2.2 test set provided by an embodiment of the present invention; Figure 8 The proportion of affirmative answers provided by an embodiment of the present invention; Figure 9 The precision analysis provided by an embodiment of the present invention; Figure 10 The recall analysis provided by an embodiment of the present invention; Figures 11a - 11e The LLM description provided by an embodiment of the present invention, including specific examples of training descriptions and inference prompts; Figure 12 The structural schematic diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts fall within the scope of protection of the present invention.

[0014] The inventors have found that rule-based task-oriented dialogue systems have poor flexibility, are unable to handle unforeseen user inputs, and the maintenance cost increases significantly as the task complexity increases. In addition, the conversations of such systems often appear rigid, reducing the user experience, and they cannot learn from interactions to improve performance. Task-oriented dialogue systems based on pre-trained language models rely on a large amount of computing resources for pre-training and fine-tuning, and there may be a risk of overfitting affecting the generalization ability. At the same time, although they can handle complex language structures, they may not be able to truly understand the user's intention in some cases. Task-oriented dialogue systems based on general large language models have strong transfer capabilities but lack specialization and may not perform optimally in specific tasks. In addition, due to the difficulty in explaining the decision-making process and the existence of hallucination problems, it is difficult to evaluate the quality and reliability of the answers generated by the model. Moreover, for model output confidence evaluation techniques, there is no unified standard, the confidence distributions of different models may vary greatly, and the stability is poor. The choice of threshold will greatly affect the classification results of the evaluation, and a suitable threshold usually needs to be set based on experience or tried multiple times on specific tasks.

[0015] The above defects are mainly caused by the following reasons: On the one hand, the inaccuracy of the system itself. Many systems based on rules and general large language models have low accuracy in specific tasks and are difficult to handle complex situations involving multi-turn conversations. On the other hand, task-oriented dialogue systems do not pay enough attention to the evaluation of output reliability. The mainstream idea for performance improvement is to directly improve the accuracy of the system by methods such as increasing the amount of training data and optimizing data quality. These methods require a large amount of time and resources. The method of improving the overall reliability of the system by evaluating the confidence of the model output is relatively less emphasized. In addition, there are limitations in the application method of confidence evaluation technology. Usually, the confidence of the output is calculated in various ways, and a threshold is selected to accept or reject the corresponding output, that is, the action is binary, which may lead to overconfidence or excessive conservatism. In fact, there may be outputs with intermediate confidence levels, that is, parts where the model is relatively uncertain. Currently, there are few methods for evaluating and processing this part.

[0016] When practitioners in this industry solve these deficiencies, they usually directly improve the performance of the dialogue system by increasing the model size, increasing the amount of training data, or improving the data quality, and reduce the error rate by improving the accuracy, thereby enhancing the overall reliability of the system. In the methods involving confidence evaluation, most systems divide the output into reliable and unreliable parts based on a selected threshold, ignoring the more detailed evaluation and processing of the "uncertain" intermediate state. The "Two-Stage Reliable Large Model Learning Framework" proposed in the embodiments of this application effectively improves the overall reliability at a relatively low additional cost by introducing the "uncertain" intermediate state and performing further slow thinking on the uncertain part in the second stage to determine the final result. At the same time, our training framework also allows the effective training of the model's ability to explicitly evaluate the confidence of the output using only a small amount of data, solving the deficiencies of traditional confidence evaluation algorithms such as "lack of a unified standard and incomparability" and "difficulty in finding an appropriate threshold". The above optimization methods across multiple dimensions are not easily conceived by the prior art.

[0017] In the embodiments of this application, the dialogue state tracking task in the traditional task-based dialogue system is redesigned from a single-stage process of "directly giving the result" to a two-stage process of "first giving a preliminary result and explicitly distinguishing the confidence of each output, and then deeply thinking about the uncertain part to determine the final result" through the "Two-Stage Reliable Large Model Learning Framework". Corresponding training methods are involved for the two stages. In the first stage, through data synthesis and optimization algorithms (such as SFT (Supervised Fine-Tuning, a method of further training a pre-trained model on a specific task dataset to improve performance) and DPO (Direct Preference Optimization, a technique for improving the model's effect by optimizing the matching degree between the preference of the model output and the user feedback)), the self-awareness ability of the open-source general large model is effectively improved, endowing the model with the ability to effectively and explicitly provide answers and corresponding confidence labels. In the second stage, a slow thinking processing model is trained using data constructed by the high-performance closed-source model GPT-4o (Generative Pre-trained Transformer 4 optimized) for the slow thinking process of the uncertain part. Through multi-dimensional slow thinking, the more difficult uncertain part in the first stage can be processed with high accuracy to obtain the final highly reliable result.

[0018] Please refer to Figure 1 , which shows the processing flow of the embodiments of this application.

[0019] 1) Prediction results and confidence labels • Description: Based on the input conversation history, the system automatically analyzes the results that should be extracted or predicted and their corresponding confidence labels. The confidence labels are divided into "definitely accepted", "uncertain", and "definitely rejected".

[0020] • Technical key points: o Result prediction: Use natural language processing technology to extract information required by the task from the input history, such as user intent, etc.

[0021] o Explicitly assign confidence labels: For each predicted result, its confidence label needs to be explicitly provided, and ensure that the confidence label can effectively reflect the true confidence state of the model.

[0022] o Reasonable classification ratio of confidence labels: Overall, the ratio of the confidence labels of the system needs to meet the preset goals and principles. In-depth slow thinking on the "uncertain" part can improve performance, but it will also introduce additional overhead. It is necessary to balance the ratio of the labels.

[0023] • Technical effects: o By analyzing the conversation history, the system can provide accurate prediction results and reliable confidence labels, enabling users to quickly understand the reliability of the results and decide on subsequent operations for each result according to the confidence labels.

[0024] o The overall classification ratio of the confidence labels of the system is reasonable, and it can achieve a balance between high overall reliability and low additional overhead.

[0025] 2) Confidence label classification • Description: For the confidence label corresponding to each predicted result, it is assigned to different processing paths for processing. For example, the "definitely accepted" is directly output as the final result, the "definitely rejected" is directly rejected and the user is told that the result is unreliable, and the "uncertain" is submitted to the slow thinking model in the next stage for further decision-making processing.

[0026] • Technical key points: o Flexible assignment of processing: For results with different confidence labels, system designers need to set corresponding processing paths according to the preset goals and principles.

[0027] • Technical effects: o For predicted results with different confidences, the system can flexibly process them, achieving high reliability of the final results and low overall additional overhead.

[0028] 3) Slow thinking model • Description: For the "uncertain" part in the previous stage, more in-depth and detailed slow thinking is carried out to determine the final processing result. This step may introduce additional overhead, but it will also improve the overall performance.

[0029] • Technical key points: o Slow thinking process design: Design the principles and logic of slow thinking for specific task requirements and design an effective thinking process.

[0030] o Slow thinking data synthesis: According to the designed thinking process, use existing data and labels, and use the high-performance closed-source model GPT-4o to synthesize data for the training of its own model.

[0031] o Slow thinking model training: Use the synthesized data to train a model that can effectively perform slow thinking through the method of SFT supervised training.

[0032] • Technical effects: o The slow thinking model can effectively think about and process "uncertain" results and has a sufficiently high accuracy rate.

[0033] 4) Output result • Description: After the task is completed, the system returns the final result to the user, and at the same time optimizes the confidence label ratio, the slow thinking model processing flow, etc. according to the feedback of the task to improve the processing effect of future tasks.

[0034] • Technical key points: o Feedback learning: The system continuously optimizes the strategies in the process by analyzing the feedback during the task execution.

[0035] o Data synthesis and optimization training: Based on the task results and expected strategy adjustments, construct optimized training data and further improve the system's adaptability through reinforcement learning.

[0036] • Technical effects: o Through continuous learning and optimization, the system can continuously improve its performance in complex tasks, making it more efficient and accurate.

[0037] o The optimized system can adapt to more complex and changeable task requirements and further improve the reliability of task execution.

[0038] During the implementation of this application, the applicant also adopted some alternative solutions and beta versions. Among them, the alternative solution is a two-stage task-based dialogue system based on normalized token probability.

[0039] Description This solution adopts a similar two-stage design. The difference is that the confidence tags in the system are calculated based on the average token probability value of the model's output. Specifically, the average token probability value of the output is a value greater than zero and less than one, which can be regarded as the confidence of the model in this output. The distribution of the average token probability values of different models may vary greatly, so it is necessary to normalize its value to unify the standards of different models. Then, according to the task requirements, two appropriate thresholds are found to divide the confidence value interval into three parts, representing "definitely accepted", "uncertain", and "clearly rejected" respectively. The subsequent process is similar to the official method we applied for.

[0040] Advantages Simple and easy to implement: For simple tasks and scenarios with clear rules, this method does not need to involve model training, is simple to implement, has low costs, and short development time.

[0041] High controllability: Since the thresholds can be set flexibly, the behavior of the system is very predictable, can be adjusted in a timely manner according to specific task requirements, and system designers can clearly understand the basis for each operation.

[0042] Disadvantages Low accuracy: The average token probability value of the model output can indeed provide certain information, but the accuracy is not high. There are samples with high probability values but actual errors, which will reduce the trust level of the system.

[0043] Poor flexibility: For different tasks, the best threshold settings need to be obtained through multiple attempts, grid search, etc., which is not flexible enough.

[0044] In the process of implementing this application, the inventor also adopted the following Beta version solution: a slow thinking model in the second stage based on user confirmation.

[0045] Description In this beta version, an automated model for slow thinking about "uncertain" results is not obtained through model training, but a strategy of seeking user confirmation is adopted. Specifically, when the system encounters uncertainty, it interacts with the user through simple user interactions to confirm details and seek correct corrections to the prediction results. This version does not fully achieve automation, but can make reasonable choices in simple scenarios.

[0046] Advantages Quick implementation: The development of this version is relatively rapid. There is no need to obtain a highly accurate automated model through complex automated thinking data design, construction, and multiple optimization iterations. A usable slow thinking module for processing uncertain results can be obtained in a short time.

[0047] User Interaction: Although the functions are relatively basic, through user interaction, some ambiguous task scenarios can be solved with high accuracy most of the time, improving the flexibility of the system.

[0048] Disadvantages Low Efficiency: Due to the use of conversational interaction, the latency of the overall system is very high, resulting in low efficiency of the entire process. The overhead for user interaction has become the main contradiction and cannot be ignored.

[0049] Limited Applicable Scenarios and Dependence on Manual Intervention: Due to the latency brought by user interaction, the system is difficult to apply to complex tasks because this usually requires multiple rounds of interaction confirmation, and its performance cost limits the generality of its applicable scenarios.

[0050] In the embodiments of the present application, by optimizing the decision-making mechanism in the task-oriented dialogue system, through the introduction of explicit three-class confidence labels and further slow thinking for uncertain results, the reliability and success rate of task execution are significantly improved, the error rate and resource waste are reduced, and the system efficiency is enhanced. The chain reaction includes enhancing user trust and experience, reducing manual intervention, lowering costs, facilitating the handling of more complex tasks, promoting the popularization of AI and automation technologies in various industries, enhancing safety and risk control capabilities, and at the same time accelerating cross-industry cooperation and innovation, providing support for the optimization and application expansion of future intelligent systems.

[0051] Please refer to Figure 2 , which shows a flowchart of an embodiment of the large model processing method for task-oriented dialogue of the present application.

[0052] As Figure 2 shown, in step 201, according to the dialogue history input by the user, the results that should be extracted or predicted and the corresponding confidence labels are automatically analyzed, where the confidence labels include definitely accepted, uncertain, and definitely rejected; In step 202, according to the confidence label corresponding to each predicted result, the predicted results are assigned to different processing paths for processing, where the definitely accepted ones are directly output as the final results, the definitely rejected ones are directly rejected and the user is told that the results are unreliable, and the uncertain ones are submitted to the slow thinking model in the next stage for further decision-making processing; In step 203, for the uncertain part in the previous stage, the slow thinking model is used for more in-depth and detailed slow thinking to determine the final processing result; In step 204, after the task is completed, the final result is returned to the user, and at the same time, the confidence label ratio is adjusted according to the feedback of the task.

[0053] In some alternative embodiments, the slow thinking model includes slow thinking process design, slow thinking data synthesis, and slow thinking model training. Among them, the slow thinking process design designs the principles and logic of slow thinking for specific task requirements to design an effective thinking process; the slow thinking data synthesis uses the existing data and labels according to the designed thinking process, and uses the high-performance closed-source model GPT-4o to synthesize data for the training of its own model; the slow thinking model training uses the synthesized data to train a model that can effectively perform slow thinking through the method of SFT supervised training.

[0054] In some alternative embodiments, the automatic analysis of the results to be extracted or predicted and the corresponding confidence labels according to the user input conversation history includes: using natural language processing technology to extract the information of task requirements from the input history; for each predicted result, an explicit corresponding confidence label needs to be provided, and it is ensured that the confidence label can effectively reflect the true confidence state of the model. Among them, the proportion of the confidence labels needs to meet the preset goals and principles.

[0055] In some alternative embodiments, for the results with different confidence labels, corresponding processing paths are set according to the preset goals and principles.

[0056] In some alternative embodiments, after the task is completed, the final result is returned to the user, and at the same time, the confidence label ratio according to the feedback of the task includes: by analyzing the feedback during the task execution process, continuously optimizing the strategies in the process; based on the task results and expected strategy adjustments, constructing optimized training data to further improve the adaptability of the system through reinforcement learning.

[0057] In some alternative embodiments, the method significantly improves the reliability and success rate of task execution, reduces the error rate and resource waste, and improves the system efficiency by optimizing the decision-making mechanism in the task-based dialogue system, introducing explicit three-class confidence labels, and further slow thinking for uncertain results.

[0058] It should be noted that the above method steps do not limit the execution order of each step. In fact, some steps may be executed simultaneously or in the reverse order of the step definition. This application has no restrictions here.

[0059] The following uses a specific embodiment to illustrate the solution of this application, so that those skilled in the art can better understand the solution of this application.

[0060] Current large language models (LLMs) often struggle to determine their boundaries of capabilities, especially when dealing with complex queries, leading to hallucinations, that is, generating seemingly reasonable but incorrect answers. This problem is particularly prominent in task-oriented information extraction, where traditional methods can extract answers in a single step and implicitly model confidence, making it complex to evaluate outputs that are beyond the model's capabilities and have low confidence, thus resulting in the unreliability of the model. Embodiments of this application introduce the Explicit Knowledge Boundary Model (EKBM), which redefines information extraction as a multi-step decision-making process. By incorporating the concepts of "certain" and "uncertain", EKBM enables the model to identify and clarify its boundaries of capabilities, thus producing more reliable outputs. Through a two-stage training process, the EKBM framework significantly improves the reliability of the model and promotes a clear understanding of its knowledge limitations. The research results highlight the potential of EKBM in improving the reliability of information extraction tasks and address the major challenge of hallucinations in LLMs.

[0061] 1. Introduction Recently, large language models (LLMs) have demonstrated impressive text generation capabilities, making them applicable to various downstream tasks. Task-oriented dialogue (TOD) systems are an example with great application potential and practical value, such as intelligent reservation systems. Some studies have investigated the application of LLMs in TOD tasks and their impacts.

[0062] However, there is ample evidence that LLMs sometimes produce hallucinations, providing untrue and misleading information. This phenomenon complicates the deployment of LLMs in real-world environments, where reliable answers are crucial and answers that deviate from the truth are almost intolerable. As Figure 3 shown, even the most advanced general-purpose LLMs, such as GPT-4 (Optimized Generative Pretrained Transformer 4 - a high-performance closed-source language model released by OpenAI), exhibit suboptimal performance in TOD applications. This highlights a significant gap between their capabilities and the requirements of actual production scenarios, and related technologies have also expressed the same concerns. Among them, Figure 3 shows the performance of the model on the test set of the dialogue state tracking task dataset MultiWOZ2.2. BART+DAIR and DFM are fine-tuned models compared with LLMs. JGA refers to the joint goal accuracy, and Slot-F1 refers to the F1 metric at the slot-value pair level.

[0063] Many studies have explored the reasons for hallucinations in LLMs and investigated mitigation strategies, including supervised fine-tuning (SFT) during training and reinforcement learning from human feedback (RLHF), as well as self-reflection and external tools during inference. Currently, the training of LLMs relies on large text corpora and matching specific human preferences. However, models often struggle to "know what they don't know" (Yin et al., 2023). When models encounter questions beyond their knowledge boundaries, they often hallucinate, leading them to generate plausible responses based on maximum likelihood estimation.

[0064] In the embodiments of this application, we introduce a novel framework - Explicit Knowledge Boundary Modeling (EKBM), as Figure 4 shown, which uses two components, "certain" and "uncertain", to demarcate the ability boundaries of the model to represent the confidence in responses. This framework aims to improve the reliability of LLM outputs in TOD systems and meet the urgent need for models that are both accurate and transparent about their limitations.

[0065] To this end, we propose a two-stage reliability training process, including reliability adaptive training and reliability preference training. The former enables the model to explicitly describe its knowledge boundaries, thereby understanding and differentiating the concepts and requirements of "certain" and "uncertain" responses. The latter aligns the model's output with the user's preferences for reliability and usefulness.

[0066] Please refer to Figure 4 , which shows the Explicit Knowledge Boundary Modeling (EKBM) framework. "Traditional Classification-based execution" refers to the direct execution of traditional classification-based methods; "Decision-making process with EKBM" refers to the decision-making execution process based on the EKBM framework; "LLM" refers to the large language model; "Reliable LLM" refers to the reliable large language model; Sure Responses refers to certain responses; Unsure Responses refers to uncertain responses; Certainty Refiner refers to the refiner that optimizes the certainty of uncertain responses; Refined Responses refers to the refined uncertain responses.

[0067] Our experimental evaluation shows that the EKBM framework and training process significantly improve the reliability of the model. Notably, the model's perfect accuracy for high-confidence positive answers exceeds 83%, representing a significant improvement compared to 72% of traditional methods. Additionally, our analysis shows that the explicit representation of knowledge boundaries only refines uncertain answers, thereby improving the performance of information extraction at a lower cost.

[0068] Our contributions are summarized as follows: - We propose the reliability requirements for information extraction tasks in TOD and introduce the principles for measuring and comparing model reliability.

[0069] We propose a new framework, Explicit Knowledge Boundary Modeling (EKBM), which introduces a method to explicitly describe the boundary of the model's capabilities (certain and uncertain responses), enabling more reliable handling of information extraction tasks. We also outline the requirements for measuring model reliability within the EKBM framework.

[0070] - We propose a corresponding training process for the EKBM framework and experimentally verify the effectiveness of this training process as well as the significance and value of the EKBM framework.

[0071] 2 Related Work 2.1 Task-Oriented Dialogue Systems Task-oriented dialogue systems (TOD) aim to assist users in completing specific tasks within a certain domain, and practical applications such as intelligent reservation systems are of great significance in the real world.

[0072] An important part of the dialogue system is dialogue state tracking (DST), which involves continuously extracting and tracking important information from the dialogue history between the user and the assistant. For example, in a reservation system, DST needs to capture details such as the user's departure location and time. Accurately tracking multi-turn conversations is a major challenge because there may be many information slots to extract.

[0073] 2.2 LLM Confidence Evaluation To reduce the hallucinations of LLMs, a large amount of research has focused on evaluating the credibility of their outputs. Effective methods include evaluating confidence through the consistency of multiple samples and calculating confidence scores based on prediction probabilities (logits).

[0074] Research has also explored confidence in the dialogue state tracking (DST) task to evaluate the uncertainty of responses. However, these methods usually involve freezing the weights of the LLM and cannot train or enhance the model's ability to evaluate its own boundary of capabilities. Therefore, this method does not inherently improve the reliability of the model output.

[0075] 3 Problem Formulation In information extraction tasks such as slot filling (SF) and dialogue state tracking (DST), the reliability of the model is crucial. Since the information extracted usually provides a reference for the execution phase, inaccuracies can lead to serious consequences.

[0076] 3.1 Reliability of the Model Reliability is not just a synonym for accuracy. As discussed in previous related technologies, a reliable model must understand its capabilities and limitations and only provide assistance within its knowledge scope. When faced with questions outside its expertise, the model should avoid providing incorrect answers. A model may have a high accuracy, but if its answers carry an opaque risk of error, especially in cases where the error can have serious implications, it will still be considered unreliable.

[0077] According to the definition of related technologies, we believe that a reliable model should provide the most help by giving as many correct answers as possible and clearly indicating when faced with unknown questions. Therefore, we define reliability as a mixture of accuracy and truthfulness, where the former represents the correctness of the information provided, and the latter ensures that the model transparently conveys its uncertainty.

[0078] 3.2 Explicit Knowledge Boundary Modeling In this section, we will formally present our proposed framework - Explicit Knowledge Boundary Modeling (EKBM), as Figure 4 shown. EKBM aims to improve the reliability of LLMs by explicitly modeling the knowledge boundaries. Our goal is to ensure a high precision of answers even when the model is faced with uncertain or unfamiliar questions, while maintaining a reasonable recall rate. The framework revolves around two different types of answers: definite answers (high confidence) and indefinite answers (low confidence).

[0079] Definite answers are characterized by high confidence and fall within the established knowledge scope of the model. Their main goal is to achieve high precision and minimize the risk of generating incorrect information, especially in critical scenarios such as intelligent reservation systems. The model should confidently provide accurate responses to ensure that users receive reliable information.

[0080] Indefinite responses can address queries outside the model's knowledge boundaries. In this category, we prioritize overall recall, allowing the model to provide useful information even when it cannot give a definite response. The aim is to provide relevant insights to the user while clearly indicating the limitations of the model, helping the user understand the context of the response and adjust their expectations accordingly.

[0081] Decision-making Process Traditionally, information extraction tasks have been regarded as classification problems where the model assigns entities to predefined slots. In our EKBM framework, we redefine this process as a decision-making framework by distinguishing between "certain" and "uncertain" responses.

[0082] The framework operates as follows 1. Clear boundary assessment: The model first evaluates whether each potential piece of information belongs to its knowledge boundary and classifies it as "certain" or "uncertain".

[0083] 2. Execute certain responses: For responses classified as "certain", the model directly enters the execution phase as a high accuracy rate is expected in this phase.

[0084] 3. Refine uncertain responses: For "uncertain" responses, the model will be further refined to reduce uncertainty. This may involve asking the user for confirmation or using a specially trained improvement model to enhance the response quality.

[0085] The EKBM framework can ensure the high accuracy of "certain" responses while effectively managing the uncertainty of "uncertain" responses, thereby improving the reliability of dialogue state tracking (DST) and other task-oriented dialogue tasks. This approach not only improves the overall performance but also enhances users' trust in the model's capabilities, thus enabling more effective and reliable interactions in practical applications.

[0086] 3.3 Optimization Objectives and Evaluation 3.3.1 Optimization Objectives In the EKBM framework we proposed, the requirements for model performance have changed. Traditionally, the main concern was accuracy; however, our framework introduces a series of more refined objectives to improve the reliability of responses.

[0087] Maximize the precision of certain responses The primary task is to ensure the precision of certain responses, which means the precision should be close to 1.

[0088] Maximize the overall recall rate (certain + uncertain) Our goal is to improve the overall recall rate of the model, ensuring that sufficient information can be provided even for queries beyond its knowledge boundary, thus building a more useful model.

[0089] To ensure the stability and sufficient helpfulness of the model, it is also crucial to control the precision of "uncertain" answers and the proportion of "certain" answers. The balance of these aspects can ensure that the model output is reliable and informative.

[0090] 3.3.2 Metrics As our requirements continue to evolve, traditional DST metrics such as Joint Goal Accuracy (JGA) are no longer applicable. We propose several new metrics suitable for our framework: Sample-level metrics We adopt a definition similar to traditional DST metrics and introduce the following metrics: - Joint Goal Precision (JGP): This metric measures the accuracy of affirmative answers. If all Sure predictions are completely accurate, the sample is considered correct, which reflects the precision of the model in this category.

[0091] - Joint Goal Recall (JGR): This metric evaluates the overall recall rate, including both "affirmative" and "uncertain" answers. If a sample captures all relevant pairs that should be extracted, the sample is considered correct, which highlights the model's ability to retrieve necessary information.

[0092] - Joint Goal Precision and Recall (JGPR): This comprehensive metric requires that samples satisfying both JGP and JGR conditions are correct, thus comprehensively evaluating the model performance. When there are no "uncertain" responses, JGPR is consistent with JGA. However, when "uncertain" responses are included, JGPR reflects the theoretical upper limit of JGA, that is, assuming full trust in certain responses and making optimal decisions based on the ideal refinement process of uncertain responses.

[0093] Pair-level metrics In addition to sample-level metrics, we also consider several pair-level metrics to provide a more refined evaluation, including the precision of the certain and uncertain parts, the overall recall rate, and the proportion of certain responses.

[0094] 4 Methods 4.1 Reliability Alignment To train a reliability model, we adopt a two-stage alignment method called "Reliability Training" pipeline, which includes "Reliability Adaptive Training" (RAT) and "Reliability Preference Training". This method aims to improve the model's ability to draw the knowledge boundary while making its output more consistent with the reliability criteria we defined.

[0095] 4.1.1 Reliability Adaptive Training In this stage, we will construct reliability adaptive data as described above. These data will be used as the labeled input for the model, and we use the Supervised Fine-Tuning (SFT) technique to effectively train the model.

[0096] The purpose of this is to enable the model to adapt to the concepts of certain and uncertain answers, thus learning how to allocate confidence based on whether the question is within its knowledge boundary. This adaptation is crucial as it enables the model to make informed decisions about its answers, thereby reducing the likelihood of inadvertently providing misinformation.

[0097] Among them, Figure 5 shows Reliability Adaptive Data Construction. Gold refers to the standard reference answer; Pred refers to the answer predicted by the model; ConstructedSample refers to the constructed sample data; Legend refers to the legend. Figure 6 shows Reliability Preference Data Construction. Gold refers to the standard reference answer; Pred refers to the answer initially predicted by the model; Refined sample refers to the iteratively optimized sample data.

[0098] 4.1.2 Reliability Preference Training The model trained in the previous stage, namely the RAT model, already has the ability to demarcate the knowledge boundary. However, its performance may not fully align with our specific requirements and goals. For example, positive answers may still contain errors, the proportion of positive answers may be too low, and so on.

[0099] To address these issues, we construct preference data according to the requirements and preferences outlined in the foregoing embodiments. Using the aforementioned data, we employ the Direct Preference Optimization (DPO) method for reliability preference training. This training process can further fine-tune the model to ensure it better meets our expectations for a reliable LLM.

[0100] 4.2 Reliability Adaptive Data Construction The construction of reliability adaptive data is specific to a model. For an unadjusted model, we perform inference on the training set of the dataset according to the target information. By defining the concepts of certain and uncertain answers through prompts, we generate an initial version of the model-specific raw data. Detailed prompt information can be found in the subsequent embodiment "A.1.1 Reliability Training Instructions".

[0101] Based on these raw data, we compare the predicted results of each sample with the gold standard results. We classify the segments correctly extracted by the model as within its knowledge boundary (True Positive, TP), while mis-extracted (False Positive, FP) or missed (False Negative, FN) segments are considered outside its capabilities.

[0102] - For TP, our confidence is positive.

[0103] - For FP, we set the confidence as uncertain and correct a portion of the errors with a predefined probability p.

[0104] - For FN, we complete the missing information, randomly assign a definite confidence with probability p, and set the rest as uncertain.

[0105] Here, p is a hyperparameter representing our expectation for improving the ability boundary of the target model. The larger p is, the better the model performance, but the greater the deviation from the original model's ability boundary.

[0106] By constructing the data in this way, we clearly demarcate the knowledge boundary of the original model using confidence.

[0107] 4.3 Construction of Reliability Preference Data The construction of reliability preference data is based on the RAT model obtained in the previous stage of reliability adaptive training. We again perform inference on the training set of the same target information extraction dataset. The results of this stage have distinguished between definite and uncertain confidences, but they may not fully meet our requirements for a reliable LLM.

[0108] Assuming the ideal result is that all the formed data are correctly classified as definite, then we use the gold labels in the original dataset to construct preference data. This method clearly guides the model to achieve our ideal result.

[0109] Specifically, we adopt an iterative optimization strategy, gradually adjusting the current sequence according to our preference principle with the gold label as the ultimate goal. This will generate a strictly increasing list of preference samples from which we randomly draw preference pairs.

[0110] The data can still be divided into correctly extracted segments (true positives, TP), mis-extracted segments (false positives, FP), and missed segments (false negatives, FN): - For TP, we retain these samples without any modification and set the confidence of all samples as positive.

[0111] - For FPs, if the original confidence is certain, we change the confidence to uncertain; if the original confidence is uncertain, we randomly assign a new confidence.

[0112] - For FNs, we complete the missing information and randomly set its confidence.

[0113] In each optimization step, we focus on one optimizable point in the entire sample set to ensure that the optimization list is strictly improved in each iteration.

[0114] 5 Experiments 5.1 Experimental Setup 5.1.1 Models and Baselines We select the 8b version of Llama3 model released by Meta as the backbone and combine the following baselines for reliability evaluation.

[0115] GPT-4 + direct&EBKM prompt GPT-4 is widely regarded as one of the most powerful LLMs currently. It is an important baseline representing the performance ceiling of unadjusted LLMs.

[0116] Llama3-8b + direct&EBKM prompt The Llama3-8b model is our backbone model, allowing us to directly observe the performance differences between adjusted and unadjusted models.

[0117] Llama3-8b + direct SFT We utilize the Llama3-8b model to directly perform supervised fine-tuning (SFT) on the original dataset without adding any reliability constraints. Training instructions are in Appendix A.1.2. Through this setup, we can clearly demonstrate the differences and optimizations between the explicit knowledge boundary modeling (EKBM) framework and traditional classification paradigms.

[0118] Llama3-8b-RAT&Llama3-8b-RPT We use the models trained with the reliability training processes introduced in the foregoing embodiments, where RAT refers to the Llama3-8b model obtained from the reliability adaptive training stage, and RPT refers to the model further trained from the reliability preference training stage.

[0119] 5.1.2 Datasets We use the dialogue state tracking (DST) task as the test scenario and select the MultiWOZ2.2 dataset for training and testing. The MultiWOZ 2.2 dataset includes 67,473 training samples and 7,372 test samples.

[0120] In the reliability adaptive training phase, we randomly sampled 10,000 samples from the training data. Then, we used the original Llama3-8b model for inference and utilized the EBKM prompt to generate the original data for the specific model (Subsequent Example A.2.2 EBKM Prompt). Subsequently, we constructed the reliability adaptive data according to the foregoing example and set the parameter p to 0.3.

[0121] In the reliability preference training phase, we again randomly sampled 10,000 samples from the training dataset. Using the trained Llama3-8b-RAT model, we constructed the reliability preference data as described in Section 4.3 and randomly selected 2,000 pairs of reliability preference pairs as the final training dataset.

[0122] Figure 7 Shows the reliability of the MultiWOZ 2.2 test set. "EKBM" indicates whether the results were obtained under the EBKM framework (i.e., with certainty / uncertainty markers) or using the direct method (i.e., without certainty / uncertainty markers). In particular, the data points marked with "*" indicate that the corresponding results were obtained using the direct method.

[0123] 5.1.3 Training Details In the reliability adaptive training phase, we used DeepSpeed-Chat. In the reliability preference training phase, we modified DeepSpeed-Chat to support direct preference optimization (DPO) training. We adopted all the default training parameters in DeepSpeed-Chat, except that for training stability, we set the learning rate in DPO training to 1e-7 and the sequence length to 1024. We trained both phases for 1 epoch and conducted all experiments using Nvidia A800 GPUs with a batch size of 2 for each GPU.

[0124] 5.2 Reliability Evaluation The main experimental results are shown in Table 2. Clearly, compared with the unadjusted models of GPT-4 and Llama3-8b, our reliability training model has significant improvements in all evaluation metrics. In addition, we also observed that the Llama3-8b model has obvious improvements under direct SFT without certainty / uncertainty labels. This highlights the effectiveness of our reliability training method.

[0125] The improvement in Joint Goal Precision (JGP) indicates that the model has become more reliable. Among the positive responses with high confidence, more than 83% of the samples are completely correct. This represents an improvement of over 14% compared to the direct SFT version, meaning our framework has avoided nearly 11% of sample errors that would be incorrect under traditional methods, making our model more reliable. Additionally, the improvement in Joint Goal Recall (JGR) indicates that our Llama3-8b-RPT model can provide more potential information in simpler classification methods.

[0126] The increase in the precision of the determined part (Precisionsure) and the overall recall (Recallcombined) further confirms these findings, indicating that our model not only improves accuracy but also expands its information retrieval capabilities.

[0127] 5.2.1 JGPR and JGA As mentioned earlier, the Joint Goal Precision and Recall (JGPR) metrics in the EKBM framework represent the theoretical upper limit of Joint Goal Accuracy (JGA), i.e., the joint goal accuracy when we fully trust and accept the determined responses while ensuring that the uncertain responses are ideally refined (e.g., seeking user confirmation for each uncertain response). From Figure 7 it can be seen that compared with the model obtained through direct SFT, the JGPR of the RPT model trained by DPO has increased significantly by approximately 13.6% (from 58.33 to 66.30). This indicates that our framework has considerable potential in improving model reliability.

[0128] When comparing JGA and JGPR, to mitigate the potential interference effect of ideal refinement, we can perform equivalent refinement on the results of the direct SFT model. In the Llama3-8b-RPT model, the proportion of uncertain responses is 13.47%, with a total of 6,276 uncertain units that need to be refined, which we call the cost. We applied the same random refinement cost to the direct SFT results (including 42,004 units) and calculated the resulting JGA, finding it to be 59.02, showing only a slight improvement compared to the JGA of 58.33 before refinement and still a large gap compared to the result of 66.30 for the Llama3-8b-RPT model.

[0129] The reason for this difference is that the direct SFT model does not explicitly learn the knowledge boundary. As a result, a large amount of refinement cost is spent on answers that are already correct. In this case, to ensure the complete accuracy of information extraction, it refines uncertain answers, thus effectively reducing information loss, while traditional methods usually require exhaustive refinement of all predictions, resulting in high costs. In contrast, in our EKBM framework work, the Llama3-8b-RPT model effectively delimits the knowledge boundary and only needs to refine a small part (13.47%) of the low-confidence answers to significantly improve the JGA. This shows that it is feasible to achieve high-reliability information extraction at low cost.

[0130] In addition, the results of GPT-4 and the unadjusted Llama3-8b model show that simply adopting the prompting method is not sufficient to describe the ability boundary of the model, thus strengthening the effectiveness of our reliability training process.

[0131] 5.2.2 Deterministic and Uncertainty Research Figure 8 The proportion of affirmative answers is shown. Among them, Proportionsure refers to the proportion of deterministic answers among all answers.

[0132] In this part of the content, we will make a detailed comparison between the "affirmative" and "uncertain" parts. As Figure 8 shown, the proportion of "deterministic" answers remains at a reasonable level, meeting our expectation that while ensuring reliability, it can also provide sufficient help to users.

[0133] Figure 9 The precision analysis is shown. Among them, Sure-only refers to only deterministic answers at this time; Sure refers to deterministic answers; Unsure refers to uncertain answers; TP refers to True Positive, that is, true positive examples; FP refers to False Positive, that is, false positive examples; Percentage refers to the proportion.

[0134] Regarding precision, we analyzed the "true positive" (TP) and "false positive" proportions of "deterministic" and "uncertain" responses, as Figure 9 shown. The results show that the performance gap between these two categories is large. Specifically, the accuracy of "affirmative" answers of the Llama3-8b-RPT model exceeds that of the traditional direct SFT version, which is consistent with the expected improvement. However, it is worth noting that since our training did not specifically optimize the accuracy of uncertain answers, the accuracy of uncertain answers is still relatively low (about 50%).

[0135] Figure 10 The recall rate analysis is shown. Among them, Sure refers to a definite answer; Unsure refers to an uncertain answer; Missing refers to a missing answer; Percentage refers to the proportion.

[0136] In terms of recall rate, we further illustrate the contributions of the "definite" and "uncertain" parts in Figure 10 . The contribution of the "uncertain" response to the overall recall rate is significant, which is consistent with our main goal in this category: to extract as much potentially relevant information as possible to improve the overall recall rate, which can be used as a reference for users or a basis for subsequent optimization.

[0137] These findings validate the effectiveness of our "reliability training" process, indicating that the model has indeed learned the distinctions and requirements related to the "definite" and "uncertain" concepts. By achieving a balance between these two categories, our framework has successfully improved the reliability and usefulness of the model, demonstrating that our method has achieved the established goal of improving task-oriented dialogue tasks.

[0138] 6 Conclusion In the embodiments of this application, we propose an Explicit Knowledge Boundary Modeling (EKBM) framework to redefine the reliability principle in the TODs information extraction task, shifting from a traditional classification problem to a novel decision-making process.

[0139] Through a two-stage reliability training process, the experimental results show that the EKBM framework significantly improves the reliability of the model, making significant progress compared with traditional methods, paving the way for building a more reliable and effective dialogue system.

[0140] A LLM Description A.1 Training Description A.1.1 Reliability Training Description Please refer to Figure 11a , and generate the dialogue state according to the given dialogue context. Ensure that the result is as reliable as possible, the "definite" part is as accurate as possible, and the overall coverage includes all relevant slots.

[0141] A.1.2 Direct SFT Training Instructions Please refer to Figure 11b , and generate the dialogue state according to the given dialogue context.

[0142] A.2 Inference Prompt A.2.1 Direct Prompt Please refer to Figure 11c, given you a conversation history between USER and SYSTEM, I want you to analyze it and generate the conversation state.

[0143] First, we use a json dict to describe the slots in each domain and their corresponding value spaces. Then, we will specify the requirements you need to follow. Finally, we will demonstrate some use cases.

[0144] ## Domain and Slot Spaces The available ontologies for the conversation state are as follows: { "Hotel": { "name": { "data_type": str, "example": "hamilton lodge" },... },... } ## Requirements 1. Analyze the conversation history carefully and fill in the relevant domains and slots.

[0145] 2.2. For slots with a specified value range, the response must be within the provided range. For slots without a specified value range, the answer must be extracted from the history.

[0146] The answer must be extracted from the history. If the user has no preference, set the value to "dontcare" (don't care).

[0147] 3. You only need to consider the domains and periods relevant to the conversation history. Do not include unrelated domains and slots in your response, and avoid empty domains or slots.

[0148] 4. Your answer should also be in single-line jsonl format and ensure the output is all lowercase. Do not provide any additional prefixes, suffixes, or any explanations.

[0149] ## Example of Output Dialog State {"Hotel": {"area": "center", "name": "alexander bed and breakfast", "parking": "yes", "type": "guesthouse"}, "Attraction": {"name": "Kambar"}} ## Shooting Example... A.2.2 EBKM Prompt Please refer to Figure 11d and Figure 11e, given you a conversation history between USER and SYSTEM, I want you to analyze it and generate the conversation state.

[0150] First, we use a json dict to describe the slots in each domain and their corresponding value spaces. Then, we will specify the requirements you need to comply with. Finally, we will demonstrate some use cases.

[0151] ## Domain and Slot Spaces The available ontologies for the conversation state are as follows: { "Hotel": { "name": { "data_type": str, "example": "hamilton lodge" },... },... } ## Requirements 1. Analyze the conversation history carefully and fill in the relevant domains and slots.

[0152] 2.2. For slots with a specified value range, the response must be within the provided range. For slots without a specified value range, the answer must be extracted from the history.

[0153] The answer must be extracted from the history. If the user has no preference, set the value to "dontcare" (don't care).

[0154] 3. You should cover as many subject and slot-value pairs related to the conversation history as possible. Mark "certain" or "uncertain" for each slot value according to your confidence. If you are very sure that the slot-value pair is indeed relevant and should be included, and you are very confident in the correctness of the slot-value, it should be marked as "certain".

[0155] "certain". Otherwise, the slot-value pair should be marked with "uncertain".

[0156] 4. Principles a. Goal one: As close as possible to 100% accuracy.

[0157] Those "certain" time slot value pairs.

[0158] b. Goal two: The "certain" part and the "uncertain" part should b. Goal two: The "certain" part and the "uncertain" part should cover all possible slots involved in the conversation history as much as possible.

[0159] As much as possible.

[0160] c. Heavy penalty: Providing incorrect slot value pairs in the "Determined" section.

[0161] Providing incorrect time slot value pairs in the "Determined" section.

[0162] d. Heavy penalty: Omitting time slot value pairs that should be extracted.

[0163] Extraction.

[0164] e. Mild penalty: Providing incorrect or redundant time slots in the "Undetermined" section value pairs.

[0165] 5. Your answer should also be in single-line jsonl format and ensure that the output is all in lowercase. Do not provide any additional prefixes, suffixes, or any explanations. For format requirements, please refer to the output dialog status example.

[0166] ## Output dialog status example {"Hotel": {"Area": {"Value": "Center", "Confidence": "Determined"}, "Name": {"Value": "alexander bed and breakfast", "Confidence": "Affirmative"}, "Parking": {"Value": "Yes", "Confidence": "Undetermined"}, "Type": {"Value": "Guesthouse", "Confidence": "Affirmative"}}, "Scenic Spot": "Affirmative"}, "Attraction": {"Name": {"Value": "Kamba", "Confidence": "Undetermined"}}} ## Shooting example... In some other embodiments, the embodiments of the present invention further provide a non-volatile computer storage medium. The computer storage medium stores computer-executable instructions, and the computer-executable instructions can execute the large model processing method for task-oriented dialogue in any of the above method embodiments; As an implementation manner, the non-volatile computer storage medium of the present invention stores computer-executable instructions, and the computer-executable instructions are set as: According to the dialogue history input by the user, automatically analyze the results that should be extracted or predicted and the corresponding confidence labels, where the confidence labels include definitely accepted, undetermined, and clearly rejected; According to the confidence labels corresponding to each prediction result, allocate the prediction results to different processing paths for processing. Among them, those determined to be accepted are directly output as the final result, those clearly rejected are directly rejected and the user is informed that the result is unreliable, and those that are uncertain are submitted to the slow-thinking model in the next stage for further decision-making processing; For the uncertain part in the previous stage, use the slow-thinking model to conduct more in-depth and detailed slow thinking to determine the final processing result; After the task is completed, return the final result to the user, and at the same time adjust the confidence label ratio according to the feedback of the task.

[0167] The non-volatile computer-readable storage medium may include a storage program area and a storage data area. Among them, the storage program area can store an operating system and application programs required for at least one function; the storage data area can store data created according to the use of the large model processing method and system for task-oriented dialogue, etc. In addition, the non-volatile computer-readable storage medium may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the non-volatile computer-readable storage medium may optionally include a memory remotely set relative to the processor, and these remote memories can be connected to the large model processing method for task-oriented dialogue through a network. Examples of the above network include but are not limited to the Internet, enterprise intranet, local area network, mobile communication network, and their combinations.

[0168] The embodiment of the present invention also provides a computer program product. The computer program product includes a computer program stored on a non-volatile computer-readable storage medium. The computer program includes program instructions. When the program instructions are executed by a computer, the computer is made to execute any one of the above-mentioned large model processing methods for task-oriented dialogue.

[0169] Figure 12 is a schematic structural diagram of the electronic device provided by the embodiment of the present invention, as Figure 12 shown, the device includes: one or more processors 1210 and a memory 1220, Figure 12 Taking one processor 1210 as an example. The device of the large model processing method and system for task-oriented dialogue may further include: an input device 1230 and an output device 1240. The processor 1210, the memory 1220, the input device 1230, and the output device 1240 can be connected through a bus or other means, Figure 12Take the bus connection as an example. The memory 1220 is the above-mentioned non-volatile computer-readable storage medium. The processor 1210 executes various functional applications and data processing of the server by running non-volatile software programs, instructions, and modules stored in the memory 1220, that is, implements the above-mentioned method embodiment of the large model processing method for task-oriented dialogue. The input device 1230 can receive input digital or character information, and generate key signal inputs related to user settings and function controls of the large model processing device for task-oriented dialogue. The output device 1240 can include display devices such as a display screen.

[0170] The above product can execute the method provided by the embodiment of the present invention, and has corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiment of the present invention.

[0171] As an implementation manner, the above electronic device is applied to a large model processing device for task-oriented dialogue, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to: Automatically analyze the results to be extracted or predicted and the corresponding confidence labels according to the dialogue history input by the user, wherein the confidence labels include definitely accepted, uncertain, and clearly rejected; According to the confidence label corresponding to each prediction result, allocate the prediction results to different processing paths for processing, wherein the definitely accepted ones are directly output as the final results, the clearly rejected ones are directly rejected and the user is told that the results are unreliable, and the uncertain ones are submitted to the slow thinking model in the next stage for further decision-making processing; For the uncertain part in the previous stage, use the slow thinking model to perform more in-depth and detailed slow thinking to determine the final processing result; After the task is completed, return the final result to the user, and at the same time adjust the confidence label ratio according to the feedback of the task.

[0172] The electronic device in the embodiment of the present application exists in various forms, including but not limited to: (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.

[0173] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDAs, MIDs, and UMPC devices, etc.

[0174] (3) Portable entertainment devices: Such devices can display and play multimedia content. This type of device includes: audio and video players, handheld game consoles, e-books, as well as smart toys and portable in-vehicle navigation devices.

[0175] (4) Servers: Devices that provide computing services. The composition of a server includes a processor, hard disk, memory, system bus, etc. Servers are similar to general computer architectures, but due to the need to provide highly reliable services, they have higher requirements in terms of processing power, stability, reliability, security, scalability, manageability, etc.

[0176] (5) Other electronic devices with data interaction functions.

[0177] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.

[0178] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course, it can also be implemented by hardware. Based on this understanding, the essence of the above technical solutions, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods of each embodiment or some parts of the embodiments.

[0179] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of each embodiment of the present invention.

Claims

1. A large model processing method for task-oriented dialogue, comprising: According to the conversation history input by the user, automatically analyze the results that should be extracted or predicted and the corresponding confidence labels, where the confidence labels include definitely accepted, uncertain, and definitely rejected; According to the confidence label corresponding to each prediction result, the prediction result is assigned to different processing paths for processing, wherein the ones that are determined to be accepted are directly output as the final result, the ones that are clearly rejected are directly rejected and the user is told that the result is unreliable, and the uncertain ones are submitted to the next stage of the slow thinking model for further decision processing; For the uncertain parts in the previous stage, use the slow thinking model to conduct more in-depth and detailed slow thinking to determine the final processing results; After the task is completed, the final result is returned to the user, and the confidence label ratio is adjusted according to the feedback of the task.

2. The method according to claim 1, wherein: The slow thinking model includes slow thinking process design, slow thinking data synthesis and slow thinking model training. The slow thinking process design is to design the principles and logic of slow thinking according to specific task requirements, and to design an effective thinking process; The slow thinking data synthesis uses existing data and labels based on a designed thinking process, and utilizes the high-performance closed-source model GPT-4o to synthesize data for own model training; The slow thinking model training uses synthetic data to train a model that can effectively perform slow thinking through the SFT supervised training method.

3. The method according to claim 1, wherein: The results that should be extracted or predicted by automatic analysis based on the conversation history input by the user and the corresponding confidence labels include: Use natural language processing techniques to extract information required by the task from the input history; For each predicted result, a corresponding confidence label needs to be explicitly provided, and it is ensured that the confidence label can effectively reflect the true confidence state of the model, wherein the confidence label ratio needs to meet the preset goals and principles.

4. The method according to claim 3, wherein: For results with different confidence labels, corresponding processing paths are set according to the preset goals and principles.

5. The method according to claim 1, wherein: After the task is completed, the final result is returned to the user, and the confidence label ratio according to the feedback of the task includes: Continuously optimize the strategies in the process by analyzing the feedback during task execution; Based on task results and expected strategy adjustments, optimized training data is constructed to further improve the system's adaptability through reinforcement learning.

6. The method according to any one of claims 1 to 5, wherein: The method optimizes the decision-making mechanism in the task-based dialogue system, introduces explicit three-category confidence labels and further slow thinking for uncertain results, significantly improves the reliability and success rate of task execution, reduces the error rate and resource waste, and improves system efficiency.

7. An electronic device, comprising: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the steps of the method described in any one of claims 1 to 6.

8. A storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Cited By

  • Method and device for training model, storage medium and electronic equipment

    CN120523957A