Task type multi-round dialogue method and system based on large language model
Through structured knowledge configuration based on large language model and dynamic prompt word template injection, combined with general intention recognition and dialogue status tracking, the accuracy and fluency of task-type multi-round dialogue system in the lack of domain data is solved, and high availability and high accuracy dialogue task completion is achieved.
Patent Information
- Application Number
- CN202510383884.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-28
- Publication Date
- 2025-07-29
AI Technical Summary
In the absence of field data, existing task-type multi-round dialogue systems are difficult to achieve high availability and high accuracy dialogue performance.
Using a method based on a large language model, through structured knowledge configuration and dynamic prompt word template injection, combined with general intention recognition and dialogue status tracking, a task-type multi-round dialogue system is built to break away from domain data dependence and improve intention recognition capabilities.
It improves the accuracy and fluency of the dialogue system, reduces dependence on domain data, and achieves high availability and high accuracy dialogue tasks.
Smart Images

Figure CN120386839A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of task-oriented multi-turn dialogue, and in particular to a task-oriented multi-turn dialogue method and system based on a large language model. Background Art
[0002] As a core technology of human-computer interaction, the task-oriented dialogue system (TOD) aims to complete specific domain tasks (such as ticket booking, product query, etc.) through continuous dialogue, and its technology development has experienced an evolution from rule-driven to data-driven. Early systems relied on manually designed grammar rules (such as Eliza), suffering from insufficient cross-domain scalability; with the rise of statistical machine learning and deep learning, the dialogue management module based on neural networks has gradually become the mainstream.
[0003] Modern TOD systems usually adopt a modular architecture, including four core components: natural language understanding (NLU), dialogue state tracking (DST), dialogue policy learning (DPL), and natural language generation (NLG). Among them, DST dynamically updates the dialogue state by integrating the current semantic understanding result and the historical dialogue context; DPL formulates response strategies based on the state information to gradually achieve the task goal. The application of joint learning technology further improves the collaborative efficiency of intent recognition and slot filling, for example, by sharing the encoding layer parameters through a multi-task learning framework.
[0004] Although large language models (such as the GPT series) have significantly enhanced dialogue fluency, the actual deployment of task-oriented systems still faces challenges: end-to-end models have problems such as poor controllability and high data dependence, while the modular architecture needs to balance generalization ability and engineering complexity. In industrial scenarios, how to achieve accurate context understanding, efficient state tracking, and robust policy optimization is still the key direction of technological breakthrough. Summary of the Invention
[0005] Aiming at the above-mentioned defects of the prior art, the purpose of the present invention is to provide a task-oriented multi-turn dialogue method and system based on a large language model, aiming to quickly build a highly available and highly accurate task-oriented multi-turn dialogue system when lacking domain data for model training.
[0006] The present invention adopts the following technical solutions to solve the technical problems:
[0007] In a first aspect, a task-oriented multi-turn dialogue method based on a large language model includes the following steps:
[0008] Step S1, customizing scenario knowledge in the form of structured knowledge configuration and initializing the system;
[0009] Step S2: Obtain the current user input and the human-machine conversation history, and use a general intent recognition large model to recognize the general intent of the current user input.
[0010] Step S3: Inject the scenario knowledge through a dynamic prompt template, input it into the scenario element extraction large model, and recognize the scenario intent and task elements of the current user input.
[0011] Step S4: Based on the recognized general intent, scenario intent, and task elements, perform dialogue state tracking and dialogue strategy generation according to the rules.
[0012] Step S5: Input the current user input, the human-machine conversation history, and the dialogue strategy into the dialogue generation large model to generate a dialogue response.
[0013] Step S6: Repeat Steps S1 to S5 until the dialogue task is completed and the dialogue ends.
[0014] Further, in Step S1, the scenario knowledge includes: scenario intent, task elements, and task logic.
[0015] Further, the element attributes include: the task to which it belongs, which is the global element by default if not specified to be displayed; the value type, including the basic string type, discrete type, and numerical type; the assignment method, including the attachable type and non-attachable type; whether to reset (whether the element is restored to the null value in future conversations after being assigned a value).
[0016] Further, in Step S2, the process of generating the general intent includes:
[0017] Obtain a task-based multi-turn dialogue dataset annotated based on the general intent system; the general intent is a general dialogue behavior formulated by humans and does not depend on a specific dialogue scenario.
[0018] Use this dataset to construct prompts and fine-tune the large language model. After fine-tuning, obtain the general intent recognition large model, and input the human-machine conversation history and the current user dialogue into the general intent recognition large model to recognize the general intent of the current user input.
[0019] Further, in Step S3, the method of injecting the scenario knowledge through a dynamic prompt template includes:
[0020] Decompose and serialize the scenario intent and task elements in the scenario knowledge into the following text form:
[0021] “{Scenario intention Figure 1}: \nIts meaning is {Scenario intention Figure 1Meaning}\nThe slots it contains are {Slot 1 affiliated with Task 1}{Meaning of Slot 1}{Value type of Slot 1}{Legal values of Slot 1 (optional)}\n{Slot 2 affiliated with Task 1}{Meaning of Slot 2}{Value type of Slot 2}{Legal values of Slot 2 (optional)}\n{Scenario meaning Figure 2}...”
[0022] Concatenate the above scenario knowledge with the human-machine dialogue history and the user's current input to form a prompt input for the un-fine-tuned element extraction large model, and obtain the scenario intention and task elements of the user's current input.
[0023] Further, in step S4, the rules for dialogue state tracking and dialogue strategy generation based on the general intention, scenario intention, and task elements include:
[0024] Update the current dialogue state according to the combination of the general intention, scenario intention, and task elements; each general intention has a corresponding state update method and possible dialogue strategies, update the state parameters of the current task, active task, and element values in the dialogue context, and the dialogue strategy is specific to certain intentions and is inserted into the dialogue process as a high-priority dialogue requirement;
[0025] Specifically, it is divided into the following situations: The general intention group contains an informing intention, and update the dialogue state according to the identified task elements; the general intention group contains a responding intention or its sub-intentions, and update the dialogue state according to the elements sent by the system in the context but not yet replied; the general intention group contains a requesting intention, judge whether the identified scenario intention is consistent with the scenario intention maintained in the current dialogue state, if not, trigger the task conversion mechanism, save the context state, and make a dialogue strategy of asking the user whether to switch tasks, if consistent, continue the current dialogue; the general intention group contains an asking intention or its sub-intentions, then query the database with the identified task elements as the content of the question, and formulate a reply strategy according to the specific sub-intentions; the general intention group contains a sub-intention of a modifying intention, verify whether the intention is consistent with the assignment method of the identified task elements, if not, make a dialogue strategy indicating that the modification fails, otherwise modify according to the assignment method; the general intention group contains a complaining intention, then make an apologetic response.
[0026] Further, in step S5, use the large model to generate a dialogue reply and construct a prompt word in the following structural order:
[0027] {Define the statement that sets the role of the large model to the customer service role corresponding to the scenario}\n{Precautions}\n{Examples}\n{Human-machine historical dialogue}\n{User's current dialogue}\n{Statement to guide the large model to use the chain of thought}\n{Database query results (optional)}\n{Dialogue strategy}.
[0028] In a second aspect, a task-based multi-turn dialogue system based on a large language model includes:
[0029] A human-computer interaction module for providing a chat interaction interface for users, with functions of content input and chat record viewing;
[0030] A scenario knowledge customization module for pre-inputting the scenario knowledge of the dialogue scenarios applied by the system, including scenario intents, task elements, and task logics;
[0031] An intent recognition module for recognizing the user's dialogue input in each turn of the dialogue, and recognizing the user's general intent, scenario intent, and task elements;
[0032] A dialogue state management module for maintaining the global dialogue state in each dialogue and recording the update status of the dialogue state in each turn;
[0033] A dialogue policy generation module for formulating the current dialogue policy based on the results of the intent recognition module and the dialogue state management module;
[0034] A dialogue response generation module for generating dialogue responses based on the dialogue policy;
[0035] A data storage module for storing scenario data in a database and implementing relevant interfaces to query the required data during the dialogue process.
[0036] In a third aspect, a terminal device includes a memory, a processor, and a program of a task-based multi-turn dialogue method based on a large language model stored in the memory and executable on the processor. When the processor executes the program of the task-based multi-turn dialogue method based on a large language model, the steps of the task-based multi-turn dialogue method based on a large language model described in any one of the above are implemented.
[0037] In a fourth aspect, a computer-readable storage medium stores a program of a task-based multi-turn dialogue method based on a large language model. When the program of the task-based multi-turn dialogue method based on a large language model is executed by a processor, the steps of the task-based multi-turn dialogue method based on a large language model described in any one of the above are implemented.
[0038] Beneficial effects:
[0039] Based on large language models and agent technology, on the basis of the modular architecture of traditional task-based multi-turn dialogue systems, the present invention presents an implementation method of task-based multi-turn dialogue systems in the era of large language models. Through the paradigm innovation of introducing large language models, we construct a zero-shot dialogue system architecture based on prompt engineering, replace model fine-tuning with structured knowledge configuration, inject scenario knowledge through dynamic prompt templates, and realize multi-turn dialogue reasoning in combination with the context management mechanism. At the same time, taking advantage of the fact that the effect of general intent recognition is independent of domain data, the general intent and scenario intent are decoupled, and the overall intent recognition ability is improved by separately enhancing the general intent recognition ability. The potential dialogue patterns in the general intent are fully explored and applied in the dialogue state tracking and dialogue policy generation modules, improving the naturalness and fluency of the generated dialogue responses. BRIEF DESCRIPTION OF THE DRAWINGS
[0040] Figure 1 is a flowchart of the task-based multi-turn dialogue method based on large language models of the present invention;
[0041] Figure 2 is a simple example diagram of the work of each module in completing a round of dialogue based on the task-based multi-turn dialogue method based on large language models of the present invention;
[0042] Figure 3 is a relationship diagram of each module of the task-based multi-turn dialogue system based on large language models of the present invention;
[0043] Figure 4 is a display diagram of the human-machine dialogue interface in the embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0044] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0045] In the specific implementation of the present invention, a task-based multi-turn dialogue method based on large language models is proposed to solve the problem of poor dialogue performance of traditional task-based multi-turn dialogue systems in the general situation of lacking domain data. This method aims to improve the accuracy and universality of task-based multi-turn dialogue systems by exploring the potential dialogue patterns in general intents that are irrelevant to specific dialogue scenarios and applying them to each module of traditional task-based multi-turn dialogue systems.
[0046] To facilitate the reader's understanding of the entire disclosure, the following terms used in this disclosure are now explained. It should be noted that the explanations of terms in this article are only for assisting the reader's understanding and do not constitute a limitation on the technical solutions of this disclosure;
[0047] Large language model: A deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of language text. Generally, it refers to a language model with scale parameters reaching billions or more;
[0048] Task-oriented multi-turn dialogue: A human-computer interaction system specifically designed to help users complete specific tasks. Different from open-domain dialogue systems, TOD systems focus on solving specific problems or performing specific tasks, such as booking a restaurant, querying the weather, purchasing an airline ticket, providing navigation guidance, etc.
[0049] Reference Figure 1 And Figure 2 , the present invention discloses a task-oriented multi-turn dialogue method based on a large language model, including the following steps:
[0050] Step S1, customize scenario knowledge in the form of structured knowledge configuration and initialize the system;
[0051] Step S2, obtain the current user input and the human-computer dialogue history, and use a general intent recognition large model to recognize the general intent of the current user input;
[0052] Step S3, inject the scenario knowledge through a dynamic prompt template, input it into the scenario element extraction large model, and recognize the scenario intent and task elements of the current user input;
[0053] Step S4, based on the recognized general intent, scenario intent and task elements, perform dialogue state tracking and dialogue strategy generation according to rules;
[0054] Step S5, input the current user input, the human-computer dialogue history and the dialogue strategy into the dialogue generation large model to generate a dialogue response;
[0055] Step S6, repeat steps S1 to S5 until the dialogue task is completed and the dialogue ends.
[0056] Further optimize the technical solution. In step S1, the scenario knowledge includes: Scenario intention; The scenario intention is an intention closely related to the conversation scenario, such as the "book a flight" intention in the airline customer service scenario and the "investment plan" intention in the financial investment customer service scenario. The scenario knowledge includes all services that the task-based dialogue system can complete for the customer service. Each scenario intention corresponds to a task. Task elements; Each element in the task element group has several element attributes. The element attributes include: The task to which it belongs, and if not specified to be hidden, it is defaulted to a global element; Value type, including basic string type, discrete type (can only select from a finite number of legal values), and numerical type (can only be a number, and its numerical range can be specified); Assignment method, including appendable type (the value of this element can add new values on the original basis) and non-appendable type; Whether to reset (whether this element is restored to an empty value in future conversations after being assigned a value). Task logic; Corresponding to each task, the main process of completing the task is set through a configuration file. The main process stipulates from the perspective of the system the user information that should be collected, actions in the face of various user responses, and other contents.
[0057] Further optimize the technical solution. In step S2, the process of generating general intentions includes:
[0058] Obtain a task-based multi-turn dialogue dataset annotated based on the general intention system; The general intention is a general dialogue behavior formulated artificially and does not depend on a specific conversation scenario;
[0059] Use this dataset to construct prompt words and fine-tune the large language model. After fine-tuning, obtain a general intention recognition large model, and input the human-machine conversation history and the current user conversation into the general intention recognition large model to recognize the general intention input by the current user.
[0060] Further optimize the technical solution. In step S3, the method of injecting scenario knowledge through a dynamic prompt word template includes:
[0061] Decompose and serialize the scenario intention and task elements in the scenario knowledge into the following text form:
[0062] “{Scenario intention Figure 1}: \nIts meaning is {Scenario intention Figure 1 's meaning}\nThe slots it contains are {Slot 1 belonging to task 1}{Meaning of slot 1}{Value type of slot 1}{Legal values of slot 1 (optional)}\n{Slot 2 belonging to task 1}{Meaning of slot 2}{Value type of slot 2}{Legal values of slot 2 (optional)}\n{Scenario intention Figure 2}...”;
[0063] The above scene knowledge is concatenated with the human-machine dialogue history and the user's current input to form a prompt input for the unfinetuned element extraction large model, so as to obtain the scene intention and task elements of the user's current input.
[0064] Further optimize the technical solution. In step S4, the rules for dialogue state tracking and dialogue policy generation based on the general intention, scene intention, and task elements are as follows:
[0065] Update the current dialogue state according to the combination of the general intention, scene intention, and task elements; each general intention has a corresponding state update method and possible dialogue policies, and update the state parameters of the current task, active tasks, element values, etc. in the dialogue context. The dialogue policy is specific to certain intentions and is inserted into the dialogue process as a high-priority dialogue requirement;
[0066] Specifically, it is divided into the following situations: If the general intention group contains an informing intention, update the dialogue state according to the recognized task elements; if the general intention group contains a responding intention or its sub-intentions, update the dialogue state according to the elements sent by the system in the context but not yet replied; if the general intention group contains a requesting intention, judge whether the recognized scene intention is consistent with the scene intention maintained in the current dialogue state. If not, trigger the task conversion mechanism, save the context state, and make a dialogue policy of asking the user whether to switch tasks. If they are consistent, continue the current dialogue; if the general intention group contains an asking intention or its sub-intentions, query the database with the recognized task elements as the content of the question, and formulate a reply policy according to the specific sub-intentions; if the general intention group contains a sub-intention of a modifying intention, verify whether the intention is consistent with the assignment method of the recognized task elements. If not, make a dialogue policy indicating that the modification fails, otherwise modify according to the assignment method; if the general intention group contains a complaining intention, make an apologetic response.
[0067] Further optimize the technical solution. In step S5, use the large model to generate a dialogue reply and construct the prompt according to the following structural order:
[0068] {Define the statement that sets the role of the large model to the customer service role corresponding to the scene}\n{Precautions}\n{Examples}\n{Human-machine historical dialogue}\n{User's current dialogue}\n{Statement to guide the large model to use the chain of thought}\n{Database query results (optional)}\n{Dialogue policy}.
[0069] In the second aspect, the present invention also provides a task-based multi-turn dialogue system based on a large language model, including:
[0070] A human-computer interaction module, which is used to provide a chat interaction interface for users and provide functions such as content input and chat record viewing;
[0071] A scenario knowledge customization module for pre-inputting the scenario knowledge of the dialogue scenarios applied by the system, including scenario intents, task elements, and task logics;
[0072] An intent recognition module for recognizing the user's dialogue input in each round of dialogue and identifying the user's general intent, scenario intent, and task elements;
[0073] A dialogue state management module for maintaining the global dialogue state in each dialogue and recording the update status of the dialogue state in each round;
[0074] A dialogue policy generation module for formulating the current dialogue policy based on the results of the intent recognition module and the dialogue state management module;
[0075] A dialogue response generation module for generating dialogue responses based on the dialogue policy;
[0076] A data storage module for storing scenario data in a database and implementing relevant interfaces to query the required data during the dialogue process.
[0077] Thirdly, the present invention further provides a terminal device, which includes a memory, a processor, and a program of the task-based multi-round dialogue method based on a large language model stored in the memory and executable on the processor. When the processor executes the program of the task-based multi-round dialogue method based on a large language model, the steps of the task-based multi-round dialogue method based on a large language model according to any one of claims 1-7 are implemented.
[0078] Fourthly, the present invention further provides a computer-readable storage medium, on which a program of the task-based multi-round dialogue method based on a large language model is stored. When the program of the task-based multi-round dialogue method based on a large language model is executed by a processor, the steps of the task-based multi-round dialogue method based on a large language model according to any one of the above are implemented.
[0079] Embodiment
[0080] This embodiment provides a task-based multi-round dialogue method based on a large language model, including the following steps:
[0081] Step S1, customizing scenario knowledge and initializing the system in the way of structured knowledge configuration;
[0082] In this step, a configuration file is used to enable dialogue system developers in each field to define scenario knowledge. To ensure that the scenario knowledge filled in by the developers meets the specifications, the system provides a front-end interface and uses a GUI (Graphical User Interface) to limit the content to be filled in.
[0083] Specifically, system developers need to fill in the following three parts of content.
[0084] The first part is the scenario intention. The scenario intention is an intention closely related to the dialogue scenario, such as the "book a flight" intention in the airline customer service scenario and the "investment planning" intention in the financial investment customer service scenario. Scenario knowledge contains all services that the task-based dialogue system can perform for the customer service. Each scenario intention corresponds to a task. In this embodiment, the system will set some default scenario intentions, including but not limited to "other", "product retrieval", etc., to handle some common tasks. For example, for "other", the system will consider that the user's current dialogue does not hit the services provided by the system; for "product retrieval", the default behavior of the system is to collect the attributes of the product to be retrieved, then initiate a retrieval request to the internal database, and return the results. Developers can disable or rewrite this intention to implement more complex logic. In the GUI, it can be represented as a list that can be infinitely added, and the name of the element is filled in by the developer himself.
[0085] The second part is the task elements. Task elements refer to the information required to complete a task, and their manifestation form in the dialogue state is key-value pairs. Each element in the task element group needs to be defined with several attributes. These attributes include: the task to which it belongs, and if not explicitly specified, it defaults to a global element. If an element is specified as a global element, it will be recognized as a legally recognizable element in each round of dialogue process, otherwise it will only be considered a recognizable element when the task to which it belongs is activated; value type, including basic string type, discrete type (can only select from a finite number of legal values), numerical type (can only be a number, and its numerical range can be specified); assignment method, including appendable type (the value of this element can be newly added on the original basis) and non-appendable type; whether to reset (whether the element is restored to an empty value after the relevant state update and policy formulation operations are completed after being assigned a value). In the GUI, it can be represented as a list that can be infinitely added. When adding a new element each time, the name and the above-defined attributes need to be filled in.
[0086] Task Logic. Corresponding to each task, the main process of completing the task is set in the form of a configuration file. The main process stipulates from the perspective of the system the user information that should be collected and the response behaviors when speaking to various users. The task logic consists of a series of ordered system behaviors. These system behaviors are general, that is, most system behaviors can be covered by selecting these actions and then specifying the parameters attached to these actions. In this embodiment, the system behaviors can be divided into the following categories: information inquiry: (element name); database retrieval; reply; custom function call; branch judgment. For the information inquiry behavior, the developer needs to fill in the element name to be inquired. The subordinate task attribute of the selected element should correspond to this task, otherwise it cannot pass the verification pre-loaded by the system. For the database retrieval behavior, the developer needs to fill in the specific table name and column name. For the reply behavior, the reply statement can be customized or generated by the dialogue reply model. For the custom function call behavior, the developer needs to specify the function name to be called and implement the function in another code file. For the branch call behavior, the developer needs to use the conditional branch statement of python to write the branch condition and branch result.
[0087] Step S2: Obtain the current user input and the human-computer dialogue history, and use the general intent recognition large model to recognize the general intent of the current user input and generate the corresponding context description;
[0088] In this embodiment, the general intent system adopted consists of 24 general intents organized in a hierarchical form. These intents are divided into six major categories: inquiry intent, reply intent, modification intent, notification intent, request intent, and other intents. Among them, the inquiry intent is divided into 4 sub-intents, namely, information inquiry intent, whether inquiry intent, comparison inquiry intent, and reason inquiry intent; the reply intent is divided into 9 sub-intents, namely, reply notification, reply affirmation, reply denial, reply ignorance, reply awareness, reply acceptance, reply rejection, reply completion, and reply incompletion; the modification intent is divided into 3 sub-intents, namely, modification override, modification supplement, and modification cancellation; the other intents include 6 intents: complaint, gratitude, greeting, apology, goodbye, and unknown.
[0089] In this embodiment, the general intent recognition large model used is fine-tuned based on the dataset labeled by the above general intent system. The open-source datasets MultiWOZ and CrossWOZ are obtained, and 40 rounds of conversations are extracted from each of them for annotation. Then fine-tuning is performed. Use the general intent recognition large model to recognize the general intent of the current user input;
[0090] Step S3: Inject the scenario knowledge through the dynamic prompt template, input it into the scenario element extraction large model, and recognize the scenario intent and task elements of the current user input;
[0091] In this embodiment, the scenario knowledge loaded in step S1 is constructed in the form of the following prompt words:
[0092] {Statement defining the role of the large language model as an expert in scenario intention recognition}\n{Task description}\n{Precautions}\n{Examples}\n{List of scenario intentions}\n{List of task elements}\n{Conversation history}\n{Current user input}\n{Output guiding words}
[0093] The scenario intentions placed in this prompt word template are all scenario intentions plus the system-built-in intentions. The task elements placed in this prompt word template are the task elements belonging to the currently ongoing or activated task and their legal values or value types.
[0094] Step S4, based on the recognized general intention, scenario intention, and task elements, perform dialogue state tracking and dialogue strategy generation according to the rules;
[0095] During the progress of the conversation, several dialogue states are maintained, including the state of each task and the filling status of each task element. The task state is divided into four states: not activated, activated, ongoing, and completed; all tasks are initially in the not activated state, indicating that the task has not been awakened by the user; until the user's scenario intention for a certain task and the general intention is a request intention, the state of the task is changed to ongoing; when the user's topic temporarily deviates from a certain task, the state of the task is changed to activated, indicating that it is currently put on hold by the user; when the task logic of a certain task is completed, the task is marked as completed, and the most recently started task is selected from the tasks in the activated state and marked as the ongoing state according to the LIFO (Last In First Out) principle.
[0096] Each general intention has a corresponding state update method and possible dialogue strategies, which update the state parameters such as the current task, activated tasks, and element values in the dialogue context. The dialogue strategy is specific to certain intentions and is inserted into the task logic as a high-priority dialogue requirement.
[0097] Specifically, it is divided into the following situations: the general intent group includes the notification intent, which updates the dialogue status according to the identified task elements; the general intent group includes the reply intent or its sub-intent, which updates the dialogue status according to the elements sent by the system in the context but not replied; the general intent group includes the request intent, which determines whether the identified scene intent is consistent with the scene intent maintained in the current dialogue status. If not, the task switching mechanism is triggered, the context state is saved, and a dialogue strategy is made to ask the user whether to switch tasks. If they are consistent, the current dialogue is continued; the general intent group includes the inquiry intent or its sub-intent, and the identified task elements are used as the content of the inquiry to query the database, and a reply strategy is formulated according to the specific sub-intent; the general intent group contains the sub-intent of the modification intent, which verifies whether the intent is consistent with the assignment method of the identified task elements. If not, a dialogue strategy is made to prompt that the modification failed, otherwise the modification is made according to the assignment method; the general intent group contains the complaint intent, and an apologetic response is made.
[0098] Step S5: Input the user's current input, the human-computer dialogue history and the dialogue strategy into the dialogue generation model to generate dialogue responses, such as Figure 4 As shown;
[0099] If there is no fixed line for a response in the task logic, the dialogue generation model will be called to generate the dialogue.
[0100] In this embodiment, the large model used to generate the dialogue response is not fine-tuned. The prompt word template used to generate the dialogue response is as follows:
[0101] {Set the role of the large language model to the role definition statement of the customer service expert in the corresponding scenario}\n{Task description}\n{Notes}\n{Example}\n{Dialogue history}\n{Current user input}\n{Context description of the general intent}\n{Dialogue strategy}\n{Output guide words}.
[0102] Step S6, repeating steps S1 to S5 until the dialogue task is completed and the dialogue ends.
[0103] like Figure 3 As shown, the present invention provides a task-based multi-turn dialogue system based on a large language model. The core functions of the system include: human-computer interaction, scenario knowledge customization, intent recognition, dialogue state management, dialogue strategy generation, dialogue response generation, data storage and other modules.
[0104] Detailed description of system modules;
[0105] The human-computer interaction module is used to provide users with a chat interaction interface, providing functions such as content input and chat history viewing;
[0106] A scenario knowledge customization module, which is used to pre-enter the scenario knowledge of the dialogue scenarios applied by the system, including scenario intents, task elements, and task logics;
[0107] An intent recognition module, which is used to recognize the user's dialogue input in each round of dialogue, and recognize the user's general intent, scenario intent, and task elements;
[0108] A dialogue state management module, which is used to maintain the global dialogue state in each dialogue and record the update status of the dialogue state in each round;
[0109] A dialogue policy generation module, which is used to formulate the current dialogue policy based on the results of the intent recognition module and the dialogue state management module;
[0110] A dialogue response generation module, which is used to generate dialogue responses based on the dialogue policy;
[0111] A data storage module, which is used to store scenario data in a database and implement relevant interfaces to query the required data during the dialogue process.
[0112] In the description of this specification, the descriptions referring to terms such as "one embodiment", "example", "specific example", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.
[0113] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A task-based multi-turn dialogue method based on large language models, characterized in that, It includes the following steps: Step S1, customize scenario knowledge in the way of structured knowledge configuration and initialize the system; Step S2, obtain the current user input and the human-computer dialogue history, and use the general intent recognition large model to recognize the general intent of the current user input; Step S3, inject the scenario knowledge through the dynamic prompt template, input it into the scenario element extraction large model, and recognize the scenario intent and task elements of the current user input; Step S4, based on the recognized general intent, scenario intent and task elements, perform dialogue state tracking and dialogue strategy generation according to the rules; Step S5, input the current user input, the human-computer dialogue history and the dialogue strategy into the dialogue generation large model to generate a dialogue response; Step S6, repeat steps S1 to S5 until the dialogue task is completed and the dialogue ends.
2. The task-based multi-round dialogue method based on a large language model according to claim 1, wherein, In step S1, the scenario knowledge includes: scenario intent, task elements and task logic.
3. The task-based multi-turn dialogue method based on a large language model according to claim 2, wherein, The element attributes include: the affiliated task, which is the global element by default if not specified; the value type, including the basic string type, discrete type, and numerical type; the assignment method, including the attachable type and non-attachable type; whether to reset (whether the element is restored to the null value in future dialogues after being assigned a value).
4. A task-based multi-turn dialogue method based on a large language model according to claim 1, characterized in that, In step S2, the process of generating the general intent includes: Obtain a task-based multi-turn dialogue dataset annotated based on the general intent system; the general intent is a general dialogue behavior formulated by humans and does not depend on a specific dialogue scenario; Use this dataset to construct a prompt and fine-tune the large language model. After fine-tuning, obtain the general intent recognition large model, and input the human-computer dialogue history and the current user dialogue input into the general intent recognition large model to recognize the general intent of the current user input.
5. A task-based multi-turn dialogue method based on a large language model according to claim 1, characterized in that, In step S3, the method of injecting the scenario knowledge through the dynamic prompt template includes: Decompose and serialize the scenario intent and task elements in the scenario knowledge into the following text form: "{Scenario Intent 1}: \nIts meaning is {the meaning of Scenario Intent 1}\nThe slots it contains are {Slot 1 affiliated with Task 1}{the meaning of Slot 1}{the value type of Slot 1}{the legal value of Slot 1 (optional)}\n{Slot 2 affiliated with Task 1}{the meaning of Slot 2}{the value type of Slot 2}{the legal value of Slot 2 (optional)}\n{Scenario Intent 2}..."; Concatenate the above scenario knowledge with the human-computer dialogue history and the current user input as a prompt and input it into the un-fine-tuned element extraction large model to obtain the scenario intent and task elements of the current user input.
6. The task-based multi-round dialogue method based on a large language model according to claim 1, wherein, In step S4, the rules for performing dialogue state tracking and dialogue strategy generation based on the general intent, scenario intent and task elements include: Update the current dialogue state according to the combination of the general intent, scenario intent and task elements; each general intent has a corresponding state update method and possible dialogue strategies, update the state parameters of the current task, active task, and element values in the dialogue context, and the dialogue strategy is specific to some intents and is inserted into the dialogue process as a high-priority dialogue requirement; Specifically, it is divided into the following situations: The general intention group contains an informing intention, and the dialogue state is updated according to the recognized task elements; The general intention group contains a responding intention or its sub-intentions, and the dialogue state is updated according to the elements sent by the system but not replied in the context; The general intention group contains a requesting intention. It is judged whether the recognized scenario intention is consistent with the scenario intention maintained in the current dialogue state. If they are inconsistent, the task conversion mechanism is triggered, the context state is saved, and a dialogue strategy of asking the user whether to switch tasks is made. If they are consistent, the current dialogue continues; The general intention group contains an asking intention or its sub-intentions, then the recognized task elements are used as the content of the query to the database, and a response strategy is formulated according to the specific sub-intentions; The general intention group contains a sub-intention of a modifying intention. It is verified whether this intention is consistent with the assignment method of the recognized task elements. If they are inconsistent, a dialogue strategy of indicating that the modification fails is made, otherwise the modification is made according to the assignment method; The general intention group contains a complaining intention, then a coping statement of apology is made.
7. The task-based multi-turn dialogue method based on a large language model according to claim 1, characterized in that, In step S5, when using the large model to generate a dialogue response, the prompt words are constructed in the following structural order: {The statement defining the role of the large model as the customer service role corresponding to the scenario}\n{Precautions}\n{Examples}\n{Human-machine historical dialogue}\n{User's current dialogue}\n{The statement guiding the large model to use the chain of thought}\n{Database query results (optional)}\n{Dialogue strategy}.
8. A task-based multi-turn dialogue system based on a large language model, characterized in that, Including: A human-computer interaction module, which is used to provide a chat interaction interface for users, and provide functions such as content input and chat record viewing; A scenario knowledge customization module, which is used to pre-enter the scenario knowledge of the dialogue scenarios applied by the system, including scenario intentions, task elements, and task logics; An intention recognition module, which is used to recognize the user's dialogue input in each round of dialogue, and recognize the user's general intention, scenario intention, and task elements; A dialogue state management module, which is used to maintain the global dialogue state in each dialogue and record the update status of the dialogue state in each round; A dialogue strategy generation module, which is used to formulate the current dialogue strategy based on the results of the intention recognition module and the dialogue state management module; A dialogue response generation module, which is used to generate a dialogue response based on the dialogue strategy; A data storage module, which is used to store scenario data in the database and implement relevant interfaces to query the required data during the dialogue process.
9. A terminal device, characterized in that, The terminal device includes a memory, a processor, and a program of the task-based multi-round dialogue method based on the large language model stored in the memory and executable on the processor. When the processor executes the program of the task-based multi-round dialogue method based on the large language model, the steps of the task-based multi-round dialogue method based on the large language model described in any one of claims 1-7 are implemented.
10. A computer-readable storage medium, characterized in that, A program of the task-based multi-round dialogue method based on the large language model is stored on the computer-readable storage medium. When the program of the task-based multi-round dialogue method based on the large language model is executed by the processor, the steps of the task-based multi-round dialogue method based on the large language model described in any one of claims 1-7 are implemented.