Efficient task type dialogue construction method and system based on large language model and program product

By using the large language model and Prompt module in the task-based dialogue system, the model is directly guided to complete specific tasks, solving the problems of traditional systems in process definition and node jumping, and achieving an efficient and smooth task-based dialogue experience.

CN120086332APending Publication Date: 2025-06-03HUA DATA TECH (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510148266.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

Traditional task-based dialogue systems require a lot of effort in the process definition process, and may cause node loops or multiple nodes to jump back and forth repeatedly, affecting the user experience.

Method used

The efficient task-based dialogue construction method based on the large language model is adopted, and the large model is directly guided to complete specific tasks through the Prompt module, achieving efficient construction of multiple rounds of task-based dialogues, avoiding dependence on process trees and status trackers.

Benefits of technology

It realizes the rapid construction of a task-based dialogue process without the need for model training and manual annotation, the process is smoother, entity recognition is more accurate, and user experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120086332A_ABST
    Figure CN120086332A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of large language models, in particular to an efficient task type dialogue construction method based on a large language model, and the method comprises the steps: jointly updating the input verbal skill of a user and the current entity information into the current Prompt; after the Prompt is updated, the large language model can give specific output according to the Prompt, then the output of the large language model is output to Response for post-processing, entity information is updated according to a post-processing result, and whether the conversation is continued or not is judged; when it is judged that the conversation does not need to be continued, conversation ending and conversation ending are carried out in the next round of conversation of the user, and content is output. By fusing the large language model and the task-based dialogue system, a task-based dialogue process can be quickly established, model training and manual annotation are not needed, a specific process tree does not need to be pre-defined either, the whole establishment process is high in speed and high in adjustability, and meanwhile the whole dialogue process is smoother.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of large language models, and particularly to an efficient task-based dialogue construction method, system and program product based on large language models. Background Art

[0002] With the rapid development of artificial intelligence technology, dialogue systems based on natural language processing have been widely used in various fields. An intelligent task-based dialogue system can usually be used to complete specific scenario tasks and support good natural language interaction. For example, a task-oriented dialogue framework such as RASA can build a complete task-based dialogue processing framework by constructing scenario stories, defining potential user intentions, identifying key entity information, designing conversation templates, etc., to support intelligent communication with users, interpret user intentions, collect effective information, and give appropriate responses.

[0003] In recent years, with the emergence of large language models, the capabilities of task-based dialogue systems have been further improved. With the powerful semantic understanding ability of large language models, the cumbersome data annotation and model training links in the past can be simplified, and dynamic entity recognition and intention recognition tasks can be completed more accurately. At the same time, the text generation ability of large language models can make task-based dialogues more fluent and natural. However, in the above methods, the development of a process tree and a state tracker for a task-based dialogue is essential. Especially in the process of predefined processes, it is necessary to preconceive the jump logic of a task-based dialogue and the information of process nodes in advance.

[0004] For traditional task-based dialogue systems, they rely on a large amount of manual data annotation and model training. Annotating data requires a large amount of human resources and professional knowledge. At the same time, it is necessary to train the models for intention recognition and entity recognition to better complete tasks. And when there are process updates or intention updates, the models also need to be retrained, consuming more time and energy. For task-based dialogue systems optimized based on large language models, although the steps of manual data annotation and model training can be omitted, a certain amount of effort is still required in the process definition link. Developers need to conceive and design the process logic of the dialogue at the initial stage of building a task-based dialogue, determine the jump conditions and status information of each process node, etc. At the same time, since the system needs to use a state tracker to track the dialogue state, there may be situations where nodes loop or multiple nodes repeatedly jump back and forth in a task-based dialogue, which will affect the user experience to a certain extent.

[0005] The present invention explores a lighter and more efficient method for constructing task-based dialogues. It uses promptengineering technology to directly guide large models to complete specific tasks, thereby realizing an efficient dialogue system that supports the completion of multiple rounds of task-based dialogues without relying on process trees and state trackers. Summary of the invention

[0006] In view of the deficiencies in the prior art, the present invention provides a method, system and program product for constructing an efficient task-based dialogue based on a large language model, which solves the problem that the traditional task-based dialogue system optimized based on a large language model needs to spend a lot of effort in the process definition link, and there will be node loops or multiple nodes jumping back and forth repeatedly, affecting the user experience.

[0007] To achieve the above objectives, the present invention is implemented through the following technical solutions: an efficient task-based dialogue construction method based on a large language model, comprising the following process:

[0008] S1. Construction of task-based dialogue process

[0009] The initial entity information is initialized to empty, and the current user input content is captured and passed to the Prompt module. The Prompt module will update the content according to the passed parameters; the updated Prompt module will be passed to the large language model for response and reply; the entity information returned by the large language model will update the previously initialized entity information; then, based on the indication of whether to end, it is determined whether the current conversation needs to continue. If it needs to continue, the next round of interaction will proceed normally. If it does not need to continue, the conversation will be concluded and the session will be ended in the next round of interaction;

[0010] S2. Construction of Prompt module

[0011] The construction of the prompt module includes seven parts: task definition, entity information to be collected, special requirements for responses, response format and examples, current entity information, other background information, and dialogue messages;

[0012] S3. Large language model call and response processing

[0013] The Prompt module constructed above will be directly passed to the large language model for output. For the output of the large language model, under the constraints of Prompt, the specified content will be generated. From the perspective of user experience, only the reply words need to be displayed as the content to be displayed in the streaming output. The output content is in json format, and the start and end characters of the reply streaming output are determined by post-processing to achieve the effect of displaying only the reply words.

[0014] Preferably, in the step S1, the content of the response of the large language model includes three parts: response words, entity information, and whether to end.

[0015] Preferably, in the step S2, the construction of the task definition is as follows:

[0016] Describe the definition of the entire task, inform the large language model of the current task scenario, and let the large model continuously design the words for the next round of asking the user based on the known information and the information still missing, and at the same time, it can also maintain the entity information required by the system, so that the large language model can have a more complete prior knowledge when completing the task-based dialogue, and can make the large language model generate words more in line with the task description defined by the user.

[0017] Preferably, in the step S2, the construction of the entity information to be collected is as follows:

[0018] Describe the entity information required in the process of task-based dialogue. The large language model has powerful semantic understanding and entity recognition capabilities. However, to better exert its capabilities, it is necessary to describe the entities to be recognized in the Prompt. This description is not only the entity name itself, but also corresponding examples.

[0019] Preferably, in the step S2, the construction of the special requirements for the response is as follows:

[0020] By setting special requirements, the requirements or styles that the large model needs to follow when generating responses are restricted. The special requirements include the response strategy when the user asks a question that is not the current topic, which can be to directly refuse to answer or to allow chatting; at the same time, it can be required that the large language model conducts the dialogue in a specific tone during the interaction.

[0021] Preferably, in the step S2, the construction of the return format and examples is as follows:

[0022] The return format and examples restrict the content constraints and format requirements of the large model's return. By restricting the content and format to be returned in the Prompt input to the large language model, leveraging the powerful semantic understanding and generation capabilities of the large language model, and using commercial models, it enables the model to output according to the system-restricted mode; this part plays a role in restricting the large model's output on the one hand, and at the same time, specific return fields are set to meet the needs of multi-round dialogue.

[0023] Preferably, the return fields include Response, ner_info, and is_complete; among them, Response is used to generate the specific words for replying to the user, and the words in this field are the content of the words shown to the user; ner_info is used to combine the known entity information and the user's current words to return the updated and maintained entity information. The function of this field is to continuously maintain the entity information collected by the current system during the multi-round conversation, so as to facilitate the large language model to have more complete known information for reference when asking questions; is_complete is used to assist the system in determining whether the current conversation process can trigger the end. The purpose of this field is to judge whether the current multi-round conversation is completed. If it is judged to be completed, the end of the entire conversation will be actively triggered.

[0024] Preferably, in the step S2, the construction of the current entity information is specifically as follows:

[0025] This part is responsible for recording and transmitting the current entity information. The purpose is to completely transmit the entity information in the current state to the large language model. Just directly replace the json format of the entity information here.

[0026] Preferably, in the step S2, the construction of other background information is specifically as follows:

[0027] This part is used to transmit background information for the large language model to have reference information when generating words or extracting entities.

[0028] Preferably, in the step S2, the construction of the dialogue messages is specifically as follows:

[0029] This part is used to transmit the historical conversation content of the past N rounds and the user's current words. The number of rounds of the historical conversation can be controlled according to the context length accepted by the large model.

[0030] Preferably, in the step S3, the specific process of only showing the reply words effect is as follows:

[0031] Monitor the streaming output of the large model in the background in real time. When the Response field in the json is recognized, it means that the output of the reply words is about to start. From this moment on, the newly generated input content of the large language model will be shown to the user side. When the background real-time monitoring monitors the ner_info or is_complete field again, it means that the output of the current Response field has been completed, and the large language model starts to output content unrelated to the reply words. The new content starting from this moment will no longer be shown to the user; the entity information and whether to end content output by the model later will be formatted after being completely obtained by the background for maintaining the entity information, and at the same time, it will be judged whether to trigger the end of the conversation in the next round of the process according to whether to end.

[0032] Preferably, the step S3 also includes the following process:

[0033] In order to improve the response speed of the entire process, the large language model is constrained to give priority to the output of the reply dialogue, and then continue to output other content; specifically: in the requirements and examples of returning json, the Response field is placed first, so that the background can monitor the dialogue content that needs to be returned to the user as early as possible, that is, the time required for the first character of the reply to appear can be shortened. After the reply dialogue is completed, the user end can see the complete dialogue, and the system background will continue to output and monitor the remaining content, so that the entire process will be smoother and the user will have a shorter waiting time.

[0034] An efficient task-oriented dialogue construction system based on a large language model, including:

[0035] A user input module, used for the user to input the conversation content;

[0036] An entity information module, which is updated in each round of the conversation to assist the system in generating more relevant and accurate responses;

[0037] A large language model, which has semantic understanding and generation capabilities and is used to generate task-based dialogue content;

[0038] Prompt module, which is used by the large language model to generate appropriate text results based on the input prompt when it is called;

[0039] Response, which is used to post-process the content output by the large language model and update the entity content.

[0040] The present invention provides a method, system and program product for constructing an efficient task-based dialogue based on a large language model. It has the following beneficial effects:

[0041] 1. The present invention integrates a large language model with a task-based dialogue system. This method can quickly build a task-based dialogue process. It does not require model training and manual annotation, nor does it require pre-definition of a specific process tree. The entire construction process is fast and highly adjustable, and only some modules in Prompt need to be optimized and modified. At the same time, the entire dialogue process will be smoother, and there will be no abrupt node jumps and changes in the content of the speech. At the same time, in the design of Prompt, entity information is maintained. This method can better perform the task of entity recognition, especially when some users modify or update the previously mentioned entities, so that the output reply speech is more accurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0042] Figure 1Schematic diagram of the task-based dialogue process of the present invention. Detailed implementation manners

[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0044] Embodiment:

[0045] As Figure 1 shown, the embodiment of the present invention provides an efficient task-based dialogue construction method based on a large language model, including the following processes:

[0046] S1. Construction of the task-based dialogue process

[0047] Initialize the initial entity information to be empty. For example, in a meeting room reservation task, it is initialized as {"meeting_room":"","meeting_people":[],"meeting_time":""}, which respectively represent the meeting room name, participants, and meeting time. At the same time, capture the current user input content and pass it to the Prompt module. The Prompt module will update the content according to the passed parameters; the updated Prompt module will be passed to the large language model for response. The content returned by the large language model includes three parts: response words, entity information, and whether to end. The entity information returned by the large language model will update the previously initialized entity information, that is, when multiple rounds of dialogue start, each round of dialogue will trigger the update of entity information. Although the update is performed in each round, whether the stored entity value will change depends on the content of the user's words in the current round and the entity recognition situation of the large language model; then, judge whether the current dialogue needs to continue according to the indication of whether to end. If it needs to continue, the next round of interaction will proceed normally. If it does not need to continue, the dialogue will be ended and the session will be terminated in the next round of dialogue interaction. Repeat this process to implement a task-based multi-round dialogue process. During the process of the process or after triggering the end, the system can obtain the currently recognized entity information at any time.

[0048] S2. Construction of the Prompt module

[0049] Prompt construction is the core module in this method. Through reasonable Prompt construction, the large model can better complete specific tasks. When the generative large language model is called, it can generate appropriate text results according to the input Prompt. The large language model has powerful semantic understanding and generation capabilities, enabling the system to complete some specific complex tasks only through the construction of Prompt. The construction of the Prompt module includes seven parts: task definition, entity information to be collected, special requirements for responses, return format and examples, current entity information, other background information, and conversation messages.

[0050] The construction of the above task definition is as follows:

[0051] The task definition describes the definition of the entire task, informing the large language model of the current task scenario, so that the large model can have a more complete prior knowledge when completing task-based conversations. It enables the large model to generate more appropriate conversation words when generating conversation words. Refer to the following example: "You are a meeting room reservation assistant. Your task is to have multiple rounds of conversations with users, collect the information required for meeting room reservation, record and maintain key slot information, and design conversation words based on the information already collected and the information still missing." Such a task definition can enable the large model to continuously design conversation words for the next round of questions to the user based on the known information and the information still missing, and at the same time maintain the entity information we need. The difference from the regular entity recognition task here is that by maintaining the entity in this way, it can better handle the situation where the user modifies the previously mentioned content. The main purpose of "You are a meeting room reservation assistant" in this example is to highlight the background of this task-based conversation. And the subsequent "Your task is to have multiple rounds of conversations with users, collect the information required for meeting room reservation, record and maintain key slot information, and design conversation words based on the information already collected and the information still missing." is a set of relatively effective Prompt conversation words for information collection task-based conversations. For different tasks, only a small amount of modification is required on this basis.

[0052] The construction of the above entity information to be collected is as follows:

[0053] This section mainly defines and describes the entity information that needs to be identified in the entire task, as well as the entity type or example reference. The main purpose of this section is to describe the entity information that needs to be collected during the task-based dialogue process. The large language model has powerful semantic understanding and entity recognition capabilities. However, in order to better exert its capabilities, the system needs to describe the entities that need to be identified in the prompt. The description is not only the entity name itself, but also requires corresponding examples. For example: "You need to collect three types of slot information, namely the meeting room name (meeting_room), the participants (meeting_people), and the meeting time (meeting_time). The specific information maintenance format is as follows: ner_info:{"meeting_room":"xxx","meeting_people":["AA","BB","CCC"],"meeting_time":"xxxx年xx月xx日xx时xx分"}". This example shows the three entity information names and corresponding examples that the system needs to capture in this scenario, such as the meeting time meeting_time, whose value is a string containing the year, month, day and time. When the description of the entity information is richer and comes with some examples, the large model tends to perform better in capturing the entity.

[0054] The special requirements of the above reply are constructed as follows:

[0055] In the process of multiple rounds of dialogue, it is hoped that the robot can interact with the user more naturally, rather than asking questions in a rigid manner. Therefore, some special requirements can be set to constrain the requirements or styles that the large model needs to follow when generating responses. The special requirements here generally include the answer strategy for users who ask questions that are not on the current topic, which can be a direct refusal to answer or allowing small talk. For example, this part can be described by the following prompt: "When the user's question is not related to the reservation of the conference room, please answer his question as briefly as possible and pull the conversation back to the process of reserving the conference room. When the entity information is fully collected, ask the user if he has any additional information, mainly for the participants to supplement and confirm, and do not confirm the confirmed information again." Through the above constraints, the response of the large language model will be more natural. At the same time, it will also answer some questions of the user while completing specific tasks, and the entire process will revolve around the specific task, and there will be no large amount of small talk. In addition, the large language model can also be required to communicate in a certain tone during the interaction, such as the tone of the ancients, the playful tone, or the serious tone.

[0056] The above return format and examples are constructed as follows:

[0057] This part mainly restricts the content and format requirements of the responses from large language models. By constraining the content and format to be returned in the Prompt input to the large language model, leveraging the powerful semantic understanding and generation capabilities of the large language model, and using powerful commercial models such as GPT4, Tongyi Qianwen, Doubao, Wenxin Yiyan, etc., the model output can be in accordance with the system-constrained pattern. For example:

[0058] "You must generate json according to my example format: {"Response": "Is there any other information that needs to be supplemented? Are there any other people who need to attend the meeting?", "ner_info": {"meeting_room": "Large meeting room", "meeting_people": ["General Manager Huang", "Wang Wei"], "meeting_time": "January 30, 2024, 14:00"}, "is_complete": false}. The Response field represents the specific words to reply / ask the user. The ner_info field represents the content of the current status of the entity information you maintain. The is_complete field represents whether to judge that the current task-based dialogue process has been completed."

[0059] The above example shows the constraints on the generated content in the Prompt passed by the system, fully expressing the content and format to be returned, giving examples, and explaining each field in detail to facilitate the large language model to better understand the meaning of each field.

[0060] This part plays a role in restricting the output of the large model and also sets some specific return fields to meet the needs of multi-round conversations. The specific return fields here refer to some other information required in addition to interacting with the user to meet some needs in multi-round task-based conversations. The role of Response is to generate the specific words to reply to the user, and the words in this field are the content shown to the user. The role of ner_info is to combine the known entity information and the user's current words to return the updated and maintained entity information. The purpose of this field is to continuously maintain the entity information we have collected during the multi-round conversation, so as to facilitate the large language model to have more complete known information for reference when asking questions. The role of is_complete is to assist the system in determining whether the current dialogue process can trigger the end. The purpose of this field is to judge whether the current multi-round conversation is completed. If it is judged to be completed, it will actively trigger the end of the entire conversation. In some other task-based dialogue scenarios, some other fields can also be added to meet specific needs.

[0061] The construction of the above current entity information is as follows:

[0062] This part is responsible for recording and transmitting the current entity information. The purpose is to completely transmit the entity information in the current state to the large language model. You can directly replace the json format of the entity information here. This way can enable the large language model to know which information has been collected and which information is still to be collected during each interaction. For example, at the beginning of the conversation, the values of the recorded entity information are all null; as the conversation progresses, some entity information will be filled, and then the large language model will ask questions and collect information about the remaining entities. This can ensure that the current task completion status is recorded and tracked throughout the conversation.

[0063] The construction of the above other background information is as follows:

[0064] This part is used to transmit background information for the large language model to have reference information when generating conversation words or extracting entities. For example, the date today, the day of the week today, the user's name, etc. can be transmitted here. After transmitting this information, when the user says "I want to book a meeting room for tomorrow", the specific date can be more accurately extracted; or when the user says "I will attend the meeting with someone", the user's name can also be extracted. The large language model is an offline model and cannot know the current date, user personal information, etc. For example, today is September 20th. When the user says "I want to book a meeting room for tomorrow", the large model can locate the date of the meeting time as September 21st. If the current date is not informed in the Prompt, the large model cannot determine the specific date referred to by "tomorrow". Therefore, supplementing some additional background information can effectively improve the effect of the entire task-based conversation.

[0065] The construction of the above conversation messages is as follows:

[0066] This part is used to transmit the historical conversation content of the past N rounds and the user's current conversation words. The number of historical conversation rounds can be controlled according to the context length accepted by the large model. Whether calling the commercial large language model api interface or the locally deployed large language model, the large language model usually has a maximum supported input length, such as eight thousand, sixteen thousand, or thirty-two thousand, etc. The specific supported length depends on the selected model and the video memory capacity of the machine on which the model is deployed, etc. For the number of historical conversation rounds that should be included in the Prompt, it depends on the remaining length after the Prompt is constructed. Subtract the length of the content already included in the Prompt from the maximum length supported by the model, and the result is the length that the context historical conversation can support. This length limit can be used to dynamically adjust the historical conversation passed to the large language model in each round. When the model can support a very long input length, the number of historical conversation rounds can be default limited to about 5 rounds, and specific adjustments and experiments can be carried out according to the actual situation to obtain the optimal effect.

[0067] Through the combination of the above seven parts, a complete large model Prompt is constructed. Under the Prompt framework, a task-based dialogue process can be constructed very efficiently. Referring to the examples listed above, each Prompt module can be spliced ​​together to obtain a complete large language model Prompt:

[0068] (1) You are a conference room fixed assistant. Your task is to have multiple rounds of conversations with users, collect the information required for conference room reservations, and record and maintain key slot information. Please design your words based on the information you have collected and the missing information.

[0069] (2) You need to collect three types of slot information, namely the meeting room name (meeting_room), meeting participants (meeting_people), and meeting time (meeting_time). The specific information maintenance format is as follows: ner_info:{"meeting_room":"xxx","meeting_people":["AA","BB","CCC"],"meeting_time":"xxxx年xx月xx日xx时xx分"}.

[0070] (3) When the user's question is not related to the conference room reservation, please answer his question as briefly as possible and bring the conversation back to the conference room reservation process. When the entity information is fully collected, ask the user if he has any additional information, mainly for the participants to supplement and confirm. Do not reconfirm the information that has been confirmed.

[0071] (4) You must generate json according to my example format:

[0072] {"Response":"Is there any other information you need to add? Is there anyone else who needs to attend the meeting?","ner_info":{"meeting_room":"Large conference room","meeting_people":["Mr. Huang","Wang Wei"],"meeting_time":"January 30, 2024 14:00"},"is_complete":False}.

[0073] The Response field indicates the current words that need to be responded to / asked by the user.

[0074] The ner_info field indicates the current state of the entity information you maintain.

[0075] The is_complete field indicates whether the current task-based dialogue process has been completed.

[0076] (5) The current entity information ner_info is:

[0077] {"meeting_room":"","meeting_people":[],"meeting_time":""}.

[0078] (6) Some background information for reference: Today’s date is October 24, 2024; today is Thursday; the current user’s name is Zhang San.

[0079] (7) Historical conversation information: [user:xxxxx,assistant:xxxxx,......].

[0080] (8) Please return the JSON content as required.

[0081] S3. Large language model call and response processing

[0082] The Prompt module constructed above will be directly passed to the large language model for output. Since the ability of the large model is needed to perform more complex tasks, this step uses a large language model with strong overall capabilities, such as some commercial model APIs, such as OpenAI, Alibaba, Byte, etc.; or open source large parameter large models, such as Tongyi Qianwen 72B, Llama70B, etc. The constructed Prompt can be completely given to the large language model as the model input, and then the output of the large language model will be received in a streaming manner.

[0083] For the output of the large language model, under the constraint of Prompt, the specified content will be generated. From the perspective of user experience, only the reply words need to be displayed as the content required for streaming output. The output content is in json format. The start and end characters of the reply streaming output are determined by post-processing to achieve the effect of only displaying the reply words. The streaming output of the large model is monitored in real time in the background. When the Response field in json is recognized, it means that the reply words are about to be output. From this moment on, the newly generated input content of the large model will be displayed to the user end. When the real-time monitoring of the background monitors the ner_info or is_complete field again, it means that the output of the current Response field has been completed, and the large language model starts to output content unrelated to the reply words. The new content starting from this moment will no longer be displayed to the user. The entity information and whether the content is terminated that are subsequently output by the model will be formatted after the background is fully obtained. On the one hand, it is used to maintain the entity information, and on the other hand, it will determine whether the next round of dialogue in the process should trigger the end of the session based on whether it is terminated.

[0084] In order to improve the response speed of the entire process, the large language model is constrained to give priority to the output of the reply dialogue, and then continue to output other content; specifically: in the requirements and examples of returning json, the Response field is placed first, so that the background can monitor the dialogue content that needs to be returned to the user as early as possible, that is, the time required for the first character of the reply to appear can be shortened. After the reply dialogue is completed, the user end can see the complete dialogue, and the system background will continue to output and monitor the remaining content. In this way, the entire process will be smoother and the user will have a shorter waiting time.

[0085] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. An efficient task-based dialogue construction method based on a large language model, characterized by: The following processes are included: S1. Construction of task-based dialogue process The initial entity information is initialized to empty, and the current user input content is captured and passed to the Prompt module. The Prompt module will update the content according to the passed parameters; the updated Prompt module will be passed to the large language model for response and reply; the entity information returned by the large language model will update the previously initialized entity information; then, based on the indication of whether to end, it is determined whether the current conversation needs to continue. If it needs to continue, the next round of interaction will proceed normally. If it does not need to continue, the conversation will be concluded and the session will be ended in the next round of interaction; S2. Construction of Prompt module The construction of the prompt module includes seven parts: task definition, entity information to be collected, special requirements for responses, response format and examples, current entity information, other background information, and dialogue messages; S3. Large language model call and response processing The Prompt module constructed above will be directly passed to the large language model for output. For the output of the large language model, under the constraints of Prompt, the specified content will be generated. From the perspective of user experience, only the reply words need to be displayed as the content to be displayed in the streaming output. The output content is in json format, and the start and end characters of the reply streaming output are determined by post-processing to achieve the effect of displaying only the reply words.

2. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In the step S1, the content of the response from the large language model includes three parts: the response words, entity information and whether it is ended.

3. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In step S2, the construction of the task definition is as follows: Describe the definition of the entire task, inform the large language model of the current task scenario, and let the large model continuously design the dialogue to continue asking users in the next round based on known information and currently missing information. At the same time, it can also maintain the entity information required by the system so that the large language model can have a more complete prior knowledge when completing task-based dialogues, allowing the large language model to generate dialogue that is more in line with the task description defined by the user.

4. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In the step S2, the entity information to be collected is constructed as follows: Describe the entity information that needs to be collected during the task-based dialogue process. The large language model has powerful semantic understanding and entity recognition capabilities. However, in order to better utilize its capabilities, it is necessary to describe the entities that need to be recognized in the prompt. The description is not only the entity name itself, but also requires corresponding examples.

5. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In the S2 step, the special requirements for the reply are constructed as follows: By setting special requirements, we can constrain the requirements or style that the large model needs to follow when generating responses. Special requirements include answering strategies when users ask questions that are not on the current topic, which can be either directly refusing to answer or allowing small talk. At the same time, we can require the large language model to communicate in a specific tone during interaction.

6. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In step S2, the return format and example construction are as follows: The return format and examples constrain the content constraints and format requirements of the big model's return. By constraining the content and format to be returned in the prompt input to the big language model, with the help of the big language model's powerful semantic understanding and generation capabilities, and using commercial models, it can output the model in a mode constrained by the system. This part not only constrains the output of the big model, but also sets specific return fields to meet the needs of multiple rounds of dialogue.

7. The method for constructing an efficient task-based dialogue based on a large language model according to claim 6, characterized in that: The return fields include Response, ner_info and is_complete; Response is used to generate specific words to reply to the user, and the words in this field are the words displayed to the user; ner_info is used to combine the known entity information and the user's current words to return the updated and maintained entity information. The function of this field is to continuously maintain the entity information collected by the current system during multiple rounds of dialogue, so that the large language model has more complete known information reference when asking questions; is_complete is used to assist the system in determining whether the current dialogue process can be triggered to end. The purpose of this field is to determine whether the current multiple rounds of dialogue are completed. If it is determined to be completed, it will actively trigger the end of the entire dialogue.

8. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In the step S2, the current entity information is constructed as follows: This part is responsible for recording and transmitting the current entity information. The purpose is to pass the entity information of the current state completely to the large language model. Simply replace the JSON format of the entity information here.

9. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In the step S2, the construction of other background information is as follows: This part is used to convey background information so that the large language model can have reference information when generating words or extracting entities.

10. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In step S2, the construction of the dialogue messages is as follows: This part is used to transmit the historical conversation content of the past N rounds and the user's current speech. The number of historical conversation rounds can be controlled according to the context length accepted by the large model.

11. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: In the S3 step, the specific process of only displaying the effect of the reply speech is as follows: The streaming output of the large model is monitored in real time in the background. When the Response field in json is recognized, it means that the reply dialogue is about to be output. From this moment on, the newly generated input content of the large language model will be displayed to the user. When the real-time monitoring of the background monitors the ner_info or is_complete field again, it means that the output of the current Response field has been completed, and the large language model begins to output content that is not related to the reply dialogue. The new content starting from this moment will no longer be displayed to the user; the entity information and whether the content is ended that are subsequently output by the model will be formatted after the background is fully obtained for maintaining the entity information. At the same time, it will determine whether the next round of dialogue in the process will trigger the end of the session based on whether it is ended.

12. The method for constructing an efficient task-based dialogue based on a large language model according to claim 1, characterized in that: The S3 step also includes the following process: In order to improve the response speed of the entire process, the large language model is constrained to give priority to the output of the reply dialogue, and then continue to output other content; specifically: in the requirements and examples of returning json, the Response field is placed first, so that the background can monitor the dialogue content that needs to be returned to the user as early as possible, that is, the time required for the first character of the reply to appear can be shortened. After the reply dialogue is completed, the user end can see the complete dialogue, and the system background will continue to output and monitor the remaining content, so that the entire process will be smoother and the user will have a shorter waiting time.

13. An efficient task-based dialogue construction system based on a large language model, characterized by: include: A user input module, used for the user to input the conversation content; An entity information module, which is updated in each round of the conversation to assist the system in generating more relevant and accurate responses; A large language model, which has semantic understanding and generation capabilities and is used to generate task-based dialogue content; Prompt module, which is used by the large language model to generate appropriate text results based on the input prompt when it is called; Response, which is used to post-process the content output by the large language model and update the entity content.

14. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein: When the processor executes the program, the method according to any one of claims 1 to 12 is implemented.

15. A computer program product comprising computer instructions, characterized in that: When the computer instructions are executed by a processor, the method according to any one of claims 1 to 12 is implemented.