Task-based dialogue method and related equipment

By using a large language model to extract the core information of dialogue statements and perform multi-dimensional retrieval in a task-based dialogue system, and determining expected actions in combination with learning examples, the problems of error accumulation, complex structure and poor scalability in existing systems are solved, and higher semantic understanding and overall experience are achieved.

CN120045659APending Publication Date: 2025-05-27VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510023763.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

The existing task-based dialogue systems have problems such as accumulation of errors, complex structure, high maintenance costs, poor scalability, poor interpretability, and poor maintenance.

Method used

By obtaining the user's current dialogue statement, inputting it into the large language model to extract core information, multi-dimensional searching based on the core information to obtain learning examples, and inputting the dialogue statements and learning examples into the large language model to determine the expected actions, realizing the user's intention correlation.

Benefits of technology

It improves the semantic understanding effect, simplifies the structure of the dialogue system, reduces maintenance costs, improves scalability and interpretability, and solves the problem of poor overall experience of traditional methods.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045659A_ABST
    Figure CN120045659A_ABST
Patent Text Reader

Abstract

The invention discloses a task-based dialogue method and related equipment, relates to the field of language models, and mainly aims to solve the problem that an existing task-based dialogue system is still not simple and accurate enough. The method comprises the steps of obtaining a current dialogue statement of a user; inputting the current dialogue statement into a large language model to extract core information of the current dialogue statement; performing multi-dimensional retrieval based on the core information to obtain a learning example; the current dialog statement and the learning example are input to a large language model to determine an expected action, where the expected action is associated with a user intent. The method is used for the task type dialogue process.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of language models, and particularly to a task-based dialogue method and related devices. Background Art

[0002] Task-based dialogue systems assist users in completing various tasks such as ordering food and tickets, media playback, encyclopedia Q&A, search and navigation, etc. through voice or text interaction. Especially in the field of intelligent cockpits, voice operation can avoid the distraction of manual operation by drivers, which has important value for ensuring driving safety and improving user experience. Currently, the common implementation methods of task-based dialogue are divided into pipeline and end-to-end methods.

[0003] However, the pipeline system has problems such as error accumulation, complex structure, high maintenance cost, and poor scalability, while the end-to-end system implemented by traditional NLP technology also has disadvantages such as poor interpretability, poor maintainability, and poor scalability. Summary of the Invention

[0004] In view of the above problems, the present invention provides a task-based dialogue method and related devices, mainly aiming to solve the problem that the current task-based dialogue system is still not concise and accurate enough.

[0005] To solve the above at least one technical problem, in a first aspect, the present invention provides a task-based dialogue method, which includes:

[0006] Obtain the user's current dialogue statement;

[0007] Input the current dialogue statement into a large language model to extract the core information of the current dialogue statement;

[0008] Perform multi-dimensional retrieval based on the core information to obtain learning examples;

[0009] Input the current dialogue statement and the learning examples into a large language model to determine the expected action, where the expected action is associated with the user's intention.

[0010] Optionally, the inputting the current dialogue statement into a large language model to extract the core information of the current dialogue statement includes:

[0011] When the current dialogue statement depends on historical dialogue statements, obtain the relevance between the current dialogue statement and the historical dialogue statements;

[0012] Extract the core information of the current dialogue statement based on the relevance;

[0013] In the case that the relevance does not meet the preset threshold, output a specific mark, where the specific mark is used to represent that the current dialogue statement cannot extract the core information.

[0014] Optionally, the inputting the current dialogue statement into a large language model to extract the core information of the current dialogue statement includes:

[0015] In the case that the current dialogue statement does not depend on historical dialogue statements, obtain the core information of the current dialogue statement.

[0016] Optionally, the inputting the current dialogue statement into a large language model to extract the core information of the current dialogue statement includes:

[0017] In the case that there are multiple core information of the current dialogue statement, split the core information of the current dialogue statement, where different core information corresponds to different user intents.

[0018] Optionally, the multi-dimensional retrieval based on the core information to obtain learning examples includes:

[0019] Retrieve learning examples of the core information based on the atomic instruction vector library;

[0020] In the case that the core information cannot be extracted, retrieve from the dialogue vector library to obtain learning examples;

[0021] Retrieve learning examples of the historical dialogue statement based on the database, where the database contains multiple rounds of user dialogue statements.

[0022] Optionally, the expected action corresponds to an action agent.

[0023] Optionally, the above method further includes:

[0024] Control the action agent to input the historical dialogue statement, the core information, and the learning examples into a large language model to output a corrected expected action.

[0025] In a second aspect, an embodiment of the present invention further provides a task-based dialogue device, including:

[0026] An acquisition unit, configured to acquire the current dialogue statement of the user;

[0027] An extraction unit, configured to input the current dialogue statement into a large language model to extract the core information of the current dialogue statement;

[0028] A retrieval unit, configured to perform multi-dimensional retrieval based on the core information to obtain learning examples;

[0029] A determination unit for inputting the current conversation statement and the learning example into a large language model to determine an expected action, where the expected action is associated with the user intention.

[0030] To achieve the above object, according to the third aspect of the present invention, there is provided a computer-readable storage medium, the computer-readable storage medium including a stored program, wherein when the above program is executed by a processor, the steps of the above task-based conversation method are implemented.

[0031] To achieve the above object, according to the fourth aspect of the present invention, there is provided an electronic device, including at least one processor and at least one memory connected to the processor; wherein the above processor is used to call the program instructions in the above memory to execute the steps of the above task-based conversation method.

[0032] By means of the above technical solutions, for the problem that the current task-based conversation system is still not concise and accurate enough, the task-based conversation method and related devices provided by the present invention obtain the user's current conversation statement; input the current conversation statement into a large language model to extract the core information of the current conversation statement; perform multi-dimensional retrieval based on the core information to obtain learning examples; input the current conversation statement and the learning example into the large language model to determine an expected action, where the expected action is associated with the user intention. In the above solution, a large language model is used to replace the traditional NLP method as the base model of NLU, and with the help of fine-tuning or retrieval enhancement technology, the intention slot is predicted in the form of natural language generation. The above large language model improves the effect of semantic understanding. This application combines large language model technology and retrieval enhancement technology to build an intelligent agent to implement complex downstream skills, and solves the problems of insufficient interpretability, difficult expansion, and poor overall experience existing in the traditional pipeline method and end-to-end method.

[0033] Correspondingly, the task-based conversation device, device, and computer-readable storage medium provided by the embodiments of the present invention also have the above technical effects.

[0034] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other objects, features, and advantages of the present invention more obvious and understandable, the specific embodiments of the present invention are specifically described below. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present invention. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0036] Figure 1 A schematic diagram showing a flow chart of a task-based dialogue method provided by an embodiment of the present invention;

[0037] Figure 2 A schematic block diagram showing the composition of a task-based dialogue device provided by an embodiment of the present invention is shown;

[0038] Figure 3 A schematic block diagram of the composition of a task-based conversational electronic device provided by an embodiment of the present invention is shown. DETAILED DESCRIPTION

[0039] The exemplary embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present invention are shown in the accompanying drawings, it should be understood that the present invention can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided in order to enable a more thorough understanding of the present invention and to enable the scope of the present invention to be fully communicated to those skilled in the art.

[0040] This application takes into account that the pipeline method is generally divided into the following steps: NLU natural language understanding, understanding user intent based on user input and extracting entity slot information; DST dialogue state tracking, determining the next multi-round dialogue state based on historical and current slot information, dialogue state definition and current state; DP dialogue strategy, determining the current dialogue strategy and action based on the dialogue state, including inquiry, confirmation, etc. DP and DST are collectively referred to as DM dialogue management; NLG natural language generation, generating feedback words for users to meet user needs or guide users to continue the dialogue. An end-to-end dialogue system generally merges one or more of NLU, DST, and DP, defines the action space of the dialogue system as a series of functions plus parameters (or intent plus merged multi-round slots), and directly predicts the next system action by modeling the dialogue context, environmental information, scene state, and dialogue state. The above-mentioned pipeline system has the following problems: error accumulation. Since there are multiple processing links in the system and the downstream is completely dependent on the upstream results, an error in any link will cause the overall result to be unavailable; the structure is complex and the maintenance cost is high. Each processing link has an independent technical system and requires independent personnel to maintain. In particular, the pipeline system often divides functions into vertical fields. Each vertical field maintains its own NLU and DM capabilities, etc., and has independent technical implementation solutions, which incurs additional maintenance costs; poor scalability. For complex functions such as multi-round, multi-intent, and cross-vertical domains, any link in the pipeline may become a bottleneck, and often needs to be implemented through external special algorithm components.

[0041] End-to-end systems implemented using traditional NLP techniques have the following problems: poor interpretability. Through the encoder, vector representation and fusion of multiple inputs are performed, and the prediction results are output through classification or sequence generation. The inference process of the model from input to output is a black box, which is not conducive to the understanding and upgrading of the system; poor maintainability. A set of end-to-end models needs to support the prediction of the entire system action space (the specific form is function plus parameters, or intent plus slots). They affect each other during the iteration process and there are boundary problems; poor scalability. New functions require retraining the model. Some complex functions may depend on additional inputs or produce additional outputs, requiring additional development and debugging costs, and may affect existing functions.

[0042] Integrating large language model technology into task-based dialogue systems is a current trend, and generally there are the following methods: 1. Using a large language model to replace traditional NLP methods as the base model of NLU, and leveraging fine-tuning or retrieval enhancement techniques to predict intent slots in a natural language generation manner. The large language model improves the effect of semantic understanding, but still has the problems of the above pipeline system. 2. Using a large language model to build an agent to achieve complex downstream skills.

[0043] To solve the problem that the current task-based dialogue system is still not concise and accurate enough, an embodiment of the present invention provides a task-based dialogue method, as Figure 1 shown, the method includes:

[0044] Determine the intent routing. The above intent routing maps the user input statement, context, and current state to the action space of the dialogue system. Specifically:

[0045] S101. Obtain the user's current dialogue statement;

[0046] S102. Input the current dialogue statement into a large language model to extract the core information of the current dialogue statement;

[0047] The steps of S102 further include S1021, S1022, and S1023:

[0048] S1021. When the current dialogue statement depends on historical dialogue statements, obtain the relevance between the current dialogue statement and the historical dialogue statements; based on the relevance, extract the core information of the current dialogue statement; when the relevance does not meet the preset threshold, output a specific marker, where the specific marker is used to indicate that the core information cannot be extracted from the current dialogue statement.

[0049] S1022. When the current dialogue statement does not depend on historical dialogue statements, obtain the core information of the current dialogue statement.

[0050] S1023. In the case where there are multiple core messages in the current dialogue statement, split the core messages of the current dialogue statement, where different core messages correspond to different user intents.

[0051] Specifically, accept the user's nth round of dialogue statement, that is, the above-mentioned current dialogue statement, and combine the historical dialogue history of the previous k rounds, that is, the user's dialogue statements and the dialogue system's response statements from the (n - 1)th round to the (n - 1 - k)th round, and input them into the large language model to understand and summarize the user's semantics. According to different inputs, the output is one of the following situations:

[0052] 1. An atomic statement independent of context rewritten from a dialogue statement that depends on the above context, which summarizes the necessary context information;

[0053] 2. A dialogue statement depends on the above context, and the context information complexity is too high to be summarized, and a specific marker "= Unable to summarize" is output;

[0054] 3. A list of atomic statements independent of context for a dialogue statement containing multiple intents, which integrates the necessary context information and reasonably splits the user's multiple intents;

[0055] 4. A dialogue statement does not depend on the above context and has a single intent, and is output as it is.

[0056] The output content is one of the above four, and can be uniformly identified in the following form: List[Tuple(atomic instruction, marker)], where the marker value is 'NULL' or 'Unable to summarize'.

[0057] S103. Perform multi-dimensional retrieval based on the core information to obtain learning examples;

[0058] The steps of the above S103 also include S1031:

[0059] S1031. Retrieve learning examples of the core information from the atomic instruction vector library; in the case where the core information cannot be extracted, retrieve from the dialogue vector library to obtain learning examples; retrieve learning examples of the historical dialogue statements from the database, where the database contains multiple rounds of user dialogue statements.

[0060] Furthermore, for each 'Tuple(atomic instruction, marker)' output above, perform multi-dimensional retrieval to obtain 'dialogue -> action (Action) -> action brief' learning examples, including the following retrieval methods:

[0061] 1. If the marker is NULL, retrieve from the'vector library - atomic instruction' through the atomic instruction;

[0062] 2. If the marker is 'Unable to summarize', retrieve from the'vector library - dialogue';

[0063] 3. For the n-1 rounds of Actions, retrieve possible downstream Actions and corresponding learning examples from the defined multi-round dialogue state database or graph structure.

[0064] S104. Input the current dialogue statement and the learning example into a large language model to determine the expected action, where the expected action is associated with the user intention.

[0065] The steps of the above S104 further include S1041:

[0066] S1041. Control the action agent to input the historical dialogue statement, the core information, and the learning example into the large language model to output a corrected expected action.

[0067] Specifically, input the dialogue history, atomic instruction, and learning example into the large language model to predict the Action for this round; for each 'Tuple (atomic instruction, tag)', according to the Action for this round, submit the request to the appropriate 'action agent' for processing.

[0068] In one embodiment, the expected action corresponds to an action agent.

[0069] 'Action agent' is divided according to the granularity of business implementation, such as 'navigation agent', 'control agent','media agent', etc.; 'Action' is the functional implementation divided by semantic granularity, such as 'navigation search', 'air conditioner adjustment','media playback', etc. One 'Action' corresponds to only one 'action agent', and one 'action agent' can undertake multiple 'Actions'.

[0070] The agent needs to correct possible Action errors within the Action selection range of this agent according to the input Tuple (atomic instruction, tag, Action), and output a reasonable Action + parameter. Agents that do not rely on specific hardware will also give results.

[0071] Further, the following shows the processing process of a typical action agent:

[0072] 1. (Optional) Secondary retrieval, retrieve relevant learning examples (dialogue / atomic instruction -> Action -> parameter description) according to the atomic instruction, tag, Action, and previous round Action.

[0073] 2. Input the dialogue history, atomic instruction, learning example, and intention routing prediction Action into the large language model, and the large language model outputs the final Action + parameter according to the prompt information, and corrects possible Action errors during the process.

[0074] A complex action agent can use additional tools or functions, and its processing process is:

[0075] (1) Planning and implementation process;

[0076] (2) Calling a specific tool to obtain execution results;

[0077] (3) Repeat 1 and 2 or give the final result according to the tool execution results.

[0078] 3. Action execution and feedback script generation

[0079] The above action execution process includes the following steps:

[0080] (1) Determine the conflict between multiple 'Action+parameters' corresponding to user statements

[0081] (2) Determine the scenario satisfaction for each Action

[0082] (3) If there is a conflict of intent or the scenario is not satisfactory, provide appropriate feedback

[0083] (4) If the conditions are met, each Action is executed and takes effect.

[0084] The feedback speech generation process generates the final feedback speech based on the number of actions executed, the results of the action execution, the feedback information, the context, the current scenario and the speech style information.

[0085] The high interpretability of the above solution is reflected in each step of the intent routing and action agent, which outputs the intermediate process and results in natural language, making it easy to check the output of each link and make targeted optimizations. Compared with the scalability problems caused by fine-tuning the unified group model in the traditional end-to-end method, this method has higher scalability:

[0086] With the above technical solution, for the problem that the task-based dialogue method provided by the present invention is still not concise and accurate enough for the current task-based dialogue system, the present invention obtains the user's current dialogue statement; inputs the current dialogue statement into a large language model to extract the core information of the current dialogue statement; performs multi-dimensional retrieval based on the core information to obtain learning examples; inputs the current dialogue statement and the learning examples into the large language model to determine the expected action, where the expected action is associated with the user's intention. In the above solution, a large language model is used to replace the traditional NLP method as the base model of NLU, and with the help of fine-tuning or retrieval enhancement technology, the intention slot is predicted in the form of natural language generation. The above large language model improves the effect of semantic understanding. The present application combines the large language model technology and the retrieval enhancement technology to build an intelligent agent to implement complex downstream skills, and solves the problems of insufficient interpretability, difficult expansion, and poor overall experience existing in the traditional pipeline method and the end-to-end method. The above intention routing uses multi-dimensional retrieval and utilizes the few-shot learning ability of the large language model. When adding new functions, only the retrieval library needs to be updated, and the large language model does not need to be retrained to achieve expansion; and new functions can be extended to the existing action intelligent agents or new intelligent agents can be added according to the function category and difficulty level, and the influence range is controlled within one intelligent agent, reducing the impact of the newly added function on the existing functions.

[0087] Further, as an implementation of the above Figure 1 shown method, the embodiment of the present invention also provides a task-based dialogue device for implementing the above Figure 1 shown method. This device embodiment corresponds to the foregoing method embodiment. For the convenience of reading, the details in the foregoing method embodiment will not be described one by one in this device embodiment, but it should be clear that the device in this embodiment can correspondingly implement all the contents in the foregoing method embodiment. As Figure 2 shown, the device includes: an acquisition unit 21, an extraction unit 22, a retrieval unit 23, and a determination unit 24, where

[0088] The acquisition unit 21 is used to acquire the user's current dialogue statement;

[0089] The extraction unit 22 is used to input the current dialogue statement into a large language model to extract the core information of the current dialogue statement;

[0090] The retrieval unit 23 is used to perform multi-dimensional retrieval based on the core information to obtain learning examples;

[0091] The determination unit 24 is used to input the current dialogue statement and the learning examples into a large language model to determine the expected action, where the expected action is associated with the user's intention.

[0092] The processor contains a kernel, which retrieves the corresponding program units from the memory. One or more kernels can be set, and by adjusting the kernel parameters, a task-based dialogue method can be implemented, which can solve the problem that the current task-based dialogue system is still not concise and accurate enough.

[0093] An embodiment of the present invention provides a computer-readable storage medium. The above computer-readable storage medium includes a stored program, and when the program is executed by a processor, the task-based dialogue method is implemented.

[0094] An embodiment of the present invention provides a processor. The processor is used to run a program, and when the program runs, the task-based dialogue method is executed.

[0095] An embodiment of the present invention provides an electronic device. The above electronic device includes at least one processor and at least one memory connected to the processor. Wherein, the above processor is used to call the program instructions in the above memory and execute the task-based dialogue method as described above.

[0096] An embodiment of the present invention provides an electronic device 30, as Figure 3 shown, the electronic device includes at least one processor 301, at least one memory 302 connected to the processor, and a bus 303. Wherein, the processor 301 and the memory 302 complete communication with each other through the bus 303. The processor 301 is used to call the program instructions in the memory to execute the above task-based dialogue method.

[0097] The intelligent electronic device in this article can be a PC, a PAD, a mobile phone, etc.

[0098] This application also provides a computer program product, which is suitable for executing a program initialized with the steps of the above task-based dialogue method when executed on a process management electronic device.

[0099] It should be noted that in the above embodiments, the descriptions of each embodiment have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0100] Those skilled in the art should understand that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0101] This application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded computer, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0102] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that implement the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0103] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or in multiple blocks.

[0104] Embodiments of the present application also provide a computer program product that includes computer software instructions. When the computer software instructions run on a processing device, the processing device is caused to execute the process of controlling the memory as in Figure 1 the corresponding embodiment.

[0105] A computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, they implement the processes or functions in accordance with the embodiments of the present application, either fully or partially. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from a website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be stored by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).

[0106] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein again.

[0107] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is only a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the couplings, direct couplings, or communication connections shown or discussed with each other may be indirect couplings or communication connections through some interfaces, devices, or units, and may be in electrical, mechanical, or other forms.

[0108] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0109] In addition, in each embodiment of the present application, the functional units may be integrated into one processing unit, or each unit may exist physically alone, or two or more units may be integrated into one unit. The above-mentioned integrated units may be implemented in the form of hardware or in the form of software functional units.

[0110] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, may be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods in the various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs that can store program codes.

[0111] The above are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present application.

Claims

1. A task-based dialogue method, characterized in that: include: Get the user's current conversation statement; Inputting the current dialogue sentence into a large language model to extract core information of the current dialogue sentence; Perform multi-dimensional retrieval based on the core information to obtain learning examples; The current dialogue sentence and the learning example are input into a large language model to determine an expected action, wherein the expected action is associated with a user intention.

2. The method according to claim 1, characterized in that The step of inputting the current dialogue sentence into a large language model to extract core information of the current dialogue sentence includes: In the case where the current dialogue sentence depends on the historical dialogue sentence, obtaining the relevance between the current dialogue sentence and the historical dialogue sentence; Extracting core information of the current dialogue sentence based on the relevance; When the correlation does not meet a preset threshold, a specific mark is output, wherein the specific mark is used to indicate that the core information cannot be extracted from the current dialogue sentence.

3. The method according to claim 1, characterized in that The step of inputting the current dialogue sentence into a large language model to extract core information of the current dialogue sentence includes: In a case where the current dialogue sentence does not depend on the historical dialogue sentences, core information of the current dialogue sentence is obtained.

4. The method according to claim 1, characterized in that: The step of inputting the current dialogue sentence into a large language model to extract core information of the current dialogue sentence includes: In the case that there are multiple pieces of core information of the current dialogue sentence, the core information of the current dialogue sentence is split, wherein different core information corresponds to different user intentions.

5. The method according to claim 2, characterized in that: The multi-dimensional search based on the core information to obtain learning examples includes: A learning example of retrieving the core information based on an atomic instruction vector library; When the core information cannot be extracted, searching from the dialogue vector library to obtain learning examples; Retrieve learning examples of the historical dialogue sentences based on a database, wherein the database contains multiple rounds of user dialogue sentences.

6. The method according to claim 1, characterized in that The expected action corresponds to an action agent.

7. The method according to claim 6, characterized in that Also includes: The action agent is controlled to input historical dialogue sentences, the core information and the learning examples into a large language model to output a revised expected action.

8. A task-based dialogue device, characterized in that: Also includes: An acquisition unit, used to acquire the current dialogue statement of the user; an extraction unit, configured to input the current dialogue sentence into a large language model to extract core information of the current dialogue sentence; A retrieval unit, used for performing multi-dimensional retrieval based on the core information to obtain learning examples; A determination unit is used to input the current dialogue sentence and the learning example into a large language model to determine an expected action, wherein the expected action is associated with the user's intention.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium includes a stored program, wherein when the program is executed by a processor, the steps of the task-based dialogue method according to any one of claims 1 to 7 are implemented.

10. An electronic device, characterized in that: The electronic device includes at least one processor and at least one memory connected to the processor; wherein the processor is used to call program instructions in the memory to execute the steps of the task-based dialogue method as described in any one of claims 1 to claim 7.

Citation Information

Cited By

  • Sample generation method, model training method, data processing method and electronic equipment

    CN121681783A