Interaction method, interaction apparatus, electronic device, storage medium, and computer program
The dialogue method addresses the challenge of understanding user needs and providing personalized responses by identifying usage scenarios, processing user queries with scenario-specific tools, and training a tool large language model, resulting in effective and personalized user interactions.
Patent Information
- Application Number
- JP2024196022
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-25
- Filing Date
- 2024-11-08
- Publication Date
- 2025-05-27
AI Technical Summary
Existing Task-Oriented Dialogue Systems (TODS) face challenges in effectively understanding user needs and providing personalized responses across diverse usage scenarios.
The proposed solution involves a dialogue method that identifies usage scenarios, obtains user data, calls tools corresponding to the scenario to process user queries, and generates answer information based on the tool execution results. This method also includes training a tool large language model using historical dialogue information and labels to improve tool labeling accuracy.
The solution enables deep understanding of user needs through tools tailored to specific scenarios, providing purposeful and personalized responses that effectively meet diverse user needs and enhance purchasing desire.
Smart Images

Figure 2025081252000001_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of artificial intelligence technology, and in particular to the fields of natural language processing and deep learning technology.
Background Art
[0002] Task-Oriented Dialogue Systems (TODS) aim to achieve specific user tasks such as ticket reservation and music playback through multiple rounds of dialogue. Usually, a modular architecture including four main modules: Natural Language Understanding (NLU), Dialogue State Tracking (DST), Dialogue Policy Learning (DPL), and Natural Language Generation (NLG) is adopted.
[0003] Natural language understanding is the process of converting a user's natural language input into a meaning representation that can be understood by the system by the NLU module. This usually relates to subtasks such as domain classification, intent recognition, and slot filling.
[0004] Dialogue state tracking is to track and update the dialogue state by analyzing the dialogue history and the current user's question by the DST module, which is the core of dialogue management.
[0005] Dialogue policy learning is to determine the next dialogue action by the DPL module based on the current dialogue state and a predefined policy model. Policy learning usually uses methods of supervised learning or reinforcement learning.
[0006] Natural language generation is to convert it into a response in natural language form by the NLG module. Usually, an end-to-end generation method is used to generate the answer.
Summary of the Invention
[0007] Embodiments of the present disclosure provide a dialogue method, a dialogue device, an electronic device, a storage medium, and a computer program.
[0008] In a first aspect, embodiments of the present disclosure include: a step of identifying a usage scenario corresponding to user query information; a step of obtaining user data in the usage scenario; a step of calling a tool corresponding to the usage scenario to process the user query information and the user data and obtaining a tool execution result; and a step of generating answer information corresponding to the user query information based on the tool execution result.
[0009] In a second aspect, embodiments of the present disclosure include: a step of obtaining a training sample including historical dialogue information of a first sample user and labels of a first sample tool corresponding to at least one usage scenario; a step of inputting the historical dialogue information of the first sample user into a large language model to obtain first predicted tool information; a step of calculating a first loss based on the first predicted tool information and the labels of the first sample tool; and a step of adjusting parameters of the large language model based on the first loss to obtain a tool large language model.
[0010] In a third aspect, embodiments of the present disclosure provide a dialogue device including: an identification module configured to identify a usage scenario corresponding to user query information; an acquisition module configured to obtain user data in the usage scenario; a call module configured to call a tool corresponding to the usage scenario to process the user query information and the user data and obtain a tool execution result; and a first generation module configured to generate answer information corresponding to the user query information based on the tool execution result.
[0011] In a fourth aspect, an embodiment of the present disclosure provides a training apparatus for a tool large language model, including: a first acquisition module configured to acquire a training sample including historical dialogue information of a first sample user and labels of a first sample tool corresponding to at least one usage scenario; a first prediction module configured to input the historical dialogue information of the first sample user into a large language model to acquire first predicted tool information; a first calculation module configured to calculate a first loss based on the first predicted tool information and the labels of the first sample tool; and a first adjustment module configured to adjust parameters of the large language model based on the first loss to obtain a tool large language model.
[0012] In a fifth aspect, an embodiment of the present disclosure provides an electronic device including at least one processor and a memory communicably connected to the at least one processor. Instructions executable by the at least one processor are stored in the memory. When the instructions are executed by the at least one processor, the at least one processor is caused to execute the dialogue method described in the first aspect or the training method of the tool large language model described in the second aspect.
[0013] In a sixth aspect, an embodiment of the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions, which cause a computer to execute the dialogue method described in the first aspect or the training method of the tool large language model described in the second aspect.
[0014] In a seventh aspect, an embodiment of the present disclosure provides a computer program that, when executed by a processor, causes the dialogue method described in the first aspect or the training method of the tool large language model described in the second aspect to be executed.
[0015] Embodiments of the present disclosure provide an interaction method, which can deeply understand the needs of users through tools corresponding to usage scenarios, appropriately satisfy those needs, and achieve a purposeful response to user queries in usage scenarios, meeting the personalized and diverse needs of users and effectively attracting the purchasing desire of users.
[0016] The key features or important features of the embodiments of the present disclosure do not limit the scope of the present disclosure. Other features of the present disclosure will be more easily understood from the following description.
Brief Description of the Drawings
[0017] Other features, objectives, and advantages of the present disclosure will become clearer by reading the detailed description of non-limiting embodiments with reference to the following drawings. The drawings are used to better understand the present disclosure and are not a limitation to the present disclosure.
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
Embodiments for Carrying Out the Invention
[0018] The following describes exemplary embodiments of the present disclosure with reference to the drawings. For the purpose of assisting understanding, various details of the embodiments of the present disclosure are described herein, but it should be understood that these are merely exemplary. Therefore, it should be understood that those skilled in the art can make various changes and modifications to the embodiments herein without departing from the scope and gist of the present disclosure. In the following description, for the sake of clarity and simplicity, descriptions of known functions and structures are omitted.
[0019] It should be noted that the embodiments of the present disclosure and the features in the embodiments can be combined with each other as long as there is no contradiction. The present disclosure will be described in detail below with reference to the drawings and embodiments.
[0020] FIG. 1 shows a flow 100 of an embodiment of a method for training a tool large language model according to the present disclosure. The method for training the tool large language model includes the following steps.
[0021] In step 101, training samples corresponding to at least one usage scenario are obtained.
[0022] In this embodiment, the execution entity of the method for training the tool large language model can obtain training samples corresponding to at least one usage scenario.
[0023] Generally, the entity that executes the training method of a large language model is a server. The server may be hardware or software. When the server is hardware, it may be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it may be realized as multiple software or software modules (for example, those for providing distributed services), or as a single software or software module. It is not particularly limited here.
[0024] Generally, historical conversation information in various usage scenarios of a large number of users can be collected and processed to obtain training samples corresponding to various usage scenarios. Here, the usage scenario may be the scene where the user is located when having a conversation. The usage scenario may include, but is not limited to, sales scenarios, question-and-answer scenarios, etc. Different tools for processing historical conversation information can be provided for different usage scenarios. For example, in a sales scenario, it may include, but is not limited to, tools such as product recommendation, product comparison, and user behavior analysis. In a question-and-answer scenario, it may include, but is not limited to, tools such as answering questions and evaluating user satisfaction. The training sample may include the historical conversation information of the first sample user and the label of the first sample tool. The historical conversation information of the first sample user may be the past conversation information of the first sample user. The label of the first sample tool may be a tool label obtained by labeling the historical conversation information of the first sample user according to the usage scenario. For example, by inputting the historical conversation information of the first sample user into a large language model, the label of the first sample tool can be obtained. As another example, a person skilled in the art can also manually label the historical conversation information of the first sample user to obtain the label of the first sample tool.
[0025] In some embodiments, depending on different usage scenarios, removable tools and corresponding description files can be provided. The installation of such removable tools has high flexibility and expandability and can adapt to various usage scenarios. Usually, the description file of a tool may include the name of the tool, its functions, basic parameters, optional parameters, etc. It may also include some usage examples. Such description files are useful for good planning, tool selection, generation of tool parameters, and execution of tools.
[0026] In step 102, the historical conversation information of the first sample user is input into the large language model to obtain the first predicted tool information.
[0027] In this embodiment, the above execution entity can input the historical conversation information of the first sample user into the large language model to obtain the first predicted tool information.
[0028] The large language model may be a trained model with a general labeling ability by tools. When the large language model cannot identify the usage scenario of the historical conversation information of the first sample user, a general label is attached to the historical conversation information of the first sample user by tools to obtain the first predicted tool information.
[0029] In step 103, a first loss is calculated based on the first predicted tool information and the label of the first sample tool.
[0030] In this embodiment, the above execution entity can calculate a first loss based on the first predicted tool information and the label of the first sample tool.
[0031] Here, an appropriate loss function can be selected. The first prediction tool information and the label of the first sample tool can be input into the loss function to obtain the first loss by calculation. Here, the first loss can be used to represent the difference between the first prediction tool information and the label of the first sample tool. The smaller the difference, the stronger the ability of the large language model to label tools for different usage scenarios, and the larger the difference, the weaker the ability of the large language model to label tools for different usage scenarios.
[0032] In step 104, the parameters of the large language model are adjusted based on the first loss to obtain a tool large language model.
[0033] In this embodiment, the above execution entity can adjust the parameters of the large language model based on the first loss to obtain a tool large language model.
[0034] Until the loss is small enough and the model converges, the parameters of the large language model can be repeatedly updated during training to obtain a tool large language model. The tool large language model can have the ability to label tools for each usage scenario.
[0035] In some embodiments, in order to ensure the accuracy and reliability of tool labeling by the tool large language model for each application scenario, the tool large language model can be tested and optimized. Specifically, it is as follows.
[0036] First, obtain test samples corresponding to at least one usage scenario.
[0037] Generally, historical conversation information in various usage scenarios of a large number of users can be collected and processed to obtain test samples corresponding to various usage scenarios. Here, the test samples may include the historical conversation information of the second sample user and the labels of the second sample tools. The historical conversation information of the second sample user may be the past conversation information of the second sample user. The labels of the second sample tools may be tool labels obtained by labeling the historical conversation information of the second sample user with tools according to the usage scenario.
[0038] After that, the historical conversation information of the second sample user is input into the tool large language model to obtain the second predicted tool information.
[0039] Next, based on the labels of the second sample tools and the second predicted tool information, the accuracy of the tool large language model is calculated.
[0040] Generally, when the difference between the labels of the second sample tools and the second predicted tool information is small, it is considered that the tool labeling based on the tool large language model is accurate. When the difference between the labels of the second sample tools and the second predicted tool information is large, it is considered that the tool labeling based on the tool large language model is inaccurate. Here, dividing the number of times of accurate labeling by the tool large language model by the total number of times of tool labeling based on the tool large language model can obtain the accuracy of the tool large language model.
[0041] Finally, it is determined whether the accuracy of the tool large language model is less than a preset accuracy threshold. If it is above the preset accuracy threshold, it is determined that the tool large language model has passed the test. If it is less than the preset accuracy threshold, a second loss is calculated based on the labels of the second sample tools and the second predicted tool information, and the parameters of the tool large language model are adjusted based on the second loss.
[0042] Embodiments of the present disclosure provide a method for training a tool large language model, which performs supervised fine-tuning on a large language model using training samples corresponding to various usage scenarios to obtain a tool large language model, and improves the labeling ability of tools based on the tool large language model for various usage scenarios.
[0043] Figure 2 shows a flowchart of supervised fine-tuning of a large language model.
[0044] In step 201, the large language model labels online conversations to obtain labeled data.
[0045] In step 202, the person in charge of labeling checks the labeled data.
[0046] In step 203, the large language model is trained using the checked labeled data.
[0047] In step 204, the trained large language model is evaluated.
[0048] In step 205, loop optimization is performed using the evaluation effect of the large language model to label data.
[0049] Figure 3 shows a flow 300 of an embodiment of the dialogue method according to the present disclosure. The dialogue method includes the following steps.
[0050] In step 301, the usage scenario corresponding to the user's query information is identified.
[0051] In this embodiment, the execution entity of the dialogue method can receive the user's query information sent by the user and identify the usage scenario corresponding to the user's query information.
[0052] Generally, the execution entity of the conversation method is a server. The server may be hardware or software. When the server is hardware, it may be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it may be realized as multiple software or software modules (for example, those for providing distributed services), or as a single software or software module. It is not particularly limited here.
[0053] Generally, the portal sites for different usage scenarios are different, and users can access the corresponding pages according to their own needs and input their query information. From the Web page where the user's query information is input, the usage scenario corresponding to the user's query information can be identified. Here, the usage scenario may be the scene where the user is located when having a conversation. The usage scenario may include, but is not limited to, a sales scene, a question-and-answer scene, etc. Also, for each usage scenario, the character of the agent interacting with the user is different. For example, in a sales scene, the agent interacting with the user may play the role of a salesperson who actively recommends products or persuades the user to purchase products. In a question-and-answer scene, the agent having a conversation with the user may play the role of a question-and-answer assistant who actively answers the user's questions and improves the user's satisfaction.
[0054] In step 302, user data in the usage scenario is acquired.
[0055] In this embodiment, the above execution entity can acquire user data in the usage scenario.
[0056] For each usage scenario, user data may be stored separately. Here, corresponding user data can be obtained according to the usage scenario. Here, user data may include, but is not limited to, the dialogue context, dialogue state, and user persona. The dialogue context may include the user's historical dialogue information and an intelligence summary regarding the user's historical dialogue. The user's historical dialogue information can record all the dialogues between the user and the agent. The intelligence summary regarding the user's historical dialogue may be an intelligence summary of the user's historical dialogue information every few rounds, which can prevent forgetting over time. The dialogue state can record the execution paths and states of the historical agent and the current agent, and can play a guiding role for the agent's next plan. The user persona may store the user's personal characteristics and a summary of the past user paths, which is important for personalized recommendations and long-term memory / understanding.
[0057] In step 303, a tool corresponding to the usage scenario is called to process the user's query information and user data to obtain a tool execution result.
[0058] In this embodiment, the above execution entity can call a tool corresponding to the usage scenario to process the user's query information and user data to obtain a tool execution result.
[0059] For each usage scenario, different tools for processing the user's query information and user data can be provided. For example, in the sales scenario, it may include, but is not limited to, tools such as product recommendation, product comparison, and user behavior analysis. In the question-and-answer scenario, it may include, but is not limited to, tools such as question-and-answer and user satisfaction evaluation.
[0060] In some embodiments, depending on different usage scenarios, detachable tools and corresponding description files can be provided. The installation of such detachable tools has high flexibility and scalability and can adapt to various usage scenarios. Usually, the description file of a tool may include the name of the tool, its functions, basic parameters, optional parameters, etc. It may also include some usage examples. Such a description file is helpful for good planning, tool selection, generation of tool parameters, and execution of the tool, etc.
[0061] In step 304, based on the tool execution result, response information corresponding to the user's query information is generated.
[0062] In this embodiment, the above execution entity can generate response information corresponding to the user's query information based on the tool execution result. By means of the tool corresponding to the usage scenario, it is possible to purposefully answer the user's query in the usage scenario.
[0063] In some embodiments, the above execution entity can process the tool execution result based on the prompt information of the usage scenario to generate response information. For each usage scenario, there may be different prompt information. For each usage scenario, in order to meet the needs of different agent characters and users, the identification of the agent character and the alignment and adjustment of the language style are carried out based on the prompt information. For example, in a sales scenario, the agent may act as a salesperson who actively recommends products or induces the user to purchase products. In a question-and-answer scenario, the agent needs to actively answer the user's questions and act as a question-and-answer assistant to improve the user's satisfaction.
[0064] In some embodiments, the above execution entity can generate guidance information for the next question based on the user's query information, response information, and user data, thereby actively prompting the user to start the next dialogue round.
[0065] In some embodiments, the execution entity may update user data based on the user's query information and answer information. After each round of interaction ends, the user data can be updated based on the user's interaction information in the current round to assist the agent in making decisions in the next round. Here, the user data may include, but is not limited to, the dialogue context, the dialogue state, and the user persona.
[0066] Embodiments of the present disclosure provide a dialogue method, which can deeply understand the user's needs through tools corresponding to the usage scenarios and appropriately satisfy those needs, and can purposefully answer the user's queries in the usage scenarios, satisfying the personalized and diverse needs of the user and effectively attracting the user's purchase desire.
[0067] FIG. 4 shows a flow 400 of another embodiment of the dialogue method according to the present disclosure. The dialogue method includes the following steps.
[0068] In step 401, identify the usage scenario corresponding to the user's query information.
[0069] In this embodiment, the execution entity of the dialogue method can receive the user's query information sent by the user and identify the usage scenario corresponding to the user's query information.
[0070] Generally, the execution entity of the dialogue method is a server. The server may be hardware or software. When the server is hardware, it may be implemented as a distributed server cluster composed of multiple servers or as a single server. When the server is software, it may be realized as multiple software or software modules (for example, for providing distributed services), or as a single software or software module. It is not particularly limited here.
[0071] Generally, the portal sites for different usage scenarios are different, and users can access the corresponding pages according to their own needs and input their query information. From the web page where the user's query information is input, the usage scenario corresponding to the user's query information can be identified. Here, the usage scenario may be the scene where the user is located when interacting. The usage scenario may include, but is not limited to, a sales scene, a question-and-answer scene, etc. Also, for each usage scenario, the character of the agent interacting with the user is different. For example, in a sales scene, the agent interacting with the user may play the role of a salesperson who actively recommends products or persuades the user to purchase products. In a question-and-answer scene, the agent interacting with the user may play the role of a question-and-answer assistant who actively answers the user's questions and improves the user's satisfaction.
[0072] In step 402, user data in the usage scenario is acquired.
[0073] In this embodiment, the above execution entity can acquire user data in the usage scenario.
[0074] For each usage scenario, user data may be stored separately. Here, corresponding user data can be obtained according to the usage scenario. Here, user data may include, but is not limited to, the dialogue context, dialogue state, and user persona. The dialogue context may include the user's historical dialogue information and an intelligence summary regarding the user's historical dialogue. The user's historical dialogue information can record all the dialogues between the user and the agent. The intelligence summary regarding the user's historical dialogue may be an intelligence summary of the user's historical dialogue information every few rounds, which can prevent forgetting over time. The dialogue state can record the execution paths and states of the historical agent and the current agent, and can play a guiding role for the agent's next plan. The user persona may store the user's personal characteristics and a summary of the past user paths, which is important for personalized recommendations and long-term memory and understanding.
[0075] In step 403, the user's query information and user data are input into the tool large language model to obtain the tool execution result.
[0076] In this embodiment, the above execution entity can input the user's query information and user data into the tool large language model to obtain the tool execution result. Among them, the user's query information and user data are input into the tool large language model, and the tool large language model is a large language model for outputting tools to solve the user's problems in the usage scenario. The tool execution result can be obtained by using the script execution tool.
[0077] The tool large language model can predict the tools used to solve the user's problems in the usage scenario. The task objective of the tool large language model may be to output all the called tools and their inputs at once. The tool large language model may be obtained by fine-tuning the large language model with supervision using the historical conversation information of sample users with tool labels, and its training process can refer to the embodiment shown in FIG. 1, which will be omitted here.
[0078] For each usage scenario, different tools for processing the user's query information and user data can be provided. For example, in the sales scenario, it may include, but is not limited to, tools such as product recommendation, product comparison, and user behavior analysis. In the question and answer scenario, it may include, but is not limited to, tools such as question and answer and user satisfaction evaluation.
[0079] In some embodiments, depending on the usage scenario, detachable tools and corresponding description files can be provided. The installation of such detachable tools has high flexibility and scalability and can adapt to various usage scenarios. Usually, the description file of the tool may include the name, function, basic parameters, optional parameters, etc. of the tool. It may also include some usage examples. Such description files are useful for good planning, tool selection, tool parameter generation, and tool execution, etc.
[0080] In step 404, the prompt information and the tool execution result are input into the character large language model to obtain the answer information.
[0081] In this embodiment, the above execution entity can input the prompt information and the tool execution result into the character large language model to obtain the answer information.
[0082] The character large language model may be a trained large language model, which can align the definition of the character based on the prompt information to answer the user's question and reject the answer to an irrelevant question. For each usage scenario, it may have different prompt information. For each usage scenario, in order to meet the needs of different agent characters and users, based on the prompt information, the identification and language style of the agent character are aligned and adjusted. For example, in a sales scenario, the agent may act as a salesman who actively recommends products or persuades users to purchase products. In a question and answer scenario, the agent needs to actively answer the user's question and act as a question and answer assistant to improve the user's satisfaction.
[0083] In step 405, the user's query information, answer information, and user data are input into the decision tree to obtain the next decision node.
[0084] In this embodiment, the above execution entity can input the user's query information, answer information, and user data into the decision tree to obtain the next decision node. Among them, the decision tree may be created based on the user's historical conversations, and the next decision node of the user can be predicted by the decision tree logic.
[0085] In step 406, the next decision node is input into the large language model to obtain the guiding information for the next question.
[0086] In this embodiment, the above execution entity can input the next decision node into the large language model to obtain the guiding information for the next question. For example, the large language model can submit guiding questions by few-shot. Finally, an active invitation question is selected by the PK battle mechanism.
[0087] Embodiments of the present disclosure provide an interaction method, which can deeply understand the needs of users and satisfy them appropriately by means of character alignment, tool inference, and active persuasion methods, and can effectively enhance the user's purchase intention. Through character alignment, it is possible to quickly switch to multiple agents. By means of quickly detachable tools, they can be applied to multiple usage scenarios. By the cooperative operation of multiple agents, it is possible to efficiently respond to the complex and diverse selection and purchase needs of users.
[0088] FIG. 5 shows a model configuration diagram of the interaction method. The model configuration of the interaction method may include a tool large language model 501, a character large language model 502, a decision tree 503, and a large language model 504.
[0089] First, input the query information and user data of the user in the usage scenario into the tool large language model 501 to obtain a tool execution result. Here, the user data may include the dialogue context, dialogue state, user persona, etc.
[0090] Then, input the prompt information corresponding to this usage scenario and the tool execution result into the character large language model 502 to obtain answer information.
[0091] Also, input the query information, answer information, and user data of the user into the decision tree 503 to obtain the next decision node. Input the next decision node into the large language model 504 to obtain the guidance information for the next question.
[0092] Here, a general agent architecture is designed, and through character alignment, it is possible to quickly switch to multiple agents. By means of quickly detachable tools, they can be applied to multiple usage scenarios. By the cooperative operation of multiple agents, it is possible to efficiently respond to the complex and diverse selection and purchase needs of users. Its advantages are mainly manifested in the following aspects.
[0093] 1. It can predict the user's selection and purchase path, improve the user experience, and promote conversion. Through the mechanism of multi-agent collaborative operation, it can deeply understand the user's needs and appropriately satisfy them, effectively enhance the user's purchase intention, promote the progress to the next stage of the selection and purchase path, and effectively improve the sales efficiency and conversion rate.
[0094] 2. Improve the flexibility and scalability of the solution. Design a general-purpose agent architecture that can be flexibly adapted to different usage scenarios through pluggable tool installation and character prompts.
[0095] 3. Reduce the operation cost. Through agent cooperation, a large number of user requests can be automatically processed, thereby greatly reducing the workload of customer service staff and saving operation costs.
[0096] There are obvious advantages in terms of user experience, sales efficiency, product flexibility and scalability, and operation cost.
[0097] Figure 6 shows a flowchart of single-agent interaction. Single-agent interaction includes the following steps.
[0098] In step 601, the user's query information and user data are input into the tool large language model to predict the tool used to solve the user's problem.
[0099] In step 602, it is determined whether it is necessary to call a tool. If "yes", step 603 is executed; if "no", step 604 is executed.
[0100] In step 603, the tool is executed.
[0101] In step 604, context information is generated.
[0102] In step 605, the prompt information and the context information are input into the character large language model.
[0103] In step 606, the answer information is output.
[0104] In step 607, the user's query information, the answer information, and the user data are input into the decision tree to obtain the next decision node.
[0105] In step 608, the next decision node is input into the large language model.
[0106] In step 609, the guiding information for the next question is generated.
[0107] In step 610, based on the answer information and the guiding information for the next question, comprehensive answer information is generated.
[0108] Figure 7 shows a flowchart of multi-agent interaction. Taking the sale of automobiles as an example, three agents, namely a purchase guidance agent, a question-and-answer agent, and a marketing agent, are designed and applied to cover different stages of user selection and purchase, and constitute the entire sales service system. Among them, the purchase guidance agent can play the role of guiding the purchase of automobiles. The question-and-answer agent can play the role of an automobile question-and-answer assistant. The marketing agent can play the role of selling automobiles.
[0109] As shown in Figure 7, the multi-agent interaction process includes the following steps.
[0110] In step 701, in the selection and purchase scenario, the purchase guidance agent is used to recommend the vehicle models that the user is interested in and clarify them until a single vehicle model is obtained.
[0111] In step 702, the purchase guidance agent determines whether the brand that the user is interested in is the only one. If the answer is "yes", step 703 is executed; if the answer is "no", the process returns to step 701 for clarification.
[0112] In step 703, in the question-and-answer scenario, or when the brand that the user is interested in is the only one, a series of questions about automobiles are answered for the user using the question-and-answer agent.
[0113] In step 704, the marketing agent is used to persuade the user to purchase a car and record personal information.
[0114] In step 705, the current user interaction is updated in the memory.
[0115] In step 706, the next flow is actively started.
[0116] Figure 8 shows a diagram of the memory update mechanism of the multi-agent. For the purchase guidance agent 801, the question-and-answer agent 802, and the marketing agent 803, the user data is updated every time the agent finishes execution. The user data may include the dialogue context 804, the dialogue state 805, and the user persona 806.
[0117] Through the single-agent interaction flow in Figure 6 and the multi-agent interaction flow in Figure 7, the cooperation of the multi-agent is realized, providing the user with a personalized automobile purchase consultation service, persuading the user to record personal information and visit the store, and enhancing the user's willingness to purchase an automobile.
[0118] In addition, the sales service system based on the cooperative multi-agent has a very wide application field. In addition to automobile sales, it can be applied to various fields such as online retail, e-commerce, customer service, online marketing, education consultation, and financial services.
[0119] Referring further to FIG. 9, as an embodiment of the method shown in each of the above figures, the present disclosure provides an embodiment of a training apparatus for a tool large language model, and the embodiment of the apparatus corresponds to the embodiment of the method shown in FIG. 1, and the apparatus can be specifically applied to various electronic devices.
[0120] As shown in FIG. 9, the training apparatus 900 for the tool large language model according to this embodiment may include a first acquisition module 901, a first prediction module 902, a first calculation module 903, and a first adjustment module 904. Here, the first acquisition module 901 is configured to acquire training samples corresponding to at least one usage scenario. The training samples include historical conversation information of a first sample user and labels of first sample tools. The first prediction module 902 is configured to input the historical conversation information of the first sample user into the large language model to obtain first predicted tool information. The first calculation module 903 is configured to calculate a first loss based on the first predicted tool information and the labels of the first sample tools. The first adjustment module 904 is configured to adjust the parameters of the large language model based on the first loss to obtain a tool large language model.
[0121] In this embodiment, in the training apparatus 900 for the tool large language model, the specific processes of the first acquisition module 901, the first prediction module 902, the first calculation module 903, and the first adjustment module 904, and the technical effects brought about by them can refer to the relevant descriptions of steps 101 to 104 in the corresponding embodiments of FIG. 1, and the description is omitted here.
[0122] In some optional embodiments of the present embodiment, the training device 900 for the tool large language model is configured to obtain a test sample including the historical dialogue information of a second sample user and the label of a second sample tool corresponding to at least one usage scenario, a second prediction module configured to input the historical dialogue information of the second sample user into the tool large language model to obtain second predicted tool information, a second calculation module configured to calculate the accuracy of the tool large language model based on the label of the second sample tool and the second predicted tool information, and a determination module configured to determine that the tool large language model has passed the test when the accuracy of the tool large language model is equal to or higher than a preset accuracy threshold.
[0123] In some optional embodiments of the present embodiment, the training device 900 for the tool large language model further includes a third calculation module configured to calculate a second loss based on the label of the second sample tool and the second predicted tool information in response to the accuracy of the tool large language model being smaller than a preset accuracy threshold, and a second adjustment module configured to adjust the parameters of the tool large language model based on the second loss.
[0124] Referring further to FIG. 10, as an implementation manner of the methods shown in the above figures, the present disclosure provides an embodiment of an interaction device, the embodiment of the device corresponding to the embodiment of the method shown in FIG. 3, and the device can be specifically applied to various electronic devices.
[0125] As shown in FIG. 10, the dialogue device 1000 of the present embodiment may include a specific module 1001, an acquisition module 1002, a call module 1003, and a first generation module 1004. The specific module 1001 is configured to identify a usage scenario corresponding to the user's query information. The acquisition module 1002 is configured to acquire user data in the usage scenario. The call module 1003 is configured to call a tool corresponding to the usage scenario to process the user's query information and user data, and obtain a tool execution result. The first generation module 1004 is configured to generate answer information corresponding to the user's query information based on the tool execution result.
[0126] In the present embodiment, in the dialogue device 1000, for the specific processing of the specific module 1001, the acquisition module 1002, the call module 1003, and the first generation module 1004 and the technical effects brought about thereby, reference can be made to the relevant descriptions of steps 301 to 304 in the corresponding embodiment of FIG. 3, and the description thereof will be omitted here.
[0127] In some optional embodiments of the present embodiment, the call module 1003 is further configured to input the user's query information and user data into a tool large language model to obtain a tool execution result, and the tool large language model is obtained by fine-tuning the large language model with supervision using the historical dialogue information of sample users with tool labels.
[0128] In some optional embodiments of the present embodiment, the first generation module 1004 includes a generation sub-module configured to process the tool execution result based on the prompt information of the usage scenario and generate answer information.
[0129] In some optional embodiments of the present embodiment, the generation sub-module is further configured to input the prompt information and the tool execution result into the character large language model to obtain the answer information.
[0130] In some optional embodiments of the present embodiment, the dialogue device 1000 further includes a second generation module configured to generate guiding information for the next question based on the user's query information, answer information, and user data.
[0131] In some optional embodiments of the present embodiment, the second generation module is further configured to input the user's query information, answer information, and user data into a decision tree to obtain the next decision node, and input the next decision node into a large language model to obtain guiding information for the next question.
[0132] In some optional embodiments of the present embodiment, the dialogue device 1000 further includes an update module configured to update user data based on the user's query information and answer information.
[0133] In the technical solution of the present disclosure, the acquisition, storage, and application of relevant user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and good customs.
[0134] According to the embodiments of the present disclosure, the present disclosure further provides an electronic device, a readable storage medium, and a computer program.
[0135] FIG. 11 shows a schematic block diagram of an exemplary electronic device 1100 that can be used to implement embodiments of the present disclosure. The electronic device represents various forms of digital computers such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Also, the electronic device can represent various forms of mobile devices such as personal digital processors, mobile phones, smartphones, wearable devices, and other similar computing devices. Note that the components shown here, their connection relationships, and their functions are merely examples and are not intended to limit the embodiments of the present disclosure described and / or claimed herein.
[0136] As shown in FIG. 11, the electronic device 1100 includes a computing unit 1101 that can execute various appropriate operations and processes by a computer program stored in a read-only memory (ROM) 1102 or a computer program loaded from a storage unit 1108 into a random access memory (RAM) 1103. Various programs and data necessary for the operation of the device 1100 can be further stored in the RAM 1103. The computing unit 1101, the ROM 1102, and the RAM 1103 are connected to each other via a bus 1104. An input / output (I / O) interface 1105 is also connected to the bus 1104.
[0137] In the electronic device 1100, a plurality of components including an input unit 1106 such as a keyboard and a mouse, an output unit 1107 such as various types of displays and speakers, a storage unit 1108 such as a magnetic disk and an optical disk, and a communication unit 1109 such as a network plug-in, a modem, and a wireless communication transceiver are connected to the I / O interface 1105. The communication unit 1109 enables the electronic device 1100 to exchange information or data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0138] The computing unit 1101 may be various general-purpose and / or dedicated processing components having processing and computing capabilities. Some examples of the computing unit 1101 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that execute machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1101 executes various methods and processes such as the above-described interaction method. For example, in some embodiments, the interaction method may be implemented as a computer software program tangibly included in a machine-readable medium such as the storage unit 1108. In some embodiments, part or all of the computer program may be loaded and / or installed into the electronic device 1100 via the ROM 1102 and / or the communication unit 1109. When the computer program is loaded into the RAM 1103 and executed by the computing unit 1101, one or more steps of the above-described interaction method can be executed. Alternatively, in other embodiments, the computing unit 1101 may be configured to execute the interaction method in any other suitable manner (e.g., via firmware).
[0139] The various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. Each of these embodiments can be implemented in one or more computer programs, which can be executed and / or interpreted in a programmable system including at least one programmable processor, which can be a dedicated or general-purpose programmable processor, and can receive data and instructions from a memory system, at least one input device, and at least one output device, and can include transmitting the data and instructions to the memory system, the at least one input device, and the at least one output device.
[0140] The program code for implementing the methods of the present disclosure may be created in any combination of one or more programming languages. These program codes can be provided to the processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing devices. When these program codes are executed by the processor or controller, the functions or operations defined in the flowchart and / or block diagram are implemented. The program code can be executed entirely on the device, partially on the device, partially on the device while being executed partially on a remote device as a stand-alone software package, or entirely on a remote device or server.
[0141] In the context of this disclosure, a machine-readable medium may be a tangible medium that includes or stores a program for use by or in combination with a command execution system, apparatus, or device. The machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. The machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination thereof. More specific examples of the machine-readable storage medium can include a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), a compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination thereof.
[0142] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having a display device (e.g., a cathode ray tube (CRT) or liquid crystal display (LCD) monitor) for displaying information to the user, and a keyboard and a pointing device (e.g., a mouse or a trackball), through which the user can provide input to the computer. Other types of devices can also be used to interact with the user. For example, the feedback provided to the user can be any form of sensing feedback, such as visual feedback, auditory feedback, or tactile feedback, and the input received from the user can be in any form, including sound input, voice input, or tactile input.
[0143] The systems and techniques described herein may be implemented in a computing system that includes backend components (e.g., a data server), or in a computing system that includes middleware components (e.g., an application server), or in a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser), where the user may interact with embodiments of the systems and techniques described herein via the graphical user interface or the web browser, or in a computing system that includes any combination of such backend components, middleware components, or frontend components. Also, the components of the system may be connected by digital data communication via any form or medium, such as a communication network. Examples of communication networks include a local area network (LAN), a wide area network (WAN), and the Internet.
[0144] The computer system may include a client and a server. The client and the server are typically located apart from each other and communicate via a communication network. The relationship between the client and the server is created by operating computer programs having a client-server relationship with each other on their respective computers. The server may be a cloud server, a server of a distributed system, or a server combined with a blockchain.
[0145] It should be understood that the order of steps can be rearranged, added, or deleted using the various forms of flow described above. For example, each step described in this disclosure may be executed in parallel, in sequence, or in a different order as long as the desired results of the technical solutions described in this disclosure can be achieved. This specification does not limit here.
[0146] The above specific embodiments do not limit the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, secondary combinations, and replacements can be made according to design requirements and other factors. Any modifications, equivalent replacements, and improvements made without departing from the spirit and principle of the present disclosure should all be included within the protection scope of the present disclosure.
Claims
1. identifying a usage scenario corresponding to the user's query information; acquiring user data in the usage scenario; calling a tool corresponding to the usage scenario, and processing the user's query information and the user data to obtain a tool execution result; generating answer information corresponding to the user's query information based on the execution result of the tool; Interaction methods including.
2. Invoking a tool corresponding to the usage scenario and processing the user's query information and the user data to obtain a tool execution result includes inputting the user's query information and the user data into a tool large-scale language model to obtain the tool execution result; The tool large-scale language model is obtained by supervised fine-tuning of the large-scale language model using historical dialogue information of a sample user with a tool label. The method of interaction according to claim 1 .
3. generating answer information corresponding to the user's query information based on the tool execution result, the step of processing the tool execution result based on prompt information of the usage scenario to generate the answer information; The method of interaction according to claim 1 .
4. The step of processing the tool execution result to generate the answer information based on the prompt information of the usage scene includes the step of inputting the prompt information and the tool execution result into a character large-scale language model to obtain the answer information. The method of interaction according to claim 1 .
5. generating next question guidance information based on the user's query information, the answer information, and the user data; The method of interaction according to claim 1 .
6. The step of generating next question guidance information based on the user's query information, the answer information, and the user data includes: inputting the user's query information, the answer information and the user data into a decision tree to obtain a next decision node; inputting the next decision node into a large-scale language model to obtain guidance information for the next question; 6. The method of claim 5, comprising:
7. updating the user data based on the query information and the answer information of the user; The method of interaction according to claim 1 .
8. Obtaining a training sample corresponding to at least one usage scenario, the training sample including historical dialogue information of a first sample user and a label of a first sample tool; inputting historical dialogue information of the first sample user into a large-scale language model to obtain first prediction tool information; Calculating a first loss based on the first prediction tool information and a label of a first sample tool; and adjusting parameters of the large-scale language model based on the first loss to obtain a tool large-scale language model. Tools for training large language models.
9. obtaining a test sample corresponding to at least one usage scenario, the test sample including historical interaction information of a second sample user and a label of a second sample tool; inputting historical dialogue information of the second sample user into the tool large-scale language model to obtain second predictive tool information; Calculating accuracy of the tool large language model based on the labels of the second sample tool and the second predictive tool information; determining that the tool large-scale language model has passed the test in response to the accuracy of the tool large-scale language model being equal to or greater than a preset accuracy threshold; The training method of claim 8 further comprising:
10. Calculating a second loss based on the label of the second sample tool and the second predicted tool information in response to the accuracy of the tool large language model being less than the preset accuracy threshold; adjusting parameters of the tool large-scale language model based on the second loss; The training method of claim 9 further comprising:
11. An identifying module configured to identify a usage scenario corresponding to the user's query information; an acquisition module configured to acquire user data in the usage scenario; A calling module configured to call a tool corresponding to the usage scenario to process the user's query information and the user data, and obtain a tool execution result; A first generation module configured to generate answer information corresponding to the user's query information based on the tool execution result.
12. The calling module further comprises: configured to input the user's query information and the user data into a tool large-scale language model to obtain the tool execution result; The tool large-scale language model is obtained by supervised fine-tuning of the large-scale language model using historical dialogue information of a sample user with a tool label.
12. An interactive device according to claim 11.
13. The first generation module: a generation submodule configured to process the tool execution result based on prompt information of the usage scenario to generate the answer information; 12. An interactive device according to claim 11.
14. The generating submodule further comprises: The prompt information and the execution result of the tool are input to a character large-scale language model to obtain the answer information.
12. An interactive device according to claim 11.
15. and a second generating module configured to generate next question guidance information based on the user's query information, the answer information, and the user data. An interactive device according to any one of claims 11 to 14.
16. The second generating module further comprises: inputting the user's query information, the answer information, and the user data into a decision tree to obtain a next decision node; configured to input the next decision node into a large-scale language model to obtain guidance information for the next question; 20. An interactive device according to claim 15.
17. The interactive device according to any one of claims 11 to 14, further comprising an update module configured to update the user data based on the query information and the answer information of the user.
18. A first acquisition module configured to acquire a training sample corresponding to at least one usage scene, the training sample including a first sample user's historical interaction information and a first sample tool's label; a first prediction module configured to input historical dialogue information of the first sample user into a large-scale language model to obtain first prediction tool information; a first calculation module configured to calculate a first loss based on the first prediction tool information and a label of a first sample tool; a first tuning module configured to adjust parameters of the large scale language model based on the first loss to obtain a tool large scale language model; A tool for training large-scale language models.
19. A second acquisition module configured to acquire a test sample corresponding to at least one usage scenario, the test sample including historical interaction information of a second sample user and a label of a second sample tool; a second prediction module configured to input historical dialogue information of the second sample user into the tool large-scale language model to obtain second predictive tool information; a second calculation module configured to calculate an accuracy of the tool large-scale language model based on the label of the second sample tool and the second predictive tool information; a determining module configured to determine that the tool large-scale language model passes a test according to the accuracy of the tool large-scale language model being equal to or greater than a preset accuracy threshold; 20. The training device of claim 18, comprising:
20. a third calculation module configured to calculate a second loss based on the label of the second sample tool and the second predicted tool information in response to the accuracy of the tool large language model being less than the preset accuracy threshold; and a second tuning module configured to tune parameters of the tool large language model based on the second loss.
20. A training device according to claim 19.
21. An electronic device comprising at least one processor and a memory communicatively connected to the at least one processor, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, cause the at least one processor to perform the dialogue method described in any one of claims 1 to 7 or the tool large-scale language model training method described in any one of claims 8 to 10.
22. A non-transitory computer-readable storage medium having computer instructions stored thereon, comprising: A non-transitory computer readable storage medium, the computer instructions causing a computer to perform the interaction method according to any one of claims 1 to 7 or the tool large scale language model training method according to any one of claims 8 to 10.
23. A computer program product which, when executed by a processor, implements the interaction method according to any one of claims 1 to 7 or the method for training a tool large-scale language model according to any one of claims 8 to 10.
Citation Information
Patent Citations
Information processing device, artificial intelligence selection method and artificial intelligence selection program
JP2019056970A
SYSTEM, APPARATUS, AND METHOD FOR PROVIDING INTENT SUGGESTIONS TO A USER IN A TEXT-BASED CONVERSATIONAL EXPERIENCE WITH USER FEEDBACK
JP2023511600A
Information platform for a virtual assitant
US20200073895A1