A method and device for processing dialogue interaction based on context environment

By obtaining and analyzing the current page information, historical dialogue information and current page information input by the user, combining the multimodal large language model and intention processing module, the free switching between dialogue robots and the execution of complex action instructions is realized, which solves the shortcomings of existing dialogue robots in interactive form switching and complex instructions execution, and improves interaction efficiency and user experience.

CN119202332BActive Publication Date: 2025-05-06SHANG FEI ZHI NENG JI SHU YOU XIAN GONG SI
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202411677465.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-22
Publication Date
2025-05-06
Estimated Expiration
2044-11-22

AI Technical Summary

Technical Problem

Existing dialogue robots cannot switch freely between multiple interaction forms, cannot perform voice interactions without text on any interactive interface, and cannot execute complex action commands, such as "Help me open the attachment on the right side of the screen".

Method used

By obtaining the current page information, historical dialogue information and current page information entered by the user, the context environment information is determined, and the user's intention analysis and execution is used to generate corresponding answer information.

Benefits of technology

Free switching between multiple interactive forms is achieved, the efficiency and accuracy of dialogue interaction is improved, complex action instructions can be executed, and the user's dialogue and interaction experience is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119202332B_ABST
    Figure CN119202332B_ABST
Patent Text Reader

Abstract

The present invention provides a method and device for processing a dialogue interaction based on a context environment, the method comprising: determining the context environment information corresponding to the current round of questions based on the current round of question information input by the user into the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; performing user intention analysis based on the current round of question information, the historical dialogue information corresponding to the user, and the context environment information of the current round of questions to obtain the user's target intention; in the case of determining that the conditions required for executing the target intention are currently met, triggering the execution of the target intention, obtaining the corresponding target intention execution result, and generating the corresponding current round answer information based on the target intention execution result. The method provided by the present invention can effectively improve the efficiency and accuracy of dialogue interaction, thereby enhancing the user's dialogue interaction experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a context-based dialogue interaction processing method, device and system, and also to an electronic device, a non-transitory computer-readable storage medium and a computer program product. Background Art

[0002] In recent years, with the rapid development of computer technology, conversational robots have been increasingly widely used in various industries. Conversational robots are mainly used in customer service dialogue, information query, task execution, etc. At present, there are several ways to implement mainstream conversational robots. One is a rule-based implementation method, which often uses keyword matching and other means to identify intent and then perform related processing. The processing process is generally in the form of a relatively fixed task chain. The second is a statistical method for training neural network models, which inputs the user's natural language into the model to obtain intent and slot values, and then performs related actions or conducts conversations. However, the conversational robots designed and implemented by the above methods often have obvious defects. First, the interaction form is generally limited to plain text. For example, conversational robots in the form of chat boxes, such as customer service robots, are often limited to plain text interactions. Even if they have graphic interaction methods such as buttons, after triggering the button, they often simulate the user sending text in the chat box first, and then complete the robot's answer through natural language understanding or rule-based methods. In essence, they are still pure text-based question-and-answer robots, and they fail to provide interactive services that are separated from text and rely only on voice on any interactive interface (especially pure graphical interface). Therefore, they cannot effectively interact with the graphical interface and cannot perform actions such as "help me open the attachment on the right side of the screen". Second, it does not support free switching between multiple forms of interaction. For example, although the personal voice assistants of modern smartphones can provide voice question-and-answer services, after the user uses interactive methods such as clicking, the personal voice assistants often close and delete the chat history. When the personal voice assistant is turned on again, a new conversation will begin. This type of interaction does not have the continuity between different forms of interaction, and different forms of interaction cannot be interspersed for continuous interaction. At the same time, in order to meet the needs of agile development and rapid configuration, the solution to the above problem cannot have a customized training or fine-tuning model process, and can only use a general large language model. Therefore, in multiple interaction scenarios (scenarios allow for but are not limited to text or voice, graphic controls and other forms of interaction), it is necessary to develop a more efficient and accurate conversational robot to meet the needs of users for multiple mixed interactions. Summary of the invention

[0003] The present invention provides a context-based dialogue interaction processing method and device, which are used to solve the defects of the dialogue interaction processing scheme of the dialogue robot in the prior art that the dialogue interaction efficiency and accuracy are poor due to high limitations.

[0004] The present invention provides a context-based dialogue interaction processing method, comprising:

[0005] Acquire current round question information input by a user into a current page of an interactive system, historical conversation information corresponding to the user, and current page information of the interactive system; determine context environment information corresponding to the current round question information based on the current round question information input by the user into the current page of the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system; perform user intent analysis based on the current round question information, the historical conversation information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the target intent of the user;

[0006] When it is determined that the conditions required to execute the target intention are currently met, the execution of the target intention is triggered, the corresponding target intention execution result is obtained, and the corresponding current round answer information is generated based on the target intention execution result.

[0007] According to a context-based dialog interaction processing method provided by the present invention, the context information corresponding to the current round of question information is determined based on the current round of question information input by the user to the current page of the interactive system, the historical dialog information corresponding to the user, and the current page information of the interactive system, specifically including:

[0008] The current round of question information input by the user into the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system are input into a preset multimodal large language model for analysis, and the key information features of the current context environment information output by the multimodal large language model are obtained, and the context environment information corresponding to the current round of question information is determined based on the key information features and the context environment information stored in the context environment database in the preset intention processing module. The multimodal large language model is a machine learning model obtained by iterative training based on sample input information and key information feature labels corresponding to the sample input information; or, the identification information set of the current page is extracted through a preset interface of the interactive system, and the identification information set of the current page is obtained based on the extraction, and the identification information set of the current page is compared and analyzed with the context environment information stored in the context environment database in the preset intention processing module to determine the context environment information corresponding to the current round of question information; wherein the context environment information is the environment in which the user is located in the interactive system, including the link and state in which the user is located, and the elements that can be touched or interacted with.

[0009] According to a context-based dialogue interaction processing method provided by the present invention, the user intention is parsed based on the current round of question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round of question information to obtain the user's target intention, specifically including:

[0010] Based on the context environment information corresponding to the question information of the current round and the corresponding relationship stored in the context environment database in the preset general intention processing module, determine the environment-specific intention corresponding to the context environment information corresponding to the question information of the current round, and perform vectorization processing based on the environment-specific intention to obtain an environment-specific intention vector; the corresponding relationship stored in the context environment database is the association relationship between each piece of context environment information and the environment-specific intention corresponding to each piece of context environment information;

[0011] Performing vectorization processing based on the general intent corresponding to the interactive system to obtain a general intent vector;

[0012] Performing vectorization processing based on the current round of question information and the historical conversation information corresponding to the user to obtain a conversation information vector;

[0013] Based on the dialogue information vector, vector matching analysis is performed with the environment-specific intent vector and the general intent vector to obtain a vector matching analysis result; and the user's target intent is selected from the environment-specific intent and the general intent according to the vector matching analysis result.

[0014] According to a context-based dialogue interaction processing method provided by the present invention, when it is determined that the conditions required for executing the target intention are currently met, triggering the execution of the target intention and obtaining the corresponding target intention execution result specifically includes:

[0015] Extracting the intent slot from the current round question information;

[0016] According to the intention slot extracted from the question information of the current round, detecting whether the conditions required for executing the target intention are currently met, and in the case of determining that the conditions required for executing the target intention are currently met, triggering the execution of the target intention to obtain the corresponding target intention execution result;

[0017] Among them, the condition required to execute the target intention is whether all slots corresponding to the target intention have been filled. When it is determined that all slots corresponding to the target intention have been filled, it is determined that the condition required to execute the target intention is met; when it is determined that all slots corresponding to the target intention have not been filled, it is determined that the condition required to execute the target intention is not met.

[0018] According to a context-based dialogue interaction processing method provided by the present invention, the generating of corresponding current round answer information based on the target intention execution result specifically includes:

[0019] Based on the target intent of the user, the preferred answer information stored in the intent database in the preset general intent processing module is queried. When the preferred answer method corresponding to the target intent of the user is queried, the target intent execution result is processed according to the preferred answer method to generate corresponding current round answer information; wherein the preferred answer method includes the layout, words and elements to be displayed of the answer; when the preferred answer method corresponding to the target intent of the user is not queried, the preset natural language large model is called to process the target intent execution result to generate corresponding current round answer information.

[0020] According to a context-based dialogue interaction processing method provided by the present invention, the target intent includes: query intent and action intent;

[0021] The triggering and executing the target intention to obtain the corresponding target intention execution result specifically includes:

[0022] When the target intent is a query-type intent, the query-type intent is triggered for execution, and the corresponding query result is returned after querying the database according to the query-type intent, and the query result is used as the execution result of the target intent; or, when the target intent is an action-type intent, the action-type intent is triggered for execution, and the execution result of the action-type intent is obtained after executing the action, and the execution result is used as the execution result of the target intent.

[0023] The present invention also provides a context-based dialogue interaction processing device, comprising:

[0024] An intention parsing unit is used to obtain the current round of question information input by the user to the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; determine the context environment information corresponding to the current round of question information based on the current round of question information input by the user to the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; perform user intention parsing based on the current round of question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round of question information to obtain the target intention of the user;

[0025] The intention execution and result return unit is used to trigger the execution of the target intention when it is determined that the conditions required to execute the target intention are currently met, obtain the corresponding target intention execution result, and generate the corresponding current round answer information based on the target intention execution result.

[0026] According to a context-based dialogue interaction processing device provided by the present invention, the intention parsing unit is specifically used to:

[0027] The current round of question information input by the user into the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system are input into a preset multimodal large language model for analysis, and the key information features of the current context environment information output by the multimodal large language model are obtained, and the context environment information corresponding to the current round of question information is determined based on the key information features and the context environment information stored in the context environment database in the preset intention processing module. The multimodal large language model is a machine learning model obtained by iterative training based on sample input information and key information feature labels corresponding to the sample input information; or, the identification information set of the current page is extracted through a preset interface of the interactive system, and the identification information set of the current page is obtained based on the extraction, and the identification information set of the current page is compared and analyzed with the context environment information stored in the context environment database in the preset intention processing module to determine the context environment information corresponding to the current round of question information; wherein the context environment information is the environment in which the user is located in the interactive system, including the link and state in which the user is located, and the elements that can be touched or interacted with.

[0028] According to a context-based dialogue interaction processing device provided by the present invention, the intention parsing unit is specifically used to:

[0029] Based on the context environment information corresponding to the question information of the current round and the corresponding relationship stored in the context environment database in the preset general intention processing module, determine the environment-specific intention corresponding to the context environment information corresponding to the question information of the current round, and perform vectorization processing based on the environment-specific intention to obtain an environment-specific intention vector; the corresponding relationship stored in the context environment database is the association relationship between each piece of context environment information and the environment-specific intention corresponding to each piece of context environment information;

[0030] Performing vectorization processing based on the general intent corresponding to the interactive system to obtain a general intent vector;

[0031] Performing vectorization processing based on the current round of question information and the historical conversation information corresponding to the user to obtain a conversation information vector;

[0032] Based on the dialogue information vector, vector matching analysis is performed with the environment-specific intent vector and the general intent vector to obtain a vector matching analysis result; and the user's target intent is selected from the environment-specific intent and the general intent according to the vector matching analysis result.

[0033] According to a context-based dialogue interaction processing device provided by the present invention, the intention execution and result return unit is specifically used to:

[0034] Extracting the intent slot from the current round question information;

[0035] According to the intention slot extracted from the question information of the current round, detecting whether the conditions required for executing the target intention are currently met, and in the case of determining that the conditions required for executing the target intention are currently met, triggering the execution of the target intention to obtain the corresponding target intention execution result;

[0036] Among them, the condition required to execute the target intention is whether all slots corresponding to the target intention have been filled. When it is determined that all slots corresponding to the target intention have been filled, it is determined that the condition required to execute the target intention is met; when it is determined that all slots corresponding to the target intention have not been filled, it is determined that the condition required to execute the target intention is not met.

[0037] According to a context-based dialogue interaction processing device provided by the present invention, the intention execution and result return unit is specifically used to:

[0038] Based on the target intent of the user, the preferred answer information stored in the intent database in the preset general intent processing module is queried. When the preferred answer method corresponding to the target intent of the user is queried, the target intent execution result is processed according to the preferred answer method to generate corresponding current round answer information; wherein the preferred answer method includes the layout, words and elements to be displayed of the answer; when the preferred answer method corresponding to the target intent of the user is not queried, the preset natural language large model is called to process the target intent execution result to generate corresponding current round answer information.

[0039] According to a context-based dialogue interaction processing device provided by the present invention, the target intent includes: query intent and action intent;

[0040] The triggering and executing the target intention to obtain the corresponding target intention execution result specifically includes:

[0041] When the target intent is a query-type intent, the query-type intent is triggered for execution, and the corresponding query result is returned after querying the database according to the query-type intent, and the query result is used as the execution result of the target intent; or, when the target intent is an action-type intent, the action-type intent is triggered for execution, and the execution result of the action-type intent is obtained after executing the action, and the execution result is used as the execution result of the target intent.

[0042] The present invention also provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the program, the context-based dialogue interaction processing method as described in any one of the above items is implemented.

[0043] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, and when the computer program is executed by a processor, the context-based dialogue interaction processing method as described in any of the above items is implemented.

[0044] The present invention also provides a computer program product, including a computer program, which, when executed by a processor, implements the context-based dialogue interaction processing method as described in any of the above items.

[0045] The context environment-based dialogue interaction processing method provided by the present invention obtains the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; determines the context environment information corresponding to the current round question information based on the current round question information input by the user to the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; performs user intention analysis based on the current round question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intention; when it is determined that the conditions required to execute the target intention are currently met, triggers the execution of the target intention to obtain the corresponding target intention execution result, and generates the corresponding current round answer information based on the target intention execution result, which can effectively improve the dialogue interaction efficiency and accuracy, thereby enhancing the user's dialogue interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0046] In order to more clearly illustrate the technical solutions in the present invention or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0047] Figure 1 This is one of the flow charts of the context-based dialogue interaction processing method provided by the present invention.

[0048] Figure 2 This is the second flow chart of the context-based dialogue interaction processing method provided by the present invention.

[0049] Figure 3 It is a structural schematic diagram of a context-based dialogue interaction processing device provided by the present invention.

[0050] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION

[0051] In order to make the purpose, technical solution and advantages of the present invention clearer, the technical solution of the present invention will be clearly and completely described below in conjunction with the drawings of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0052] Combine the following Figure 1-Figure 4 The context-based dialogue interaction processing method and device of the present invention are described, and the embodiments thereof are described in detail.

[0053] The following is a detailed description of an embodiment of the context-based dialogue interaction processing method of the present invention. Figure 1 As shown, it is one of the flow charts of the context-based dialogue interaction processing method provided by the present invention, and the specific implementation process includes the following steps:

[0054] Step 101: obtain current round question information input by a user into a current page of an interactive system, historical conversation information corresponding to the user, and current page information of the interactive system; determine context environment information corresponding to the current round question information based on the current round question information input by the user into the current page of the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system; perform user intent analysis based on the current round question information, the historical conversation information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intent.

[0055] In an embodiment of the present invention, after the conversational robot obtains the current round of question information input by the user into the current page of the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system, the current round of question information input by the user into the current page of the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system can be input into a preset multimodal large language model for analysis, and the key information features of the current context environment information output by the multimodal large language model are obtained, and based on the comparison and analysis of the key information features with the context environment information stored in the context environment database in the preset intention processing module, the context environment information corresponding to the current round of question information is determined. Contextual environment information; wherein the multimodal large language model is a machine learning model obtained by iterative training based on sample input information and key information feature labels corresponding to the sample input information; or, extracting the identification information set of the current page through the preset interface of the interactive system, based on the extraction of the identification information set of the current page, and based on the identification information set of the current page and the contextual environment information stored in the contextual environment database in the preset intention processing module, determining the contextual environment information corresponding to the question information of the current round; wherein the contextual environment information is the environment in which the user is located in the interactive system, including the link, state, and elements that can be touched or interacted. The conversational robot runs in a computer device.

[0056] It should be noted that the conversational robot in the present invention includes the following structures: a conversation management module and an intention processing module, wherein the intention processing module includes an intention executor, a context environment database, and an intention database. Figure 2As shown. The intention processing module includes an intention executor, a context environment database, and an intention database. Among them, the context environment database and the intention database are both created and maintained by the user in advance. The intention executor: When the prerequisite for the execution of a certain intention is met, it is responsible for executing the intention according to the input parameters, or returning the action parameters of the intention. For example, for query-type intentions, query and calculate relevant data and then return the results. For action-type intentions, calculate action parameters and return results. The context environment database: is responsible for storing possible intentions under each context (environment). When a user is in a different context environment, he may have a corresponding intention. For example, when a user is in the start interface, he may only ask some relatively macro questions and not specific questions. When the user has entered a specific environment, such as a detailed form of a project, he may ask the specific content of the project. At the same time, if the user has experienced multiple rounds of dialogue before, even if the interface seen by the user has not changed, he may still be in a different scene. For example, a user asks information questions for different areas in the same interface. After the user specifies different areas, different areas obviously have different information, and the questions that can be asked will also be different. In other words, the user's current intentions may be related to the current context. It should be noted that there are also some general intentions, such as "exit the current interface", "return to the main menu", etc. Therefore, the user's current intentions may include environment-specific intentions, as well as general intentions. The intention database: responsible for storing parameters of each intent, including but not limited to intent, intent slot, example questions, preferred answer methods, actions to be performed, and other information.

[0057] In addition, if Figure 2As shown, the dialogue management module includes a dialogue manager. The dialogue manager is responsible for interacting with the user and opening up interactive services to the outside world (i.e., a dialogue interaction processing method based on a context environment). The method includes three links: the intention analysis link corresponding to the intention analysis unit: This link is responsible for analyzing the user's target intention. If the user's target intention can be identified, proceed to the next link. Otherwise, initiate a counter-question. The input of this link is the current round of question information, historical dialogue information, and context environment information input by the user into the current page of the interactive system. The output is the target intention (for the next link) or the prompt information when the target intention is unclear (returned to the user). In this link, the context environment information is first obtained. This context environment information is given by the interactive system used by the user or by analyzing the current state of the interactive system and the user's historical dialogue records. There are two ways to perform this analysis. One is to input the current page information of the interactive system (i.e., the screenshot of the current page or its element information) and the current round of question information and historical dialogue information input by the user into the current page of the interactive system into a multimodal large language model, and obtain key information (i.e., key information features) through multimodal large language model analysis, and then compare the key information with the preset context environment information stored in the context environment database in several intention processing modules to obtain the current context environment information (i.e., the context environment information corresponding to the current round of question information). The other is to preset several rules and judge the current context environment information by a rule-based method, such as extracting the identification information set of the current page through the preset interface of the interactive system, obtaining the identification information set of the current page based on the extraction, and comparing and analyzing the identification information set of the current page with the context environment information stored in the preset context environment database in the intention processing module to determine the context environment information corresponding to the current round of question information. Since obtaining contextual information does not rely on a specific input, the input sources are wide, including but not limited to user interface information, user questions in the current round, and historical conversation information. This information (especially user interface information, etc.) always exists and will not disappear due to reasons such as closing the dialog box. Therefore, the conversation robot can be seamlessly intervened at any time. At the same time, accurate contextual information can be determined through the fusion analysis of multiple information, and then the types of operations that the user may perform can be known through the contextual information, so that a more accurate judgment of user intentions can be obtained. In addition, since it does not rely on a specific type of input, it allows users to perform multiple types of input, including but not limited to voice / text input, button clicks, sensor input, etc.

[0058] After obtaining the context information corresponding to the current round of question information, the context database is queried to find the possible environment-specific intent under the context information, and the user's current round of questions and historical dialogue information are analyzed through vector matching, natural language understanding, etc., and the user's target intent is selected from the environment-specific intent + general intent, and the target intent is input into the next link. If the user's target intent is unclear, a counter-question can be initiated. For example, the vector matching process includes: in the process of parsing the user's intention based on the current round of question information, the historical conversation information corresponding to the user, and the context environment information corresponding to the current round of question information, and obtaining the user's target intention, the conversation robot can determine the environment-specific intention corresponding to the context environment information corresponding to the current round of question information based on the context environment information corresponding to the current round of question information and the corresponding relationship stored in the context environment database in the preset general intention processing module, and perform vectorization processing based on the environment-specific intention to obtain an environment-specific intention vector; the corresponding relationship stored in the context environment database is the association relationship between each context environment information and the environment-specific intention corresponding to each context environment information; perform vectorization processing based on the general intention corresponding to the interactive system to obtain a general intention vector; perform vectorization processing based on the current round of question information and the historical conversation information corresponding to the user to obtain a conversation information vector; perform vector matching analysis based on the conversation information vector with the environment-specific intention vector and the general intention vector to obtain a vector matching analysis result; select the user's target intention from the environment-specific intention and the general intention according to the vector matching analysis result. The context environment information (i.e. Figure 2 Context (environment) information in the interactive system): refers to the environment in which the user is in the interactive system, including the link and status, the elements that can be touched or interacted with, the context, etc.

[0059] The intent execution link corresponding to the intent execution and result return unit: responsible for processing the call process. First, check whether the conditions required to execute the target intent are met. Generally speaking, it extracts parameters (intent slots) from the user's question and checks whether the parameters are sufficient. If not, a counter-question is initiated to ask the missing parameters. Then, if the conditions required to execute the target intent are met (for example, the intent slots are filled), the intent is executed. For query intents (common intents for question-and-answer dialogue robots), query the database and return the query results. For action intents (common intents for task-based dialogue robots), calculate the action parameters and return the action parameters, or return the execution results after executing the action. The result return link corresponding to the intent execution and result return unit: responsible for organizing the answer. If the intent has a preferred answer method in the intent database, the answer is given according to the preferred answer method, including but not limited to the layout of the answer, the words, and the elements to be displayed. The display form of the element includes but is not limited to fixed text, text filled in according to the template, charts (such as bar charts, line charts, pie charts, waterfall charts, subway maps, card charts, etc.), links, etc. The layout of the element includes the overall layout, the location, etc. Each form will be pre-defined in the conversational robot program. Users only need to fill in the elements they want and the layout, data and other attributes of the elements in the intent database. Since there are many types of elements, the output is diverse. If there is no preferred answer method, let the large model organize the answer by itself. Depending on the situation, intermediate results or original execution results of the intent can also be given. It should be noted that all the above processes only require the use of a (multimodal) general large language model, and there is no need for secondary pre-training or fine-tuning of the model (no targeted training is required).

[0060] The present invention can realize the seamless intervention of natural language input of the dialogue robot at any time in a multi-hybrid interactive scenario (the scenario allows including but not limited to text or voice, graphic controls and other interactive forms); the dialogue robot combines user questions, historical records and the current environment (such as the current stop page, executed instructions, etc.) through multiple rounds of dialogue to accurately judge the user's intention; no additional model training is required, only a general large language model is needed; highly configurable intention database and context (environment) database can realize the development and configuration of dialogue robots that meet user needs. The present invention provides a large model interaction mode based on context environment, which greatly improves the scene understanding ability and execution accuracy of the dialogue robot. Its core process is: construct a context environment database and an intention database, analyze the context environment where the user is currently located in the multi-hybrid interactive process in real time according to multiple information, obtain the user's possible intention and intention parameters in the context environment, execute the user's intention and give multiple return results, and realize the interactive function of the dialogue robot that intervenes at any time in each scenario.

[0061] Step 102: When it is determined that the conditions required to execute the target intention are currently met, trigger the execution of the target intention, obtain the corresponding target intention execution result, and generate the corresponding current round answer information based on the target intention execution result.

[0062] In an embodiment of the present invention, the intent slot can be specifically extracted from the current round of question information; based on the intent slot extracted from the current round of question information, it is detected whether the conditions required to execute the target intent are currently met; if it is determined that the conditions required to execute the target intent are currently met, the execution of the target intent is triggered to obtain the corresponding target intent execution result. Wherein, the condition required to execute the target intent is whether all slots corresponding to the target intent have been filled; if it is determined that all slots corresponding to the target intent have been filled, it is determined that the conditions required to execute the target intent are met; if it is determined that all slots corresponding to the target intent have not been filled, it is determined that the conditions required to execute the target intent are not met. Wherein, the target intent includes: query type intent and action type intent. Correspondingly, the triggering and executing the target intent and obtaining the corresponding target intent execution result specifically include: when the target intent is a query intent, triggering and executing the query intent, and returning the corresponding query result after querying the database according to the query intent, and using the query result as the target intent execution result; or, when the target intent is an action intent, triggering and executing the action intent, obtaining the execution result of the action intent after executing the action, and using the execution result as the target intent execution result. The intent slot refers to a parameter position in an intent with parameters. An intent can have one or more intent slots. For example, in the intent "user wants to book a flight", possible intent slots are "time", "departure place", and "destination". Intent slots can be assigned values ​​in specific user questions. For example. In the sentence "I want to depart from place A to place B today", the slot values ​​of the intent "user wants to book a flight" are "today", "place A", and "place B" respectively.

[0063] In addition, in the process of generating corresponding current round answer information based on the target intent execution result, the preferred answer information stored in the intent database in the preset general intent processing module can be queried based on the user's target intent; when the preferred answer method corresponding to the user's target intent is queried, the target intent execution result is processed according to the preferred answer method to generate corresponding current round answer information; wherein, the preferred answer method includes the layout, words and elements to be displayed of the answer; when the preferred answer method corresponding to the user's target intent is not queried, the preset natural language large model is called to process the target intent execution result to generate corresponding current round answer information.

[0064] The context-based conversational interaction processing method described in the present invention can be implemented by a conversational robot (conversational program or software) to realize multi-faceted mixed interactions in a specified scenario, and different interaction processes can be freely interspersed without customized model training for the specified scenario. The process of the conversational robot processing user questions includes: parsing user intentions, checking and executing intentions, and answering user questions.

[0065] For example, in a complete embodiment of the present invention: the context database and intention database set by the user are as follows.

[0066] Table 1 Context database

[0067]

[0068] Table 2 Intent database

[0069]

[0070] The user is currently in the context of the "personnel query interface" and asks "I want to check the details of Zhang San." Now the process of parsing user intent, checking and executing intent, and answering user questions begins. In the process of parsing user intent, the interface information shows that the user is in the context of the "personnel query interface". By querying the context database, it can be learned that the unique intents of the current context are "check attendance rate" and "check the details of a specific person", and the general intent is "return to the start interface". Through prompt engineering, vectorized matching and other means of analysis, it can be learned that among these three intents, the user's accurate intent is "check the details of a specific person". The intent is clear, and the next step is entered. In the process of checking and executing intent, the execution premise of the intent is first checked, and the premise is that all necessary slots need to be filled. There is only one slot for this intent, which is name, which means "the name to be queried" and is a necessary slot. Using prompt engineering and other means, it can be learned that in the user's question, the value of this slot is "Zhang San". Therefore, the slot is filled and the intent can be executed. Then the target intent is executed. The action that this intent needs to execute is to call the relevant API. Therefore, the conversation robot calls the relevant API, and the parameter is name=Zhang San. Get the result. Assume that the result is {age: 30, gender: "male"}. Enter the result into the next step. In the process of answering user questions, the target intent does not prefer the return form, so the big model is used to organize the answer by itself. Through the big model prompt engineering and other methods, the above execution results are organized into natural language answers. For example, it can be organized into "Zhang San, male, 30 years old." The organization is completed and the result is returned to the user. The whole process ends.

[0071] The context environment-based dialogue interaction processing method provided by the present invention obtains the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; determines the context environment information corresponding to the current round question information based on the current round question information input by the user to the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; performs user intention analysis based on the current round question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intention; when it is determined that the conditions required to execute the target intention are currently met, triggers the execution of the target intention to obtain the corresponding target intention execution result, and generates the corresponding current round answer information based on the target intention execution result, which can effectively improve the dialogue interaction efficiency and accuracy, thereby enhancing the user's dialogue interaction experience.

[0072] The following is a description of a conversation interaction processing device based on a context environment provided by the present invention. The conversation interaction processing device based on a context environment described below and the conversation interaction processing method based on a context environment described above can be referred to in correspondence with each other. Figure 3 As shown, it is a schematic diagram of the structure of the context-based dialogue interaction processing device provided by the present invention. The context-based dialogue interaction processing device of the present invention specifically includes the following parts:

[0073] The intention parsing unit 301 is used to obtain the current round question information of the current page input by the user to the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system; determine the context environment information corresponding to the current round question information based on the current round question information of the current page input by the user to the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system; perform user intention parsing based on the current round question information, the historical conversation information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intention.

[0074] The intention execution and result return unit 302 is used to trigger the execution of the target intention when it is determined that the conditions required to execute the target intention are currently met, obtain the corresponding target intention execution result, and generate the corresponding current round answer information based on the target intention execution result.

[0075] According to a context-based dialogue interaction processing device provided by the present invention, the intention parsing unit is specifically used to:

[0076] The current round of question information input by the user into the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system are input into a preset multimodal large language model for analysis, and the key information features of the current context environment information output by the multimodal large language model are obtained, and the context environment information corresponding to the current round of question information is determined based on the key information features and the context environment information stored in the context environment database in the preset intention processing module. The multimodal large language model is a machine learning model obtained by iterative training based on sample input information and key information feature labels corresponding to the sample input information; or, the identification information set of the current page is extracted through a preset interface of the interactive system, and the identification information set of the current page is obtained based on the extraction, and the identification information set of the current page is compared and analyzed with the context environment information stored in the context environment database in the preset intention processing module to determine the context environment information corresponding to the current round of question information; wherein the context environment information is the environment in which the user is located in the interactive system, including the link and state in which the user is located, and the elements that can be touched or interacted with.

[0077] According to a context-based dialogue interaction processing device provided by the present invention, the intention parsing unit is specifically used to:

[0078] Based on the context environment information corresponding to the question information of the current round and the corresponding relationship stored in the context environment database in the preset general intention processing module, determine the environment-specific intention corresponding to the context environment information corresponding to the question information of the current round, and perform vectorization processing based on the environment-specific intention to obtain an environment-specific intention vector; the corresponding relationship stored in the context environment database is the association relationship between each piece of context environment information and the environment-specific intention corresponding to each piece of context environment information;

[0079] Performing vectorization processing based on the general intent corresponding to the interactive system to obtain a general intent vector;

[0080] Performing vectorization processing based on the current round of question information and the historical conversation information corresponding to the user to obtain a conversation information vector;

[0081] Based on the dialogue information vector, vector matching analysis is performed with the environment-specific intent vector and the general intent vector to obtain a vector matching analysis result; and the user's target intent is selected from the environment-specific intent and the general intent according to the vector matching analysis result.

[0082] According to a context-based dialogue interaction processing device provided by the present invention, the intention execution and result return unit is specifically used to:

[0083] Extracting the intent slot from the current round question information;

[0084] According to the intention slot extracted from the question information of the current round, detecting whether the conditions required for executing the target intention are currently met, and in the case of determining that the conditions required for executing the target intention are currently met, triggering the execution of the target intention to obtain the corresponding target intention execution result;

[0085] Among them, the condition required to execute the target intention is whether all slots corresponding to the target intention have been filled. When it is determined that all slots corresponding to the target intention have been filled, it is determined that the condition required to execute the target intention is met; when it is determined that all slots corresponding to the target intention have not been filled, it is determined that the condition required to execute the target intention is not met.

[0086] According to a context-based dialogue interaction processing device provided by the present invention, the intention execution and result return unit is specifically used to:

[0087] Based on the target intent of the user, the preferred answer information stored in the intent database in the preset general intent processing module is queried. When the preferred answer method corresponding to the target intent of the user is queried, the target intent execution result is processed according to the preferred answer method to generate corresponding current round answer information; wherein the preferred answer method includes the layout, words and elements to be displayed of the answer; when the preferred answer method corresponding to the target intent of the user is not queried, the preset natural language large model is called to process the target intent execution result to generate corresponding current round answer information.

[0088] According to a context-based dialogue interaction processing device provided by the present invention, the target intent includes: query intent and action intent;

[0089] The triggering and executing the target intention to obtain the corresponding target intention execution result specifically includes:

[0090] When the target intent is a query-type intent, the query-type intent is triggered for execution, and the corresponding query result is returned after querying the database according to the query-type intent, and the query result is used as the execution result of the target intent; or, when the target intent is an action-type intent, the action-type intent is triggered for execution, and the execution result of the action-type intent is obtained after executing the action, and the execution result is used as the execution result of the target intent.

[0091] The context environment-based dialogue interaction processing device provided by the present invention obtains the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; determines the context environment information corresponding to the current round question information based on the current round question information input by the user to the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; performs user intention analysis based on the current round question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intention; when it is determined that the conditions required to execute the target intention are currently met, triggers the execution of the target intention to obtain the corresponding target intention execution result, and generates the corresponding current round answer information based on the target intention execution result, which can effectively improve the dialogue interaction efficiency and accuracy, thereby enhancing the user's dialogue interaction experience.

[0092] Figure 4 An example of a physical structure diagram of an electronic device is shown in FIG. Figure 4 As shown, the electronic device (i.e., computer device) may include: a processor (processor) 401, a communication interface (Communications Interface) 404, a memory (memory) 402 and a communication bus 403, wherein the processor 401, the communication interface 404, and the memory 402 communicate with each other through the communication bus 403. The processor 401 can call the logic instructions in the memory 402 to execute a context-based dialogue interaction processing method, which includes: obtaining the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; determining the context environment information corresponding to the current round question information based on the current round question information input by the user to the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; performing user intent analysis based on the current round question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intent; when it is determined that the conditions required to execute the target intent are currently met, triggering the execution of the target intent, obtaining the corresponding target intent execution result, and generating the corresponding current round answer information based on the target intent execution result.

[0093] In addition, the logic instructions in the above-mentioned memory 402 can be implemented in the form of a software functional unit and can be stored in a computer-readable storage medium when it is sold or used as an independent product. Based on this understanding, the technical solution of the present application can be essentially or partly embodied in the form of a software product that contributes to the prior art. The computer software product is stored in a storage medium, including several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, etc. Various media that can store program codes.

[0094] On the other hand, the present application also provides a computer program product, which includes a computer program, which can be stored on a computer-readable storage medium. When the computer program is executed by a processor, the computer can execute the context-based dialogue interaction processing method provided by the above methods, the method including: obtaining the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; based on the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system, determining the context environment information corresponding to the current round question information; based on the current round question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round question information, performing user intent analysis to obtain the user's target intent; when it is determined that the conditions required to execute the target intent are currently met, triggering the execution of the target intent, obtaining the corresponding target intent execution result, and generating the corresponding current round answer information based on the target intent execution result.

[0095] On the other hand, the present application also provides a computer-readable storage medium, the computer-readable storage medium includes a stored program, wherein the program, when running, executes the context-based dialogue interaction processing method provided by the above-mentioned methods, the method comprising: obtaining the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; based on the current round question information of the current page input by the user to the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system, determining the context environment information corresponding to the current round question information; performing user intent analysis based on the current round question information, the historical dialogue information corresponding to the user, and the context environment information corresponding to the current round question information to obtain the user's target intent; when it is determined that the conditions required to execute the target intent are currently met, triggering the execution of the target intent, obtaining the corresponding target intent execution result, and generating the corresponding current round answer information based on the target intent execution result.

[0096] The device embodiments described above are merely illustrative, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative labor.

[0097] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0098] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for processing dialog interaction based on context environment, characterized in that: include: Acquire current round question information input by a user into a current page of an interactive system, historical conversation information corresponding to the user, and current page information of the interactive system; Determining context information corresponding to the current round of question information based on the current round of question information input by the user into the current page of the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system; Analyzing the user's intention based on the current round of question information, historical conversation information corresponding to the user, and context information corresponding to the current round of question information to obtain the user's target intention; The performing user intention analysis based on the current round of question information, the historical conversation information corresponding to the user, and the context environment information corresponding to the current round of question information to obtain the user's target intention specifically includes: Based on the context environment information corresponding to the question information of the current round and the corresponding relationship stored in the context environment database in the preset general intention processing module, determine the environment-specific intent corresponding to the context environment information corresponding to the question information of the current round, perform vectorization processing based on the environment-specific intent, and obtain an environment-specific intent vector; when the context environment information is a start interface, the environment-specific intent corresponding to the context environment information is to query personnel-related information or to query project-related information; when the context environment information is a personnel query interface, the environment-specific intent corresponding to the context environment information is to view attendance rate or to view detailed information of a specific person; when the context environment information is a project query interface, the environment-specific intent corresponding to the context environment information is to view overdue projects or to view the overall progress of the project; The corresponding relationship stored in the context environment database is the association relationship between each context environment information and the environment-specific intention corresponding to each context environment information; the context environment information is the environment in which the user is located in the interactive system, including the link, state, and elements that can be touched or interacted with; Performing vectorization processing based on the general intent corresponding to the interactive system to obtain a general intent vector; the general intent is to exit the current interface or return to the main menu; Performing vectorization processing based on the current round of question information and the historical conversation information corresponding to the user to obtain a conversation information vector; Performing vector matching analysis on the dialogue information vector, the environment-specific intent vector and the general intent vector respectively to obtain a vector matching analysis result; and selecting the target intent of the user from the environment-specific intent and the general intent according to the vector matching analysis result; When it is determined that the conditions required to execute the target intention are currently met, the execution of the target intention is triggered, the corresponding target intention execution result is obtained, and the corresponding current round answer information is generated based on the target intention execution result.

2. The context-based dialogue interaction processing method according to claim 1, characterized in that: The determining, based on the current round of question information input by the user into the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system, context information corresponding to the current round of question information specifically includes: The current round of question information input by the user into the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system are input into a preset multimodal large language model for analysis, and the key information features of the current context environment information output by the multimodal large language model are obtained, and the context environment information corresponding to the current round of question information is determined based on the key information features and the context environment information stored in the context environment database in the preset intention processing module. The multimodal large language model is a machine learning model obtained by iterative training based on sample input information and key information feature labels corresponding to the sample input information; or, the identification information set of the current page is extracted through a preset interface of the interactive system, and the identification information set of the current page is obtained based on the extraction, and the identification information set of the current page is compared and analyzed with the context environment information stored in the context environment database in the preset intention processing module to determine the context environment information corresponding to the current round of question information.

3. The context-based dialogue interaction processing method according to claim 1, characterized in that: When it is determined that the conditions required for executing the target intention are currently met, triggering the execution of the target intention and obtaining the corresponding target intention execution result specifically includes: Extracting the intent slot from the current round question information; According to the intention slot extracted from the question information of the current round, detecting whether the conditions required for executing the target intention are currently met, and in the case of determining that the conditions required for executing the target intention are currently met, triggering the execution of the target intention to obtain the corresponding target intention execution result; Among them, the condition required to execute the target intention is whether all slots corresponding to the target intention have been filled. When it is determined that all slots corresponding to the target intention have been filled, it is determined that the condition required to execute the target intention is met; when it is determined that all slots corresponding to the target intention have not been filled, it is determined that the condition required to execute the target intention is not met.

4. The context-based dialogue interaction processing method according to claim 1, characterized in that: The generating corresponding current round answer information based on the target intention execution result specifically includes: Based on the target intent of the user, the preferred answer information stored in the intent database in the preset general intent processing module is queried. When the preferred answer method corresponding to the target intent of the user is queried, the target intent execution result is processed according to the preferred answer method to generate corresponding current round answer information; wherein the preferred answer method includes the layout, words and elements to be displayed of the answer; when the preferred answer method corresponding to the target intent of the user is not queried, the preset natural language large model is called to process the target intent execution result to generate corresponding current round answer information.

5. The context-based dialogue interaction processing method according to claim 3, characterized in that: The target intent includes: query intent and action intent; The triggering and executing the target intention to obtain the corresponding target intention execution result specifically includes: When the target intent is a query-type intent, the query-type intent is triggered for execution, and the corresponding query result is returned after querying the database according to the query-type intent, and the query result is used as the execution result of the target intent; or, when the target intent is an action-type intent, the action-type intent is triggered for execution, and the execution result of the action-type intent is obtained after executing the action, and the execution result is used as the execution result of the target intent.

6. A conversation interaction processing device based on context environment, characterized in that: include: An intention parsing unit, used to obtain the current round of question information input by the user into the current page of the interactive system, the historical dialogue information corresponding to the user, and the current page information of the interactive system; Determining context information corresponding to the current round of question information based on the current round of question information input by the user into the current page of the interactive system, the historical conversation information corresponding to the user, and the current page information of the interactive system; Analyzing the user's intention based on the current round of question information, historical conversation information corresponding to the user, and context information corresponding to the current round of question information to obtain the user's target intention; The intention parsing unit is specifically used to: determine the environment-specific intention corresponding to the context environment information corresponding to the question information of the current round based on the context environment information corresponding to the question information of the current round and the corresponding relationship stored in the context environment database in the preset general intention processing module, and perform vectorization processing based on the environment-specific intention to obtain an environment-specific intention vector; when the context environment information is the start interface, the environment-specific intention corresponding to the context environment information is the query person related information or the query item related information; When the context environment information is a personnel query interface, the environment-specific intent corresponding to the context environment information is to view attendance rate or detailed information of a specific person; when the context environment information is a project query interface, the environment-specific intent corresponding to the context environment information is to view overdue projects or to view the overall progress of projects; the corresponding relationship stored in the context environment database is the association relationship between each context environment information and the environment-specific intent corresponding to each context environment information; the context environment information is the environment in which the user is located in the interactive system, including the link and status, and the elements that can be touched or interacted with; vectorization processing is performed based on the general intent corresponding to the interactive system to obtain a general intent vector; the general intent is to exit the current interface or return to the main menu; Performing vectorization processing based on the current round of question information and the historical conversation information corresponding to the user to obtain a conversation information vector; Performing vector matching analysis on the dialogue information vector, the environment-specific intent vector and the general intent vector to obtain a vector matching analysis result; selecting the user's target intention from the environment-specific intention and the general intention according to the vector matching analysis result; The intention execution and result return unit is used to trigger the execution of the target intention when it is determined that the conditions required to execute the target intention are currently met, obtain the corresponding target intention execution result, and generate the corresponding current round answer information based on the target intention execution result.

7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and running on the processor, characterized in that: When the processor executes the program, the context-based dialogue interaction processing method as described in any one of claims 1 to 5 is implemented.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the context-based dialogue interaction processing method as described in any one of claims 1 to 5 is implemented.

9. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the context-based dialogue interaction processing method as described in any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Natural language understanding processing method, device and equipment based on context inference

    CN112632961A

  • Conversation method and model based on interactive interface

    CN114895999A

  • Intelligent dialogue system and method based on large model and electronic equipment

    CN117453899A

  • Method, server and system for answering questions based on text

    CN118013981A

  • Dynamic multi-intention semantic understanding method and device, computer equipment and readable storage medium

    CN118467680A