Human-computer dialogue methods, dialogue central control server, dialogue engine and storage medium
By introducing a cache mechanism in the human-computer dialogue system, using the query information input by the user to find matching cache data in the cache model processing result data of the dialogue engine, the problem of high GPU resource cost and slow response speed in the large-model human-computer dialogue system is solved, and more efficient GPU resource utilization and response speed are achieved.
Patent Information
- Application Number
- PCT/CN2024/117831
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-12-01
- Filing Date
- 2024-09-09
- Publication Date
- 2025-06-05
AI Technical Summary
The human-computer dialogue system based on large models has not been effectively solved in the problems of high GPU resource costs and slow response speed, resulting in the inability to return reply information in real-time in real-time applications.
By introducing a cache mechanism in the dialogue central control server and dialogue engine, the query information entered by the user is used to find matching cache data in the dialogue engine's cache model processing result data, thereby reducing the call of the dialogue model, reducing GPU resource costs and improving response speed.
Through the cache mechanism, the call of dialogue models is reduced in the human-computer dialogue system, the cost of GPU resources is reduced, the response speed is improved, and the system's concurrent processing capability and overall response speed are improved.
Smart Images

Figure CN2024117831_05062025_PF_FP_ABST
Abstract
Description
Human-computer dialogue method, dialogue control server, dialogue engine and storage medium
[0001] This disclosure claims priority to the Chinese patent application filed with the China Patent Office on December 1, 2023, with application number 202311646407.7 and application name “Human-computer dialogue method, dialogue central control server, dialogue engine and storage medium”, the entire content of which is incorporated by reference into this disclosure. Technical Field
[0002] The present disclosure relates to computer technology, and in particular to a human-computer dialogue method, a dialogue control server, a dialogue engine, and a storage medium. Background Art
[0003] With the rapid development of large model technology, it is being applied in an increasing number of fields. Large language models (LLMs), among other large models, are also being used in human-computer interaction scenarios, such as intelligent conversational robots. When large models are applied to human-computer interaction systems, they must call upon the large model to generate responses during each round of conversation with the user.
[0004] However, the applicant found that the following technical problems were encountered in the process of applying large model technology to human-computer dialogue systems: First, the resource cost of the Graphics Processing Unit (GPU) is high. Since large models have a large number of parameters, their training and reasoning processes need to rely on a large number of GPU resources. Second, the response speed is slow. In many real-time applications (such as dialogue systems, intelligent assistants, etc.), the response time requirements for applications are very high, usually at the millisecond level. Since large models have a large number of parameters, their reasoning process takes a long time (usually at the second level), and the response speed is slow, resulting in the inability to return reply information in real time.
[0005] Summary of the Invention
[0006] The present disclosure provides a human-computer dialogue method, a dialogue central control server, a dialogue engine, and a storage medium to solve the problems of high GPU resource cost and slow response speed in a human-computer dialogue system based on a large model.
[0007] In a first aspect, the present disclosure provides a human-computer dialogue method, which is applied to a dialogue control server, including: based on first query information input by a user, searching for cached data matching the first query information in model processing result data cached by at least one dialogue engine, and obtaining cached search results of each dialogue engine for the first query information; according to the cached search results of each dialogue engine for the first query information, selecting one of the dialogue engines as the target engine, and controlling the target engine to generate reply information according to the cached search results of the first query information.
[0008] In a second aspect, the present disclosure provides a human-computer dialogue method, which is applied to a dialogue engine, and the method includes: receiving a first query message sent by a dialogue central control server; searching for cached data matching the first query message in the cached model processing result data to obtain a cached search result of the first query message; returning the cached search result of the first query message to the dialogue central control server; and generating a reply message based on the cached search result of the first query message based on the control information of the dialogue central control server.
[0009] In a third aspect, the present disclosure provides a human-computer dialogue method, which is applied to a dialogue control server in an intelligent customer service system, the method comprising: receiving a first question that a user needs to consult, searching for cached data related to the first question in model processing result data cached by at least one dialogue engine, and obtaining cached search results of each dialogue engine for the first question; based on the cached search results of each dialogue engine for the first question, selecting one of the dialogue engines as a target engine, and controlling the target engine to generate reply information based on the cached search results for the first question.
[0010] In a fourth aspect, the present disclosure provides a dialogue control server, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the dialogue control server to execute the method described in the first or third aspect above.
[0011] In a fifth aspect, the present disclosure provides a dialogue engine comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the dialogue engine to execute the method described in the second aspect above.
[0012] In a sixth aspect, the present disclosure provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method described in the first aspect, the second aspect, or the third aspect is implemented.
[0013] In a seventh aspect, the present disclosure provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in the first aspect, the second aspect, or the third aspect.
[0014] The present disclosure provides a human-computer dialogue method, a dialogue control server, a dialogue engine, and a storage medium. When conducting a human-computer dialogue, the dialogue control server searches for cached data matching the first query information in the model processing result data cached by each dialogue engine based on the first query information input by the user, and obtains the cached search results of each dialogue engine for the first query information; based on the cached search results of each dialogue engine for the first query information, one of the dialogue engines is selected as the target engine, and the target engine is controlled to generate reply information based on the cached search results for the first query information. By caching the model processing result data of historical query information in each dialogue engine, when encountering the same or similar queries in the subsequent human-computer dialogue process, there is no need to call the dialogue model, and the reply information is generated based on the cached data of the model processing results hit by the current query information (i.e., the first query information). This can reduce the calling of the dialogue model, reduce the GPU resource cost of the human-computer dialogue, and improve the response speed of the human-computer dialogue, thereby improving the concurrent processing capability of the human-computer dialogue system and the overall response speed of the human-computer dialogue system. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present disclosure and, together with the description, serve to explain the principles of the present disclosure.
[0016] FIG1 is a schematic diagram of an example system architecture applicable to the present disclosure;
[0017] FIG2 is an exemplary framework diagram of a two-stage scheduling solution for a dialogue system provided by an exemplary embodiment of the present disclosure;
[0018] FIG3 is a flow chart of a human-computer dialogue method provided by an exemplary embodiment of the present disclosure;
[0019] FIG4 is a flowchart of generating reply information based on cache search results provided by an exemplary embodiment of the present disclosure;
[0020] FIG5 is a flow chart of a human-computer dialogue method provided by another exemplary embodiment of the present disclosure;
[0021] FIG6 is an exemplary framework diagram of a two-stage scheduling solution for a dialogue system provided by another exemplary embodiment of the present disclosure;
[0022] FIG7 is a flowchart of a human-computer dialogue of an intelligent customer service system provided by an exemplary embodiment of the present disclosure;
[0023] FIG8 is a schematic diagram of the structure of a dialogue control server provided in an embodiment of the present disclosure;
[0024] FIG9 is a schematic diagram of the structure of a conversation engine provided by an embodiment of the present disclosure.
[0025] The above drawings illustrate specific embodiments of the present disclosure, which will be described in more detail below. These drawings and textual descriptions are not intended to limit the scope of the present disclosure in any way, but rather to illustrate the concepts of the present disclosure to those skilled in the art by reference to specific embodiments. DETAILED DESCRIPTION
[0026] Exemplary embodiments will be described in detail herein, with examples illustrated in the accompanying drawings. In the following description, when referring to the drawings, identical numerals in different figures represent identical or similar elements, unless otherwise indicated. The embodiments described in the following exemplary embodiments are not intended to represent all possible embodiments consistent with the present disclosure. Rather, they are merely examples of apparatus and methods consistent with certain aspects of the present disclosure, as detailed in the appended claims.
[0027] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this disclosure are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant laws, regulations and standards, and corresponding operation entrances must be provided for users to choose to authorize or refuse.
[0028] First, the terms involved in this disclosure are explained:
[0029] Large Language Model (LLM): Also known as a large-scale language model or large language model, LLM is a model based on machine learning and natural language processing technologies. It is trained on large amounts of text data to learn the ability to support human language understanding and generation. The core concept of LLM is to learn the patterns and linguistic structures of natural language through large-scale unsupervised training, which can simulate the human language cognition and generation process to a certain extent. Compared with traditional natural language processing (NLP) models, LLM can better understand and generate natural text, while also demonstrating certain logical thinking and reasoning capabilities.
[0030] Human-computer dialogue system: A human-computer dialogue system based on natural language understanding technology and dialogue management technology, which is widely used in scenarios such as intelligent customer service, intelligent outbound calls, online education, office software, and e-commerce.
[0031] Dialogue Central Control Server: Also known as the Dialogue Central Control. Dialogue systems need to support multiple dialogue engines, including but not limited to engines for task-based dialogue, frequently asked questions (FAQ), graph Q&A, table Q&A, document Q&A, and casual chat. These dialogue engines support different use cases. Task-based dialogue engines primarily target multi-round, task-based Q&A scenarios, such as checking the weather, booking hotels, and recharging data, which can be abstracted into intents and slots. FAQ engines support dialogue scenarios based on question-and-answer pairs. Graph Q&A engines support dialogue scenarios based on knowledge graph data. Table Q&A engines support dialogue scenarios based on tabular data. Document Q&A engines support dialogue scenarios based on document reading comprehension. Chat supports dialogue scenarios based on open domains with less business relevance, such as small talk and casual conversation. Human-computer dialogue often requires the coordination of multiple dialogue engines. The Dialogue Central Control, which implements the scheduling and ordering capabilities of dialogue engines in a multi-engine dialogue system, is a core module in this system and plays a critical role in improving dialogue capabilities.
[0032] Cache: refers to a memory that can perform high-speed data exchange. In layman's terms, it is to increase access speed by storing data in memory in advance.
[0033] Multimodal tasks: refers to downstream tasks whose input and output data involve multiple modal data such as images and text, such as visual question answering tasks, image description tasks, visual implication tasks, referential expression and understanding tasks, image generation tasks, etc.
[0034] Multimodal pre-trained model: refers to a pre-trained model whose input and output data involve multiple modal data such as images and text. After fine-tuning and training, it can be applied to multimodal task processing.
[0035] Large models refer to deep learning models with large-scale model parameters, typically containing hundreds of millions, tens of billions, or even hundreds of billions of model parameters. Large models, also known as foundation models (FMs), are pre-trained on large-scale unlabeled corpora, producing pre-trained models with parameters exceeding 100 million. These models are adaptable to a wide range of downstream tasks and exhibit good generalization capabilities. Examples include large language models (LLMs) and multi-modal pre-training models.
[0036] When large models are used in practice, only a small number of samples are needed to fine-tune the pre-trained model and it can be applied to different tasks. Large models can be widely used in natural language processing, computer vision and other fields. Specifically, they can be applied to computer vision tasks such as visual question answering (VQA), image caption (IC), and image generation, as well as natural language processing tasks such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, and intelligent design.
[0037] With the rapid development of large model technology, it is being applied in an increasing number of fields. Large language models (LLMs), among other large models, are also being used in human-computer interaction scenarios, such as intelligent conversational robots. When large models are applied to human-computer interaction systems, they must call upon the large model to generate responses during each round of conversation with the user.
[0038] However, the applicant found that the following technical problems were encountered in the process of applying large model technology to human-computer dialogue systems: First, the resource cost of the Graphics Processing Unit (GPU) is high. Since large models have a large number of parameters, their training and reasoning processes need to rely on a large number of GPU resources. Second, the response speed is slow. In many real-time applications (such as dialogue systems, intelligent assistants, etc.), the response time requirements for applications are very high, usually at the millisecond level. Since large models have a large number of parameters, their reasoning process takes a long time (usually at the second level), and the response speed is slow, resulting in the inability to return reply information in real time.
[0039] The present disclosure provides a human-computer dialogue method. When conducting a human-computer dialogue, the dialogue control server searches for cached data matching the first query information in the model processing result data cached by each dialogue engine based on the first query information input by the user, and obtains the cached search results of each dialogue engine for the first query information; according to the cached search results of each dialogue engine for the first query information, one of the dialogue engines is selected as the target engine, and the target engine is controlled to generate reply information according to the cached search results for the first query information; by caching the model processing result data of historical query information in each dialogue engine, when encountering the same or similar query in the subsequent human-computer dialogue process, there is no need to call the dialogue model, and the reply information is generated based on the cached data of the model processing result hit by the current query information (i.e., the first query information). This can reduce the calling of the dialogue model, reduce the GPU resource cost of the human-computer dialogue, improve the response speed of the human-computer dialogue, and thereby improve the concurrent processing capability of the human-computer dialogue system, and improve the overall response speed of the human-computer dialogue system.
[0040] Figure 1 is a schematic diagram of an example system architecture applicable to the present disclosure. As shown in Figure 1, the system architecture includes a human-computer dialogue system that provides human-computer dialogue services and a terminal device used by a user.
[0041] Among them, the user terminal device can be various terminal devices that can obtain user input, wherein the user input can be in the form of text or voice. The user terminal device can be a device with a screen or a device with voice interaction function. It includes but is not limited to: smart mobile terminals, smart home devices, wearable devices, personal computers (PCs), etc. Among them, smart mobile devices can include mobile phones, tablets, laptops, Internet cars, etc. Smart home devices can include smart home appliances such as smart TVs, smart air conditioners, smart refrigerators, etc. Wearable devices can include smart watches, smart glasses, smart bracelets, virtual reality (VR) devices, augmented reality (AR) devices, mixed reality devices (i.e., devices that can support virtual reality and augmented reality), etc. The user can obtain the query information input by the user through the user terminal device and send the query information (Query) input by the user to the dialogue control server.
[0042] The human-computer dialogue system includes a dialogue control server and at least one dialogue engine. These dialogue engines support different use cases, including but not limited to engines for implementing various dialogue tasks, such as task-based dialogue, FAQs, graph Q&A, table Q&A, document Q&A, and casual chatting. The dialogue engines run dialogue models that implement the corresponding dialogue capabilities. These dialogue models can be implemented using large models such as large language models, or other smaller models, without specific limitations here. For example, the task-based dialogue engine runs a task-based multi-turn dialogue model, primarily targeting multi-turn task-based Q&A scenarios such as checking the weather, booking hotels, and recharging data, which can be abstracted into intents and slots. The FAQ engine runs a knowledge Q&A model to implement dialogue scenarios based on question and answer pairs. The graph Q&A engine runs a knowledge graph Q&A model to implement dialogue scenarios based on knowledge graph data. The table Q&A engine runs a table Q&A model to implement dialogue scenarios based on table data. The document Q&A engine runs a knowledge Q&A model to implement dialogue scenarios based on document reading comprehension. Other conversation engines, such as small talk, can also be used to support open-domain conversation scenarios with less business relevance, such as small talk and small talk. The conversation engines supported by the conversation system can be configured and adjusted based on the needs of the actual use case, and are not specifically limited here.
[0043] Human-computer dialogue often requires the coordination of multiple dialogue engines. The dialogue control server, a core module of a multi-engine dialogue system, implements the scheduling and ranking capabilities of each dialogue engine. It plays a crucial role in improving dialogue capabilities. The dialogue control server receives user-entered query information, schedules and ranks multiple dialogue engines, selects a dialogue engine as the target engine for the current round of dialogue, and controls the target engine to generate a response. The dialogue control server then returns the response to the user's device.
[0044] The above-mentioned dialogue control server, dialogue engine, etc. can be set up on a single server, or on a server group consisting of multiple servers, or on a cloud server.
[0045] It should be understood that the number of end-side devices, dialogue control servers, and dialogue engines used by users in Figure 1 is only illustrative. Depending on implementation requirements, the human-computer dialogue system can have any number of end-side devices, dialogue control servers, and dialogue engines.
[0046] Based on the system architecture of Figure 1, when conducting a human-computer dialogue, the end-side device sends the first query information input by the user to the dialogue control server. The dialogue control server schedules at least one dialogue engine based on the first query information input by the user, searches for cached data that matches the first query information in the model processing result data cached by at least one dialogue engine, and obtains the cached search results of each dialogue engine for the first query information. Furthermore, the dialogue control server selects one of the dialogue engines as the target engine based on the cached search results of each dialogue engine for the first query information, and controls the target engine to generate reply information based on the cached search results for the first query information. Furthermore, the dialogue control server returns the reply information to the end-side device.
[0047] In scenarios where the dialogue engine uses large models such as LLM, the dialogue control server in the dialogue system adopts a two-stage scheduling scheme for multiple dialogue engines: the first stage is the diversion stage, and the second stage is the execution stage. During the diversion stage, the dialogue control server calls the diversion interfaces of multiple dialogue engines in parallel, causing the dialogue engines to run the diversion interface implementation program, perform the diversion stage processing on the current query information, and obtain the diversion result. Furthermore, the dialogue control server ranks the diversion results of each dialogue engine according to a preset diversion ranking strategy to obtain the diversion ranking result. The diversion interface is a natural language understanding (NLU) interface provided by the dialogue engine to the dialogue control server, which is used to implement the diversion processing of query information and obtain the diversion result.
[0048] During the execution phase, the conversation control server selects a conversation engine as the target engine based on the sorting results. It then calls the target engine's execution interface, causing it to run the implementation of the execution interface and generate a response based on the results of the query. Furthermore, the conversation control server notifies each conversation engine of the execution results of the current round of conversations, which may include the query information, the selected target engine for diversion, and the generated response information.
[0049] For example, Figure 2 is a schematic diagram of a two-stage scheduling scheme for a dialog system provided by this embodiment. As shown in Figure 2, in this two-stage scheduling scheme for a dialog system, the dialog control server performs parallel traffic diversion to each dialog engine based on the query and context information input by the user in the current round. Specifically, the dialog control server concurrently calls the diversion interface of each dialog engine and transmits the query and context information to each dialog engine as input parameters of the diversion interface. Each dialog engine performs the diversion phase based on the query and context information, obtains a diversion result, and returns the diversion result to the dialog control server. The dialog control server then sorts the diversion results of each dialog engine based on a given sorting strategy, decomposes the sorted diversion results, and selects a dialog engine as the target engine. Furthermore, the dialog control server schedules the target engine to perform the execution phase and generates a reply based on the intermediate results of the diversion phase. Specifically, the conversation control server can call the target engine's execution interface and transmit the intermediate results of the diversion phase as input parameters to the target engine, causing the target engine to run the implementation program of the execution interface to generate response information based on the intermediate results of the diversion phase and return the response information to the conversation control server. Furthermore, the conversation control server updates the context information stored in the conversation control server based on the query information and response information of the current round of conversation. The conversation control server can also transmit the execution results of the current round of conversation (including but not limited to query information and response information) to each conversation engine other than the target engine, so that each conversation engine can update the context information stored within the engine.
[0050] The dialogue control server and each dialogue engine respectively store and maintain a copy of context information. The context information stored by the dialogue control server is called global context information, and the context information stored by each dialogue engine is called single-engine context information (or local context information).
[0051] It should be noted that, based on the needs of actual application scenarios, the dialogue system can be configured to support multiple different dialogue engines. The number and types of dialogue engines supported by the dialogue system can vary for different application scenarios. Figure 2 uses the dialogue system's support for task-based multi-turn dialogue engines, table question-and-answer engines, knowledge question-and-answer engines (which can simultaneously support FAQ and document question-and-answer), and open domain dialogue engines (which support casual conversation) as examples to illustrate the process framework of the two-stage scheduling solution for multiple engines in the dialogue system. There is no limit on the number and types of dialogue engines that the dialogue system can support. The purpose of open domain dialogue is not to try to complete a task, but to chat with users without task or domain restrictions, and is usually data-driven.
[0052] In addition, different dialogue engines may perform different processing during the diversion phase, and may also perform different processing during the execution phase. In the field of intelligent customer service, most dialogue robots perform questions and answers based on existing static knowledge data, and can use knowledge question and answer engines, such as document question and answer engines, FAQ engines based on question and answer pairs, etc. The diversion result is a candidate set recalled based on query information. For document question and answer engines, the candidate set contains at least one document, and the knowledge type corresponding to the diversion result is a document (denoted as Doc type). For FAQ engines, the candidate set contains at least one question and answer pair, and the knowledge type corresponding to the diversion result is a question and answer pair (denoted as FAQ type). In the diversion phase, candidate sets (such as candidate documents, candidate question and answer pairs, etc.) are obtained based on query information and context information. In the execution phase, knowledge question and answer models (such as document question and answer LLM, FAQ question and answer LLM) are used to generate reply information based on query information, context information, and candidate sets.
[0053] For task-based multi-turn dialogue scenarios, such as question-and-answer scenarios for process-related knowledge, a task-based multi-turn dialogue engine can be used. In the diversion phase, a task-based multi-turn dialogue model (such as LLM) is used to perform natural language understanding (NLU) based on query information and context information to obtain the intent and / or entity information of this round of dialogue. The knowledge types corresponding to the diversion results can be divided into the knowledge type that directly outputs the result after hitting the preset intent (denoted as Ds non-other or Ds non-Anything Else), and the knowledge type that results after hitting "other" without hitting the intent (denoted as Ds other or Ds Anything Else). In the execution phase, response information is generated based on the query information, context information, intent and / or entity information.
[0054] For table question and answer scenarios for table data, a table question and answer engine can be used. In the diversion stage, a table question and answer model (such as LLM) is used to convert the query information into a structured query statement (Structured Query Language, referred to as SQL) based on the query information and context information. The diversion result is SQL, and the knowledge type corresponding to the diversion result is recorded as a table query (Table). In the execution stage, the SQL generated in the diversion stage is used to query the table data to obtain the SQL query result, and the reply information is generated according to the query information, context information and SQL query result. In some optional embodiments, in the execution stage, the table question and answer engine can also use a large model (such as LLM) to generate reply information according to the query information, context information and SQL query results.
[0055] For open-domain conversation scenarios like small talk, an open-domain conversation engine can be used. This engine can be independent of parallel splitting. Based on the splitting results of each conversation engine in parallel splitting, if no other conversation engines match, the open-domain conversation engine is selected as the target engine. During the execution phase, an open-domain search is performed based on the query and context information to obtain open-domain search results. An open-domain conversation model (such as an LLM) is used to generate a response based on the query, context, and open-domain search results.
[0056] Based on the priority of the knowledge type corresponding to the diversion results of each dialogue engine, the result confidence, etc., the diversion results of each dialogue engine are sorted, and a dialogue engine is selected as the target engine according to the sorting result of the diversion results.
[0057] Exemplarily, the priority of each knowledge type can be configured as: FAQ>Ds Non-Anything Else>Table>Doc>Ds Anything Else>Open Domain Search. In addition, the priority of the knowledge type corresponding to the diversion results returned by each dialogue engine in the diversion stage can be configured and adjusted according to the needs of the actual application scenario, which is not specifically limited here. Optionally, in the preset diversion sorting strategy, the diversion results can be sorted according to the priority of the knowledge type corresponding to the diversion results returned by each dialogue engine, with the knowledge type with a higher priority being ranked first and the knowledge type with a lower priority being ranked last.
[0058] The confidence level of the diversion result is the confidence level given by the dialogue engine for the diversion result, reflecting the reliability of the diversion result. The dialogue control server can also set a preset diversion sorting strategy based on the priority of the knowledge types corresponding to the diversion results of each dialogue engine and the confidence level of the result. This strategy can be configured and adjusted based on the actual application of the dialogue system and is not specifically limited here.
[0059] In some embodiments, a two-stage scheduling scheme based on multiple dialogue engines is used. The dialogue control server is responsible for parallel offloading and execution of each dialogue engine. Each dialogue engine caches model processing results and is responsible for recalling and writing to the cache. Considering that some dialogue engines, such as task-based multi-turn dialogue engines and table-based question-answering engines, invoke dialogue models during the offload phase, if the dialogue control server unconditionally offloads dialogue models to each dialogue engine in parallel, that is, invokes each dialogue engine's offload phase in parallel, regardless of which target engine's cache is ultimately hit, invalid calls to dialogue models will inevitably occur during this parallel offload phase, affecting dialogue efficiency. For example, if both a task-based multi-turn dialogue engine and a table-based question-answering engine invoke a large model during the offload phase, and if a subsequent hit occurs in the task-based multi-turn dialogue engine's cache, the table-based question-answering engine's call to the large model during the offload phase will be invalid and should be avoided.
[0060] To minimize the number of conversation model calls, some embodiments incorporate a caching phase before the diversion phase. The conversation control server caches the diversion results of each conversation engine for query information. For conversation engines that invoke conversation models during the diversion phase, the diversion results include model processing data. During the caching phase, the cache is searched (i.e., recalled) based on the current first query information. Based on the cache search results, if a match is found in a conversation engine's cache (i.e., a match is found in the diversion result cached by that conversation engine), the corresponding conversation engine is scheduled for the post-diversion execution phase based on the matching diversion result.
[0061] In this scenario, if the cache hit is a diversion result returned by a conversational engine with strong recall capabilities (such as a table Q&A engine or a document Q&A engine), and a higher-priority conversational engine (such as an FAQ engine or a task-based multi-turn conversational engine) is configured with relevant knowledge information, it is easy for the engine to misinterpret the result. For example, in one session (denoted as session 1), the user inputs "A insurance." The document Q&A engine is diverted and selects the response, "A insurance is a good insurance company." In another session (denoted as session 2), the user in the first round inputs "apply for insurance," and the system outputs "What kind of insurance do you want to apply for?" The user in the second round inputs "A insurance." The user input in the second round is exactly the same as in session 1, so the document Q&A engine cache for session 1 is hit, and based on the cache hit, the response from session 1 is output: "A insurance is a good insurance company." This clearly does not match the user's intention to apply for insurance. Based on the user's actual intention, the dialogue system should hit the task-based multi-turn dialogue model to generate a response. There is a situation where it is expected to hit the cache of the task-based multi-turn dialogue model, but actually hits the cache of the document question-answering engine.
[0062] In an embodiment of the present disclosure, a two-stage scheduling scheme is based on multiple dialogue engines, wherein each dialogue engine caches model processing result data of a dialogue model, and a caching stage is added before the diversion stage. In the caching stage, the dialogue control server searches for cached data matching the first query information in the model processing result data cached by at least one dialogue engine based on the first query information input by the user, and obtains the cached search results of each dialogue engine for the first query information. Based on the cached search results of each dialogue engine for the first query information, two-stage scheduling of multiple dialogue engines is performed, one of the dialogue engines is selected as the target engine, and the target engine is controlled to generate reply information based on the cached search results for the first query information, so as to reduce unnecessary dialogue model calls, improve the response speed of human-computer dialogue, and thereby improve the concurrent processing capability of the human-computer dialogue system, and improve the overall response speed of the human-computer dialogue system.
[0063] The following detailed description of the technical solution of the present disclosure and how the technical solution of the present disclosure solves the above-mentioned technical problems is provided with specific embodiments. The following specific embodiments may be combined with each other, and the same or similar concepts or processes may not be described in detail in some embodiments. The embodiments of the present disclosure will be described below in conjunction with the accompanying drawings.
[0064] FIG3 is a flowchart of a human-computer dialogue method provided by an exemplary embodiment of the present disclosure. The execution subject of this embodiment is the dialogue control server in the aforementioned system architecture. As shown in FIG3, the specific steps of the method are as follows:
[0065] Step S301: Based on the first query information input by the user, cached data matching the first query information is searched in the model processing result data cached by at least one dialogue engine to obtain the cached search results of each dialogue engine for the first query information.
[0066] The first query information refers to the query information currently input by the user. The user input can be in the form of text or voice. The first query information refers to the text content input by the user. The text input by the user can be used directly as the first query information or after being rewritten as the first query information. The voice input by the user can be converted into text and used as the first query information.
[0067] In traditional dialogue-based control scheduling solutions, regardless of the selected dialogue engine, the dialogue engine must invoke the corresponding dialogue model (such as LLM) for inference to generate responses. However, especially when using large models (such as LLM), inference takes a long time (typically seconds) and relies on a large amount of GPU resources, resulting in slow response times in human-machine dialogues.
[0068] In this embodiment, each dialogue engine caches model processing result data for query information. During a human-computer dialogue, based on the first query information input by the user, the dialogue control server first searches the model processing result data cached by at least one dialogue engine for cached data that matches the first query information before scheduling the dialogue engine to run the dialogue model. This server then obtains cached search results for the first query information from each dialogue engine.
[0069] Among them, different caching strategies are configured for different dialogue engines based on the knowledge data characteristics of each dialogue engine. For example, the key of the cache of a task-based multi-round dialogue engine can include query information and the response information of the previous round, and the value can include the processing results of the query information by the task-based multi-round dialogue model, such as intent recognition results, entity recognition results, etc. By adding the response information of the previous round to the cache key, the accuracy of the cache is improved by adding necessary contextual information. For example, the key of the cache of a table question-answering engine can include query information and the SQL of the previous round, and the value includes the processing results of the query information by the table question-answering model, such as the SQL converted into the query information of this round. For example, the key of the cache of a document question-answering engine can include query information and candidate sets, and the value is the response information of the query information.
[0070] The cache search result of the first query information includes a cache hit status and the knowledge type of the hit model processing result. Optionally, the cache search result of the first query information may also include the hit model processing result, confidence level, etc. The information included in the cache search result is not specifically limited here.
[0071] Optionally, the dialogue control server may search for cached data that matches the first query information in the model processing result data of all dialogue engines; or, the dialogue control server may select some dialogue engines for cache search based on factors such as the priority of each dialogue engine, the inference time using the dialogue model, and the amount of GPU resources used, and search for cached data that matches the first query information in the model processing result data cached by at least one selected dialogue engine.
[0072] Step S302: Select one of the dialogue engines as a target engine based on the cache search results of the first query information of each dialogue engine, and control the target engine to generate reply information based on the cache search results of the first query information.
[0073] The dialog control server sorts the cached search results of each dialog engine for the first query based on a preset sorting strategy. Based on the sorting results, the dialog engine with the highest cached search result is selected as the target engine. The target engine is then controlled to generate a response message based on the cached search result of the first query. If a cache hit is found, the target engine uses the model processing result of the hit to generate a response message. If a cache miss is found, the target engine obtains the model processing result of the first query by calling the corresponding dialog model, and then generates a response message based on the model processing result of the first query.
[0074] The method of this embodiment adds a caching mechanism, whereby each dialogue engine caches the model processing result data of query information. During a human-computer dialogue, the dialogue control server searches for cached data matching the first query information among the model processing result data cached by each dialogue engine based on the first query information input by the user, obtaining the cached search results of each dialogue engine for the first query information. Based on the cached search results of each dialogue engine for the first query information, one dialogue engine is selected as the target engine, and the target engine is controlled to generate a response message based on the cached search results for the first query information. By caching the model processing result data of historical query information in each dialogue engine, when the same or similar query is encountered in a subsequent human-computer dialogue, there is no need to call the dialogue model. Instead, a response message is generated based on the cached data of the model processing result matched by the current query information (i.e., the first query information). This can reduce the call of the dialogue model, reduce the GPU resource cost of the human-computer dialogue, and improve the response speed of the human-computer dialogue, thereby improving the concurrent processing capability of the human-computer dialogue system and the overall response speed of the human-computer dialogue system.
[0075] In an optional embodiment, step S301 can be implemented as follows: the conversation control server sends the first query information input by the user to each conversation engine, controls each conversation engine to search cached model processing result data for cached data matching the first query information, and obtains a cached search result for the first query information. The conversation control server receives the cached search result for the first query information returned by each conversation engine.
[0076] Specifically, each dialogue engine provides a cache search state interface to the dialogue control server. The dialogue control server can call the cache search state interface of any dialogue engine and transmit the first query information as an input parameter of the interface to the dialogue engine. This interface controls the dialogue engine to search for cached data matching the first query information in the cached model processing result data, obtain the dialogue engine's cache search result for the first query information, and return the dialogue engine's cache search result for the first query information to the dialogue control server. The dialogue control server receives the cache search results for the first query information returned by each dialogue engine.
[0077] Optionally, when the dialog control server calls the cache search status interface of any dialog engine, it transmits the first query information and the word segmentation result of the first query information to the dialog engine as input parameters of the interface. Based on the first query information and the word segmentation result of the first query information, the dialog engine searches the cached model processing result data for cached data that matches the first query information, thereby obtaining the dialog engine's cache search result for the first query information.
[0078] In another optional embodiment, the dialogue control server may cache the received query information. Step S301 may also be implemented as follows: the dialogue control server searches for similar query information to the first query information entered by the user in the cached historical query information. The dialogue control server sends the first query information and similar query information to each dialogue engine, and controls each dialogue engine to search for cached data matching the first query information or similar query information in the cached model processing result data, thereby obtaining a cached search result for the first query information, thereby improving the cache hit rate of each dialogue engine. The dialogue control server receives the cached search result for the first query information returned by each dialogue engine.
[0079] Specifically, each dialogue engine provides a cache search state interface to the dialogue control server. The dialogue control server calls the cache search state interface of any dialogue engine and transmits the first query information and similar query information as input parameters of the interface to the dialogue engine to control the dialogue engine to run the implementation program of the cache search state interface, thereby searching for cached data that matches the first query information or similar query information in the cached model processing result data, obtaining the cache search result of the first query information, and returning the cache search result of the dialogue engine for the first query information to the dialogue control server. The dialogue control server receives the cache search result for the first query information returned by each dialogue engine. The cache search result returned by the dialogue engine may include the identification information of the dialogue engine. The dialogue control server distinguishes the cache search results returned by different dialogue engines based on the identification information of the dialogue engine.
[0080] Optionally, when the dialog control server calls the cache search status interface of any dialog engine, it transmits the first query information and its word segmentation result, and similar query information and its word segmentation result, as input parameters of the interface to the dialog engine. Based on the first query information and its word segmentation result, and the similar query information and its word segmentation result, the dialog engine searches the cached model processing result data for cached data that matches the first query information or the similar query information, and obtains the cache search result for the first query information.
[0081] Optionally, when the dialogue control server searches for similar query information of the first query information entered by the user in the cached historical query information, it can use the ES (Elastic Search) algorithm, or any other search algorithm or retrieval algorithm based on the similarity of query information, to recall similar query information of the first query information from the cached historical query information. This embodiment is not specifically limited here.
[0082] Optionally, when the dialogue control server searches for similar query information of the first query information input by the user in the cached historical query information, it first uses the ES (Elastic Search) algorithm to recall the first similar query information of the first query information from the cached historical query information; then, it calculates the similarity between the first query information and the first similar query information recalled by ES, sorts the first similar query information based on the similarity, and uses a preset number (such as TOP N) of first similar query information with a higher similarity to the first query information as the similar query information of the first query information determined by the dialogue control server. Wherein, N is a preset number, N is a positive integer, which can be specifically configured according to the actual application scenario. For example, N can take a value of 1, 3, 5, etc., which is not specifically limited here.
[0083] In addition, the dialog control server may also determine second similar query information whose similarity to the first query information is greater than or equal to a similarity threshold as similar query information to the first query information. The similarity threshold can be configured and adjusted according to the needs of the actual application scenario and is not specifically limited here.
[0084] The method of this embodiment searches for multiple similar query information of the current first query information from the cached historical query information through the dialogue central control server, and searches for the model processing result data cached by each dialogue engine in parallel based on the first query information and its similar query information. This can maximize the utilization rate of the model processing result data cache of each dialogue engine, thereby minimizing the call of the dialogue model and improving the response speed of the human-computer dialogue.
[0085] In an optional embodiment, as shown in FIG4 , to further reduce the call of dialogue models, a two-stage scheduling strategy based on multiple dialogue engines is used. Step S302 selects one dialogue engine as a target engine based on the cached search results of each dialogue engine for the first query information, and controls the target engine to generate a response message based on the cached search results for the first query information. This can be specifically implemented using the following steps:
[0086] Step S401: sort the cache search results of the first query information of each dialogue engine according to a preset sorting strategy, and determine the preferred engine according to the sorting results.
[0087] In this embodiment, the cache search results of each dialogue engine for the first query information are sorted according to a preset sorting strategy, and a dialogue engine ranked higher (such as first) is selected as the preferred engine for this cache hit based on the sorting results.
[0088] The preset sorting strategy is configured based on factors such as the knowledge type of model processing results cached by each dialogue engine, the priority of each knowledge type, and the potential for false intercepts due to differences in recall capabilities between different dialogue engines. This strategy rationally sorts the cached search results returned by each dialogue engine to select the dialogue engine that will hit the cache. This reduces dialogue model calls while ensuring cache hit accuracy.
[0089] For example, for task-based multi-round dialogue scenarios, such as question-answering scenarios for process-related knowledge, a task-based multi-round dialogue engine can be used. In the cache of the task-based multi-round dialogue engine, the key can include query information and the previous round of reply information, and the value can include the processing results of the query information by the task-based multi-round dialogue model, such as intent recognition results, entity recognition results, etc. The knowledge type of the cached model processing results can be divided into the knowledge type of the direct result after hitting the preset intent (denoted as Ds non-other or Ds non-Anything Else type), and the knowledge type of the result after missing the intent but hitting "other" (denoted as Ds other or Ds Anything Else type).
[0090] For table-based Q&A scenarios involving tabular data, a table-based Q&A engine can be used. In the table-based Q&A engine's cache, the key can include the query information and the previous round of SQL, while the value includes the table-based Q&A model's processing results for the query information, such as the SQL converted from the current round of query information. The knowledge type of the cached model processing results is designated as table query (Table type).
[0091] In the field of intelligent customer service, conversational robots often answer questions based on existing static knowledge data, such as document Q&A and FAQs based on question-and-answer pairs. Knowledge Q&A engines can be used for this purpose. In the cache of a knowledge Q&A engine, the key can include query information and candidate sets, while the value is the response to the query. Depending on the type of knowledge in the candidate set, the knowledge type of the model processing results cached by the knowledge Q&A engine can be categorized as either documents (denoted as Doc) or question-and-answer pairs (denoted as FAQ).
[0092] For example, the priorities of the aforementioned knowledge types may be configured as: FAQ>Ds Non-Anything Else>Table>Doc>. The preset sorting strategy may sort the knowledge types of the model processing results cached by each dialogue engine according to their priorities.
[0093] In addition, the cache search results returned by each conversation engine may include a confidence level, indicating the reliability of the cache search result hit. The conversation control server may also combine the priority and confidence levels of the knowledge types corresponding to the cache search results returned by each conversation engine to set a preset sorting strategy for the cache search results of each conversation engine. This can be configured and adjusted based on the actual application of the conversation system and is not specifically limited here.
[0094] In an optional embodiment, after sorting the cache search results for the first query information returned by each dialogue engine and selecting the preferred engine, if the cache hit status of the cache search results for the first query information of the preferred engine is cache hit, the dialogue control server uses the preferred engine as the target engine, and serially calls the diversion interface and execution interface of the target engine, so that the target engine performs the diversion stage and execution stage processing in serial based on the cache search results of the first query information to generate reply information for the first query information.
[0095] Considering that if the preferred engine is a dialogue engine with strong recall capabilities, such as a table question-answering engine (the knowledge type hitting the cache is Table type) or a document question-answering engine (the knowledge type hitting the cache is Doc type), when a dialogue engine with a higher priority (such as an FAQ engine or a task-based multi-round dialogue engine) is configured with related knowledge information, it is easy to cause false interception problems between dialogue engines.
[0096] In this embodiment, a first engine set and a second engine set are configured, wherein the first engine set includes preset dialogue engines with higher priorities, such as FAQ engines, task-based multi-round dialogue engines, etc. The second engine set includes preset dialogue engines with stronger recall capabilities, such as table question-and-answer engines, document question-and-answer engines, etc. For any dialogue engine with stronger recall capabilities in the second engine set, the dialogue engine is configured to use dialogue engines that may be mistakenly blocked. For example, for a document question-and-answer model with stronger recall capabilities, the dialogue engine that may be mistakenly blocked is a task-based multi-round dialogue engine. The first engine set and the second engine set can be configured based on actual application scenarios and experience, and are not specifically limited here.
[0097] After sorting the cache search results for the first query information returned by each dialogue engine and selecting the preferred engine, if the preferred engine hits the cache and is a dialogue engine with a higher priority, through steps S402-S403, the preferred engine is directly used as the target engine, and the diversion phase and execution phase of the target engine are serially called, and the target engine generates reply information based on the cache search results.
[0098] If the preferred engine hits the cache and is a dialogue engine with strong recall capabilities, through steps S404-S407, the preferred engine and the dialogue engines that may be mistakenly blocked by the preferred engine are added to the candidate engine list for parallel diversion. The candidate engines in the candidate engine list are parallelly diverted, and the target engine is selected based on the diversion results. The execution phase of the target engine is called to generate reply information, which can solve the problem of mistaken blocking of dialogue engines with strong recall capabilities.
[0099] If the preferred engine does not hit the cache, through steps S408-S411, each dialogue engine is added to the candidate engine list for parallel diversion, and the candidate engines in the candidate engine list are diverted in parallel. The target engine is selected based on the diversion result, and the execution phase of the target engine is called to generate the reply information.
[0100] Step S402: If the cache hit status of the cache search result of the preferred engine for the first query information is cache hit, and the preferred engine belongs to the first engine set, the preferred engine is used as the target engine.
[0101] In this embodiment, if the cache hit status of the preferred engine in the cache search result of the first query information is a cache hit, and the preferred engine belongs to the first engine set, that is, the preferred engine that hits the cache this time is a dialogue engine with a higher priority, the dialogue control server directly uses the preferred engine that hits the cache as the target engine.
[0102] Step S403: Control the target engine to perform processing in the diversion phase and the execution phase based on the cache search result of the first query information, and generate reply information for the first query information.
[0103] In this step, the dialogue control server serially calls the diversion interface and execution interface of the target engine, so that the target engine serially executes the processing of the diversion stage and the execution stage based on the cached search result of the first query information to generate the reply information of the first query information, and returns the reply information of the first query information to the dialogue control server.
[0104] In actual applications, different conversation engines call conversation models at different stages: some call models during the diversion stage, some during the execution stage, and some call different models in both stages. For example, a task-based multi-turn conversation engine calls a conversation model during the diversion stage to generate a diversion result containing intent recognition or entity recognition results; it does not call the model during the execution stage, but instead generates a response based on the diversion result. A table question-answering engine calls a table question-answering model during the diversion stage to convert query information into SQL as the diversion result; it calls another large model during the execution stage to generate a response based on the SQL query result data, query information, and contextual information. An FAQ engine does not call a model during the diversion stage, but only recalls a candidate set for the query information. It calls the model during the execution stage to generate a response based on the query information, candidate set, and contextual information. Therefore, different conversation engines cache different model processing result data, and the cached data corresponds to different stages.
[0105] The dialog control server serially calls the target engine's diversion interface and execution interface. Based on the cached search results of the first query information, the target engine serially executes the diversion and execution phases. During any phase of processing, if a model needs to be called, if the cached data hit by the first query information contains the model processing result, the model is not called again, and the model processing result is directly obtained from the cached data hit by the first query information. Only if the model processing result does not exist in the cached data is the model called, and the model processing result is generated through model inference.
[0106] In this step, since the cache hit status in the target engine's cache search result for the first query information is cache hit, the target engine does not need to call the model during the diversion stage and execution stage processing of the first query information. It directly obtains the model processing result from the cache data hit by the first query information, and determines the reply information based on the model processing result, which can reduce model calls and improve the response speed of human-computer dialogue.
[0107] Step S404: If the cache hit status of the cache search result of the preferred engine for the first query information is cache hit, and the preferred engine belongs to the second engine set, the preferred engine and the dialog engine that would be mistakenly blocked by the preferred engine are selected as candidate engines.
[0108] In this embodiment, if the cache hit status in the cache search result of the preferred engine for the first query information is a cache hit, and the preferred engine belongs to the second engine set, that is, the preferred engine hits the cache and is a dialogue engine with strong recall capability, then the dialogue control server will take the preferred engine and the dialogue engine that the preferred engine may mistakenly block as candidate engines, add them to the candidate engine list for parallel diversion, perform parallel diversion on the candidate engines in the candidate engine list, select the target engine based on the diversion result, call the execution phase of the target engine to generate reply information, which can solve the problem of mistaken blocking of dialogue engines with strong recall capability.
[0109] Step S405: Control each candidate engine to perform a diversion phase based on the cached search result of the first query information to obtain a diversion result of each candidate engine.
[0110] In this step, based on the candidate engines determined in the previous step, the dialog control server uses the first query information as the input parameter of the diversion interface and calls the diversion interface of each candidate engine in parallel. Each candidate engine processes the first query information in the diversion stage, obtains the diversion results, and returns the diversion results to the dialog control server.
[0111] It should be noted that the cache hit status of the cache query results for the first query information may vary between different candidate engines. That is, the first query information may hit the cached data in some candidate engines, but not in others. In addition, the cached data hits may also vary between different candidate engines.
[0112] For candidate engines that call models in the diversion stage (such as table question-and-answer engines, task-based multi-round dialogue engines), if the candidate engine does not hit the cached data, during the processing of the diversion stage for the first query information, the candidate engine calls the model normally and generates the model processing results through model reasoning. If the candidate engine hits the cached data, during the processing of the diversion stage for the first query information, the candidate engine does not need to call the model, but directly obtains the model processing results from the cached data hit by the first query information, and determines the reply information based on the model processing results, which can reduce model calls and improve the response speed of human-computer dialogue. For candidate engines that do not need to call models in the diversion stage (such as document question-and-answer engines, FAQ engines), the target engine can ignore the cache, perform the diversion stage processing normally, and generate diversion results based on the first query information.
[0113] Step S406: Select a candidate engine as the target engine based on the diversion results of each candidate engine.
[0114] The conversation control server sorts the diversion results returned by each candidate engine according to a preset diversion ranking strategy. Based on the sorted diversion results, it selects the candidate engine with the highest ranking (e.g., first place) as the target engine. The preset diversion ranking strategy can be configured and adjusted based on the priority of the knowledge type corresponding to the diversion results of each conversation engine, the confidence level of the result, and other factors, and is not specifically limited here.
[0115] Optionally, in the preset diversion sorting strategy, the diversion results returned by each dialogue engine can be sorted according to the priority of the knowledge type corresponding to the diversion results, with the corresponding knowledge type with a higher priority being ranked at the front, and the corresponding knowledge type with a lower priority being ranked at the back. Exemplarily, the priority of each knowledge type can be configured as: FAQ>Ds Non-Anything Else>Table>Doc>Ds Anything Else>Open Domain Search. In addition, the priority of the knowledge type corresponding to the diversion results returned by each dialogue engine in the diversion stage can be configured and adjusted according to the needs of the actual application scenario, and is not specifically limited here.
[0116] Optionally, the confidence level of the diversion result is the confidence level given by the dialogue engine for the diversion result, reflecting the reliability of the diversion result. The dialogue control server can also set a preset diversion sorting strategy based on the priority of the knowledge type corresponding to the diversion results of each dialogue engine and the confidence level of the result. The specific configuration and adjustment can be based on the actual application of the dialogue system and are not specifically limited here.
[0117] Step S407: Control the target engine to perform execution processing on the cache search result of the first query information to generate reply information of the first query information.
[0118] After determining the target engine, the dialog control server invokes the target engine's execution interface and transmits the first query, context, and diversion results as input parameters. The target engine then performs execution based on the first query, context, and diversion results, obtains a response to the first query, and returns the response to the dialog control server.
[0119] It should be noted that the cache hit status of the cache query results for the first query information may vary between different candidate engines. That is, the first query information may hit cache data in some candidate engines, while it may not hit cache data in others. Furthermore, the cache data hits may also differ between different candidate engines. Therefore, the cache hit status of the cache query results for the first query information in the target engine selected from the candidate engines may be either a cache hit or a cache miss.
[0120] In the case where the target engine calls the model during the execution phase (such as a document question-and-answer engine or FAQ engine), if the target engine does not hit the cached data, the target engine will call the model normally during the execution phase of processing the first query information and generate the model processing result through model inference. If the target engine hits the cached data, during the execution phase of processing the first query information, the target engine does not need to call the model, but directly obtains the model processing result from the cached data hit by the first query information, and determines the response information based on the model processing result. This can reduce model calls and improve the response speed of human-computer dialogue.
[0121] For the case where the target engine does not call the model during the execution phase (such as a table question-and-answer engine), the target engine can ignore the cache and perform normal processing during the execution phase, generating reply information for the first query information based on the first query information, context information, and diversion results.
[0122] Step S408: If the cache hit status of the cache search result of the preferred engine for the first query information is cache miss, each dialogue engine is used as a candidate engine.
[0123] In this embodiment, if the cache hit status in the cache search result of the preferred engine for the first query information is a cache miss, it means that the first query information does not hit the cache of any dialogue engine. In this case, the dialogue control server will treat each dialogue engine as a candidate engine, add it to the candidate engine list, and perform parallel diversion on the candidate engines in the candidate engine list. Based on the diversion result, the target engine is selected, and the execution phase of the target engine is called to generate reply information.
[0124] Step S409: Control each candidate engine to perform diversion processing based on the corresponding dialogue model to obtain diversion results of each candidate engine.
[0125] In this step, the dialog control server uses the first query information as the input parameter of the diversion interface and calls the diversion interface of each candidate engine in parallel. Each candidate engine processes the first query information in the diversion stage, obtains the diversion result, and returns the diversion result to the dialog control server.
[0126] It should be noted that since the first query information in each candidate engine does not hit the cached data, each candidate engine performs the diversion stage processing normally based on the dialogue model. When the model needs to be called, the model is called directly without searching the cached data.
[0127] Step S410: Select a candidate engine as a target engine based on the diversion results of each candidate engine.
[0128] This step is similar to the implementation of the aforementioned step S406, except that the number and type of candidate engines may be different. For details, please refer to the relevant content of step S406, which will not be repeated here.
[0129] Step S411: Control the target engine to perform execution phase processing based on the corresponding dialogue model to generate reply information for the first query information.
[0130] After determining the target engine, the dialog control server invokes the target engine's execution interface and transmits the first query, context, and diversion results as input parameters. The target engine then performs execution based on the first query, context, and diversion results, obtains a response to the first query, and returns the response to the dialog control server.
[0131] It should be noted that since the first query information in each candidate engine does not hit the cached data, each candidate engine performs normal execution phase processing based on the corresponding session model. When the model needs to be called, it directly calls the model without searching the cached data.
[0132] In the case of missing the cache of any dialogue engine, the dialogue control server can implement the two-stage scheduling scheme for multiple dialogue engines by using any existing two-stage scheduling scheme for multiple dialogue engines, which is not specifically limited here.
[0133] Furthermore, if the first query information does not hit any of the conversation engines' caches, the conversation control server writes the first query information to the cache. Exemplarily, the conversation control server may write the first query information to the cache via a message queue (MQ). Optionally, regardless of whether the query information hits any of the conversation engines' caches, the conversation control server may write the first query information to the cache and perform deduplication on the query information in the cache.
[0134] In an optional embodiment, after selecting and determining the target engine, the dialogue control server sends the diversion result of the current round of dialogue to each dialogue engine to notify each dialogue engine of the target engine for generating the reply information of the current round.
[0135] In this embodiment, a caching phase is added before the diversion phase in the two-stage scheduling scheme for multiple dialogue engines in the dialogue control server. Each dialogue engine caches model processing results based on its own knowledge and data characteristics. During the caching phase, the cache search results of the first query information in each dialogue engine are queried in parallel, sorted using a reasonable cache search result sorting strategy, and the preferred engine for this round of hits is selected. If the preferred engine hits the cache and has a higher priority, the diversion phase and execution phase of the preferred engine are executed serially based on the cache search results, which can reduce model calls for lower-priority dialogue engines. If the preferred engine hits the cache and has a strong recall capability, the preferred engine and engines that may have been mistakenly blocked by the preferred engine are selected as candidate engines. Each candidate engine is scheduled in a two-stage manner to generate response information. While improving the cache hit rate, the problem of false blocking between engines due to different recall capabilities is minimized, ensuring the accuracy of cache hits.
[0136] The following describes the processing flow of the dialogue engine during the human-computer dialogue process from the perspective of the dialogue engine. Figure 5 is a flow chart of the human-computer dialogue method provided by an exemplary embodiment of the present disclosure. The execution subject of this embodiment is any dialogue engine in the aforementioned system architecture. As shown in Figure 5, the specific steps of this method are as follows:
[0137] Step S501: Receive first query information sent by the dialogue control server.
[0138] The first query information refers to the query information currently input by the user (Query). User input can be in the form of text or voice. The first query information refers to the text content input by the user. For text input by the user, it can be used directly as the first query information or after being rewritten as the first query information. For voice input by the user, the user input voice can be converted into text and used as the first query information.
[0139] In traditional dialogue-based control scheduling solutions, regardless of the selected dialogue engine, the dialogue engine must invoke the corresponding dialogue model (such as LLM) for inference to generate responses. However, especially when using large models (such as LLM), inference takes a long time (typically seconds) and relies on a large amount of GPU resources, resulting in slow response times in human-machine dialogues.
[0140] In this embodiment, each dialogue engine caches the model processing results of query information and provides a cache search status interface to the dialogue control server. During a human-computer dialogue, based on the first query information input by the user, the dialogue control server uses the first query information as an input parameter before scheduling a dialogue engine to run the dialogue model. The dialogue engine receives the first query information from the dialogue control server.
[0141] Step S502: Search cached data matching the first query information in the cached model processing result data to obtain a cached search result of the first query information.
[0142] In this step, the dialog engine executes the implementation program of the cache search state interface, searches for cache data matching the first query information in the cached model processing result data, and obtains the cache search result of the first query information.
[0143] Optionally, when the dialog control server calls the cache search status interface of any dialog engine, it transmits the first query information and the word segmentation result of the first query information to the dialog engine as input parameters of the interface. Based on the first query information and the word segmentation result of the first query information, the dialog engine searches the cached model processing result data for cached data that matches the first query information, thereby obtaining the dialog engine's cache search result for the first query information.
[0144] In an optional embodiment, the dialogue control server can cache the received query information. Based on the first query information input by the user, the dialogue control server searches for one or more similar query information of the first query information in the cached historical query information. When calling the cache search status interface of each dialogue engine, the dialogue control server uses the first query information and similar query information as input parameters and calls the cache search status interface of each dialogue engine in parallel. The dialogue engine receives the first query information and similar query information sent by the dialogue control server. The dialogue engine executes the implementation program of the cache search status interface, searches for cached data that matches the first query information or similar query information in the cached model processing result data, and obtains the cache search result of the first query information.
[0145] Optionally, the conversation engine sorts the similar query information based on their similarity to the first query information, first searches for cached data based on the first query information, and if cached data matching the first query information is found, the cached data matching the first query information is used as cached data that hit the first query information, thereby obtaining a cached search result for the first query information. If no cached data matching the first query information is found, the conversation engine searches for cached data matching the similar query information in sequence based on the sorted similar query information until cached data matching the similar query information is found, then stops searching, and uses the cached data that matches the similar query information found as cached data that hit the first query information, thereby obtaining a cached search result for the first query information; or until it is determined that no cached data matching the first query information or similar query information is found, it is determined that the first query information does not hit cached data, thereby obtaining a cached search result for the first query information.
[0146] Optionally, the dialogue engine can respectively search for cached data and the first confidence level that hits the first query information and similar query information, perform a weighted calculation on the first confidence level of the hit cached data based on the degree of similarity between the similar query information and the first query information, and obtain the second confidence level of the hit cached data; based on the second confidence level, select one of the cached data with a higher confidence level as the cached data that hits the first query information, and obtain the cache search result for the first query information.
[0147] Step S503: Return the cache search result of the first query information to the dialogue control server.
[0148] Step S504: Based on the control information of the dialog control server and the cache search result of the first query information, a reply message is generated.
[0149] Among them, the control information of the dialogue control server can be the call of the interface provided by the dialogue control server to the dialogue engine, the instructions and requests sent to the dialogue engine, etc. The dialogue control server can realize the scheduling and control of the dialogue engine by calling the interface and sending instructions / requests. This embodiment does not specifically limit the interactive method of the dialogue control server for scheduling and controlling the dialogue engine.
[0150] In this embodiment, the dialogue engines cache the model result processing data. The dialogue control server selects one of the dialogue engines as the target engine based on the cached search results of the first query information returned by each dialogue engine, and controls the target engine to generate a response message based on the cached search results of the first query information.
[0151] In this step, the dialogue engine serving as the target engine generates reply information based on the control information of the dialogue control server and the cached search result of the first query information.
[0152] Specifically, based on the control information of the dialogue control server, when the cache hit status in the cache search result of the first query information is cache hit, the dialogue engine performs diversion stage and / or execution stage processing based on the cache search result of the first query information to process the result data according to the model hit by the first query information, and generate reply information to reduce the call to the model and improve the response speed of the human-computer dialogue.
[0153] When the cache hit status in the cache search result of the first query information is a cache miss, the diversion stage and / or execution stage are processed based on the corresponding dialogue model to process the first query information through the corresponding dialogue model, and generate reply information based on the obtained model processing result.
[0154] For example, for a conversation engine whose model call is in the diversion phase, when the conversation control server calls the conversation engine's diversion interface, if the cache hit status in the cache search result for the first query information is cache hit, the conversation engine will no longer call the model during the diversion phase. Instead, it will obtain the corresponding model processing result from the cache data and return the diversion result to the conversation control server. This can reduce model calls and improve the conversation engine's response speed during the diversion phase. If the cache hit status in the cache search result for the first query information is cache miss, the diversion phase is performed based on the corresponding conversation model to process the first query information through the corresponding conversation model. The obtained model processing result is returned to the conversation control server. When the conversation control server calls the conversation engine's execution interface, since the execution phase does not require calling the model, the conversation engine performs the execution phase processing normally, generates reply information, and returns the reply information to the conversation control server.
[0155] For example, for a dialogue engine whose model call is in the execution phase, when the dialogue control server calls the dialogue engine's diversion interface, since the diversion phase does not require calling the model, the dialogue engine performs the diversion phase processing normally, obtains the diversion result, and returns the diversion result to the dialogue control server. When the dialogue control server calls the dialogue engine's execution interface, if the cache hit status in the cache search result of the first query information is a cache hit, the dialogue engine no longer calls the model during the execution phase processing, but obtains the corresponding model processing result from the cache data, which can reduce model calls and improve the response speed of the dialogue engine in the execution phase. If the cache hit status in the cache search result of the first query information is a cache miss, the execution phase processing is performed based on the corresponding dialogue model to process the first query information and the diversion result through the corresponding dialogue model, obtain the model processing result (reply information), and return the reply information to the dialogue control server.
[0156] For example, for a conversation engine whose model call is simultaneously in the diversion phase and the execution phase, when the conversation control server calls the conversation engine's diversion interface, if the cache hit status in the cache search result for the first query information is cache hit, the conversation engine will no longer call the model during the diversion phase processing. Instead, it will obtain the corresponding model processing result from the cache data and return the diversion result to the conversation control server. This can reduce model calls and improve the conversation engine's response speed during the diversion phase. If the cache hit status in the cache search result for the first query information is cache miss, the diversion phase processing is performed based on the corresponding conversation model to process the first query information through the corresponding conversation model, and the obtained model processing result is returned to the conversation control server.
[0157] When the dialogue control server calls the execution interface of the dialogue engine, if the cache hit status in the cache search result of the first query information is a cache hit, the dialogue engine will no longer call the model during the execution phase, but will obtain the corresponding model processing result from the cache data, which can reduce model calls and improve the response speed of the dialogue engine in the execution phase. In the case that the cache hit status in the cache search result of the first query information is a cache miss, the execution phase is processed based on the corresponding dialogue model to process the first query information and the diversion result through the corresponding dialogue model, and the obtained model processing result (reply information) is returned to the dialogue control server. Among them, for the dialogue engine whose model call is in the diversion phase and the execution phase at the same time, the dialogue model used in the two phases can be the same model or two different models, which is not specifically limited here.
[0158] Furthermore, after processing the first query information using the dialogue model, the dialogue engine caches the model processing result data of the first query information.
[0159] In actual applications, different caching strategies can be configured for different dialogue engines based on the knowledge data characteristics of each dialogue engine, so that each dialogue engine can cache different data according to its own knowledge data characteristics to improve the cache hit rate and minimize the number of model calls while ensuring cache accuracy.
[0160] For example, for a dialog engine used for table Q&A (i.e., a table Q&A engine), the corresponding dialog model is the table Q&A model. During the diversion phase, the table Q&A engine uses the table Q&A model to convert query information into structured query statements (SQL). During the execution phase, it obtains the query results from the SQL query information and generates response information based on the query results, query information, and contextual information.
[0161] The model processing result data cached by the Tableau Q&A engine includes a cache key and a cache value. The cache key includes the query information and the previous round of SQL, while the cache value is the SQL for the current round converted from the query information. By adding the previous round of SQL to the cache key, you can improve the accuracy of cache hits.
[0162] When the dialogue engine searches for cached data that matches the first query information, it first obtains the previous round of SQL for the first query information, and matches the first query information and the previous round of SQL with the cache key in the cache data to determine the cached data that is hit by the first query information and the previous round of SQL as the cached data that matches the first query information.
[0163] Optionally, when matching the first query information and the previous round of SQL with the cache key in the cache data, a consistency match can be performed. When the first query information and the previous round of SQL are consistent with the query information in the cache key and the previous round of SQL, it is determined that the first query information and the previous round of SQL hit the cache data, which can improve the accuracy of cache hits.
[0164] In addition, a fuzzy match can be performed based on the first query information and the previous round of SQL with the cache key in the cache data. When the matching degree between the first query information and the previous round of SQL and the cache key in the cache data reaches a first matching degree threshold, the cache data hit by the first query information and the previous round of SQL is determined. The first matching degree threshold can be configured and adjusted according to the needs of the actual application scenario and is not specifically limited here. In addition, other cache hit methods can be used to search for cache data hit by the first query information and the previous round of SQL, which is not specifically limited here.
[0165] Among them, the previous round of SQL of the first query information can be obtained from the context information recorded by the dialogue engine, or provided by the dialogue control server to the dialogue engine, which is not specifically limited in this embodiment.
[0166] Exemplarily, the dialogue engine's corresponding dialogue model is a knowledge question-and-answer model, such as a document question-and-answer model or FAQ model. During the diversion phase, the dialogue engine obtains a knowledge candidate set for the query information. During the execution phase, the knowledge question-and-answer model generates a response based on the query information, its context, and the knowledge candidate set. The cache key in the model processing result data cached by the dialogue engine includes the query information and the knowledge candidate set, and the cache value is the response information.
[0167] For example, the conversation engine can be a document question-and-answer engine, and the conversation model used is a document question-and-answer model. The document question-and-answer engine obtains a document candidate set for the query information in the diversion stage, and the document candidate set contains at least one document that matches the query information. In the execution stage, the document question-and-answer engine uses the document question-and-answer model to generate response information based on the query information, its context, and the document candidate set. The model processing result data cached by the document question-and-answer engine includes a cache key and a cache value, wherein the cache key includes the query information and the document candidate set for the query information, and the cache value is the response information for the query information. By adding the document candidate set for the query information to the cache key, the accuracy of the document question-and-answer engine hitting the cache can be improved.
[0168] When the document question-and-answer engine searches for cached data that matches the first query information, it first obtains the document candidate set of the first query information, and matches the first query information and the document candidate set with the cache key in the cached data to determine the cached data that hits the first query information and the document candidate set as the cached data that matches the first query information.
[0169] Optionally, when matching the first query information and document candidate set with the cache key in the cache data, consistency matching can be performed. When the first query information and document candidate set are consistent with the query information and document candidate set in the cache key, it is determined that the first query information and document candidate set hit the cache data, which can improve the accuracy of cache hits.
[0170] In addition, fuzzy matching can be performed based on the first query information and the document candidate set and the cache key in the cache data. When the matching degree between the first query information and the document candidate set and the cache key in the cache data reaches a second matching degree threshold, the cache data hit by the first query information and the document candidate set is determined. The second matching degree threshold can be configured and adjusted according to the needs of the actual application scenario, and is not specifically limited here. In addition, other cache hit methods can be used to find the cache data hit by the first query information and the document candidate set, and are not specifically limited here. The document candidate set of the first query information can be recalled by the document question and answer engine from the candidate documents of the knowledge base according to the first query information in the caching stage to obtain the document candidate set. The document question and answer engine can directly use the recalled document candidate set in the diversion stage without repeatedly recalling the document candidate set.
[0171] For example, the conversation engine may be a FAQ engine, and the conversation model used may be a FAQ model. The FAQ engine obtains a candidate set of question-answer pairs for the query information in the diversion stage, and the candidate set of question-answer pairs contains at least one question-answer pair that matches the query information. In the execution stage, the FAQ engine uses the FAQ model to generate reply information based on the query information, its context, and the candidate set of question-answer pairs. The model processing result data cached by the FAQ engine includes a cache key and a cache value, wherein the cache key includes the query information and the candidate set of question-answer pairs for the query information, and the cache value is the reply information for the query information. By adding the candidate set of question-answer pairs for the query information to the cache key, the accuracy of the FAQ engine hitting the cache can be improved.
[0172] When the FAQ engine searches for cached data that matches the first query information, it first obtains the candidate set of question-answer pairs for the first query information, and matches the first query information and the candidate set of question-answer pairs with the cache key in the cached data to determine the cached data that hits the first query information and the candidate set of question-answer pairs as the cached data that matches the first query information.
[0173] Optionally, when matching the first query information and the candidate set of question and answer pairs with the cache key in the cache data, a consistency match can be performed. When the first query information and the candidate set of question and answer pairs are consistent with the query information and the candidate set of question and answer pairs in the cache key, it is determined that the first query information and the candidate set of question and answer pairs hit the cache data, which can improve the accuracy of the cache hit.
[0174] In addition, fuzzy matching can be performed based on the first query information and the candidate set of question and answer pairs and the cache key in the cache data. When the matching degree between the first query information and the candidate set of question and answer pairs and the cache key in the cache data reaches a third matching degree threshold, the cache data hit by the first query information and the candidate set of question and answer pairs is determined. Among them, the third matching degree threshold can be configured and adjusted according to the needs of the actual application scenario, and is not specifically limited here. In addition, other cache hit methods can be used to find the cache data hit by the first query information and the candidate set of question and answer pairs, and are not specifically limited here. Among them, the candidate set of question and answer pairs for the first query information can be recalled by the FAQ engine from the candidate documents of the knowledge base according to the first query information in the caching stage to obtain the candidate set of question and answer pairs. The FAQ engine can directly use the recalled candidate set of question and answer pairs in the diversion stage, without the need to repeatedly recall the candidate set of question and answer pairs.
[0175] Exemplarily, the dialogue engine can be a task-based multi-turn dialogue engine, and the corresponding dialogue model is a task-based multi-turn dialogue model. The dialogue engine performs natural language processing on the query information through the task-based multi-turn dialogue model in the diversion stage, and generates reply information according to the query information, the context of the query information and the natural language processing results in the execution stage.
[0176] The model processing result data cached by the task-based multi-turn dialogue engine consists of a cache key and a cache value. The cache key includes the query information and the previous round's response information, while the cache value is the current round's response information. By adding the previous round's response information to the cache key, the accuracy of the task-based multi-turn dialogue engine's cache hit rate can be improved.
[0177] When the task-based multi-round dialogue engine searches for cached data that matches the first query information, it first obtains the reply information of the previous round, and matches the cache key in the cached data based on the first query information and the reply information of the previous round to determine the cached data hit by the first query information and the reply information of the previous round as the cached data that matches the first query information.
[0178] Optionally, when matching the first query information and the reply information of the previous round with the cache key in the cache data, a consistency match can be performed. When the first query information and the reply information of the previous round are consistent with the query information in the cache key and the reply information of the previous round, it is determined that the first query information and the reply information of the previous round hit the cache data, which can improve the accuracy of hitting the cache of the task-based multi-round dialogue engine.
[0179] In addition, fuzzy matching can also be performed based on the first query information and the previous round of reply information and the cache key in the cache data. When the matching degree between the first query information and the previous round of reply information and the cache key in the cache data reaches the fifth matching degree threshold, the cache data hit by the first query information and the previous round of reply information is determined. Among them, the fifth matching degree threshold can be configured and adjusted according to the needs of the actual application scenario, and is not specifically limited here. In addition, other cache hit methods can be used to find the cache data hit by the first query information and the previous round of reply information, which is not specifically limited here. Among them, the previous round of reply information of the first query information can be obtained by the task-based multi-round dialogue engine from the recorded context information, or it can be provided by the dialogue control server to the task-based multi-round dialogue engine, which is not specifically limited here in this embodiment.
[0180] In another optional embodiment, the cache key in the cache data of each dialogue engine contains query information, and each dialogue engine caches different cache values according to its own knowledge data characteristics. When searching the cache, the cache data hit by the first query information is determined by matching the cache key in the cache data based on the current first query information, or the first query information and similar query information.
[0181] In this embodiment, a two-stage scheduling scheme for multiple dialogue engines on a dialogue control server adds a caching stage before the diversion stage. Each dialogue engine caches model processing results based on its own knowledge and data characteristics, improving the accuracy of cache hits for each dialogue engine. During the caching stage, cache search results for the first query information in each dialogue engine are queried in parallel, sorted using a reasonable cache search result sorting strategy, and the preferred engine for this round of hits is selected. If a preferred engine hits the cache and has a higher priority, the diversion and execution stages of the preferred engine are executed serially based on the cache search results, reducing model calls for lower-priority dialogue engines. If a preferred engine hits the cache and has a strong recall capability, the preferred engine, along with engines that may have been mistakenly blocked by the preferred engine, are selected as candidate engines. These candidate engines are then scheduled in a two-stage process to generate response information. This improves cache hit rates while minimizing the risk of false blocking between engines due to varying recall capabilities, ensuring cache hit accuracy.
[0182] Figure 6 is a schematic diagram of a two-stage scheduling scheme for a dialog system, provided by an exemplary embodiment of the present disclosure. As shown in Figure 6, the two-stage scheduling scheme provided by this embodiment includes a cache stage before the diversion stage. Based on the query information (query) entered by the user in the current round, the dialog control server searches for similar query information to the user's first query information in its locally cached historical query information. Based on the first query information and similar query information, the dialog control server concurrently calls the cache search state interface of each dialog engine, causing each dialog engine to search its cached model processing result data for cached data that matches the first query information, thereby obtaining a cache search result for the first query information. The dialog control server sorts the cache search results for the first query information returned by each dialog engine and selects a preferred engine based on the sorting results. If the preferred engine matches the cache, the dialog control server serially calls the diversion interface and execution interface of the preferred engine, causing the preferred engine to sequentially perform the diversion and execution stages. During the diversion and execution stages, when the preferred engine needs to call a model, it retrieves the model processing result corresponding to the first query information from the cached data, bypassing the model call, thereby reducing the number of model calls. When the preferred engine misses the cache, the dialog control server performs parallel diversion on each dialog engine, sorts the diversion results, selects the target engine, and calls the execution phase of the target engine.
[0183] Optionally, to avoid the problem of false blocking caused by the different recall capabilities of different engines, after selecting the preferred engine, if the preferred engine hits the cache, if the preferred engine is a conversation engine with strong recall capabilities (such as a table question-and-answer engine or a document question-and-answer engine), the preferred engine and the conversation engines that the preferred engine may have falsely blocked are added to the candidate engine list for parallel diversion. The candidate engines in the candidate engine list are parallel diverted, and the target engine is selected based on the diversion results. The execution phase of the target engine is called to generate reply information. This can solve the problem of false blocking of conversation engines with strong recall capabilities. In the case of a preferred engine hit in the cache, if the preferred engine is a conversation engine with a higher priority (such as an FAQ engine or a task-based multi-round conversation engine), the preferred engine is directly used as the target engine, and the diversion phase and execution phase of the target engine are serially called. The reply information is generated by the target engine based on the cache search results.
[0184] Figure 7 is a flowchart of a human-computer dialogue in an intelligent customer service system provided by an exemplary embodiment of the present disclosure. The execution subject of this embodiment is the dialogue control server. As shown in Figure 7, when applied to an intelligent customer service scenario, the specific process of the human-computer dialogue is as follows:
[0185] Step S701: The dialogue control server receives a first question that the user needs to ask, searches for cached data corresponding to the first question in the model processing result data cached by at least one dialogue engine, and obtains cached search results of each dialogue engine for the first question.
[0186] The first question refers to the question the user wants to ask in the intelligent customer service system, that is, the first query information currently entered by the user. The first question refers to the text content entered by the user. The text entered by the user can be used as the first question directly or after being rewritten. The voice input by the user can be converted into text and used as the first question.
[0187] In an optional embodiment, in this step, the dialogue control server can search for similar questions to the first question in the cached historical question information; send the first question and similar questions to each dialogue engine, control each dialogue engine to search for cached data matching the first question or similar questions in the cached model processing result data, and obtain the cached search result for the first question; receive the cached search result for the first question returned by each dialogue engine.
[0188] In this step, the dialogue control server searches for cached data matching the first question in the model processing result data cached by at least one dialogue engine, and obtains the cached search results of each dialogue engine for the first question. This is consistent with the implementation method of the dialogue control server searching for cached data matching the first query information in the model processing result data cached by at least one dialogue engine in the aforementioned step S301, and obtaining the cached search results of each dialogue engine for the first query information. The first question is the first query information, the historical question information is the historical query information, and similar questions to the first question are similar query information to the first query information. For details, please refer to the relevant content of the aforementioned embodiment, which will not be repeated here.
[0189] Step S702: The dialog control server selects one of the dialog engines as the target engine based on the cache search results of the dialog engines for the first question, and controls the target engine to generate a reply message based on the cache search results for the first question.
[0190] In this step, the dialog control server sorts the cached search results for the first question by each dialog engine based on a preset sorting strategy. Based on the sorting results, the dialog engine with the highest cached search result is selected as the target engine. The target engine is then instructed to generate a response message based on the cached search result for the first question. If a cache hit is found, the target engine uses the model processing result that hit the cache. If a cache miss is found, the target engine invokes the corresponding dialog model to obtain the model processing result for the first question, and then generates a response message based on the model processing result for the first question.
[0191] Specifically, the dialogue control server sorts the cache search results of each dialogue engine for the first question according to a preset sorting strategy, and determines the preferred engine based on the sorting results; if the cache hit status of the cache search result of the preferred engine for the first question is a hit cache, and the preferred engine belongs to the first engine set, the preferred engine is used as the target engine; the target engine is controlled to perform processing in the diversion stage and the execution stage based on the cache search result of the first question to generate reply information for the first question.
[0192] If the cache hit status of the cache search result of the preferred engine for the first question is cache hit, and the preferred engine belongs to the second engine set, the preferred engine and the dialogue engine that will be mistakenly blocked by the preferred engine will be used as candidate engines; control each candidate engine to perform diversion stage processing based on the cache search result of the first question to obtain the diversion results of each candidate engine; select a candidate engine as the target engine based on the diversion results of each candidate engine; control the target engine to perform execution stage processing on the cache search result of the first question to generate reply information for the first question.
[0193] If the cache hit status of the cache search result of the preferred engine for the first question is a cache miss, each dialogue engine is used as a candidate engine; each candidate engine is controlled to perform diversion stage processing based on the corresponding dialogue model to obtain the diversion results of each candidate engine; based on the diversion results of each candidate engine, one candidate engine is selected as the target engine; the target engine is controlled to perform execution stage processing based on the corresponding dialogue model to generate response information for the first question.
[0194] In this embodiment, the meaning and configuration of the preset sorting strategy, the first engine set, and the second engine set are similar to those described in step S302 of the aforementioned embodiment and are not repeated here. The specific implementation of step S702 is similar to that of step S302, and is specifically described in steps S401-S411, which are not repeated here.
[0195] This embodiment also provides a human-computer dialogue method for a dialogue engine applied to an intelligent customer service system. It should be noted that when applied to an intelligent customer service system, the processing performed by each dialogue engine under the scheduling of the dialogue control server is similar to the processing performed by the dialogue engine in any of the aforementioned method embodiments. Please refer to the detailed content of the aforementioned method embodiments for details, which will not be repeated here.
[0196] The solution of this embodiment is applied to intelligent customer service scenarios. In the two-stage scheduling scheme for multiple dialogue engines in the dialogue control server, a caching stage is added before the diversion stage. Each dialogue engine caches the model processing results based on its own knowledge and data characteristics. In the caching stage, the cache search results of the first question the user wants to ask are queried in parallel in each dialogue engine. A reasonable cache search result sorting strategy is used to sort the results, and the preferred engine that hits the target in this round is selected. If the preferred engine hits the cache and has a higher priority, the diversion stage and execution stage of the preferred engine are executed serially based on the cache search results, which can reduce the model calls of the lower-priority dialogue engines. If the preferred engine hits the cache and has a strong recall capability, the preferred engine and engines that may have been mistakenly blocked by the preferred engine are selected as candidate engines. Each candidate engine is scheduled in two stages to generate reply information. While improving the cache hit rate, the problem of false blocking between engines due to different recall capabilities of each engine is minimized to ensure the accuracy of cache hits.
[0197] FIG8 is a schematic diagram of the structure of a dialogue control server provided in an embodiment of the present disclosure. As shown in FIG8 , the dialogue control server includes: a memory 801 and a processor 802. The memory 801 is used to store computer-executable instructions and can be configured to store various other data to support operations on the dialogue control server. The processor 802 is communicatively connected to the memory 801 and is used to execute the computer-executable instructions stored in the memory 801 to implement the technical solution provided by any of the above-mentioned method embodiments. The specific functions and technical effects that can be achieved are similar and will not be repeated here.
[0198] Optionally, as shown in Figure 8, the dialog control server also includes other components such as a firewall 803, a load balancer 804, a communication component 805, and a power supply component 806. Figure 8 only schematically illustrates some components, which does not mean that the dialog control server only includes the components shown in Figure 8.
[0199] Figure 9 is a structural diagram of a dialogue engine provided by an embodiment of the present disclosure. The dialogue engine can run on a cloud server. As shown in Figure 9, the dialogue engine includes: a memory 901 and a processor 902. The memory 901 is used to store computer-executable instructions and can be configured to store various other data to support operations on the cloud server. The processor 902 is communicatively connected to the memory 901 and is used to execute the computer-executable instructions stored in the memory 901 to implement the technical solution provided by any of the above-mentioned method embodiments. Its specific functions and achievable technical effects are similar and will not be repeated here. Optionally, as shown in Figure 9, the dialogue engine also includes: a firewall 903, a load balancer 904, a communication component 905, a power supply component 906 and other components. Figure 9 only schematically shows some components, which does not mean that the dialogue engine only includes the components shown in Figure 9.
[0200] The embodiments of the present disclosure also provide a computer-readable storage medium, in which computer-executable instructions are stored. When a processor executes the computer-executable instructions, the method flow executed by the dialogue control server or the dialogue engine in any of the aforementioned embodiments is implemented. The specific functions and technical effects that can be achieved are not repeated here.
[0201] The present disclosure also provides a computer program product, including a computer program, which, when executed by a processor, implements the method of any of the aforementioned embodiments. The computer program is stored in a readable storage medium, and at least one processor of an electronic device can read the computer program from the readable storage medium. The at least one processor executes the computer program, causing the electronic device to execute the method flow executed by the dialog control server or dialog engine in any of the aforementioned method embodiments. The specific functions and technical effects achieved are not further described here.
[0202] The present disclosure provides a chip comprising: a processing module and a communication interface. The processing module is capable of executing the technical solution of the electronic device described in the aforementioned method embodiments. Optionally, the chip further comprises a storage module (e.g., a memory) configured to store instructions, and the processing module configured to execute the instructions stored in the storage module. Execution of the instructions stored in the storage module causes the processing module to execute the method flow executed by the dialog control server or dialog engine described in any of the aforementioned method embodiments.
[0203] It should be understood that the above-mentioned processor can be a processing unit (Central Processing Unit, referred to as CPU), or it can be other general-purpose processors, digital signal processors (Digital Signal Processor, referred to as DSP), application specific integrated circuits (Application Specific Integrated Circuit, referred to as ASIC), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in the application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor. The memory may include high-speed random access memory (Random Access Memory, referred to as RAM), and may also include non-volatile storage, such as at least one disk storage, and can also be a USB flash drive, a mobile hard disk, a read-only memory, a disk or an optical disk, etc.
[0204] The above-mentioned memory may be an object storage service (OSS). The above-mentioned memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0205] The communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on communication standards, such as mobile hotspots (WiFi), second-generation mobile communication systems (2G), third-generation mobile communication systems (3G), fourth-generation mobile communication systems (4G) / Long Term Evolution (LTE), fifth-generation mobile communication systems (5G), or a combination thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near-field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared technology, ultra-wideband (UWB) technology, Bluetooth technology, and other technologies. The power supply component provides power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.
[0206] The above-mentioned storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or special-purpose computer. An exemplary storage medium is coupled to a processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be an integral part of the processor. The processor and the storage medium can be located in a dedicated integrated circuit. Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a main control device.
[0207] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.
[0208] The order of the above-mentioned embodiments of the present disclosure is for description only and does not represent the advantages and disadvantages of the embodiments. In addition, in some of the processes described in the above-mentioned embodiments and the accompanying drawings, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or in parallel. They are only used to distinguish between different operations, and the sequence number itself does not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent the order of precedence, nor do they limit "first" and "second" to different types. The meaning of "multiple" is more than two, unless otherwise clearly and specifically defined.
[0209] Through the description of the above implementation methods, those skilled in the art can clearly understand that the above-mentioned embodiment methods can be implemented by means of software plus the necessary general hardware platform, and of course can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present disclosure, or the part that contributes to the prior art, can be embodied in the form of a software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present disclosure.
[0210] Other embodiments of the present disclosure will readily occur to those skilled in the art after considering the specification and practicing the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of the present disclosure that follow the general principles of the present disclosure and include common knowledge or customary techniques in the art not disclosed herein.
[0211] The above are only preferred embodiments of the present disclosure and are not intended to limit the patent scope of the present disclosure. Any equivalent structure or equivalent process transformation made using the contents of the present disclosure and the drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present disclosure.
Claims
1. A human-computer dialogue method, wherein: Applied to the dialogue central control server, the method includes: Based on first query information input by a user, searching for cached data matching the first query information in model processing result data cached by at least one dialogue engine, and obtaining cached search results of each dialogue engine for the first query information; According to the cache search results of each dialogue engine for the first query information, one of the dialogue engines is selected as the target engine, and the target engine is controlled to generate reply information according to the cache search results of the first query information.
2. The method according to claim 1, wherein: The method of searching cached data matching the first query information in the model processing result data cached by at least one dialogue engine based on the first query information input by the user, and obtaining the cache search results of each dialogue engine for the first query information, includes: Sending the first query information input by the user to each dialogue engine, controlling each dialogue engine to search for cached data matching the first query information in the cached model processing result data, and obtaining the cached search result of the first query information; Receive cache search results of the first query information returned by each dialogue engine.
3. The method according to claim 1, wherein: The method of searching cached data matching the first query information in the model processing result data cached by at least one dialogue engine based on the first query information input by the user, and obtaining the cache search results of each dialogue engine for the first query information, includes: Searching the cached historical query information for similar query information to the first query information input by the user; Sending the first query information and similar query information to each dialogue engine, controlling each dialogue engine to search for cached data matching the first query information or the similar query information in the cached model processing result data, and obtaining a cached search result of the first query information; Receive cache search results of the first query information returned by each dialogue engine.
4. The method according to any one of claims 1 to 3, wherein: The selecting one of the dialogue engines as the target engine according to the cache search results of each dialogue engine for the first query information, and controlling the target engine to generate reply information according to the cache search results of the first query information, includes: According to a preset sorting strategy, sorting the cache search results of each dialogue engine for the first query information, and determining a preferred engine according to the sorting results; If the cache hit status of the cache search result of the preferred engine for the first query information is cache hit, and the preferred engine belongs to the first engine set, the preferred engine is used as the target engine; The target engine is controlled to perform processing in a diversion phase and an execution phase based on a cache search result of the first query information, and generate reply information of the first query information.
5. The method according to claim 4, wherein: After determining the preferred engine according to the ranking results, the method further includes: If the cache hit status of the cache search result of the preferred engine for the first query information is cache hit, and the preferred engine belongs to the second engine set, the preferred engine and the dialog engine that will be mistakenly blocked by the preferred engine are used as candidate engines; Controlling each candidate engine to perform a diversion phase process based on a cached search result of the first query information to obtain a diversion result of each candidate engine; According to the diversion results of each candidate engine, select a candidate engine as the target engine; The target engine is controlled to perform execution phase processing on the cache search result of the first query information to generate reply information of the first query information.
6. The method according to claim 4, wherein: After determining the preferred engine according to the ranking results, the method further includes: If the cache hit status of the cache search result of the preferred engine for the first query information is a cache miss, then each of the dialogue engines is used as a candidate engine; Control each candidate engine to perform the diversion phase based on the corresponding dialogue model to obtain the diversion results of each candidate engine; According to the diversion results of each candidate engine, select a candidate engine as the target engine; The target engine is controlled to perform processing in the execution phase based on the corresponding dialogue model to generate reply information for the first query information.
7. A human-computer dialogue method, wherein: Applied to a dialogue engine, the method comprises: Receiving first query information sent by the dialog control server; Searching cached data matching the first query information in the cached model processing result data to obtain a cached search result of the first query information; Returning the cache search result of the first query information to the dialog control server; Based on the control information of the central control server in the dialogue, reply information is generated according to the cache search result of the first query information.
8. The method according to claim 7, wherein: The step of searching the cached model processing result data for cached data matching the first query information to obtain a cached search result of the first query information includes: Receiving similar query information of the first query information sent by the dialog control server; In the cached model processing result data, cached data matching the first query information or the similar query information is searched to obtain a cached search result of the first query information.
9. The method according to claim 7, wherein: The generating of reply information according to the control information of the central control server in the dialogue and the cache search result of the first query information comprises: Based on the control information of the dialog central control server, when the cache hit status in the cache search result of the first query information is cache hit, performing processing in the diversion stage and / or the execution stage based on the cache search result of the first query information, so as to generate reply information according to the model processing result data hit by the first query information; In the case where the cache hit status in the cache search result of the first query information is a cache miss, a diversion stage and / or an execution stage is performed based on the corresponding dialogue model to process the first query information through the corresponding dialogue model, and a reply information is generated according to the obtained model processing result.
10. The method according to claim 9, wherein: After the first query information is processed by the corresponding dialogue model, the method further includes: The model processing result data of the first query information is cached.
11. The method according to any one of claims 7 to 10, wherein: The dialogue model corresponding to the dialogue engine is a table question-answering model. The dialogue engine converts the query information into a structured query statement through a table question-answering model in the diversion phase, obtains the query result of the structured query statement of the query information in the execution phase, and generates reply information based on the query result, the query information and the context information; The cache key in the model processing result data cached by the dialogue engine includes query information and the structured query statement of the previous round, and the cache value is the structured query statement of the current round converted from the query information.
12. The method according to any one of claims 7 to 10, wherein: The dialogue model corresponding to the dialogue engine is a knowledge question-answering model. The dialogue engine obtains a knowledge candidate set of query information in the diversion phase, and generates reply information according to the query information, the context of the query information and the knowledge candidate set through the knowledge question-answering model in the execution phase; The cache keys in the model processing result data cached by the dialogue engine include query information and knowledge candidate sets. The cache value is the reply information.
13. A human-computer dialogue method, wherein: The method is applied to a dialog control server in an intelligent customer service system, and comprises: Receive a first question that a user needs to ask, search for cached data related to the first question in model processing result data cached by at least one dialogue engine, and obtain cached search results of each dialogue engine for the first question; According to the cache search results of each dialogue engine for the first question, one of the dialogue engines is selected as the target engine, and the target engine is controlled to generate reply information according to the cache search results for the first question.
14. The method according to claim 13, wherein: The step of searching cached data related to the first question in the model processing result data cached by at least one dialogue engine to obtain cached search results of each dialogue engine for the first question includes: Searching for similar questions to the first question in cached historical question information; Sending the first question and similar questions to each dialogue engine, controlling each dialogue engine to search for cached data matching the first question or the similar questions in the cached model processing result data, and obtaining a cached search result for the first question; Receive the cache search results of the first question returned by each dialogue engine.
15. A dialogue control server, wherein: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; Wherein, the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the dialogue control server to execute the method described in any one of claims 1-6 and 13-14.
16. A conversation engine, wherein: include: at least one processor; as well as a memory communicatively coupled to the at least one processor; The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the dialogue engine to perform the method described in any one of claims 7-12.
17. A computer-readable storage medium, wherein: The computer-readable storage medium stores computer-executable instructions. When a processor executes the computer-executable instructions, the method according to any one of claims 1 to 14 is implemented.
18. A computer program product, comprising a computer program, wherein when the computer program is executed by a processor, the method according to any one of claims 1 to 14 is implemented.
Citation Information
Patent Citations
Question answering processing method and apparatus
CN109033229A
Intelligent question and answer method and device
CN113127619A
Data searching device and method for data system
CN116628016A
Man-machine dialogue method, dialogue central control server, dialogue engine and storage medium
CN117633181A
System and method for a cognitive conversation service
US20220358295A1