Man-machine conversation method, electronic equipment and storage medium
By obtaining association information in enterprise business data and generating prompt information to enter the dialogue model, the problem of inaccurate responses of large language models in enterprise question-and-answer scenarios is solved, achieving more efficient and flexible reply processing.
Patent Information
- Application Number
- CN202410064449.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-16
- Publication Date
- 2025-07-18
AI Technical Summary
When the prior art applies large language models to enterprise business Q&A scenarios, it is impossible to accurately generate reply information, and the dialogue model is updated inefficiently and has high training costs, resulting in poor reply effects and less controllable.
By obtaining user input information, calling the dialogue engine to obtain related information in the business data of the enterprise, generating the first prompt information, and inputting it into the dialogue model for processing, improving the accuracy of the reply information.
It improves the dialogue model's understanding of user input information, enhances the accuracy and flexibility of replying information, and reduces the resource consumption and response time of the dialogue model.
Smart Images

Figure CN120336451A_ABST
Abstract
Description
Technical Field
[0001] This application relates to computer technology, and in particular to a human-machine dialogue method, an electronic device, and a storage medium. Background Art
[0002] With the rapid development of large model technology, more and more fields have started to apply large model technology, and large models such as large language models (LLMs) have also been applied to human-machine dialogue scenarios such as intelligent dialogue robots. When applying a large language model to a human-machine dialogue system, in each round of dialogue between the human-machine dialogue system and the user, the large language model needs to be called to generate a response message.
[0003] However, the applicant has found that when applying large language model technology to enterprise business Q&A scenarios, it is necessary to first train the large language model with the enterprise's business data before use. At present, the method adopted is to directly input the user input information into the large language model to obtain a response message, and this method cannot accurately generate a response message. Summary of the Invention
[0004] This application provides a human-machine dialogue method, an electronic device, and a storage medium, which are used to solve the problem that the current human-machine dialogue system cannot generate an accurate response message based on the user input information.
[0005] In a first aspect, this application provides a human-machine dialogue method, including: obtaining user input information; calling at least one dialogue engine, obtaining information associated with the user input information in the business data corresponding to the dialogue engine to obtain associated information; generating a first prompt message according to the user input information and the associated information; inputting the first prompt message into the called dialogue model for processing to obtain a response message corresponding to the user input information.
[0006] In a second aspect, this application provides a human-machine dialogue method, which is applied to a task-based dialogue engine. The method includes: receiving user input information; determining at least one target dialogue process in the dialogue process according to the user input information; if a call request carrying the target dialogue process and the user input information is received, generating a first prompt message according to the target dialogue process and the user input information; wherein, the first prompt message is used to indicate that the first prompt message is input into the called dialogue model for processing to obtain a response message corresponding to the user input information.
[0007] In a third aspect, this application provides a human-machine dialogue method, which is applied to a terminal device. The method includes: obtaining user input information and sending the user input information to a server; receiving the response message corresponding to the user input information sent by the server and displaying the response message.
[0008] Fourthly, the present application provides a human-machine dialogue method, which is applied to the server of an enterprise intelligent customer service. The method includes: obtaining user input information; calling at least one dialogue engine, obtaining information associated with the user input information from the enterprise business data corresponding to the dialogue engine to obtain enterprise-associated information; generating a first prompt message according to the user input information and the enterprise-associated information; inputting the first prompt message into the called dialogue model for processing to obtain a reply message corresponding to the user input information.
[0009] Fifthly, the present application provides an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores first instructions executable by the at least one processor, and when the first instructions are executed by the at least one processor, the processor is enabled to execute the methods of the foregoing first aspect, second aspect, third aspect or fourth aspect.
[0010] Sixthly, the present application provides a computer-readable storage medium, in which computer-executable first instructions are stored. When the processor executes the computer-executable first instructions, the methods of the first aspect, second aspect, third aspect or fourth aspect are implemented.
[0011] The human-machine dialogue method, electronic device and storage medium provided by the present application include: obtaining user input information;
[0012] calling at least one dialogue engine, obtaining information associated with the user input information from the business data corresponding to the dialogue engine to obtain associated information; generating a first prompt message according to the user input information and the associated information; inputting the first prompt message into the called dialogue model for processing to obtain a reply message corresponding to the user input information. Based on the user input information and the associated information, the present application generates a first prompt message and inputs the first prompt message into the dialogue model, which can enable the dialogue model to better understand the user input information based on the associated information, thereby improving the accuracy of generating the reply message. BRIEF DESCRIPTION OF THE DRAWINGS
[0013] The drawings herein are incorporated into the specification and constitute a part of the specification, showing embodiments consistent with the present application, and are used together with the specification to explain the principles of the present application.
[0014] Figure 1 It is a schematic diagram of an exemplary human-machine dialogue system applicable to the present application;
[0015] Figure 2 It is a flowchart of the first human-machine dialogue method provided by an exemplary embodiment of the present application;
[0016] Figure 3 It is a schematic diagram of the human-machine dialogue method provided by an exemplary embodiment of the present application;
[0017] Figure 4 The first flowchart for generating the first prompt message provided by an exemplary embodiment of the present application;
[0018] Figure 5 The second flowchart for generating the first prompt message provided by an exemplary embodiment of the present application;
[0019] Figure 6 The third flowchart for generating the first prompt message provided by an exemplary embodiment of the present application;
[0020] Figure 7 The flowchart of the second human - machine dialogue method provided by an exemplary embodiment of the present application;
[0021] Figure 8 The structural block diagram of the first human - machine dialogue device provided by an embodiment of the present application;
[0022] Figure 9 The structural block diagram of the second human - machine dialogue device provided by an embodiment of the present application;
[0023] Figure 10 The structural block diagram of the third human - machine dialogue device provided by an embodiment of the present application;
[0024] Figure 11 The structural block diagram of an electronic device provided by an embodiment of the present application.
[0025] Through the above - mentioned drawings, specific embodiments of the present application have been shown, and there will be more detailed descriptions hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed implementation manners
[0026] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.
[0027] It should be noted that the user information (including but not limited to user device information, user attribute information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in the present application are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant information need to comply with relevant laws, regulations, and standards, and corresponding operation entrances are provided for users to select authorization or rejection.
[0028] First, the terms involved in this application are explained as follows:
[0029] The dialogue model is a large language model (LLM), also known as a large-scale language model or a large language model. It is a model based on machine learning and natural language processing technologies. It learns the ability to serve human language understanding and generation by training on a large amount of text data. The core idea of the LLM is to learn the patterns and language structures of natural language through large-scale unsupervised training, which can to a certain extent simulate the human language cognition and generation process. Compared with traditional natural language processing (NLP) models, the LLM can better understand and generate natural text, and at the same time can also show certain logical thinking and reasoning abilities.
[0030] The human-machine dialogue system is a human-machine dialogue system implemented based on natural language understanding technology and dialogue management technology, and has wide applications in scenarios such as intelligent customer service, intelligent outbound calls, online education, office software, and e-commerce.
[0031] The dialogue control server, also known as the dialogue control center. There are multiple dialogue engines that a human-machine dialogue system needs to support. Among them, the dialogue engines include document dialogue engines, table dialogue engines, task-based dialogue engines, frequently-asked questions (FAQ) engines, chit-chat engines, etc. These dialogue engines support different usage scenarios respectively. Among them, the task-based dialogue engine is mainly aimed at multi-turn task-based question-and-answer scenarios such as querying the weather, booking hotels, and recharging traffic that can be abstracted into intents and slots; the FAQ engine is used to support dialogue scenarios based on question-and-answer pairs; the table dialogue engine is used to support dialogue scenarios based on table data in the database; the document dialogue engine is used to support dialogue scenarios based on document reading comprehension; the chit-chat engine is used to support open-domain dialogue scenarios with weak business relevance such as greetings and chit-chat. In the process of human-machine dialogue, the cooperation of multiple dialogue engines is often required. The dialogue control server realizes the scheduling and sorting capabilities of each dialogue engine in a multi-dialogue engine dialogue system, is the core module of the multi-dialogue engine dialogue system, and plays a very crucial role in improving the dialogue ability.
[0032] The enterprise's business data refers to the enterprise's unique business knowledge, including the enterprise's unique business dialogue processes, database data, documents, etc.
[0033] The prompt is an input data of a dialogue model, which is used to indicate how the dialogue model should process and / or generate output data when performing a specific task.
[0034] Large models refer to deep learning models with large-scale model parameters, usually containing hundreds of millions, tens of billions, or even hundreds of billions of model parameters. Large models can also be referred to as Foundation Models (FM). Through pre-training of large models with large-scale unlabeled corpora, pre-trained models with over hundreds of millions of parameters are produced. Such models can adapt to a wide range of downstream tasks and have good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0035] When large models are applied in practice, only a small number of samples are needed to fine-tune the pre-trained model for use in different tasks. Large models can be widely applied in the fields of natural language processing, computer vision, etc. Specifically, they can be applied to tasks in the field of computer vision such as Visual Question Answering (VQA), Image Caption (IC), image generation, etc., as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, machine translation, etc. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0036] The business scenarios of enterprises are generally very complex, and the knowledge of enterprises is also very complex, stored in the file system or database in the form of documents or structured data. How to build intelligent dialogue robots based on this enterprise knowledge to meet the needs of enterprise user services and intelligent outbound calls has been the direction of technical research in recent years.
[0037] Among them, before the emergence of large language model technology, the performance of NLP (Natural Language Processing) technology in document-based question answering was far lower than expected. Therefore, generally, enterprises would sort out common user questions and corresponding answers into QA pairs (question-answer pairs), and find the answer with the highest matching degree to the user question through the matching of QA pairs. This method has a high sorting cost and relatively low knowledge coverage, and cannot directly apply the richest document resources in enterprises. Secondly, in the aspect of question answering based on structured data, NL2SQL (a technology that converts natural language acquisition into structured acquisition language) performs well in simple acquisition scenarios, but its performance in complex acquisition scenarios also fails to meet the requirements of actual business applications. Finally, in the task-based question answering scenario, the dialogue flow is generally constructed according to the business process, and the processing nodes in the dialogue flow are found based on the NLU (intent recognition, entity recognition) results and context, and then the business is processed or a reply is obtained. The overall dialogue effect of this method is average, and it is also difficult to optimize.
[0038] After the emergence of the dialogue model, due to its high processing capabilities in the directions of natural language understanding, logical reasoning, response generation, etc., it has received increasing attention from more and more enterprises. It is necessary to apply the dialogue model and combine enterprise knowledge to build a reliable and controllable enterprise-level dialogue robot.
[0039] However, the applicant has found that the following technical problems are encountered in the current process of applying the dialogue model technology to the human-machine dialogue system: The current method is to directly train the enterprise's business data into the dialogue model and directly use the dialogue model to support the dialogue. The problem with this method is that the enterprise's knowledge is dynamically changing, the update efficiency of new knowledge in the dialogue model is low, the training cost is high, and the requirements for the capabilities of the dialogue model itself are very high, resulting in poor and less controllable effects when using the dialogue model for responses.
[0040] This application provides a human-machine dialogue method. When conducting a human-machine dialogue, by obtaining user input information; calling at least one dialogue engine, obtaining information associated with the user input information from the business data corresponding to the dialogue engine to obtain associated information; generating a first prompt message based on the user input information and the associated information; inputting the first prompt message into the called dialogue model for processing to obtain a response message corresponding to the user input information, it is realized that the dialogue model better understands the user input information based on the associated information, thereby improving the accuracy of the generated response message.
[0041] Figure 1 To show a schematic diagram of the interaction between the human-machine dialogue system and the terminal device. As Figure 1 shown, the human-machine dialogue system includes a dialogue control server and each dialogue engine.
[0042] Among them, the terminal device can be various devices that can obtain user input. The user input can be in the form of text or voice, etc. The terminal device can be a device with a screen or a device with voice interaction capabilities. Including but not limited to: intelligent mobile terminals, smart home devices, wearable devices, personal computers (abbreviated as PC), etc. Among them, intelligent mobile devices can include, for example, mobile phones, tablets, laptops, Internet cars, etc. Smart home devices can include smart home appliances, such as smart TVs, smart air conditioners, smart refrigerators, etc. Wearable devices can include, for example, smart watches, smart glasses, smart bracelets, virtual reality (abbreviated as VR) devices, augmented reality (abbreviated as AR) devices, mixed reality devices (i.e., devices that can support virtual reality and augmented reality), etc. Users can obtain user input information through the terminal device and send the user input information (Query) to the dialogue control server of the human-machine dialogue system.
[0043] The human-machine dialogue system includes: a dialogue control server and at least one dialogue engine. These dialogue engines respectively support different usage scenarios, including but not limited to engines for implementing various dialogue tasks such as task-based dialogue, frequently asked questions (FAQ), table question answering, document question answering, and chit-chat. Which dialogue engines the dialogue system supports can be configured and adjusted according to the actual application scenario requirements, and no specific limitation is made here.
[0044] During the human-machine dialogue process, the cooperation of multiple dialogue engines is often required. The dialogue control server's ability to implement the scheduling and sorting of each dialogue engine in a multi-dialogue engine dialogue system is the core module of the human-machine dialogue system and plays a very crucial role in enhancing the dialogue ability. The dialogue control server is responsible for obtaining the user input information, calling the dialogue engines in two stages. In the first stage, it can call each dialogue engine to generate associated information. In the second stage, it calls one of the target dialogue engines to generate a reply message. Then, the dialogue control server can return the reply message to the user terminal.
[0045] The above human-machine dialogue system can be set up on a single server, or the dialogue control server and dialogue engines can be respectively set up on a single server, or it can be set up in a server group composed of multiple servers, or it can also be a cloud server, and no limitation is imposed on this.
[0046] It should be understood that Figure 1 the numbers of the terminal devices, dialogue control servers, and dialogue engines in are merely illustrative. According to the implementation requirements, there can be any number of terminal devices, dialogue control servers, and dialogue engines.
[0047] Based on Figure 1 the human-machine dialogue system, when conducting a human-machine dialogue, the terminal device sends the user input information to the dialogue control server of the human-machine dialogue system. The dialogue control server, based on the user input information, calls at least one dialogue engine, obtains the information associated with the user input information in the business data corresponding to the dialogue engine, and gets the associated information; generates a first prompt message according to the user input information and the associated information; further, inputs the first prompt message into the called dialogue model for processing to obtain a reply message. Then, the dialogue returns the reply message to the terminal device.
[0048] Figure 2 The flowchart of the human-machine dialogue method provided by an exemplary embodiment of this application. The execution subject of this embodiment is Figure 1 the human-machine dialogue system in Figure 2 As shown in , the specific steps of this method are as follows:
[0049] S201, obtain the user input information.
[0050] Among them, the user input information refers to the user input information currently input by the user. The content input by the user can be in the form of text or voice, etc. The user input information refers to the text corresponding to the content input by the user. For the text content input by the user, it can be directly used as the user input information or used as the user input information after being rewritten. For the voice content input by the user, the voice content can be converted into text and used as the user input information.
[0051] S202, call at least one dialogue engine, and obtain information associated with the user input information from the business data corresponding to the dialogue engine to obtain associated information.
[0052] In the embodiments of the present application, referring to Figure 3 , multiple dialogue engines can be set. Each dialogue engine has corresponding types of business data, and the business data is the business data of the enterprise, such as the business data of the enterprise's documents corresponding to the document dialogue engine, such as the product description document of the product. The table dialogue engine corresponds to the business data of the tables in the enterprise database, such as the product information of the product. The task-based dialogue engine corresponds to the business data of the enterprise's dialogue flow, such as the SOP of the business. In addition, the business data also includes FAQ data, etc.
[0053] Specifically, the dialogue control server sends the user input information to each dialogue engine, and the dialogue engine obtains information associated with the user input information from the corresponding business data according to the user input information to obtain associated information. Further, if the dialogue engine fails to obtain information associated with the user input information, it can return a null value indicating no associated information.
[0054] In the embodiments of the present application, if none of the dialogue engines obtain associated information, the dialogue control server can return to the terminal device "unable to reply to the user input information", thereby avoiding the subsequent call of the dialogue model and improving the reply efficiency.
[0055] In addition, the associated information can be one or more. For example, if the document dialogue engine obtains an associated information, the table dialogue engine obtains an associated information, and the task-based dialogue engine does not obtain associated information, the obtained associated information is two.
[0056] Further, generate a first prompt message according to the user input information and the associated information, including: determine a target dialogue engine from at least one dialogue engine according to the associated information; call the target dialogue engine, and generate a first prompt message according to the user input information and the associated information.
[0057] Among them, determining a target dialogue engine from at least one dialogue engine according to the associated information includes: determining a target dialogue engine from at least one dialogue engine according to the number of the associated information.
[0058] Specifically, according to the number of associated information, a target dialogue engine is determined in at least one dialogue engine, including: if the number of associated information is 1, the dialogue engine that outputs the associated information in at least one dialogue engine is determined as the target dialogue engine; if the number of associated information is at least two, according to a preset priority strategy, the associated information with the highest priority is determined as the target associated information among at least two pieces of associated information; the dialogue engine that outputs the target associated information in at least one dialogue engine is determined as the target dialogue engine.
[0059] Referring to Figure 3 , at least one dialogue engine includes at least one of a document dialogue engine, a table dialogue engine, and a task-based dialogue engine. Among them, the dialogue control server sends the user input information to different dialogue engines respectively. If only one dialogue engine among all dialogue engines returns associated information.
[0060] For example, the document dialogue engine returns associated information A, or the table dialogue engine returns associated information B, or the task-based dialogue engine returns associated information C. Among them, the dialogue engine that returns the associated information is determined as the target dialogue engine. For example, if only associated information A is returned, the document dialogue engine is determined as the target dialogue engine. If only associated information B is returned, the table dialogue engine is determined as the target dialogue engine. If only associated information C is returned, the task-based dialogue engine is determined as the target dialogue engine.
[0061] In the embodiments of the present application, a preset priority strategy of the dialogue engine can be set in advance. Specifically, the preset priority strategy is the priority of the dialogue engine set according to the specific business needs of the enterprise, and this priority is adjustable according to the business. For example, the priority of the task-based dialogue engine is set to be greater than the priority of the table dialogue engine which is greater than the priority of the document dialogue engine.
[0062] Furthermore, if a dialogue engine has a high priority, the associated information output by it also has a high corresponding priority. For example, the priority of the associated information C output by the task-based dialogue engine is greater than the priority of the associated information B output by the table dialogue engine which is greater than the priority of the associated information A output by the document dialogue engine.
[0063] Exemplarily, if in the parallel scheduling stage, the document dialogue engine returns associated information A, the table dialogue engine returns associated information B, and the task-based dialogue engine returns a null value. Then associated information A and associated information B are obtained. Then according to the preset priority strategy, it is determined that the priority of the table dialogue engine is higher than the priority of the document dialogue engine. Furthermore, it can be determined that the priority of associated information B is higher than the priority of associated information A. Then the table dialogue engine that outputs associated information B can be determined as the target dialogue engine, and the table dialogue engine is called to process the associated information.
[0064] Furthermore, referring to Figure 3, after obtaining at least two pieces of associated information through parallel scheduling, it is also possible to first determine whether there is associated information C among the at least two pieces of associated information. If there is, the task-based dialogue engine is determined as the target dialogue engine. If not, it is determined whether there is associated information B among the at least two pieces of associated information. If there is, the form dialogue engine is determined as the target dialogue engine. If not, it is determined whether there is associated information A among the at least two pieces of associated information. If there is, the document dialogue engine is determined as the target dialogue engine. If not, the preset open-domain dialogue engine can be called to process the user input information. The open-domain dialogue engine serves as a fallback, which can improve the response rate of the dialogue and enhance the user experience.
[0065] Among them, the open-domain dialogue engine does not participate in the parallel scheduling in the first stage. In the scheduling process of the second stage, when none of the other dialogue engines are scheduled, the open-domain dialogue engine is used as the target dialogue engine. The context information of the user input information can be obtained based on the user input information, and then an open-domain search is performed according to the user input information and the context information to obtain the open-domain search result. The open-domain dialogue model (such as LLM) is used to generate the reply information of the user input information based on the user input information, the context information, and the open-domain search result.
[0066] S203, generating the first prompt information according to the user input information and the associated information.
[0067] Among them, the generated first prompt information includes: the first instruction, the user input information, and the associated information. The first instruction is used to instruct the dialogue model to reply to the user input information.
[0068] In an alternative embodiment, the first instruction, the user input information, and the associated information can be concatenated to obtain the first prompt information. For example, the first instruction is to reply to the user input information, the user input information is: Which wealth management products have a yield rate greater than 10%, and the associated information is that the yield rate of product A is 15% and the yield rate of product B is 20%.
[0069] The first prompt information is as follows: Suppose you are an intelligent assistant, please reply to the user input information according to the user input information: Which wealth management products have a yield rate greater than 10%, and the associated information: The yield rate of product A is 15% and the yield rate of product B is 20%.
[0070] In the embodiments of the present application, the concatenation method of the first instruction, the user input information, and the associated information is not limited.
[0071] In another alternative embodiment, it is also possible to process the user input information and the associated information to obtain other relevant information, and concatenate the first instruction, the user input information, the associated information, and the other relevant information to obtain the first prompt information.
[0072] Optionally, context information of the user input information can also be obtained, and based on the user input information, the context information, and the associated information, a target dialogue engine is called to generate a first prompt message. The first prompt message (prompt) further includes: the context information of the user input information, where the context information is the multi-round dialogue information between the user and the human-machine dialogue system before the current user input information.
[0073] S204, input the first prompt message into the called dialogue model for processing to obtain a reply message corresponding to the user input information.
[0074] In this application, the first prompt message spliced with the associated information obtained from the business data is input into the dialogue model. The first instruction can instruct the dialogue model to reply to the user input information according to the associated information to obtain a reply message. Therefore, the dialogue model of this application does not need to be pre-trained according to the business data. Therefore, even if the business data is flowing and of various types, it can be obtained according to the user input information and then input into the dialogue model, so that the dialogue model can learn accurate reply messages from the associated information.
[0075] In an alternative embodiment, referring to Figure 3 , after obtaining the user input information, it further includes: calling a frequently asked questions answering engine to reply to the user input information to obtain a reply result or a null value; where, if a reply result is obtained and the confidence level of the reply result is greater than or equal to the confidence level threshold, the reply result is returned to the user; if a null value is obtained, or a reply result is obtained and the confidence level of the reply result is less than the confidence level threshold, the step of understanding the user input information to obtain the first prompt message is executed.
[0076] In the embodiment of this application, referring to Figure 3 , the frequently asked questions answering engine (FAQ engine) only needs a matching model, and this matching model is a non-generative language model, where the non-generative language model is a small model with a small number of model parameters. The frequently asked questions answering engine does not need to call the dialogue model. When the user input information is obtained, if the frequently asked questions answering engine can be called first to obtain a reply result with a confidence level greater than the confidence level threshold, the subsequent dialogue engine does not need to be called, which can improve the reply efficiency of the user input information.
[0077] Specifically, enterprises will precipitate some high-quality FAQ data according to actual business data in history. These FAQ data can be directly used for the question and answer task after being processed. In the embodiment of this application, the FAQ data includes multiple question-answer pairs. The questions in the multiple question-answer pairs are vectorized to obtain question vectors, and these question vectors are stored in a vector library. At the same time, after the questions are tokenized, they are stored in a database in an inverted index manner.
[0078] In the dialogue process, when the FAQ engine receives the user input information sent by the dialogue control server, it tokenizes and vectorizes the user input information to obtain a target question vector. Then, through multi-channel retrieval (vector retrieval and inverted index retrieval), it obtains the question vectors similar to the target question vector in the vector library, as well as the similar questions of the FAQ data corresponding to these similar question vectors. Next, it determines the semantic similarity between the retrieved similar questions and the user input information, sorts the semantic similarity, and obtains the target similar question with the maximum semantic similarity. Then, it determines whether the confidence level of the target similarity is greater than or equal to the confidence threshold. If so, it finds the answer to the target similar question in the FAQ data as the reply result. If not, it returns a null value, indicating that there is no answer. In addition, if no similar questions are obtained during multi-channel retrieval, a null value can also be returned.
[0079] It can be understood that the FAQ engine is called first. First, it is judged whether the FAQ engine outputs a reply result with a high confidence level. If a reply result with a high confidence level is output, this reply result can be directly used to reply to the user input information. If not, the subsequent process shown in Figure 3 is carried out. Among them, the FAQ engine does not need to use the dialogue model and has a fast response time. Therefore, it is placed at the forefront of the dialogue process. If the FAQ engine outputs a reply result with a high confidence level, no dialogue model needs to be called.
[0080] It can be understood that if the enterprise's business scenario is relatively simple and only involves one of documents, databases, and task types, the single-engine solution can be directly used. If the enterprise's business scenario is relatively complex and involves multiple knowledge types, the multi-engine scheduling solution of the present application is adopted to organically combine the dialogue engines of various knowledge types. Moreover, the inference cost of the dialogue model is very high. Therefore, the present application adopts a smaller number of calls to the dialogue model, which can reduce unnecessary calls to the dialogue model in the dialogue execution link and improve the dialogue processing ability.
[0081] In the embodiments of the present application, first, the present application has good scalability and can reply to the user input information for various types of data (such as FAQ, document type, table type, and task type) without training a single-task dialogue model for each type of data. The user only needs to input the user input information without distinguishing which dialogue engine needs to be used for processing, improving the user experience. Second, it is not necessary to directly train the enterprise's business data into the dialogue model, nor is it necessary to train different dialogue models for different types of data. The dialogue model can be directly called to support the dialogue. Finally, the dialogue model does not need to be frequently scheduled, that is, it does not need to call multiple dialogue models in parallel or serially, reducing resource consumption and response time.
[0082] In summary, in the embodiments of the present application, by generating the first prompt information, the fluency of the conversation can be improved, and for the user input information of all business data types, the conversation model can be used for processing and reply, improving the reply effect. In addition, by adopting the multi-conversation engine scheduling method, it has high generality and flexibility. Combining the enterprise's business data during the conversation, the credibility of the reply information is high.
[0083] Furthermore, the calling times of the conversation model of the present application are few, improving the system's concurrent processing ability. And it can support flexible jumping between engines, with good conversation effects. In addition, it can be used as a general conversation platform solution to support the implementation of multiple services. Further, the present application has low requirements for the capabilities of the conversation model. Each call is a single task, and it does not require the conversation model to support particularly complex tasks, making it more feasible to implement. Finally, the present application has good scalability. It is not necessary to train enterprise knowledge into the conversation model, and the newly added enterprise knowledge can be directly used for question answering.
[0084] In an alternative embodiment, referring to Figure 4 , if the conversation engine is a document conversation engine, the business data includes: multiple document fragment vectors generated based on document data, and the associated information includes the target document fragment vector, then the document conversation engine performs the following steps:
[0085] Among them, most of the enterprise's business data is stored in documents. The present application can parse the documents into document texts that can be edited and processed. The document text may include: one or more of the title, paragraphs, document hierarchical structure, or the body text. The specific implementation process of parsing the documents to obtain the document text in the present application is not limited. Then, the document text is segmented into document fragments according to the segmentation strategy, where the segmentation strategy includes segmentation by a fixed length or by paragraphs. Then, the vectors of the document fragments are calculated to obtain the document fragment vectors, and the document fragment vectors are stored in the vector library. Further, it is also possible to calculate the vectors of the question-answer pairs in the FAQ data as the document fragment vectors and store them in the vector library as well. Then, the multiple document fragment vectors of the present application are generated based on the documents and the FAQ. After the processing of the document data is completed in advance, referring to Figure 4 , based on the document conversation engine, the following steps of the human-machine conversation method can be performed:
[0086] S401, determining the target question vector corresponding to the user input information.
[0087] Among them, when the document conversation engine receives the user input information sent by the conversation control server, it can generate the corresponding target question vector based on the user input information.
[0088] S402, obtaining multiple document fragment vectors, and determining the target document fragment vector similar to the obtained vector among the multiple document fragment vectors.
[0089] Then, based on multiple document fragment vectors of the target question vector in the vector library, retrieve the document fragment vectors similar to the target question vector. Further, calculate the similarity between the similar document fragment vectors and the target question vector, and then put the document fragment vectors with a correlation greater than the correlation threshold into the candidate set and sort them in descending order of correlation.
[0090] S403, obtain the context information of the user input information.
[0091] Among them, the context information of the user input information refers to the multi-round dialogue information before the user input information during this conversation.
[0092] S404, generate the first prompt information according to the context information, the user input information, and the target document fragment vector.
[0093] Specifically, determine the first instruction and the constraint information. The first instruction is used to instruct the dialogue model to generate a reply message according to the document. The constraint information is used to constrain the length and / or freedom degree of the generated reply message. Then, splice the first instruction, the context information, the user input information, the target document fragment vector, and the constraint information to obtain the first prompt information.
[0094] Then, after generating the first prompt information, the document dialogue engine can send the first prompt information to the dialogue control server, and the dialogue control server calls the dialogue model to process the first prompt information to generate a reply message.
[0095] In the embodiment of the present application, when using the document dialogue engine to process and generate a reply message, the dialogue model only needs to be called once, which can reduce resource consumption and response time, and the dialogue model does not need to be pre-trained with the enterprise's document data. The present application can use real-time business data to generate a reply message, which can improve the accuracy of the reply message.
[0096] In an alternative embodiment, for the table dialogue engine, the dialogue engine is a table dialogue engine, the business data includes: multiple tables in the database, and the association information includes: at least one target table among the multiple tables.
[0097] In the embodiment of the present application, an enterprise has a large amount of business data stored in the database. There are some typical scenarios that are very suitable for answering questions with NL2SQL. For example, the BI data (business intelligence data) and / or financial product introduction data stored in the database. Among them, the BI data is such as the monthly sales amount and / or the year-on-year growth rate or year-on-year decrease rate of the monthly sales amount. The financial product introduction data is such as the annualized rate of return of the financial product. Refer to Figure 5 , based on the table dialogue engine, the following steps of the human-computer dialogue method can be executed:
[0098] S501. Obtain information associated with the user input information from multiple tables in the database based on the user input information to obtain associated information, and send the associated information to the dialogue control server.
[0099] Specifically, NER (entity) recognition can be performed on the user input information to obtain multiple entities. Then, target tables having the multiple entities are obtained from multiple tables in the database according to the multiple entities.
[0100] In practical applications, the number of tables in the database is large, and it is impossible to input all the data in the tables into the dialogue model. Therefore, this application locates the target table related to the user input information among multiple tables.
[0101] Specifically, after identifying the entities in the user input information, search for the entities in the table name, table header, and table values of the table. If any, determine the table containing the entity as the target table.
[0102] S502. Obtain the context information of the user input information, and generate a second prompt message according to the user input information, context information, and at least one target table.
[0103] Specifically, the second prompt message includes a second instruction. After obtaining the target table, determine the second instruction. Among them, the second instruction is used to instruct the dialogue model to generate a database query statement. Specifically, the second instruction is used to instruct the dialogue model to generate a database query statement (SQL) according to the user input information.
[0104] Furthermore, extract the table fields and table field types of the target table, as well as the text information in some rows of the target table, and extract the relevant user input information related to the user input information in the database. Then, splice multiple items of the second instruction, user input information, dialogue context, table fields, table field types, text information in some rows of the target table, context information of the user input information, and relevant user input information into the second prompt message.
[0105] S503. Input the second prompt message into the dialogue model for processing to obtain a database query statement.
[0106] Among them, in this step, the dialogue model is called to process the second prompt message to obtain a database query statement (SQL). This database query statement is used to obtain the target data related in the target table.
[0107] S504. Obtain at least one target table according to the database query statement to obtain the target data in at least one target table.
[0108] Among them, relevant target data in the target table is obtained according to the database query statement. Since the amount of data in the target table is very large and cannot all be passed to the language, the step of target data in this step is also very important.
[0109] S505, generate a first prompt message according to the target data, the first instruction and the user input information.
[0110] Specifically, splice the first instruction, the user input information and the target data into the first prompt message.
[0111] In the embodiment of the present application, when using the data dialogue engine to process and generate the reply message, only need to call the dialogue model twice, which can reduce the resource consumption and response time, and the dialogue model does not need to be pre-trained with the table data in the enterprise database. The present application uses the table data in the real-time database to generate the reply message, which can improve the accuracy of the reply message.
[0112] In an alternative embodiment, for the task-based dialogue engine, the dialogue engine is a task-based dialogue engine, the business data includes: multiple dialogue processes, and the association information includes at least one target dialogue process.
[0113] Among them, in addition to the data based on documents and tables, enterprises also have some business processes, which are used to help users handle business, such as booking hotels, checking balances, etc. The same business will have some differences in business processes in different enterprises. These business processes need to be controllable, and over time, the business processes will change. In the embodiment of the present application, dialogue processes are used to represent these business processes. The dialogue processes will define the corresponding trigger intents and example question forms of the intents. Then, in the dialogue process, the dialogue process of the specified business process can be found according to the user input information of the user to complete the dialogue.
[0114] Specifically, referring to Figure 6 , the human-machine dialogue method based on the task-based dialogue engine specifically includes the following steps:
[0115] S601, receive the user input information.
[0116] Among them, the user input information is sent by the dialogue control server. Specifically, the content input by the user can be in the form of text, voice, etc. The user input information refers to the text corresponding to the content input by the user. For the text content input by the user, it can be directly used as the user input information or used as the user input information after being rewritten. For the voice content input by the user, the voice content can be converted into text and used as the user input information.
[0117] S602, determine at least one target dialogue process in the dialogue process according to the user input information.
[0118] Among them, according to the user input information, at least one target dialogue flow is determined in the dialogue process, including: obtaining the context information of the user input information, and identifying the intent and entity in the user input information and the context information; generating a third prompt information according to the user input information, the context information, the intent and the entity, where the third prompt information includes a third instruction for instructing the dialogue model to identify the target intent and the target entity according to the user input information and the context information; inputting the third prompt information into the dialogue model for processing to obtain the target intent and the target entity; determining at least one target dialogue flow from multiple dialogue flows according to the target intent and / or the target entity
[0119] Among them, the context information includes: the multi-round dialogue information closest to the user input information or the retrieved multi-round dialogue information related to the user input information.
[0120] Furthermore, both the user input information and the context information are identified to obtain at least one intent and / or at least one entity.
[0121] Among them, the third prompt information includes a third instruction for instructing the dialogue model to identify the target intent and the target entity according to the user input information and the context information.
[0122] Specifically, if the number of identified intents is small, all intents can be used to splice the third prompt information; if the number of identified intents is large, some intents with higher relevance are determined according to the question and the context information for splicing the third prompt information.
[0123] In the embodiment of the present application, relevant user input information with a similar way of asking as the user input information can also be obtained. The third instruction, the user input information, the context information, the intent, the relevant user input information, and the entity are spliced into the third indication information.
[0124] In this step, the dialogue model is called, and the third prompt information is input into the dialogue model to obtain the natural language understanding result, which includes: the target intent and the target entity.
[0125] Furthermore, if the target intent is a trigger intent, the corresponding dialogue flow is found according to the trigger intent as the target dialogue flow. If the trigger intent is not recognized, the target dialogue flow is searched according to the currently recognized target entity. If the corresponding dialogue flow cannot be found according to the target intent and / or the target entity, the nearest uncompleted dialogue flow in the historical dialogue is found. If the dialogue flow cannot be found according to the target intent and / or the target entity and there is no uncompleted dialogue flow, a null value is returned.
[0126] S603, if a call request carrying association information and user input information from the conversation control server is received, a first prompt message is generated according to the association information and the user input information.
[0127] Specifically, the first instruction, the user input information, the context information, and the target conversation flow are concatenated to obtain the first prompt message. The first prompt message is used to indicate that the first prompt message is input into the called conversation model for processing to obtain a reply message corresponding to the user input information.
[0128] In the embodiment of the present application, when using the task conversation engine to process and generate the reply message, the conversation model only needs to be called twice, which can reduce resource consumption and response time. And this conversation model does not need to be pre-trained with the enterprise's conversation flow, and can generate the reply message using the real-time task flow, which can improve the accuracy of the reply message.
[0129] In addition, since the number of stored conversation flows is large, it is impossible to call the conversation model once to process all the conversation flows. Therefore, the target conversation flow is first located and then the conversation model is called for processing. In this embodiment, the conversation model needs to be called twice, and the scalability requirement for the conversation model is relatively low.
[0130] Refer to Figure 7 , the embodiment of the present application also provides a human-computer conversation method, which is applied to a terminal device. The human-computer conversation method includes the following steps:
[0131] S701, obtain user input information and send the user input information to the server.
[0132] In the embodiment of the present application, a human-computer conversation system is deployed in the server. It is the terminal device that obtains the user input information and then sends the user input information to the server.
[0133] Specifically, the content input by the user to the terminal device can be in the form of text or voice, etc. The user input information refers to the text corresponding to the content input by the user. For the text content input by the user, it can be directly used as the user input information or used as the user input information after being rewritten. For the voice content input by the user, the voice content can be converted into text and used as the user input information.
[0134] Among them, obtaining the user input information includes: obtaining the initial information input by the user, performing intent understanding on the initial information to obtain multiple understanding information; displaying the multiple understanding information; and determining the user input information according to one of the understanding information selected by the user.
[0135] Exemplarily, if the initial information input by the user is "Please recommend wealth management products with a yield greater than 15% and a yield greater than 10%", the terminal device can obtain two understanding messages by understanding this initial information. One is "Do you need to recommend wealth management products with a yield greater than 15%?", and the other understanding message is "Do you need to recommend wealth management products with a yield greater than 10%?". These two understanding messages can be displayed on the terminal device for the user to select one of them. The understanding message selected by the user is determined as the user input information. If the understanding message selected by the user is "Do you need to recommend wealth management products with a yield greater than 15%?", then the user input information determined according to this understanding message is "Please recommend wealth management products with a yield greater than 15%".
[0136] S702, receive the reply message corresponding to the user input information sent by the server, and display the reply message.
[0137] Furthermore, after the terminal device displays the reply message to the user, it can further obtain the user input information input by the user based on this reply message, and loop the above steps.
[0138] In the embodiment of the present application, the user can directly input the user input information that he wants to consult. The terminal device sends this user input information to the server, and the server can quickly generate a reply message based on different types of business data. The user does not need to instruct the server to generate a reply message based on which type of business data, improving the user experience.
[0139] Furthermore, the embodiment of the present application also provides a human-machine dialogue method, which is applied to the server of an enterprise intelligent customer service. The human-machine dialogue method includes: obtaining user input information; calling at least one dialogue engine, and obtaining information associated with the user input information in the enterprise business data corresponding to the dialogue engine to obtain enterprise association information; generating a first prompt message according to the user input information and the enterprise association information; inputting the first prompt message into the called dialogue model for processing to obtain a reply message corresponding to the user input information.
[0140] Among them, generating a first prompt message according to the user input information and the enterprise association information includes: determining a target dialogue engine in at least one dialogue engine according to the enterprise association information; calling the target dialogue engine, and generating a first prompt message according to the user input information and the enterprise association information.
[0141] In the embodiment of the present application, the human-machine dialogue system shown is deployed in the server of the enterprise intelligent customer service Figure 1 for accurately replying to the user input information based on the enterprise's business data. The specific implementation process refers to the above embodiment and will not be elaborated here.
[0142] Figure 8Schematic diagram of the structure of a human-machine dialogue device 80 provided by an embodiment of the present application. The human-machine dialogue device 80 includes:
[0143] An acquisition module 801, configured to acquire user input information;
[0144] A call module 802, configured to call at least one dialogue engine, and acquire information associated with the user input information from the service data corresponding to the dialogue engine to obtain associated information;
[0145] A generation module 803, configured to generate a first prompt message according to the user input information and the associated information;
[0146] A processing module 804, configured to input the first prompt message into the called dialogue model for processing to obtain a reply message corresponding to the user input information.
[0147] In an alternative embodiment, the generation module 803 is specifically configured to: determine a target dialogue engine from at least one dialogue engine according to the associated information; call the target dialogue engine, and generate a first prompt message according to the user input information and the associated information.
[0148] In an alternative embodiment, the at least one dialogue engine includes: a document dialogue engine, and the service data corresponding to the document dialogue engine includes: a plurality of document fragment vectors generated based on document data, and the associated information includes a target document fragment vector. When the generation module 803 acquires information associated with the user input information from the service data corresponding to the dialogue engine to obtain associated information, it is specifically configured to: determine a target question vector corresponding to the user input information; acquire a plurality of document fragment vectors, and determine a target document fragment vector similar to the acquired vector from the plurality of document fragment vectors.
[0149] In an alternative embodiment, the target dialogue engine is a document dialogue engine. When the generation module 803 generates a first prompt message according to the user input information and the associated information, it is specifically configured to: acquire context information of the user input information; generate a first prompt message according to the context information, the user input information, and the target document fragment vector.
[0150] In an alternative embodiment, at least one dialogue engine includes: a table dialogue engine. The service data corresponding to the table dialogue engine includes: multiple tables in a database. The association information includes: at least one target table among the multiple tables. When generating the first prompt message according to the user input information and the association information, the generating module 803 is specifically configured to: obtain the context information of the user input information, and generate a second prompt message according to the user input information, the context information, and at least one target table; input the second prompt message into a dialogue model for processing to obtain a database query statement; obtain at least one target table according to the database query statement to obtain target data in at least one target table; and generate the first prompt message according to the target data, the first instruction, and the user input information.
[0151] In an alternative embodiment, at least one dialogue engine includes: a task-oriented dialogue engine. The service data corresponding to the task-oriented dialogue engine includes: multiple dialogue flows. The association information includes at least one target dialogue flow. When the calling module 802 obtains the information associated with the user input information in the service data corresponding to the dialogue engine to obtain the association information, it is specifically configured to: obtain the context information of the user input information, and identify the intent and entity in the user input information and the context information; generate a third prompt message according to the user input information, the context information, the intent, and the entity. The third prompt message includes a third instruction for instructing the dialogue model to identify the target intent and the target entity according to the user input information and the context information; input the third prompt message into the dialogue model for processing to obtain the target intent and the target entity; and determine at least one target dialogue flow among the multiple dialogue flows according to the target intent and / or the target entity.
[0152] In an alternative embodiment, the first prompt message includes: a first instruction, the user input information, and the association information.
[0153] In an alternative embodiment, when determining the target dialogue engine in at least one dialogue engine according to the association information, the generating module 803 is specifically configured to: determine the target dialogue engine in at least one dialogue engine according to the quantity of the association information.
[0154] In an alternative embodiment, when determining the target dialogue engine in at least one dialogue engine according to the quantity of the association information, the generating module 803 is specifically configured to: if the quantity of the association information is 1, determine the dialogue engine that outputs the association information in at least one dialogue engine as the target dialogue engine; if the quantity of the association information is at least two, determine the association information with the highest priority among the at least two association information according to a preset priority strategy as the target association information; and determine the dialogue engine that outputs the target association information in at least one dialogue engine as the target dialogue engine.
[0155] The specific implementation process of the human-machine dialogue device 80 provided in this application refers to the corresponding method embodiment above, and will not be elaborated here.
[0156] Figure 9 FIG. 4 is a schematic structural diagram of a human-machine dialogue device 80 provided in an embodiment of the present application, which is applied to a task-based dialogue engine. The human-machine dialogue device 90 includes:
[0157] A receiving module 901, configured to receive user input information;
[0158] A determining module 902, configured to determine at least one target dialogue process in the dialogue process according to the user input information;
[0159] A generating module 903, configured to, if a call request carrying the target dialogue process and the user input information is received, generate a first prompt message according to the target dialogue process and the user input information; wherein, the first prompt message is used to indicate that the first prompt message is input into the called dialogue model for processing to obtain a reply message corresponding to the user input information.
[0160] In an alternative embodiment, the determining module 902 is specifically configured to: obtain context information of the user input information, and identify the intent and entity in the user input information and the context information; generate a third prompt message according to the user input information, the context information, the intent and the entity, where the third prompt message includes a third instruction, and the third instruction is used to indicate that the dialogue model identifies the target intent and the target entity according to the user input information and the context information; input the third prompt message into the dialogue model for processing to obtain the target intent and the target entity; determine at least one target dialogue process in multiple dialogue processes according to the target intent and / or the target entity.
[0161] The specific implementation process of the human-machine dialogue device 90 provided in this application refers to the corresponding method embodiment above, and will not be elaborated here.
[0162] Figure 10 FIG. 5 is a schematic structural diagram of a human-machine dialogue device 80 provided in an embodiment of the present application, which is applied to a terminal device. The human-machine dialogue device 100 includes:
[0163] An obtaining module 101, configured to obtain user input information and send the user input information to the server;
[0164] A receiving module 102, configured to receive a reply message corresponding to the user input information sent by the server and display the reply message.
[0165] In an alternative embodiment, the obtaining module 101 is specifically configured to obtain initial information input by the user, perform intent understanding on the initial information to obtain multiple understanding messages; display the multiple understanding messages; determine the user input information according to one of the understanding messages selected by the user.
[0166] The specific implementation process of the human-machine dialogue device 100 provided in this application refers to the corresponding method embodiments described above, and will not be elaborated here.
[0167] Figure 11 It is a schematic structural diagram of an electronic device provided in an embodiment of this application. The electronic device can be on a cloud server. As Figure 11 shown, the electronic device includes: a memory 111 and a processor 112. The memory 111 is used to store a first computer-executable instruction and can be configured to store various other data to support operations on the cloud server. The processor 112 is communicatively connected to the memory 111 and is used to execute the first computer-executable instruction stored in the memory 111 to implement the technical solutions provided in any of the above method embodiments. Its specific functions and achievable technical effects are similar and will not be elaborated here. Optionally, as Figure 11 shown, the electronic device further includes: a firewall 113, a load balancer 114, a communication component 115, a power supply component 116, and other components. Figure 11 Only some components are schematically shown in the figure, and it does not mean that the dialogue engine only includes Figure 11 the components shown in the figure.
[0168] An embodiment of this application also provides a computer-readable storage medium. A first computer-executable instruction is stored in the computer-readable storage medium. When the processor executes the first computer-executable instruction, the method processes executed by the dialogue control server or the dialogue engine in any of the foregoing embodiments are implemented. Its specific functions and achievable technical effects will not be elaborated here.
[0169] An embodiment of this application also provides a computer program product, including a computer program. When the computer program is executed by the processor, the method in any of the foregoing embodiments is implemented. The computer program is stored in a readable storage medium. At least one processor of the electronic device can read the computer program from the readable storage medium, and at least one processor executing the computer program causes the electronic device to execute the method processes executed by the dialogue control server or the dialogue engine in any of the above method embodiments. Its specific functions and achievable technical effects will not be elaborated here.
[0170] An embodiment of this application provides a chip, including: a processing module and a communication interface. The processing module can execute the technical solutions of the electronic device in any of the foregoing method embodiments. Optionally, the chip further includes a storage module (such as a memory). The storage module is used to store a first instruction, and the processing module is used to execute the first instruction stored in the storage module, and the execution of the first instruction stored in the storage module causes the processing module to execute the method processes executed by the dialogue control server or the dialogue engine in any of the foregoing method embodiments.
[0171] It should be understood that the above-mentioned processor may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The steps of the method disclosed in combination with the application can be directly embodied as being executed and completed by a hardware processor, or by a combination of hardware and software modules in the processor. The memory may include high-speed random access memory (RAM), and may also include non-volatile storage, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disc, etc.
[0172] The above-mentioned memory may be an object storage service (OSS). The above-mentioned memory may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disc.
[0173] The above communication component is configured to facilitate communication, either wired or wireless, between the device where the communication component is located and other devices. The device where the communication component is located can access wireless networks based on communication standards, such as mobile hotspots (WiFi), second-generation mobile communication systems (2G), third-generation mobile communication systems (3G), fourth-generation mobile communication systems (4G) / Long Term Evolution (LTE), fifth-generation mobile communication systems (5G), etc. mobile communication networks, or combinations thereof. In an exemplary embodiment, the communication component receives broadcast signals or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component further includes a Near Field Communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on Radio Frequency Identification (RFID) technology, infrared technology, Ultra Wide Band (UWB) technology, Bluetooth technology, and other technologies. The above power supply component provides power to various components of the device where the power supply component is located. The power supply component may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device where the power supply component is located.
[0174] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk. The storage medium can be any available medium that can be accessed by a general-purpose or dedicated computer. An exemplary storage medium is coupled to the processor so that the processor can read information from and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an application-specific integrated circuit. Of course, the processor and the storage medium can also exist as discrete components in an electronic device or a master device.
[0175] It should be noted that in this article, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0176] The order of the above embodiments of the present application is only for description and does not represent the superiority or inferiority of the embodiments. Additionally, in some of the processes described in the above embodiments and the accompanying drawings, there are multiple operations that appear in a specific order. However, it should be clearly understood that these operations may not be executed in the order in which they appear herein or may be executed in parallel. They are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. Additionally, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions such as "first" and "second" in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequence, nor do they limit that "first" and "second" are of different types. The meaning of "a plurality" is more than two, unless otherwise specifically defined.
[0177] From the description of the above embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disc) and includes several first instructions for causing a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods of the various embodiments of the present application.
[0178] After considering the specification and practicing the invention disclosed herein, those skilled in the art will readily conceive of other embodiments of the present application. The present application is intended to cover any variations, uses, or adaptations of the present application that follow the general principles of the present application and include common general knowledge or conventional technical means in the technical field not disclosed in the present application.
[0179] The above are only the preferred embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural or equivalent process transformations made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, are equally included in the patent protection scope of the present application.
Claims
1. A human-machine dialogue method, characterized in that, The method includes: Obtaining user input information; Invoking at least one dialogue engine, and obtaining information associated with the user input information from the service data corresponding to the dialogue engine to obtain associated information; Generating a first prompt message according to the user input information and the associated information; Inputting the first prompt message into the invoked dialogue model for processing to obtain a reply message corresponding to the user input information.
2. The human-machine dialogue method according to claim 1, characterized in that The generating the first prompt message according to the user input information and the associated information includes: Determining a target dialogue engine from the at least one dialogue engine according to the associated information; Invoking the target dialogue engine, and generating a first prompt message according to the user input information and the associated information.
3. The human-machine dialogue method according to claim 2, wherein The at least one dialogue engine includes: a document dialogue engine, and the service data corresponding to the document dialogue engine includes: a plurality of document fragment vectors generated based on document data, and the associated information includes a target document fragment vector. The obtaining information associated with the user input information from the service data corresponding to the dialogue engine to obtain associated information includes: Determining a target question vector corresponding to the user input information; Obtaining the plurality of document fragment vectors, and determining a target document fragment vector similar to the obtained vector from the plurality of document fragment vectors.
4. The human-machine dialogue method according to claim 3, wherein The target dialogue engine is the document dialogue engine. The generating the first prompt message according to the user input information and the associated information includes: Obtaining context information of the user input information; Generating the first prompt message according to the context information, the user input information, and the target document fragment vector.
5. The human-machine dialogue method according to claim 2, characterized in that, The at least one dialogue engine includes: a table dialogue engine, and the service data corresponding to the table dialogue engine includes: a plurality of tables in a database, and the associated information includes at least one target table among the plurality of tables. The generating the first prompt message according to the user input information and the associated information includes: Obtaining context information of the user input information, and generating a second prompt message according to the user input information, the context information, and the at least one target table; Inputting the second prompt message into the dialogue model for processing to obtain a database query statement; Obtaining the at least one target table according to the database query statement to obtain target data in the at least one target table; Generating the first prompt message according to the target data, the first instruction, and the user input information.
6. The human-machine dialogue method according to claim 5, wherein The at least one dialogue engine includes: a task-based dialogue engine, and the service data corresponding to the task-based dialogue engine includes: a plurality of dialogue processes, and the associated information includes at least one target dialogue process. The obtaining information associated with the user input information from the service data corresponding to the dialogue engine to obtain associated information includes: Obtaining context information of the user input information, and identifying intents and entities in the user input information and the context information; Generate a third prompt message based on the user input information, the context information, the intent, and the entity. The third prompt message includes a third instruction for instructing the dialogue model to identify a target intent and a target entity based on the user input information and the context information; Input the third prompt message into the dialogue model for processing to obtain the target intent and the target entity; Determine the at least one target dialogue flow among the multiple dialogue flows according to the target intent and / or the target entity; 7. The human-machine dialogue method according to any one of claims 1 to 6, characterized in that The first prompt message includes: a first instruction, the user input information, and the associated information.
8. The human-machine dialogue method according to any one of claims 1 to 6, characterized in that, The determining a target dialogue engine from the at least one dialogue engine according to the associated information includes: Determine a target dialogue engine from the at least one dialogue engine according to the quantity of the associated information.
9. The human-machine dialogue method according to claim 8, wherein The determining a target dialogue engine from the at least one dialogue engine according to the quantity of the associated information includes: If the quantity of the associated information is 1, determine the dialogue engine that outputs the associated information among the at least one dialogue engine as the target dialogue engine; If the quantity of the associated information is at least two, determine the associated information with the highest priority among the at least two associated information as the target associated information according to a preset priority strategy; determine the dialogue engine that outputs the target associated information among the at least one dialogue engine as the target dialogue engine.
10. The human-machine dialogue method according to any one of claims 1 to 6, characterized in that, After obtaining the user input information, it further includes: Call a frequently asked questions answering engine to reply to the user input information to obtain a reply result or a null value; If the reply result is obtained and the confidence of the reply result is greater than or equal to the confidence threshold, return the reply result to the user; If the null value is obtained, or the reply result is obtained and the confidence of the reply result is less than the confidence threshold, perform the step of understanding the user input information to obtain the first prompt message.
11. A human-machine dialogue method, characterized in that, Applied to a task-based dialogue engine, the method includes: Receive user input information; Determine at least one target dialogue flow in the dialogue flow according to the user input information; If a call request carrying the target dialogue flow and the user input information is received, generate a first prompt message according to the target dialogue flow and the user input information; wherein, the first prompt message is used to instruct to input the first prompt message into the called dialogue model for processing to obtain the reply information corresponding to the user input information.
12. The human-machine dialogue method according to claim 11, wherein, The determining at least one target dialogue flow in the dialogue flow according to the user input information includes: Obtain the context information of the user input information, and identify the intent and entity in the user input information and the context information; Generate a third prompt message based on the user input information, the context information, the intent, and the entity. The third prompt message includes a third instruction for instructing the dialogue model to identify a target intent and a target entity based on the user input information and the context information; Input the third prompt message into the dialogue model for processing to obtain a target intent and a target entity; Determine the at least one target dialogue flow from the multiple dialogue flows according to the target intent and / or the target entity.
13. A human-machine dialogue method, characterized in that, Applied to a terminal device, the method includes: Obtain user input information and send the user input information to the server; Receive the reply information corresponding to the user input information sent by the server and display the reply information.
14. The human-machine dialogue method according to claim 13, characterized in that, The obtaining of the user input information includes: Obtain the initial information input by the user and perform intent understanding on the initial information to obtain multiple understanding messages; Display the multiple understanding messages; Determine the user input information according to one of the understanding messages selected by the user.
15. A human-machine dialogue method, characterized in that, Applied to a server of an enterprise intelligent customer service, the method includes: Obtain user input information; Call at least one dialogue engine, and obtain information associated with the user input information from the enterprise business data corresponding to the dialogue engine to obtain enterprise associated information; Generate a first prompt message according to the user input information and the enterprise associated information; Input the first prompt message into the called dialogue model for processing to obtain the reply information corresponding to the user input information.
16. The human-machine dialogue method according to claim 15, wherein The generating of the first prompt message according to the user input information and the enterprise associated information includes: Determine a target dialogue engine from the at least one dialogue engine according to the enterprise associated information; Call the target dialogue engine and generate a first prompt message according to the user input information and the enterprise associated information.
17. An electronic device, characterized in that, Includes: At least one processor; And A memory communicatively connected to the at least one processor; Wherein, the memory stores first instructions executable by the at least one processor, and the first instructions are executed by the at least one processor to enable the processor to execute the method according to any one of claims 1-16.
18. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executed first instructions, and when the processor executes the computer-executed first instructions, the method according to any one of claims 1-16 is implemented.