Intelligent dialogue method and system
By obtaining user input and dialogue context in the intelligent dialogue system, providing it to multiple dialogue engines and sorting it, the problem of multi-engine collaborative work is solved, achieving a more reasonable dialogue process and higher dialogue effect.
Patent Information
- Application Number
- CN202210515926.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-12
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2042-05-12
AI Technical Summary
In an intelligent dialogue system, how multiple dialogue engines work together to make the dialogue process more reasonable has become a technical difficulty.
By obtaining the content and dialogue context input by the user in this round of conversation, it is provided to multiple different dialogue engines, and sorting the reply text based on confidence, selecting the reply text that meets the preset requirements as the first reply text.
The collaborative work of multiple dialogue engines is realized, making the generated dialogue more reasonable, improving the dialogue effect, and reducing system operation and management costs.
Smart Images

Figure CN114860910B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of artificial intelligence technology, and in particular to an intelligent dialogue method and system. Background Art
[0002] With the continuous development of science and technology, smart devices are gradually penetrating into our lives. Among them, intelligent dialogue systems are increasingly used in various smart devices, such as smart speakers, smart TVs, smart car navigation, smart phones, and so on. When applying intelligent dialogue systems in various products, different dialogue engines are often required to support different service types and knowledge types. For example, the task-based dialogue engine supports task-based question-and-answer scenarios such as checking the weather, booking hotels, and recharging traffic; the FAQ (Frequently Asked Questions) dialogue engine supports dialogue scenarios based on question-and-answer pairs. The graph question-and-answer engine supports dialogue scenarios based on knowledge graph data; the table question-and-answer engine supports dialogue scenarios based on table data; the document question-and-answer engine supports dialogue scenarios based on document reading comprehension; the chat dialogue engine supports open domain dialogue scenarios based on greetings, chats, and other services that are not strongly related.
[0003] In the process of human-computer dialogue, multiple dialogue engines are often required to cooperate. For users, the existence of multiple dialogue engines in the intelligent dialogue system is not perceived. In fact, multiple dialogue engines are often used back and forth during the dialogue process. For example, in the following scenario, the user inputs the following voices in multiple rounds of dialogue:
[0004] "What outlets does China Mobile have in Chaoyang District, Beijing?" - Suitable for table question answering engines;
[0005] "What is a high-data package?" - suitable for knowledge question answering engines;
[0006] "What 5G packages are available?" — Suitable for graph question-answering engines;
[0007] "Help me get a 5G unlimited package?" - Suitable for task-based dialogue engines to support.
[0008] Therefore, how multiple dialogue engines work together to make the dialogue process more reasonable becomes a technical difficulty in intelligent dialogue systems. Summary of the invention
[0009] In view of this, the present application provides an intelligent dialogue method and system to achieve the collaborative work of multiple dialogue engines, making the generated dialogue more reasonable.
[0010] This application provides the following solutions:
[0011] According to a first aspect, an intelligent dialogue method is provided, comprising:
[0012] Get the content and conversation context of the user's current conversation;
[0013] Providing the content and conversation context of the user's current conversation input to N different conversation engines, where N is a positive integer greater than 1;
[0014] Sorting the reply texts obtained by the N different dialogue engines based on confidence, and selecting a reply text whose sorting meets a preset requirement as the first reply text;
[0015] The first reply text is used to generate content returned to the user in this round.
[0016] According to an achievable manner in an embodiment of the present application, before providing the content and conversation context input by the user in this round of conversation to N different conversation engines, the method further includes:
[0017] Determine whether the content input by the user in this round of dialogue hits the preset pre-dialogue strategy. If so, execute the hit pre-dialogue strategy; otherwise, continue to execute the step of providing the content input by the user in this round of dialogue and the dialogue context to N different dialogue engines.
[0018] According to an achievable method in the embodiment of the present application, the pre-dialogue strategy includes at least one of the following:
[0019] Re-listening strategy: if the content input by the user in this round of dialogue has the intention of re-listening, the content returned to the user in the previous round of dialogue is returned to the user again;
[0020] Emotion strategy: if the content input by the user in this round of dialogue is identified as a preset emotion type, then perform the pre-processing corresponding to the identified emotion type;
[0021] Sensitive word strategy: if the content input by the user in this round of dialogue is identified as a preset sensitive word, a second reply text is generated using the preset first speech technique, and the second reply text is returned to the user;
[0022] Manual transfer strategy: If the content input by the user in this round of dialogue is identified as being intended to be transferred to manual processing, the user will be connected to the manual processing system.
[0023] According to an achievable manner in an embodiment of the present application, providing the content and conversation context input by the user in this round of dialogue to N different dialogue engines includes: scheduling the natural language understanding NLU interfaces of the N different dialogue engines in parallel, and passing the content and conversation context input by the user in this round of dialogue as parameters to each dialogue engine;
[0024] Using the first reply text to generate content returned to the user in this round includes: taking the dialogue engine corresponding to the first reply text among the N different dialogue engines as the target dialogue engine for this round of dialogue; calling the execution interface of the target dialogue engine to trigger the target dialogue engine to execute the first reply text.
[0025] According to an achievable method in an embodiment of the present application, the reply texts obtained by the N different dialogue engines are sorted based on confidence, and the reply text whose sorting meets the preset requirements is selected as the first reply text, including:
[0026] The confidences of the reply texts obtained by the N different dialogue engines are sorted in combination with the weights or priorities of the dialogue engines, and the M reply texts with the highest rankings are selected as the first reply texts, where M is a positive integer;
[0027] The priority of each dialogue engine is preset according to the type of the dialogue engine.
[0028] According to an achievable method in an embodiment of the present application, the confidences of the reply texts obtained by the N different dialogue engines are sorted in combination with the weight or priority of each dialogue engine, and the M reply texts with the highest sorting are selected as the first reply text, including:
[0029] If the maximum confidence of the reply texts obtained by the N different dialogue engines is greater than or equal to the reply text of the preset first confidence threshold, then the reply text with the highest priority ranking of the corresponding dialogue engine is selected as the first reply text from the reply texts with confidence greater than or equal to the preset first confidence threshold; or,
[0030] If the maximum confidence of the reply texts obtained by the N different dialogue engines is less than a preset first confidence threshold and greater than a preset second confidence threshold, the reply texts with a confidence less than the preset first confidence threshold and greater than the preset second confidence threshold are scored based on the confidence and the dialogue engine weight, and the reply text with the highest score is selected as the first reply text; or,
[0031] If the reply texts obtained by the N different dialogue engines are clarification rhetoric, then M reply texts are selected as the first reply text according to the priority of each dialogue engine, wherein the clarification rhetoric is generated when the dialogue engine cannot generate a reply text with a confidence greater than or equal to a preset third confidence threshold;
[0032] The first confidence threshold is greater than the second confidence threshold, and the second confidence threshold is greater than or equal to the third confidence threshold.
[0033] According to an achievable manner in the embodiment of the present application, using the first reply text to generate content returned to the user in this round includes:
[0034] If the first reply text and the content returned to the user in the previous round come from different dialogue engines and the content returned to the user in the previous round is a clarifying rhetorical question, then the content of the previous round of user dialogue is used to generate a third reply text, and the first reply text and the third reply text are used to generate the content returned to the user; otherwise, the first reply text is used to generate the content returned to the user in this round.
[0035] According to an achievable manner in the embodiment of the present application, the method further includes: if the N different dialogue engines do not obtain a reply text, performing the following post-processing:
[0036] Generate a fourth reply text for guiding the user into the historical conversation scene based on the preset second speech and the historical conversation content, and use the fourth reply text to generate the content returned to the user in this round; or,
[0037] Based on the preset third speech and hot issue information, a fifth reply text is generated to prompt the user to consult the hot issue, and the fifth reply text is used to generate the content returned to the user in this round; or,
[0038] generating a sixth reply text for informing the user of the hotspot service based on the preset fourth speech and the hotspot service information, and using the sixth reply text to generate the content returned to the user in this round; or,
[0039] If the number of rounds in which the N different dialogue engines have not obtained a reply text reaches a preset round threshold, the user is connected to a manual processing system.
[0040] According to an achievable manner in the embodiment of the present application, the method further includes:
[0041] Each dialogue engine uses the sentence vector of the content input by the user in this round of dialogue to match in the intervention strategy set, where the intervention strategy set includes a pre-configured input sentence and its corresponding reply text or intent type;
[0042] If the sentence vector hits an input sentence in the intervention strategy set, the reply text corresponding to the hit input sentence is used as the reply text obtained by the dialogue engine, or the hit intent type is provided to the model in the dialogue engine to obtain the reply text; otherwise, the content input by the user in this round of dialogue and the dialogue context are input into the model in the dialogue engine to obtain the reply text.
[0043] According to an achievable manner in the embodiment of the present application, the input sentences in the intervention strategy set include high-frequency questions or bad example sentences;
[0044] If the first reply text obtained for the content input by the user cannot meet the expectation, the content input by the user is determined to be a bad example sentence.
[0045] According to an achievable manner in the embodiment of the present application, the method further includes:
[0046] The content input by the user in this round of dialogue and the content returned to the user in this round are recorded, and the recorded dialogue context is updated.
[0047] According to a second aspect, an intelligent dialogue system is provided, comprising:
[0048] A content acquisition unit, configured to acquire the content and conversation context input by the user in this round of conversation;
[0049] The engine scheduling unit is configured to provide the content and conversation context input by the user in this round of conversation to N different conversation engines, where N is a positive integer greater than 1; sort the reply texts obtained by the N different conversation engines based on confidence, and select the reply text whose sorting meets the preset requirements as the first reply text;
[0050] The content generation unit is configured to use the first reply text to generate content returned to the user in this round.
[0051] According to a third aspect, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods described in the first aspect are implemented.
[0052] According to the specific embodiments provided in this application, this application discloses the following technical effects:
[0053] 1) This application performs unified scheduling of multiple different dialogue engines, provides the content and dialogue context of the user's current dialogue input to each dialogue engine, and after obtaining the reply texts of multiple different dialogue engines, performs unified sorting to obtain the content returned to the user in this round, thereby realizing context reuse and collaborative work of multiple dialogue engines, making the generated dialogue more reasonable.
[0054] 2) Execute the pre-dialogue strategy before scheduling the dialogue engine, and execute the post-processing after the scheduled dialogue engine obtains the first reply text, which can uniformly implement some common dialogue capabilities and greatly improve the dialogue effect. The pre-dialogue strategy and post-dialogue strategy also facilitate the flexible expansion of the dialogue strategy of the intelligent dialogue system, and will not affect each dialogue engine. Each dialogue engine only needs to focus on the dialogue capability of its own engine.
[0055] 3) The sorting method provided by this application based on the confidence of the reply text combined with the priority and / or weight of the dialogue engine ensures that the management personnel only need to maintain the knowledge of each dialogue engine, and there is no need to perform additional data annotation for unified sorting, which greatly reduces the system operation and management costs. In addition, it is also convenient to upgrade and expand each dialogue engine.
[0056] 4) In this application, a two-stage scheduling method is adopted for the dialogue engine. First, the NLU interface of each dialogue engine is called to trigger each dialogue engine to obtain a reply text based on the content input by the user in this round of dialogue and the dialogue context. After the reply texts are uniformly sorted to obtain the first reply text, the execution interface of the dialogue engine from which the first reply text comes is scheduled, so that the dialogue engine initiates interaction with the service layer and outputs the first reply text. This method can effectively avoid the occurrence of erroneous execution of other dialogue engines.
[0057] 5) In this application, by setting the intervention strategy in the dialogue engine, each dialogue engine can use the sentence vector of the content input by the user in this round of dialogue to quickly match the intervention strategy to obtain the reply text, thereby effectively improving the processing efficiency of the dialogue engine and facilitating the management of users to effectively intervene in bad cases.
[0058] Of course, the implementation of any technical solution or product of the present application does not necessarily require achieving all of the advantages described above at the same time. BRIEF DESCRIPTION OF THE DRAWINGS
[0059] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0060] Figure 1 An exemplary system architecture diagram to which the embodiments of the present application can be applied is shown;
[0061] Figure 2 A main method flow chart provided for an embodiment of the present application;
[0062] Figure 3 A flow chart of a preferred method provided in an embodiment of the present application;
[0063] Figure 4 A schematic block diagram of the intelligent dialogue system according to an embodiment is shown;
[0064] Figure 5 A schematic diagram of the architecture of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0065] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments in the present application belong to the scope of protection of this application.
[0066] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments, and are not intended to limit the present invention. The singular forms "a", "said" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates other meanings.
[0067] It should be understood that the term "and / or" used in this article is only a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist at the same time, and B exists alone. In addition, the character " / " in this article generally indicates that the associated objects before and after are in an "or" relationship.
[0068] The word "if" as used herein may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to determining" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)", depending on the context.
[0069] Figure 1 An exemplary system architecture to which the embodiments of the present application can be applied is shown. Figure 1 As shown in , the system architecture mainly includes: a user terminal device, a central control unit and multiple dialogue engines, where multiple refers to more than one.
[0070] The user terminal device can be a variety of terminal devices that can obtain user input, where the user input can be in the form of text or voice. The user terminal device can be a screen device or a device with voice interaction function. Including but not limited to: smart mobile terminals, smart home devices, wearable devices, PCs (personal computers), etc. Among them, smart mobile devices can include such as mobile phones, tablets, laptops, PDAs (personal digital assistants), Internet cars, etc. Smart home devices can include smart home appliances, such as smart TVs, smart air conditioners, smart refrigerators, etc. Wearable devices can include such as smart watches, smart glasses, smart bracelets, virtual reality devices, augmented reality devices, mixed reality devices (i.e. devices that can support virtual reality and augmented reality), etc. The user can obtain the content input by the user through the user terminal device and send the content input by the user to the central control unit.
[0071] The central control unit is the abbreviation of the intermediate control device. The intelligent dialogue method provided in the embodiment of the present application is mainly executed by the central control unit, which is responsible for obtaining the content input by the user, and then scheduling multiple dialogue engines using the method provided in the embodiment of the present application to obtain appropriate reply text for generating content returned to the user.
[0072] Multiple dialogue engines are developed for different service scenarios and different types of knowledge. They can all use pre-established templates or models to output corresponding reply texts for input texts when inputting texts. Among them, multiple dialogue engines may include but are not limited to: task-based dialogue engines, FAQ dialogue engines, graph question-and-answer engines, table question-and-answer engines, document question-and-answer engines, and chat dialogue engines. Among them, the task-based dialogue engine supports task-based question-and-answer scenarios such as checking the weather, booking hotels, and recharging traffic; the FAQ dialogue engine supports dialogue scenarios based on question-and-answer pairs. The graph question-and-answer engine supports dialogue scenarios based on knowledge graph data; the table question-and-answer engine supports dialogue scenarios based on table data; the document question-and-answer engine supports dialogue scenarios based on document reading comprehension; and the chat dialogue engine supports open domain dialogue scenarios based on greetings, small talk, and other services that are not strongly related.
[0073] The central control unit, the dialogue engine, etc. can be set up in a single server, or in a server group consisting of multiple servers, or in a cloud server. A cloud server, also known as a cloud computing server or cloud host, is a host product in the cloud computing service system to solve the defects of difficult management and weak service scalability in traditional physical hosts and virtual private servers (VPS) services.
[0074] It should be understood that Figure 1The number of user terminal devices, central control units and dialogue engines in the embodiment is only for illustration. Any number of user terminal devices, central control units and dialogue engines may be provided according to implementation requirements.
[0075] Figure 2 This is a main flow chart of the intelligent dialogue method provided in the embodiment of the present application. The method can be Figure 1 The central control unit in the system architecture shown in the figure is executed. Figure 2 As shown in , the method may include the following steps:
[0076] Step 202: Obtain the content and conversation context of the user's current conversation input.
[0077] Step 204: Provide the content input by the user in this round of dialogue and the dialogue context to N different dialogue engines, where N is a positive integer greater than 1.
[0078] Step 206: sort the reply texts obtained by N different dialogue engines based on confidence, and select the reply text whose sorting meets the preset requirements as the first reply text.
[0079] Step 208: Use the first reply text to generate content returned to the user in this round.
[0080] It can be seen that in the above method provided in the embodiment of the present application, multiple different dialogue engines are uniformly scheduled, the content input by the user in this round of dialogue and the dialogue context are provided to each dialogue engine, and after obtaining the reply texts of multiple different dialogue engines, they are uniformly sorted to obtain the content returned to the user in this round, thereby realizing context reuse and collaborative work of multiple dialogue engines, making the generated dialogue more reasonable.
[0081] The above steps are described in detail below in conjunction with the embodiments. First, the above step 202, namely "obtaining the content and conversation context input by the user in this round of conversation", is described in detail.
[0082] In this step, the content input by the user in this round of dialogue can be input in the form of text or voice. The voice input can also obtain the corresponding text after voice recognition processing for subsequent processing by the central control unit and the dialogue engine. Since voice recognition processing is a relatively mature technology in this field, it will not be described in detail here. Figure 1 The system shown does not specifically show a voice recognition device.
[0083] The conversation context refers to the historical conversation content with the user under the same connection before the current conversation. The historical conversation content includes the content input by the user in the previous round and the content returned to the user. The central control unit will record the content input by the user and the content returned to the user in each round, and provide it to each conversation engine together with the user input content as the conversation context in the next round of user input. This cross-engine reuse of the conversation context enables each conversation engine to refer to the conversation context to give a more accurate and reasonable reply text. If the user's current conversation is the first round, the conversation context can be empty.
[0084] For example, suppose the user has the following conversation turns:
[0085] Round 1: The user inputs text A, and the content returned to the user is the reply text a generated by the FAQ dialogue engine;
[0086] Second round: The user inputs text B, and the content returned to the user is the reply text b generated by the task-based dialogue engine;
[0087] Round 3: User enters text C.
[0088] In this case, the content input by the user in this round is text C, so the conversation context consists of text A, reply text a, text B and reply text b.
[0089] As a more preferred embodiment, Figure 3 As shown in , before step 204, step 302, namely "pre-processing the content input by the user in this round of dialogue", can be further included.
[0090] The above-mentioned pretreatment may include but is not limited to at least one of the following treatments:
[0091] The first preprocessing: de-colloquialization processing.
[0092] Often, people tend to use some colloquial expressions during input, due to their expression habits and ways of expression, especially during voice input. These colloquial expressions may have an adverse effect on the understanding of the dialogue engine, thereby affecting the accuracy of the reply text generated by the dialogue engine. Therefore, the user input content can be de-colloquialized first.
[0093] For example, if a user inputs "Can you tell me what the weather is like today?", after de-colloquialization, the result is "How is the weather today?". Another example is "Uh-huh, do you know how high the temperature is today?", after de-colloquialization, the result is "What is the temperature today?", and so on.
[0094] When de-colloquializing, it can be implemented based on dictionaries, templates, or pre-trained models. For example, a dictionary including colloquial words is pre-set, the content input by the user is matched with the dictionary, and the matched colloquial words are deleted. For another example, an end-to-end model is pre-trained based on annotated text, the content input by the user is input into the pre-trained end-to-end model, and the end-to-end model performs de-colloquialization prediction on the content input by the user to obtain the de-colloquialized content. This part of the content can be implemented using existing technologies and will not be described in detail here.
[0095] The second type of pretreatment refers to digestion treatment.
[0096] Co-reference resolution is an important part of natural language processing. It is to determine which noun or phrase a pronoun in a text refers to. When users express themselves, they often use pronouns to refer to a noun or phrase due to their habit of expression or to express themselves more concisely. This often causes the dialogue engine to be unable to accurately understand the user's reference meaning, affecting the accuracy of the dialogue engine's generated reply text. Therefore, the content input by the user can be processed for co-reference resolution during the preprocessing process.
[0097] For example, if a user inputs "I have been to the business hall, but it is closed. Do you know what time it opens?", "there" refers to the "business hall". Another example is that a user inputs "I contacted customer service, but he did not respond to me", "he" refers to the "customer service", and so on.
[0098] When processing reference resolution, a method based on syntactic analysis or a method based on a pre-trained machine learning model can be used. Since this part can be implemented using existing technologies, it will not be described in detail here.
[0099] The third preprocessing: sentence vector calculation.
[0100] For the content input by the user in this round of dialogue, we can segment it into words to get the word vector of each word, and then perform bag of words-based processing, TF-IDF (term frequency-inverse document ratio)-based processing on the word vector of each word to get the sentence vector corresponding to the content input by the user in this round of dialogue.
[0101] The sentence vector corresponding to the content input by the user in this round of dialogue can also be obtained based on a pre-trained sentence vector calculation model. The sentence vector calculation model can be implemented based on an encoder, BERT (Bidirectional Encoder Representation from Transformers), etc.
[0102] The obtained sentence vector is mainly used by each dialogue engine before using the model to generate a reply text. It can use the sentence vector to directly match the pre-set intervention strategy set, quickly hit the intervention strategy to obtain the reply text, effectively improve the processing efficiency of the dialogue engine, and also facilitate the management user to effectively intervene in bad cases. This part will be described in detail in the subsequent embodiments.
[0103] After the above preprocessing, the preprocessing result can be provided to each dialogue engine in the subsequent step 204. For example, the content after de-colloquialization and / or reference resolution can be provided to each dialogue engine, the obtained sentence vector can be provided to each dialogue engine together with the content input by the user in this round of dialogue, the obtained sentence vector can be provided to each dialogue engine together with the content after de-colloquialization and / or reference resolution, and so on.
[0104] As a more preferred embodiment, Figure 3 As shown in, before step 204, step 304 may be further included, namely "determining whether the content input by the user in this round of dialogue hits the preset pre-dialogue strategy". If so, the hit pre-dialogue strategy is executed, otherwise, step 204 is continued.
[0105] It should also be noted that the above steps 302 and 304 can be performed either one or in sequence. Figure 3 In the illustrated embodiment, step 302 and step 304 are performed sequentially.
[0106] In a multi-dialogue engine system, different dialogue engines are developed to suit different service scenarios and different knowledge types. However, some general dialogue strategies are independent of the knowledge type. General system-level dialogue strategies can be abstracted and provided uniformly by the central control unit, so that each dialogue engine only needs to focus on solving the dialogue effect in its own field. The abstracted general system-level dialogue strategy is used as the pre-dialogue strategy before scheduling the dialogue engine.
[0107] The pre-dialogue strategy may include at least one of the following strategies:
[0108] The first strategy: re-listening strategy.
[0109] In the human-machine dialogue scenario, the user often fails to hear clearly what the robot is saying due to factors such as noise and external environment. At this time, the user may enter "Sorry, I didn't hear clearly" and other similar content. These contents represent that the user wants to listen again, that is, has the intention to re-listen. In view of this, in the embodiment of the present application, the intention to re-listen can be identified for the content input by the user in this round of dialogue. If the intention to re-listen is identified, the content returned to the user in the previous round of dialogue can be returned to the user again, without the need to call the dialogue engine, that is, direct pre-processing is performed.
[0110] Among them, the recognition of hard of hearing intention can be achieved based on a preset dictionary, speech template or a pre-trained machine learning model, etc., which will not be elaborated here.
[0111] The second strategy: emotional strategy.
[0112] During the conversation, the user's emotions often play a crucial role in the effectiveness of the intelligent conversation service. The ability to process different user emotions is an important dimension to measure the effectiveness of the intelligent conversation system. In the embodiment of the present application, emotion recognition can be performed on the content input by the user in this round of conversation. If a preset emotion type is recognized, the pre-processing corresponding to the recognized emotion type can be performed.
[0113] For example, if negative emotions are identified from the content input by the user in this round of dialogue, the user can be given soothing words in advance, such as "I understand your situation very much", "Please don't get excited", etc. The user can also be connected to the manual processing system for manual processing.
[0114] For another example, if the user's positive emotions are identified from the content input by the user in this round of conversation, preset marketing content can be returned to the user, and so on.
[0115] The addition of the above-mentioned emotion recognition and corresponding pre-processing greatly improves the humanization and harmony of the intelligent dialogue service.
[0116] The above-mentioned emotion recognition can also be implemented based on a preset dictionary, speech template or a pre-trained machine learning model, which will not be elaborated here.
[0117] The third strategy: sensitive word strategy.
[0118] In order to ensure compliance during the service process in human-computer dialogue, sensitive words can be identified for the content input by the user in this round of dialogue. If sensitive words are identified, the preset first speech script can be used to generate a second reply text and return the second reply text to the user. In this case, there is no need to schedule multiple dialogue engines later, and it is directly processed in the front-end link.
[0119] For example, if the content input by the user in this round of dialogue contains uncivilized words, a second reply text "Please use civilized language" may be returned to the user.
[0120] For another example, if the content input by the user in this round of conversation contains privacy, illegal content, etc., a second reply text such as "This question is not convenient to answer" can be returned to the user.
[0121] The above-mentioned sensitive word recognition is mainly implemented based on a preset dictionary, and can also be implemented by combining a dictionary with a speech template, etc., which will not be elaborated here.
[0122] It should be noted that the expressions "first", "second" and the like in the embodiments of the present application do not have restrictions on order, size and quantity, and are only used to distinguish in name. For example, "first reply text", "second reply text", "third reply text" and the like are used to distinguish reply texts in name.
[0123] The fourth strategy: switch to manual strategy.
[0124] If the user's input in this round of dialogue is recognized as an intention to be transferred to a human, the user will be connected to the human processing system. For example, if the user's input in this round is "I really don't understand, let me transfer it to a human" or "This is too outrageous, I want to complain", it can be transferred to a human.
[0125] The above-mentioned conversion to manual intent can be implemented based on a preset dictionary, speech template or a pre-trained machine learning model, which will not be described in detail here.
[0126] The above-mentioned pre-dialogue strategy can adopt the pipeline mode, which can be easily expanded to meet the customized needs of different services and has good scalability.
[0127] The above-mentioned step 204, i.e., "providing the content and conversation context input by the user in this round of conversation to N different conversation engines" and step 206, i.e., "ranking the reply texts obtained by the N different conversation engines based on confidence, and selecting the reply text whose ranking meets the preset requirements as the first reply text" are described in detail below in conjunction with the embodiments.
[0128] This step is the core capability of the central control unit, that is, how to reasonably combine multiple dialogue engines so that they work together to produce optimal results. In the embodiment of the present application, the central control unit uses a parallel scheduling method for multiple dialogue engines, parallel scheduling N different dialogue engines, and inputs the content and dialogue context input by the user in this round of dialogue into each dialogue engine, so that each dialogue engine can obtain a reply text based on a unified dialogue context.
[0129] As a preferred implementation, the central control unit may schedule the dialogue engine in two stages:
[0130] In the first stage, the NLU (Natural Language Understanding) interfaces of N different dialogue engines are scheduled in parallel, and the content of the user's current dialogue input and the dialogue context are passed as parameters to each dialogue engine. The reply texts returned by each dialogue engine are uniformly sorted, and the first reply text is selected based on the sorting results.
[0131] After receiving the content and conversation context of the current round of conversation input, each conversation engine can generate a reply text using the model adopted by the engine. The model adopted by each conversation engine is described in detail and limited in the embodiments of this application. However, in the conversation engine, the management user can configure an intervention strategy set, which includes a pre-configured input statement and its corresponding reply text or intent type.
[0132] Before using the model to obtain the reply text, the sentence vector obtained by preprocessing in step 302 can first be used to match it in the intervention strategy set, that is, the sentence vector corresponding to the content input by the user in this round of dialogue and the sentence vector corresponding to the input sentence in the intervention strategy set are calculated based on distance similarity. If an input sentence with a similarity greater than or equal to a preset similarity threshold is determined, the input sentence is used as the matched input sentence. The reply text corresponding to the input sentence can be directly used as the reply text obtained by this dialogue engine without further obtaining the reply text based on the model. Alternatively, the intent type corresponding to the input sentence is provided to the model in the dialogue engine to obtain the reply text. If no matching input sentence is obtained, the content input by the user in this round of dialogue and the dialogue context are input into the model in the dialogue engine to obtain the reply text.
[0133] As one of the feasible ways, frequently appearing questions can be used as input sentences to configure their corresponding reply texts, so that the corresponding reply texts can be quickly obtained through sentence vectors.
[0134] As another feasible way, the bad case that appears can be used as an input sentence to configure the corresponding reply text. For example, suppose that for the question "How much money is left in the card?", the task-based dialogue engine should be the most suitable one to check the balance. However, the task-based dialogue engine did not recognize the intention of checking the balance. Instead, the FAQ engine gave a reply text with high confidence. In other words, the first reply text obtained for the user's input content cannot meet expectations, which results in a bad case, that is, the user's input content is judged to be a bad case. In this case, the management user can configure an intervention strategy set in the task-based dialogue engine, taking "How much money is left in the card" as the input text and checking the balance as the corresponding intent type. When the user enters a similar statement as "How much money is left in the card" in this round of dialogue, the input text is obtained based on the sentence vector matching, thereby determining the intention of checking the balance and triggering the model to check the balance.
[0135] When each dialogue engine obtains a reply text, it will also obtain the confidence level corresponding to the reply text. As one of the feasible methods, the central control unit can sort the confidence levels of each reply text in combination with the weight or priority level of each dialogue engine, and select the M reply texts with the highest ranking as the first reply text, where M is a positive integer.
[0136] For example, if the maximum confidence of the reply text obtained by N different dialogue engines is greater than or equal to the reply text of the preset first confidence threshold, then the reply text with the highest priority ranking of the corresponding dialogue engine is selected from the reply texts with confidence greater than or equal to the preset first confidence threshold as the first reply text. In other words, if there is a reply text with high confidence, the reply text generated by the dialogue engine with high priority is preferentially selected from the reply texts with high confidence as the first reply text. Among them, the priority of each dialogue engine can be pre-set according to the type of dialogue engine. For example, the priority of the task-based dialogue engine can be set to the highest, the priority of the knowledge question and answer engine can be set to the second highest, the priority of the table question and answer engine can be set to the third highest, and so on. In addition, the setting of the priority of each dialogue engine can also be set according to the actual scenario, the service field to which the intelligent dialogue system belongs, the user's pre-configuration of the priority of each dialogue engine, the user's usage preference for each dialogue engine, etc.
[0137] If the maximum confidence of the reply texts obtained by N different dialogue engines is less than the preset first confidence threshold and greater than the preset second confidence threshold, the reply texts with confidence less than the preset first confidence threshold and greater than the preset second confidence threshold are scored based on the confidence and the dialogue engine weight, and the reply text with the highest score is selected as the first reply text. Among them, the first confidence threshold is greater than the second confidence threshold. In other words, if there is no high-confidence reply text, but there is a higher-confidence reply text, then these higher-confidence reply texts can be scored in combination with the confidence of the reply text and the weight of the corresponding dialogue engine, for example, the value obtained by multiplying the confidence of the reply text by the weight of the dialogue engine is used as the score, and then the reply text with the highest score is selected as the first reply text.
[0138] If the reply texts obtained by N different dialogue engines are clarification rhetoric, then M reply texts are selected as the first reply text according to the priority of each dialogue engine, where the clarification rhetoric is generated when the dialogue engine cannot generate a reply text with a confidence greater than or equal to the preset third confidence threshold, and the second confidence threshold is greater than or equal to the third confidence threshold. In other words, if there is no reply text with high confidence or relatively high confidence, for the dialogue engine, if the reply texts are all of medium confidence, the clarification rhetoric will be generated as the reply text.
[0139] For example, when the user inputs "financial management?" in this round of dialogue, each dialogue engine cannot accurately understand the intention of the content. The task-based dialogue engine may give a clarifying question "Do you want to apply for financial management services?", and the graph-based dialogue engine may give a clarifying question "Do you want to learn about financial management products?" In this case, based on the priority of the dialogue engine, the clarification question of the high-priority task-based dialogue engine can be used as the first reply text.
[0140] This sorting mechanism ensures that managers only need to maintain the knowledge of each conversation engine, without the need to perform additional data annotation for unified sorting, which greatly reduces system operation and management costs. It also makes it easier to upgrade and expand each conversation engine.
[0141] In addition to the above-mentioned sorting methods, other sorting methods may be used based on the confidence of the reply text, the priority and weight of the dialogue engine. For example, the confidence of the reply text, the priority and weight of the dialogue engine are used to calculate the score of each reply text, and the reply text is sorted according to the score.
[0142] In the second stage, the dialogue engine corresponding to the first reply text is used as the target dialogue engine for this round of dialogue; the execution interface of the target dialogue engine is called to trigger the target dialogue engine to execute the first reply text.
[0143] In the first stage, only the NLU interface of each dialogue engine is called, and the interaction between each dialogue engine and the service layer is not triggered. After the reply texts of each dialogue engine are uniformly sorted and the first reply text is selected, the execution interface of the dialogue engine corresponding to the first reply text is called in the second stage, so that the dialogue engine initiates the interaction with the service layer and outputs the first reply text.
[0144] In order to reflect the advantages of the two-stage processing in the above preferred embodiment, an example is given below for comparison of one-stage processing and two-stage processing:
[0145] One-stage processing method: After the content of the user's current dialogue input and the dialogue context are sent to each dialogue engine, the task-based dialogue engine recognizes the user's intention to apply for a certain service, and sends a text message related to the service to the user through interaction with the service layer. However, after sorting the reply texts of each dialogue engine, the reply text with the highest confidence is generated by the FAQ dialogue engine, and the reply text of the FAQ dialogue engine is finally determined and returned to the user. In this way, the user may receive the answer returned by the FAQ after receiving the text message, which will cause confusion and poor user experience.
[0146] Two-stage processing method: Separate the NLU interface and the execution interface, and the execution interface is called by the central control unit alone. The content input by the user in this round of dialogue and the dialogue context are sent to each dialogue engine by calling the NLU interface of each dialogue engine. Each dialogue engine returns a reply text but does not interact with the service layer. After sorting the reply texts of each dialogue engine, the reply text with the highest confidence is generated by the FAQ dialogue engine. The central control unit calls the execution interface of the FAQ dialogue engine to trigger the FAQ dialogue engine to output the reply text as the first reply text. In this case, the user will not receive other redundant messages.
[0147] The above step 208, ie, "using the first reply text to generate the content returned to the user in this round" is described in detail below in conjunction with an embodiment.
[0148] As one of the feasible ways, the first reply text may be directly used to generate readable content and return it to the user. The readable content may be generated based on a template or a pre-trained generative model, etc., which will not be described in detail here.
[0149] As another possible way to achieve this, Figure 3 As shown in , at least one of the following post-processing can be performed in step 308.
[0150] The first post-processing: cross-engine dialogue guidance.
[0151] If the first reply text and the content returned to the user in the previous round come from different conversation engines and the content returned to the user in the previous round is a clarifying question, then the content of the previous round of user conversation is used to generate the third reply text, and the first reply text and the third reply text are used to generate the content returned to the user; otherwise, the first reply text is used to generate the content returned to the user in this round.
[0152] That is to say, two consecutive rounds of dialogues are realized across engines, and the content returned to the user in the previous round of dialogue is a clarification question, which means that the previous round of dialogue has not been processed by the corresponding dialogue engine and is still in the clarification stage. Therefore, as a preferred implementation, while returning the first reply text to the user in the current round of dialogue, the user can be guided to continue to return to the previous round of dialogue. This method will shorten the dialogue cost of meeting user needs and give users a better experience.
[0153] For example:
[0154] In the first round of dialogue, the user inputs "Help me book a hotel in Beijing", and the task-based dialogue engine returns the reply text "What hotel do you want to book?".
[0155] In the second round of dialogue, the user inputs "How is the weather in Beijing?". After sorting, it is determined that the first reply text is the reply text of the FAQ dialogue engine "The weather in Beijing is sunny." However, because it is cross-engine and the content returned to the user in the first round of dialogue is a clarification question, the content of the previous round of dialogue can be used to generate the third reply text "Shall I continue to help you book a hotel?", and then the first reply text and the third reply text are used to generate the content returned to the user "The weather in Beijing is sunny, shall I continue to help you book a hotel?", thereby guiding the user to return to the previous round of dialogue.
[0156] The second type of post-processing: personalized rejection processing.
[0157] In some cases, due to problems expressed by the user, or some environmental factors, etc., it may happen that N different dialogue engines do not receive a reply text. This situation is called "rejection of recognition" for short. In the embodiment of the present application, the central control unit can perform unified post-processing for this situation, which may include but is not limited to the following methods:
[0158] Method 1: Generate a fourth reply text based on the preset second speech and historical conversation content to guide the user into the historical conversation scene, and use the fourth reply text to generate the content returned to the user in this round.
[0159] For example, the user enters "I want to register" in the first round of dialogue, and the task-based dialogue engine generates a reply text "What do you want to register?". If the user's input in the second round of dialogue is unclear and no dialogue engine gets a reply text, then a fourth reply text "I didn't hear what you said clearly, what do you want to register?" can be generated based on the preset second dialogue script and historical dialogue content.
[0160] Method 2: Based on the preset third speech and hot issue information, a fifth reply text is generated to prompt the user to consult the hot issue, and the fifth reply text is used to generate the content returned to the user in this round.
[0161] For example, if the user's input in the first round of dialogue is unclear and all dialogue engines do not receive a reply text, a fifth reply text "I didn't hear what you said clearly. Are you asking how to check the balance? " can be generated based on the preset third speech and hot issue information.
[0162] Method 3: Generate a sixth reply text for informing the user of the hotspot service based on the preset fourth speech and hotspot service information, and use the sixth reply text to generate the content returned to the user in this round.
[0163] For example, if the user inputs unclear content in the first round of dialogue and all dialogue engines do not get a reply text, a sixth reply text "I didn't hear what you said clearly, do you want to check your balance?
[0164] Method 4: If the number of rounds in which N different dialogue engines have not received a reply text reaches a preset round threshold, the user is connected to the manual processing system.
[0165] For example, if the dialogue engine does not receive a reply text for three consecutive rounds, the user can be connected to the manual processing system and the case can be transferred to manual processing to avoid the user's bad experience caused by not understanding their needs for a long time.
[0166] For the reply text obtained by the above post-processing, readable content can also be generated and returned to the user.
[0167] It should also be noted that when returning content to the user, it can be returned in the form of text, or the text can be synthesized into speech and the corresponding speech can be returned to the user, and so on.
[0168] Furthermore, if Figure 3 As shown in , step 310 may also be included to update the dialogue state, including recording the content input by the user in this round of dialogue and the content returned to the user in this round, and updating the recorded dialogue context.
[0169] This step can record the content of each round of human-computer dialogue and update the dialogue context, so that the dialogue context can be sent to each dialogue engine in a unified manner in subsequent rounds of dialogue, thereby realizing cross-engine context reuse, realizing complex dialogue scenarios, and improving dialogue effects. In addition, the recording of dialogue status can also serve as a reference for subsequent rounds of reference resolution, dialogue guidance, clarification and rhetorical questions, etc.
[0170] In addition, the context of this round of dialogue, the reply text of each dialogue engine, the content of the user's input in this round of dialogue, the content returned to the user in this round, etc. can also be sent to the dialogue analysis report platform and annotation platform through the asynchronous message queue. The dialogue analysis report platform can be used to analyze the effect of the intelligent dialogue system, and the operator can also continuously optimize the dialogue effect of the intelligent dialogue system through the annotation platform.
[0171] The above is a description of a specific embodiment of the specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in an order different from that in the embodiments and still achieve the desired results. In addition, the processes depicted in the drawings do not necessarily require the specific order or continuous order shown to achieve the desired results. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0172] According to an embodiment of another aspect, an intelligent dialogue is provided. Figure 4 A schematic block diagram of the intelligent dialogue system according to an embodiment is shown. The system is arranged at Figure 1 The central control unit in the system architecture shown. The system can be an application located on the server side, or it can also be a plug-in or software development kit (SDK) and other functional units in the application located on the server side, or it can be deployed on a cloud server or terminal device, and this application does not limit it. Figure 4 As shown, the system 400 may include: a content acquisition unit 402, an engine scheduling unit 404 and a content generation unit 406, and may further include: a pre-processing unit 408, a pre-strategy processing unit 410, a post-processing unit 412 and a conversation recording unit 414. The main functions of each component unit are as follows:
[0173] The content acquisition unit 402 is configured to acquire the content and conversation context input by the user in this round of conversation.
[0174] The engine scheduling unit 404 is configured to provide the content and conversation context input by the user in this round of conversation to N different conversation engines, where N is a positive integer greater than 1; sort the reply texts obtained by the N different conversation engines based on confidence, and select the reply text whose sorting meets the preset requirements as the first reply text.
[0175] The content generating unit 406 generates content returned to the user in this round by using the first reply text.
[0176] Furthermore, the preprocessing unit 408 is configured to preprocess the content input by the user in this round of dialogue, wherein the preprocessing includes at least one of de-colloquialization processing, reference resolution processing and sentence vector calculation.
[0177] Accordingly, the engine scheduling unit 404 may provide the result of preprocessing the content input by the user in this round of dialogue and the dialogue context to N different dialogue engines.
[0178] Furthermore, the pre-dialogue strategy processing unit 410 is configured to determine whether the content input by the user in this round of dialogue hits the preset pre-dialogue strategy. If so, the hit pre-dialogue strategy is executed; otherwise, the engine scheduling unit 404 is triggered to continue to execute the processing of providing the content input by the user in this round of dialogue and the dialogue context to N different dialogue engines.
[0179] The above-mentioned pre-dialogue strategy may include at least one of the following:
[0180] Re-listening strategy: If the content input by the user in this round of conversation has the intention of re-listening, the content returned to the user in the previous round of conversation will be returned to the user again;
[0181] Emotion strategy: If the content input by the user in this round of dialogue is recognized as a preset emotion type, the pre-processing corresponding to the recognized emotion type is performed;
[0182] Sensitive word strategy: If the content input by the user in this round of dialogue is identified as a preset sensitive word, the preset first speech script is used to generate a second reply text, and the second reply text is returned to the user;
[0183] Manual transfer strategy: If the content input by the user in this round of conversation is recognized as being transferred to a manual process, the user will be connected to the manual processing system.
[0184] As one of the achievable ways, the engine scheduling unit 404 may be specifically configured to: schedule the natural language understanding NLU interfaces of N different dialogue engines in parallel, and pass the content input by the user in this round of dialogue and the dialogue context as parameters to each dialogue engine;
[0185] Correspondingly, the content generation unit 406 uses the dialogue engine corresponding to the first reply text as the target dialogue engine of this round of dialogue; and calls the execution interface of the target dialogue engine to trigger the target dialogue engine to execute the first reply text.
[0186] As one of the achievable methods, the engine scheduling unit 404 may be specifically configured to sort the confidences of the reply texts obtained by N different dialogue engines in combination with the weight or priority of each dialogue engine, and select the M reply texts with the highest ranking as the first reply text, where M is a positive integer.
[0187] Among them, the priority of each dialogue engine can be preset according to the type of dialogue engine.
[0188] For example: if the maximum confidence of the reply texts obtained by N different dialogue engines is greater than or equal to the reply text of the preset first confidence threshold, then the reply text with the highest priority ranking of the corresponding dialogue engine is selected as the first reply text from the reply texts with confidence greater than or equal to the preset first confidence threshold; or,
[0189] If the maximum confidence of the reply texts obtained by N different dialogue engines is less than a preset first confidence threshold and greater than a preset second confidence threshold, the reply texts with a confidence less than the preset first confidence threshold and greater than the preset second confidence threshold are scored based on the confidence and the dialogue engine weight, and the reply text with the highest score is selected as the first reply text; or,
[0190] If the reply texts obtained by N different dialogue engines are clarification rhetoric, then M reply texts are selected as the first reply texts according to the priority of each dialogue engine, wherein the clarification rhetoric is generated when the dialogue engine cannot generate a reply text with a confidence greater than or equal to a preset third confidence threshold;
[0191] The first confidence threshold is greater than the second confidence threshold, and the second confidence threshold is greater than or equal to the third confidence threshold.
[0192] As one achievable manner, the content generation unit 406 may be configured as follows: if the first reply text and the content returned to the user in the previous round come from different conversation engines and the content returned to the user in the previous round is a clarifying rhetorical question, then a third reply text is generated using the content of the previous round of user conversation, and the first reply text and the third reply text are used to generate the content returned to the user; otherwise, the first reply text is used to generate the content returned to the user in this round.
[0193] Furthermore, the post-processing unit 412 may be configured to: if the N different dialogue engines do not obtain a reply text, perform the following post-processing:
[0194] Generate a fourth reply text for guiding the user into the historical dialogue scene based on the preset second speech and the historical dialogue content, and use the fourth reply text to generate the content returned to the user in this round; or,
[0195] Based on the preset third speech and hot issue information, a fifth reply text is generated to prompt the user to consult the hot issue, and the fifth reply text is used to generate the content returned to the user in this round; or,
[0196] generating a sixth reply text for informing the user of the hotspot service based on the preset fourth speech and the hotspot service information, and using the sixth reply text to generate the content returned to the user in this round; or,
[0197] If the number of rounds in which N different dialogue engines have not received a reply text reaches a preset round threshold, the user is connected to the manual processing system.
[0198] As one of the achievable methods, each dialogue engine can use the sentence vector calculated by the sentence vector preprocessing unit 408 to match in the intervention strategy set, which includes a pre-configured input sentence and its corresponding reply text or intent type. If the sentence vector hits the input sentence in the intervention strategy set, the reply text corresponding to the hit input sentence is used as the reply text obtained by the dialogue engine, or the hit intent type is provided to the model in the dialogue engine to obtain the reply text; otherwise, the content input by the user in this round of dialogue and the dialogue context are input into the model in the dialogue engine to obtain the reply text.
[0199] As one possible implementation, the input sentences in the above intervention strategy set may include high-frequency questions or bad example sentences; if the first reply text obtained for the content input by the user fails to meet expectations, the content input by the user is a bad example sentence.
[0200] Furthermore, the conversation recording unit 414 may be configured to record the content input by the user in this round of conversation and the content returned to the user in this round, and update the recorded conversation context.
[0201] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the scheme described in this article within the scope permitted by applicable laws and regulations, provided that the applicable laws and regulations of the country are met (for example, with the user's explicit consent, effective notification to the user, etc.).
[0202] In addition, an embodiment of the present application further provides a computer-readable storage medium on which a computer program is stored, and when the program is executed by a processor, the steps of any one of the methods in the aforementioned method embodiments are implemented.
[0203] An embodiment of the present application also provides an electronic device, comprising: one or more processors; and a memory associated with the one or more processors, wherein the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the aforementioned method embodiments are executed.
[0204] in, Figure 5 The architecture of the electronic device is shown as an example, which may include a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, and a memory 520. The processor 510, the video display adapter 511, the disk drive 512, the input / output interface 513, the network interface 514, and the memory 520 may be communicatively connected via a communication bus 530.
[0205] The processor 510 may be implemented by a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solutions provided in this application.
[0206] The memory 520 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage device, dynamic storage device, etc. The memory 520 can store an operating system 521 for controlling the operation of the electronic device 500, and a basic input and output system (BIOS) 522 for controlling the low-level operation of the electronic device 500. In addition, a web browser 523, a data storage management system 524, and an intelligent dialogue system 525, etc. can also be stored. The above-mentioned intelligent dialogue system 525 can be an application program that specifically implements the operations of the aforementioned steps in the embodiment of the present application. In short, when the technical solution provided by the present application is implemented by software or firmware, the relevant program code is stored in the memory 520 and is called and executed by the processor 510.
[0207] The input / output interface 513 is used to connect the input / output module to realize information input and output. The input / output module can be configured in the device as a component (not shown in the figure), or it can be externally connected to the device to provide corresponding functions. The input device may include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output device may include a display, a speaker, a vibrator, an indicator light, etc.
[0208] The network interface 514 is used to connect to a communication module (not shown) to realize communication interaction between the device and other devices. The communication module can realize communication through a wired mode (such as USB, network cable, etc.) or a wireless mode (such as mobile network, WIFI, Bluetooth, etc.).
[0209] The bus 530 comprises a pathway for transmitting information between the various components of the device (eg, the processor 510, the video display adapter 511, the disk drive 512, the input / output interface 513, the network interface 514, and the memory 520).
[0210] It should be noted that, although the above device only shows a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, a memory 520, a bus 530, etc., in the specific implementation process, the device may also include other components necessary for normal operation. In addition, it can be understood by those skilled in the art that the above device may also only include components necessary for implementing the solution of the present application, and does not necessarily include all the components shown in the figure.
[0211] It can be known from the description of the above implementation methods that those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application can be essentially or partly contributed to the prior art in the form of a software product, which can be stored in a storage medium such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present application or certain parts of the embodiments.
[0212] Each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can refer to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can refer to the partial description of the method embodiment. The system and system embodiments described above are merely schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without creative work.
[0213] The technical solution provided by the present application is described in detail above. The principle and implementation method of the present application are described in detail using specific examples. The description of the above embodiments is only used to help understand the method and core idea of the present application. At the same time, for those skilled in the art, according to the idea of the present application, there will be changes in the specific implementation method and application scope. In summary, the content of this specification should not be understood as limiting the present application.
Claims
1. An intelligent dialogue method, comprising: Get the content and conversation context of the user's current conversation; Providing the content and conversation context of the user's current conversation input to N different conversation engines, where N is a positive integer greater than 1; If the maximum confidence of the reply texts obtained by the N different dialogue engines is greater than or equal to the reply text of the preset first confidence threshold, then the reply text with the highest priority ranking of the corresponding dialogue engine is selected as the first reply text from the reply texts with confidence greater than or equal to the preset first confidence threshold; or, If the maximum confidence of the reply texts obtained by the N different dialogue engines is less than a preset first confidence threshold and greater than a preset second confidence threshold, the reply texts whose confidence is less than the preset first confidence threshold and greater than the preset second confidence threshold are scored based on the confidence and the dialogue engine weight, and the reply text with the highest score is selected as the first reply text; or, If the reply texts obtained by the N different dialogue engines are clarification rhetoric, then M reply texts are selected as the first reply text according to the priority of each dialogue engine, wherein the clarification rhetoric is generated when the dialogue engine cannot generate a reply text with a confidence greater than or equal to a preset third confidence threshold; If the first reply text and the content returned to the user in the previous round come from different conversation engines and the content returned to the user in the previous round is a clarification rhetorical question, then the content of the previous round of user conversation is used to generate a third reply text, and the first reply text and the third reply text are used to generate the content returned to the user; Otherwise, using the first reply text to generate content returned to the user in this round; Providing the content and the conversation context input by the user in this round of conversation to N different conversation engines includes: scheduling the natural language understanding (NLU) interfaces of the N different conversation engines in parallel, and passing the content and the conversation context input by the user in this round of conversation as parameters to each conversation engine; Using the first reply text to generate content returned to the user in this round includes: using the dialogue engine corresponding to the first reply text among the N different dialogue engines as the target dialogue engine of this round of dialogue; calling the execution interface of the target dialogue engine to trigger the target dialogue engine to execute the first reply text; The first confidence threshold is greater than the second confidence threshold, the second confidence threshold is greater than or equal to the third confidence threshold, and the priority of each dialogue engine is pre-set according to the type of the dialogue engine.
2. The method according to claim 1, before providing the content and conversation context input by the user in the current conversation to N different conversation engines, further comprising: Determine whether the content input by the user in this round of dialogue matches the preset pre-dialogue strategy, and if so, execute the pre-dialogue strategy that matches; Otherwise, continue to execute the step of providing the content input by the user in this round of dialogue and the dialogue context to N different dialogue engines.
3. The method according to claim 2, wherein: The pre-dialogue strategy includes at least one of the following: Re-listening strategy: if the content input by the user in this round of dialogue has the intention of re-listening, the content returned to the user in the previous round of dialogue is returned to the user again; Emotion strategy: if the content input by the user in this round of dialogue is identified as a preset emotion type, then perform the pre-processing corresponding to the identified emotion type; Sensitive word strategy: if the content input by the user in this round of dialogue is identified as a preset sensitive word, a second reply text is generated using the preset first speech technique, and the second reply text is returned to the user; Manual transfer strategy: If the content input by the user in this round of dialogue is identified as being intended to be transferred to manual processing, the user will be connected to the manual processing system.
4. The method according to claim 1, further comprising: If the N different dialogue engines do not obtain a reply text, the following post-processing is performed: Generate a fourth reply text for guiding the user into the historical conversation scene based on the preset second speech and the historical conversation content, and use the fourth reply text to generate the content returned to the user in this round; or, Based on the preset third speech and hot issue information, a fifth reply text is generated to prompt the user to consult the hot issue, and the fifth reply text is used to generate the content returned to the user in this round; or, generating a sixth reply text for informing the user of the hotspot service based on the preset fourth speech and the hotspot service information, and using the sixth reply text to generate content returned to the user in this round; or, If the number of rounds in which the N different dialogue engines have not obtained a reply text reaches a preset round threshold, the user is connected to a manual processing system.
5. The method according to claim 1, further comprising: Each dialogue engine uses the sentence vector of the content input by the user in this round of dialogue to match in the intervention strategy set, where the intervention strategy set includes a pre-configured input sentence and its corresponding reply text or intent type; If the sentence vector hits an input sentence in the intervention strategy set, the reply text corresponding to the hit input sentence is used as the reply text obtained by the dialogue engine, or the hit intent type is provided to a model in the dialogue engine to obtain a reply text; Otherwise, the content input by the user in this round of dialogue and the dialogue context are input into the model in the dialogue engine to obtain a reply text.
6. The method according to claim 5, wherein: The input sentences in the intervention strategy set include high-frequency questions or bad example sentences; If the first reply text obtained for the content input by the user cannot meet the expectation, the content input by the user is determined to be a bad example sentence.
7. The method according to any one of claims 1 to 6, further comprising: The content input by the user in this round of dialogue and the content returned to the user in this round are recorded, and the recorded dialogue context is updated.
8. An intelligent dialogue system, comprising: A content acquisition unit, configured to acquire the content and conversation context input by the user in this round of conversation; The engine scheduling unit is configured to select, from the reply texts whose confidence is greater than or equal to the preset first confidence threshold, a reply text with the highest priority ranking of the corresponding dialogue engine as the first reply text if the maximum confidence of the reply texts obtained by the N different dialogue engines is greater than or equal to the reply text of the preset first confidence threshold; Alternatively, if the maximum confidence of the reply texts obtained by the N different dialogue engines is less than a preset first confidence threshold and greater than a preset second confidence threshold, the reply texts whose confidence is less than the preset first confidence threshold and greater than the preset second confidence threshold are scored based on the confidence and the dialogue engine weight, and the reply text with the highest score is selected as the first reply text; Alternatively, if the reply texts obtained by the N different dialogue engines are clarifying rhetorical questions, M reply texts are selected as the first reply texts according to the priority of each dialogue engine, wherein the clarifying rhetorical questions are generated when the dialogue engine cannot generate a reply text with a confidence greater than or equal to a preset third confidence threshold; wherein the first confidence threshold is greater than the second confidence threshold, the second confidence threshold is greater than or equal to the third confidence threshold, and the priority of each dialogue engine is preset according to the type of the dialogue engine; A content generating unit, configured to generate content returned to the user in this round by using the first reply text; When the engine scheduling unit is configured to provide the content and the conversation context input by the user in this round of conversation to N different conversation engines, it is specifically configured as follows: Parallel scheduling of the natural language understanding (NLU) interfaces of the N different dialogue engines, and passing the content of the user's current dialogue input and the dialogue context as parameters to each dialogue engine; The content generation unit is specifically configured as follows: If the first reply text and the content returned to the user in the previous round come from different dialogue engines and the content returned to the user in the previous round is a clarifying rhetorical question, then the third reply text is generated using the content of the previous round of user dialogue, and the first reply text and the third reply text are used to generate the content returned to the user; otherwise, the dialogue engine corresponding to the first reply text among the N different dialogue engines is used as the target dialogue engine for this round of dialogue; and the execution interface of the target dialogue engine is called to trigger the target dialogue engine to execute the first reply text.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.
Citation Information
Patent Citations
Man-machine interaction method and system based on artificial intelligence
CN105068661A
Intelligent robot integrating chat, knowledge and task questions and answers
CN113515613A