Dialogue method, dialogue system and electronic device

By extracting key information of user statements in the voice interaction system and relating them to historical dialogues, and dynamically updating the knowledge graph, the problem that existing systems cannot effectively understand user intentions is solved, and a more intelligent and appropriate dialogue reply is achieved.

CN120216648APending Publication Date: 2025-06-27VOYAH AUTOMOBILE TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510303266.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

The existing voice interaction system cannot effectively understand and associate the user's intentions, resulting in the content of the reply being unrelated to the previous conversation, resulting in the dialogue fragmentation and information omission.

Method used

By receiving the user's to be replied to, extracting key information, and determining whether there is a contextual relationship between it and the historical dialogue statement. If there is, update the key information and knowledge graph based on the historical dialogue content to generate relevant replies; if there is no, build a new knowledge graph based on the initial key information to answer.

Benefits of technology

It improves the context understanding ability, semantic analysis ability and dynamic knowledge expansion ability of the dialogue system, ensures that the content of the reply is consistent with the user's intention, and improves the intelligence level of user experience and voice interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120216648A_ABST
    Figure CN120216648A_ABST
Patent Text Reader

Abstract

The invention discloses a dialogue method, a dialogue system and an electronic device, first key information is extracted based on a to-be-replied statement currently input by a user, and when it is determined that context association exists between the to-be-replied statement and a historical dialogue statement, the first key information is received based on the historical dialogue statement. Updating a first keyword and a first user intention in the first key information to obtain second key information; and based on the second key information and a preset knowledge base, expanding and updating a preset first knowledge graph to obtain a second knowledge graph, and outputting the dialogue content according to the second key information and the second knowledge graph. The method can be applied to multiple rounds of dialogues, and the context understanding capability, the semantic analysis capability and the knowledge dynamic expansion capability of a dialogue system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to Intelligent Conversation the technical field, and in particular to a dialogue method, a dialogue system, and an electronic device. Background Art

[0002] With the development of artificial intelligence technology, voice interaction, as one of the important forms of human-computer interaction, has been widely used in fields such as intelligent vehicles, smart homes, and mobile devices.

[0003] In the dialogue scenarios of non-task-based voice interaction, such as planning a travel route, interest guidance, emotional communication, or multi-topic discussion, etc., the user's questions are often not limited to a single task, but a continuous and progressive process. The user's needs may be gradually expressed in multiple dialogue turns, while traditional voice interaction systems usually cannot effectively understand and associate the user's intentions, and the replied content may be irrelevant to the previous dialogue content, resulting in the fragmentation of the dialogue and information omission. Summary of the Invention

[0004] The present application provides a dialogue method, a dialogue system, and an electronic device to solve the technical problem that the existing voice interaction system cannot effectively understand and associate the user's intentions, resulting in the replied content being irrelevant to the previous dialogue content.

[0005] In view of the above problems, the present application is proposed to provide a dialogue method, a dialogue system, and an electronic device that overcome the above problems or at least partially solve the above problems.

[0006] In a first aspect, a dialogue method is provided, including:

[0007] Receiving a to-be-replied statement currently input by a user, and based on the to-be-replied statement, extracting first key information, where the first key information includes a first keyword and a first user intention;

[0008] Judging whether there is a context association between the to-be-replied statement and historical dialogue statements. If so, updating the first keyword and the first user intention based on the historical dialogue statements to obtain second key information;

[0009] Based on the second key information and a preset knowledge base, expanding and updating a preset first knowledge graph to obtain a second knowledge graph, where the first knowledge graph is a historical knowledge graph generated based on historical dialogue content;

[0010] Outputting dialogue content according to the second key information and the second knowledge graph.

[0011] Optionally, after judging whether there is a context association between the to-be-replied statement and historical dialogue statements, the method further includes:

[0012] If there is no context association between the statement to be replied and the historical conversation statements, a third knowledge graph is generated based on the first key information and a preset knowledge base;

[0013] Output the conversation content according to the first key information and the third knowledge graph.

[0014] Optionally, output the conversation content according to the second key information and the second knowledge graph, including:

[0015] Perform entity extraction and relationship construction on the knowledge content of the second knowledge graph according to the second key information to obtain target entity information;

[0016] Perform logical reasoning based on the target entity information and preset reasoning rules, generate an answer corresponding to the target entity information, and integrate and optimize the answer to obtain and output the conversation content.

[0017] Optionally, expand and update the preset first knowledge graph based on the second key information and the preset knowledge base to obtain the second knowledge graph, including:

[0018] Extract key entity information from the knowledge base based on the second key information; wherein, the second key information includes a second keyword and a second user intention, and the key entity information includes entities, relationships, and / or attributes related to the second keyword, and entities, relationships, and / or attributes related to the second user intention;

[0019] Add, modify, or delete entity information in the first knowledge graph based on the key entity information to obtain the second knowledge graph.

[0020] Optionally, determine whether there is context association between the statement to be replied and the historical conversation statements, including:

[0021] Compare the first key information with the historical key information extracted from the historical conversation content. If the relevance is greater than or equal to the preset threshold, it is determined that there is context association between the statement to be replied and the historical conversation statements; if the relevance is less than the preset threshold, it is determined that there is no context association between the statement to be replied and the historical conversation statements.

[0022] Optionally, the statement to be replied input by the user currently is a voice statement; extract the first key information based on the statement to be replied, including:

[0023] Convert the voice statement into a text statement;

[0024] Perform semantic analysis on the text statement to obtain the first key information.

[0025] Optionally, based on historical conversation statements, update the first keyword and the first user intention to obtain second key information, including:

[0026] Update the first keyword based on historical keywords extracted from historical conversation statements to obtain a second keyword;

[0027] Update the first user intention based on the association between historical keywords and / or historical conversation statements and the statement to be replied, to obtain a second user intention, and the second user intention and the second keyword constitute second key information.

[0028] In a second aspect, the present application also provides a dialogue system, including:

[0029] An extraction unit, configured to receive the statement to be replied currently input by a user, and based on the statement to be replied, extract first key information, where the first key information includes a first keyword and a first user intention;

[0030] A determination and update unit, configured to determine whether there is a context association between the statement to be replied and historical conversation statements, and if so, based on the historical conversation statements, update the first keyword and the first user intention to obtain second key information;

[0031] An update unit, configured to expand and update a preset first knowledge graph based on the second key information and a preset knowledge base to obtain a second knowledge graph, where the first knowledge graph is a historical knowledge graph generated based on historical conversation content;

[0032] A dialogue output unit, configured to output dialogue content according to the second key information and the second knowledge graph.

[0033] In a third aspect, the present application also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device is caused to execute the method provided in the first aspect.

[0034] In a fourth aspect, the present application also provides a computer-readable storage medium, where the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the computer is caused to execute the method provided in the first aspect.

[0035] In a fifth aspect, the present application also provides a computer program product, including a computer program, and when the computer program is run by a computer, the computer is caused to execute the method provided in the first aspect.

[0036] The technical solution provided by the present application has at least the following technical effects or advantages:

[0037] The dialogue method, dialogue system, and electronic device provided by this application determine whether there is a context association between the to-be-replied statement currently input by the user and the historical dialogue statements; if so, based on the to-be-replied statement and a preset knowledge base, the historical knowledge graph generated based on the historical dialogue content is expanded and updated, and the to-be-replied statement is answered according to the updated knowledge graph. This application can be applied to improve the context understanding ability, semantic parsing ability, and knowledge dynamic expansion ability of the dialogue system in multi-round conversations, solve the problem that the existing voice interaction system cannot effectively understand and associate the user's intention, resulting in the problem that the replied content is irrelevant to the previous dialogue content, and at the same time improve the user experience and the intelligent level of voice interaction.

[0038] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the following specifically illustrates the specific implementation manners of this application. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of this application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0040] Figure 1 It is a flowchart of a dialogue method in an example provided by an embodiment of this application;

[0041] Figure 2 It is a flowchart of a dialogue method in another example provided by an embodiment of this application;

[0042] Figure 3 It is a flowchart of a dialogue method in another example provided by an embodiment of this application;

[0043] Figure 4 It is a flowchart of a dialogue method in another example provided by an embodiment of this application;

[0044] Figure 5 It is a block diagram of a dialogue system provided by an embodiment of this application;

[0045] Figure 6 It is a flowchart of an example of voice multi-round interaction in an embodiment of this application;

[0046] Figure 7 It is a schematic diagram of an electronic device in an embodiment of this application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0047] Exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings.

[0048] Various structural schematic diagrams according to embodiments of the present application are shown in the accompanying drawings. These figures are not drawn to scale, where for the purpose of clear expression, some details are enlarged and some details may be omitted. The shapes of various regions and layers shown in the figures, as well as their relative sizes and positional relationships, are merely exemplary and may deviate in practice due to manufacturing tolerances or technical limitations. Those skilled in the art can design regions / layers with different shapes, sizes, and relative positions according to actual needs.

[0049] For the convenience of clearly describing the technical solutions of the embodiments of the present application, in the embodiments of the present application, terms such as "first" and "second" are used to distinguish identical or similar items with basically the same functions and roles. For example, the first value and the second value are merely used to distinguish different values and do not limit their order. Those skilled in the art can understand that terms such as "first" and "second" do not limit the quantity and execution order, and "first", "second", etc. do not necessarily mean different.

[0050] It should be noted that in the present application, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.

[0051] In the present application, "at least one" means one or more, and "a plurality" means two or more. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, or B exists alone, where A and B may be singular or plural. The character " / " generally represents an "or" relationship between the associated objects before and after. "At least one (item)" or its similar expression refers to any combination of these items, including any combination of single item (item) or plural items (items). For example, at least one (item) of a, b, or c may represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, c may be single or multiple.

[0052] To better understand the above technical solutions, the following will elaborate on the above technical solutions in combination with specific implementation manners. It should be understood that the embodiments of the present disclosure and the specific features in the embodiments are detailed descriptions of the technical solutions of the present application, rather than limitations on the technical solutions of the present application. Without conflict, the technical features in the embodiments of the present application and the embodiments can be combined with each other.

[0053] With the development of artificial intelligence technology, voice interaction, as one of the important forms of human-computer interaction, has been widely applied in fields such as intelligent vehicles, smart homes, and mobile devices. The voice interaction system provides a more convenient and natural interaction method for users through technologies such as automatic speech recognition (ASR), natural language processing (NLP), and text-to-speech (TTS). However, the current voice interaction system answers questions based on the questions raised in each round of conversation, and sometimes it cannot ensure the context coherence of the conversation. This results in the inability to achieve context understanding during multi-round interactions by users, and the content replied by the voice system may be irrelevant to the previous conversation content. There is a problem of insufficient context relevance.

[0054] Taking non-task-based conversations as an example, in non-task-based conversations, users' questions often do not focus on a single task, but rather a continuous and progressive process, and users' needs may be gradually expressed in multiple conversation turns. The system needs to be able to understand and adapt to changes in users' intentions, and dynamically adjust responses according to the context of multi-round conversations. For non-task-based conversation scenarios, such as planning a travel route, interest guidance, emotional communication, or multi-topic discussions, traditional voice interaction systems usually cannot effectively understand and associate users' intentions, resulting in fragmented conversations and information omission. In the process of handling multi-round conversations, the existing technologies usually face the following problems:

[0055] 1. Lack of relevance in conversation context

[0056] Existing voice interaction systems usually respond based on the input of a single-round conversation, lacking adaptability to changes in users' intentions in multi-round conversations. Especially in non-task-based conversations, users may put forward multiple topics or emotional communication content without clearly expressing a specific task. Existing systems cannot flexibly capture these potential needs from multiple conversation turns, resulting in the inability to meet users' actual intentions in non-task-based conversations. For example, when a user asks "What are the tourist attractions in Beijing?" in the first round and then asks "What are the nearby hotels?" the existing system usually recommends hotels based on the current geographical location, rather than associating with the previous context of "tourist attractions in Beijing", resulting in the inability to meet the user's true intention.

[0057] 2. Insufficient coping ability in complex scenarios

[0058] In complex scenarios such as multi-topic switching or correcting the system's answer, the capabilities of existing systems are particularly weak. For example, when the user inputs "What you just said is not the scenic spot I'm looking for", the existing system has difficulty understanding that the user's intention is to correct the previous answer and cannot dynamically adjust the recommended results, easily leading to a decline in the user experience.

[0059] 3. Limited semantic understanding accuracy

[0060] Existing systems usually rely on pre-trained models or rule-based semantic parsing methods. When dealing with multi-turn semantic associations, these methods are restricted by training data and rule design and are difficult to accurately judge the intentions expressed by users in multi-turn conversations. In addition, for complex tasks involving multiple entities and relationships, such as travel planning and education consulting scenarios, existing technologies lack efficient information organization and semantic reasoning mechanisms, resulting in the difficulty of fully understanding and meeting user needs.

[0061] Currently, some voice interaction systems in the market attempt to improve the user experience through dialogue context information, and there are mainly the following implementation schemes:

[0062] 1. Rule-based multi-turn dialogue model

[0063] This method relies on manually written dialogue rules or templates to associate user inputs and system answers according to predefined logic, and is applicable to dialogue scenarios of single tasks, such as ordering takeout and hotel reservation. Although effective in simple scenarios, it lacks flexibility when facing non-task-based conversations. For example, when the user queries a tourist attraction and then asks "What are the nearby hotels?", the system may directly answer by matching the rules. However, the limitation of this method is that the rule design needs to cover a large number of scenarios and it is difficult to adapt to the rapid changes in user intentions and complex situations.

[0064] 2. Neural network-based context dialogue model

[0065] Based on deep learning methods such as recurrent neural networks (RNNs) or transformers, the system's adaptability to multi-turn conversations has been improved to a certain extent. However, when facing the rapid switching of multiple topics or emotions, there are still problems with low semantic association accuracy, especially in non-task-based conversations, lacking the speculation and understanding of complex emotions and interests.

[0066] 3. Method based on dialogue state tracking (DST)

[0067] This method helps the system understand the context relationship in multi-turn conversations by tracking the dialogue state and is commonly used in task-oriented conversations. However, the DST method usually relies on a clear task process design and lacks the ability to flexibly handle open-ended dialogue scenarios. Especially in non-task-oriented conversations involving multi-topic switching or emotional correction, it cannot fully play its role.

[0068] Deficiencies of the prior art

[0069] The above prior art has many deficiencies in multi-turn conversations:

[0070] (1) Lack of dynamic knowledge expansion ability and inability to construct associated information in real time according to user needs.

[0071] (2) Limited ability to understand complex contexts and difficult to accurately handle multi-turn semantic associations.

[0072] (3) Poor adaptability to scenarios of dialogue interruption, correction, or topic switching.

[0073] These problems seriously affect the user experience of the voice interaction system. Especially in complex tasks that require multi-turn conversations to complete, the prior art is difficult to meet the actual needs of users. Therefore, a new technical solution is needed that can dynamically associate dialogue contexts, enhance semantic understanding capabilities, and provide efficient and accurate voice services in complex multi-turn interaction scenarios.

[0074] In view of this, the embodiments of the present application provide a dialogue method. Please refer to Figure 1 , Figure 1 which is the flowchart of the dialogue method in the embodiments of the present application. The method includes the following operations:

[0075] S101. Receive the statement to be replied currently input by the user, and based on the statement to be replied, extract the first key information, where the first key information includes the first keyword and the first user intention;

[0076] S102. Determine whether there is a context association between the statement to be replied and the historical dialogue statements. If so, update the first keyword and the first user intention based on the historical dialogue statements to obtain the second key information;

[0077] S103. Based on the second key information and the preset knowledge base, expand and update the preset first knowledge graph to obtain the second knowledge graph, where the first knowledge graph is a historical knowledge graph generated based on the historical dialogue content;

[0078] S104. Output the dialogue content according to the second key information and the second knowledge graph.

[0079] The dialogue method provided by the embodiments of the present application extracts relevant information in real time according to user input, dynamically constructs and updates a knowledge graph, ensures the real-time nature of semantic understanding, correlates multi-round dialogues, improves the coherence of user intention parsing, and can also be applied to integrate and reason about multi-domain knowledge, support multi-domain scenarios such as tourism, transportation, and food, generate answers that meet user expectations through cross-domain knowledge reasoning, and can provide accurate answers in complex scenarios such as multi-topic switching and emotion correction.

[0080] In the operation of S101, the text statement can be analyzed through a semantic parsing and intention recognition module to identify the core needs of the user (i.e., keywords and user intentions).

[0081] During the dialogue process, the user can communicate with the machine through voice input and may also communicate with the machine through text input. When the statement to be replied to for the user's current input is a voice statement, the operation of S101 specifically includes the following steps:

[0082] First step, convert the voice statement into a text statement.

[0083] Second step, perform semantic analysis based on the text statement to obtain the first key information.

[0084] In the steps of the above first-step operation, the user's currently input voice statement can be converted into a text statement through automatic speech recognition (ASR, Automatic Speech Recognition) technology. The working process of automatic speech recognition includes steps such as audio input, preprocessing, feature extraction, speech recognition, language modeling, and output generation. Its application scenarios are extensive, such as voice assistants for smart phones, in-vehicle intelligent voice interaction systems, and smart home devices.

[0085] In the steps of the above second-step operation, a semantic analysis model can be used to analyze the grammar and semantic structure of the text statement. For example, dependency syntax analysis can determine the dependency relationships between words in the text statement, find the core words and modifiers, etc., and extract the first keyword. At the same time, by understanding the overall semantic framework of the text statement, such as judging whether the sentence is describing a scenario, expressing an opinion, or asking a question, etc., the first user intention is inferred.

[0086] For example, when the voice input by the user is to query "What are the tourist attractions in Beijing?", in the operation of S101, after converting the user's voice input into a text statement through an automatic speech recognition system, a semantic analysis model is used to extract keywords such as "Beijing" and "tourist attractions" from the text statement, as well as the user intention of "hoping to get recommendations on Beijing attractions", so as to constitute the first key information.

[0087] Among them, the semantic analysis model can be a rule-based semantic analysis model, a statistic-based semantic analysis model, a deep learning-based semantic analysis model, or a knowledge graph-based semantic analysis model. The embodiments of the present application do not limit the type of the semantic analysis model.

[0088] In the operation of S102, it is determined whether there is a context association between the current round of conversation and the historical conversation. If it is determined to give a reasonable answer in subsequent operations, specifically, a method based on rule matching, a method based on statistics, a method based on deep learning, or a method based on semantic understanding and knowledge graph can be used to determine whether there is a context association between the statement to be replied and the historical conversation statement.

[0089] The method based on rule matching can specifically adopt operations such as keyword matching or semantic rule matching. In some examples, in S102, the first key information is compared with the historical key information extracted from the historical conversation content. If the relevance is greater than or equal to a preset threshold, it is determined that there is a context association between the statement to be replied and the historical conversation statement; if the relevance is less than the preset threshold, it is determined that there is no context association between the statement to be replied and the historical conversation statement.

[0090] The historical key information includes multiple historical keywords and multiple historical user intents extracted from the historical conversation statements. Specifically, the relevance between the first keyword and the multiple historical keywords can be compared. If there is a historical keyword that at least partially overlaps with the first keyword, it is considered whether there is a context association between the statement to be replied and the historical conversation statement. For another example, when the first keyword includes multiple keywords, each keyword can also be compared with the historical keywords, and the relevance between the two can be determined according to the number, degree, and frequency of the overlap between the first keyword and the historical keywords.

[0091] In this way, by comparing the first key information with the historical key information, if the key information is consistent or there is a logical continuation, it can be determined that there is a context association between the statement to be replied and the historical conversation statement.

[0092] It is also possible to determine whether there is a context association between the statement to be replied and the historical conversation statement through a method based on semantic understanding, such as a vector space model or a deep learning model.

[0093] Exemplarily, by using pre-trained language models such as BERT (Bidirectional Encoder Representations from Transformers) and GPT (Generative Pre-trained Transformer), etc., the currently input reply statement and historical dialogue statements are mapped to a low-dimensional semantic space to obtain vector representations, and the vector similarity is calculated to determine the degree of context relevance. The semantic understanding ability of the model can also be utilized to perform tasks such as semantic role labeling and relation extraction, and analyze whether there is context relevance between the currently input reply statement and the historical dialogue statements in terms of semantic roles and semantic relations.

[0094] In some alternative embodiments, the operation of updating the first keyword and the first user intention based on the historical dialogue statements in S102 to obtain the second key information specifically includes the following steps:

[0095] First step, update the first keyword based on the historical keywords extracted from the historical dialogue statements to obtain the second keyword.

[0096] Second step, update the first user intention based on the association between the historical keywords and / or the historical dialogue statements and the reply statement to obtain the second user intention, and the second user intention and the second keyword constitute the second key information.

[0097] If it is determined that there is context relevance between the reply statement and the historical dialogue statements, it indicates that the currently input reply statement by the user is the subsequent round of a multi-turn dialogue. In a multi-turn dialogue, the user's topic and intention change continuously. Based on the association, that is, the logical relationship, between the historical keywords and / or the historical dialogue statements and the reply statement, the information in the user's historical input dialogue statements and the currently input reply statement is combined to fully understand the user's true intention, making the updated second user intention more in line with the user's needs, so as to facilitate giving an answer that fits the question raised by the user in subsequent operations and provide a more natural and fluent dialogue experience for the user.

[0098] Therefore, the operation in S102 updates the first key information to the second key information based on context relevance, more accurately analyzes the user's true needs, and further optimizes the interaction experience.

[0099] In some alternative embodiments, as Figure 2 shown, the operation of S103 specifically includes the following steps:

[0100] S201. Extract key entity information from the knowledge base based on the second key information. The key entity information includes entities, relationships, and / or attributes related to the second keyword, as well as entities, relationships, and / or attributes related to the second user intention. The preset knowledge base includes web pages, documents, databases, social media, etc.

[0101] S202. Add, modify, or delete entity information in the first knowledge graph based on the key entity information to obtain a second knowledge graph.

[0102] The second key information is more in line with the user's true needs. The operations of S201 and S202 above update the first knowledge graph to the second knowledge graph based on the second key information, correct and supplement errors and omissions in the information related to the user's true intention, and improve the accuracy and integrity of the knowledge content in the knowledge graph in real time, so as to understand the user's question more accurately. Therefore, in subsequent operations, it is easier to understand the user's search intention and needs, provide more accurate search results and personalized recommendations, and improve the user experience.

[0103] In some alternative embodiments, as Figure 3 shown, the operation of S104 specifically includes the following steps:

[0104] S301. Perform entity extraction and relationship construction on the knowledge content of the second knowledge graph according to the second key information to obtain target entity information.

[0105] S302. Perform logical reasoning based on the target entity information and preset reasoning rules, generate an answer corresponding to the target entity information, and integrate and optimize the answer to obtain and output the conversation content.

[0106] The operations of S301 and S302 above use the knowledge in the second knowledge graph for reasoning to provide a comprehensive answer to the user. Among them, the integration operation in the second step can organize the answer involving multiple information segments to make it conform to logical and language expression habits. The optimization operation can polish the generated answer in terms of language, supplement details, and adjust the format, so that its expression is more fluent, natural, and accurate, making the answer richer, more in-depth, and improving the readability of the answer.

[0107] In some alternative embodiments, after generating the answer corresponding to the target entity information, synthesize the answer into a voice response and feedback it to the user.

[0108] It can be understood that during the process of a user's conversation with a machine, the statement to be replied to input by the user can be a text statement or a voice statement, and the answer output by the machine can be a text statement or a voice statement. For example, when the user inputs a statement to be replied to in text format through a keyboard, the machine can output an answer statement in text format or an answer statement in voice format.

[0109] In the operations of S102 and S103 above, based on the association relationship between the historical conversation statement and the statement to be replied to, the first key information is updated to the second key information, and the first knowledge graph is updated to the second knowledge graph. Therefore, in the operation of S104, an answer corresponding to the statement to be replied to currently input by the user is output based on the updated second key information and the second knowledge graph, ensuring the accuracy and context coherence of the answer.

[0110] In this way, when the user's question is vague, the true intention of the user is supplemented through context information and graph reasoning, not only filling in the gaps based on the existing knowledge graph information, but also generating a personalized and accurate answer that meets the user's needs through semantic reasoning.

[0111] In some other feasible embodiments, after determining whether there is a context association between the statement to be replied to and the historical conversation statement, as Figure 4 shown, the method further includes:

[0112] S401. If it is determined that there is no context association between the statement to be replied to and the historical conversation statement, then a third knowledge graph is generated based on the first key information and a preset knowledge base.

[0113] S402. Output the conversation content according to the first key information and the third knowledge graph.

[0114] If there is no context association between the statement to be replied to and the historical conversation statement, it proves that the current conversation is not a subsequent round of the historical conversation. At this time, a third knowledge graph corresponding to the current conversation is constructed based on the first key information and the preset knowledge base, and then the most suitable response is generated based on the third knowledge graph. Among them, the preset knowledge base includes web pages, documents, databases, and social media, etc.

[0115] In some examples, the operation of S401 is specifically as follows: First, according to the first keyword, relevant data is collected from a preset knowledge base, and the application target is determined according to the user's intention, such as for intelligent question answering, information retrieval, recommendation systems, etc. Second, entities related to the first keyword and the first user intention are identified from the collected data, such as personal names, place names, organization names, product names, etc. Specifically, methods based on rules, statistical models or deep learning can be used for entity extraction. Then, the relationships between entities are determined, such as "belong to", "contain", "associate", "causal" and other relationships. Relationship extraction usually needs to be realized by combining syntactic analysis, semantic understanding and machine learning algorithms. Finally, the attribute information of the entities is extracted, such as the age, gender, occupation of a person, the price, specifications, brand of a product, etc. Attribute extraction can help to more comprehensively describe the characteristics of entities. Based on the identified entities, relationships and attributes, the third knowledge graph can be constructed.

[0116] In some examples, the operation of S402 is specifically as follows:

[0117] In the first step, according to the first key information, entity extraction and relationship construction are performed on the knowledge content of the third knowledge graph to obtain target entity information.

[0118] In the second step, logical reasoning is performed based on the target entity information and preset reasoning rules to generate an answer corresponding to the target entity information, and the answer is integrated and optimized to obtain and output the dialogue content.

[0119] The operations of the above first step and second step realize reasoning using the knowledge in the third knowledge graph to provide a comprehensive answer for the user to the current dialogue when there is no context association between the user's currently input sentence to be replied and the historical dialogue sentences.

[0120] Next, an example of a machine having an intelligent dialogue with a user using the dialogue method provided in the above embodiments is used to illustrate the dialogue method provided in the embodiments of the present application.

[0121] Example 1: Multi-turn dialogue with context association

[0122] (1) User's question: "What are the tourist attractions in Beijing?" The system identifies the entities of "Beijing" and "tourist attractions" through semantic parsing and constructs relevant entity relationships in the knowledge graph.

[0123] (2) The user then asks: "What are the nearby hotels?" The system detects that this question is contextually related to the previous question of "tourist attractions in Beijing". Therefore, the system associates "hotels" with "tourist attractions", updates the knowledge graph based on the already constructed knowledge graph, and at the same time returns relevant hotel information such as "It is recommended that you stay at the Beijing Hotel near Beijing Happy Valley".

[0124] Example 2: Complex Conversation with Corrected Answer

[0125] (1) User's question: "What are the most famous scenic spots in Shanghai?", and the system generates an answer: "The Bund and the Oriental Pearl Tower."

[0126] (2) User corrects: "No, I want to ask about the scenic spots in Hangzhou.", and the system dynamically updates the knowledge graph, replaces the original "Shanghai scenic spots" information, and generates a new and accurate answer based on the updated graph.

[0127] The system detects that the user has modified the previous information. By comparing the context and the previous conversation, it dynamically updates the knowledge graph, adjusts the relevant entities (such as replacing "Shanghai" with "Hangzhou"), and corrects it to: "It is recommended to visit the West Lake and Lingyin Temple." The system can quickly process the corrected content and prevent incorrect answers from continuing to exist.

[0128] Example 3: Multi-turn Conversation without Context Association

[0129] (1) User's question: "What are the tourist attractions in Beijing?", and the system identifies the entities of "Beijing" and "tourist attractions" through semantic parsing and constructs relevant entity relationships in the knowledge graph.

[0130] (2) The user then asks: "What hotels are near the West Lake and Lingyin Temple?", and the system detects that this question has no context association with the previous question of "Beijing tourist attractions". Therefore, the system associates "the West Lake", "Lingyin Temple" with "hotels", builds a new knowledge graph, and returns relevant hotel information.

[0131] Based on the same inventive concept, an embodiment of the present application also provides a dialogue system, as Figure 5 shown. The dialogue system 500 includes:

[0132] An extraction unit 501, configured to receive a to-be-replied statement currently input by a user, and extract first key information based on the to-be-replied statement, where the first key information includes a first keyword and a first user intention;

[0133] A judgment and update unit 502, configured to judge whether there is a context association between the to-be-replied statement and historical conversation statements. If so, update the first keyword and the first user intention based on the historical conversation statements to obtain second key information;

[0134] An update unit 503, configured to expand and update a preset first knowledge graph based on the second key information and a preset knowledge base to obtain a second knowledge graph, where the first knowledge graph is a historical knowledge graph generated based on historical conversation content;

[0135] The dialogue output unit 504 is configured to output dialogue content according to the second key information and the second knowledge graph.

[0136] The dialogue system 500 provided by the embodiments of the present application can be used to implement the dialogue method provided by the above embodiments. It realizes voice multi-round interaction based on dynamic update of the knowledge graph. By dynamically constructing and applying the knowledge graph, it enhances the semantic understanding ability of the dialogue system for context in multi-round voice conversations, and solves the deficiencies of existing voice interaction systems in complex dialogue scenarios.

[0137] When the dialogue system 500 provided by the embodiments of the present application is applied to a machine for voice multi-round interaction, its system architecture design mainly includes the following modules. Each module cooperates to perform specific tasks, so as to effectively respond to user needs:

[0138] (1) Automatic Speech Recognition module (ASR): The machine converts the user's voice input into a text statement, ensuring that the voice information can be accurately converted into a form that can be processed by the machine, and preparing for subsequent semantic parsing.

[0139] (2) Semantic Parsing and Intent Recognition module: By analyzing the user input, the machine recognizes the keywords, user intent, and sentiment of the statement and performs context correlation analysis to more accurately understand the user's needs.

[0140] (3) Knowledge Graph Construction and Dynamic Update module: The machine parses the entities and relationships in the input, and combines the context information to integrate this information into the knowledge graph in real time, promoting the dynamic update of the graph and ensuring the continuous improvement of the relationships between entities. To ensure the accuracy, relevance, and fluency of the dialogue.

[0141] (4) Knowledge Reasoning and Semantic Generation module: Based on the constructed knowledge graph, the machine can generate answers related to the user input through reasoning and adjust the answer content according to the user's dynamic needs to ensure the fluency of multi-topic and sentiment-changing conversations.

[0142] (5) Text-to-Speech module (TTS): Converts the generated text answer into voice to complete the voice feedback.

[0143] The following combines Figure 6 to introduce an example of voice multi-round interaction of a voice multi-round interaction system based on a dynamic knowledge graph. Its core technical process can be divided into the following five steps:

[0144] Step 1: User Voice Input and Preliminary Processing

[0145] When receiving user voice input, the system first converts it into a text statement through the speech recognition module. Then, the machine analyzes the text through the semantic parsing and intent recognition module to identify the keywords representing the user's core needs and the user intent, and prepares to enter the next step of processing.

[0146] Step 2: Semantic Parsing and Context Modeling

[0147] In this step, the machine extracts key information such as the user intent and keywords from the user input through the semantic parsing and intent recognition module, and determines whether there is a context association between the current input and the historical conversation content.

[0148] Step 3: Knowledge Graph Construction and Dynamic Update

[0149] If there is a context association, the current input is a subsequent round of the multi-round conversation. The machine will update the user intent and keywords in combination with the context, and update the semantic model.

[0150] Based on the user's first input and the semantic parsing result, the machine extracts relevant entities and relationships from the multi-domain knowledge base in real time to construct a preliminary knowledge graph, and continuously updates and expands the graph according to new inputs in subsequent conversations to ensure that all information in the conversation process is accurately captured. It can be understood that when determining whether there is a context association between the current input and the historical conversation content, the key information is updated (specifically performing the operations of the above-mentioned embodiment S102), and the knowledge graph is dynamically updated (specifically performing the operations of the above-mentioned embodiment S103).

[0151] In a multi-round conversation, whenever the user asks a new question, the system checks whether the current question has a context association with the previous questions. If there is a context association, it is dynamically updated based on the existing knowledge graph. The specific construction and update operations of the knowledge graph are as follows:

[0152] 1. Construct a preliminary knowledge graph: When the user asks a question for the first time, the system identifies the key entities (such as place names, things, etc.) and relationships (such as "is", "belongs to", etc.) in the question through semantic parsing, extracts this information and constructs a preliminary knowledge graph.

[0153] 2. Context recognition: The system analyzes whether the current question involves information in the previous conversation (such as the user first asks about tourist attractions and then asks related questions about hotels, etc.).

[0154] 3. Dynamic Expansion: If the current question is related to the previous conversation content, the system will extract and expand relevant entities in the existing knowledge graph. For example, when the user asks "What are the tourist attractions in Beijing?", the system extracts entities such as "Beijing" and "tourist attractions" from the tourism knowledge base, and associates relationships such as the location of attractions, transportation methods, and popularity, forming a partial knowledge graph. If the user then asks a related question in a subsequent round, such as "What hotels are nearby?", the system will dynamically update the knowledge graph and associate new information such as "Beijing tourist attractions" and "hotels" in the existing graph to ensure that the answer aligns with the user's intention. Figure 1 Consistent.

[0155] Through the updated knowledge graph, the system can handle cross-domain questions: When the user's question spans multiple domains (such as transitioning from tourist attractions to hotel recommendations), the system not only extracts relevant entities from each domain but also integrates knowledge from various domains through graph reasoning, ensuring that cross-domain relationships are dynamically updated and that the mutual relationships between multiple domains are fully considered when generating responses.

[0156] Through the updated knowledge graph, the system can achieve dynamic reasoning: The system uses the reasoning mechanism in the graph to generate reasonable answers based on the context and user input, and reasons about the answer content through the knowledge graph to ensure the coherence and accuracy of the answers.

[0157] Step 4: Knowledge Reasoning and Semantic Generation

[0158] Based on the updated knowledge graph, the system performs semantic reasoning through the knowledge reasoning and semantic generation module, generating reasonable answers by combining the context and user input.

[0159] The system can first determine the clarity of the user's question, and the specific operations are as follows:

[0160] 1. Clear question handling: When the user asks a clear question, the system directly extracts the most relevant entities and relationships from the latest knowledge graph and generates an accurate answer through reasoning to ensure the accuracy and context coherence of the answer.

[0161] 2. Vague question handling: When the user's question is vague, the machine supplements the user's true intention through context information and graph reasoning. The system not only fills in the gaps based on the existing graph information but also generates personalized and accurate answers that meet the user's needs through semantic reasoning.

[0162] For example: When the user asks "What hotels are nearby?", the system generates an answer by associating the entities "Beijing tourist attractions" and "hotels" in the knowledge graph, such as "We recommend you stay at the Beijing Hotel near Beijing Happy Valley".

[0163] Step 5: Answer Generation and Voice Output

[0164] The generated answer is passed to the text-to-speech synthesis module in text form, and the synthesized voice answer can be output as the voice answer. It can be understood that the voice interaction system feeds back the answer to the user in voice form.

[0165] In summary, the dialogue method and system provided by the embodiments of the present application significantly improve the context understanding ability and response accuracy of the voice interaction system in multi-round dialogues by introducing the knowledge graph and semantic parsing technology, and overcome the limitations of the prior art in complex dialogue scenarios. When the dialogue system and dialogue method provided by the embodiments of the present application are applied to multi-round voice interaction, the following

[0166] Beneficial effects:

[0167] 1. Significantly improve the context understanding ability of the voice interaction system

[0168] The present invention dynamically constructs a knowledge graph, associates the semantic information in each round of the user's dialogue, and realizes accurate modeling of the context. The system can dynamically update the knowledge graph based on the user's multi-round dialogue, and combine the changes in the user's emotions and interests to ensure that each answer is consistent with the user's context intention, solving the problems of dialogue fragmentation and context loss in the prior art.

[0169] For example: in the case where the user continuously asks "What are the tourist attractions in Beijing?" and "What are the nearby hotels?", the system of the present invention can preferentially retrieve answers from the knowledge graph related to "Beijing tourist attractions" instead of relying solely on the current geographical location to generate answers.

[0170] 2. Enhance the semantic parsing and intention recognition ability in complex contexts

[0171] The technical solution of the present invention effectively solves the deficiencies of the existing system in dealing with scenarios such as correction, topic switching, and supplementary information in complex multi-round dialogues:

[0172] The system can dynamically infer the user's true intention based on the entities and relationships in the knowledge graph, combined with the user's emotional changes and interest speculation.

[0173] When the user corrects or supplements the system answer, the system can quickly adjust the semantic parsing logic and reconstruct the associated knowledge graph, so as to provide an answer that better meets the user's needs.

[0174] For example: when the user says "What you just said is wrong. I want to ask about the attractions in Hangzhou", the system can adjust the answer based on the context and dynamic graph without misunderstanding.

[0175] 3. Realize dynamic knowledge expansion and improve the knowledge adaptability of the dialogue system

[0176] By constructing and updating a dynamic knowledge graph in real time, the present invention solves the problems of the existing system relying on a fixed corpus and having limited knowledge coverage:

[0177] The system can extract new knowledge from the knowledge base in real time according to the user input and the context information in the multi-round conversation, and integrate it into the knowledge graph of the current conversation to achieve dynamic expansion.

[0178] The dynamic update of the knowledge graph and the entity association mechanism enable the system to adapt to multi-domain and multi-topic conversation scenarios, such as travel planning, traffic query, restaurant recommendation, etc.

[0179] Actual effect: When the user involves multiple domains in the multi-round conversation (such as switching from "scenic spot recommendation" to "transportation method"), the system can quickly adapt to the user's needs through the knowledge expansion function, ensuring the comprehensiveness and accuracy of the answer.

[0180] 4. Improve the accuracy and coherence of the answers in voice interaction

[0181] Based on the traditional voice interaction system, the present invention integrates knowledge reasoning and dynamic semantic modeling technologies, greatly improving the accuracy and coherence of the answers:

[0182] By combining the context association in the knowledge graph and the dynamic input of the user, the system can generate answers highly relevant to the actual needs, avoiding irrelevant or disjointed answers, and ensuring coherence especially in multi-topic and emotion-changing conversations.

[0183] In semantic ambiguity scenarios (such as the user's question contains fuzzy expressions or dual intentions), the present invention analyzes the user's true needs through graph reasoning to further optimize the interaction experience.

[0184] Example: When the user asks "What are the nearby hotels?", the system can not only provide recommendations for the current geographical location, but also preferentially recommend "hotels near tourist attractions in Beijing" through the context.

[0185] 5. Improve the user interaction experience and enhance the intelligence level of the dialogue system

[0186] The system of the present invention enables the user not to need to repeatedly input key information or correct the system answer through context understanding and the construction of a dynamic knowledge graph, thus significantly reducing the interaction cost and improving user satisfaction.

[0187] In practical applications, the present invention can simulate the experience of having a conversation with a human consultant, and combine sentiment analysis and interest speculation to make the voice interaction more natural, intelligent and user-friendly.

[0188] Scenario Application: In scenarios such as user travel planning, shopping consultation, and online education that require multi-round interactions, the system can provide personalized and precise services, enhancing user trust.

[0189] 6. Improve the robustness to open multi-round dialogue scenarios

[0190] The present invention effectively enhances the system's adaptability to open multi-round dialogue scenarios by introducing knowledge graphs and semantic reasoning technologies:

[0191] The system can handle complex scenarios such as multi-topic dialogue switching, emotion correction, and fuzzy expressions, avoiding the problems that traditional systems are prone to failure in open dialogues. For example: when the user continuously switches multiple topics (such as from "querying tourist attractions" to "recommending restaurants" to "booking hotels"), the system can dynamically integrate knowledge in different fields to ensure that the answers are coherent and meet the user's needs.

[0192] 7. Scalability and practical application value of the technical solution

[0193] The technical solution of the present invention has good scalability and applicability, especially in scenarios such as intelligent vehicles, smart homes, and online customer service that require multi-round conversations and emotion speculation, and has broad application value. In the field of intelligent vehicles: through voice interaction, it can meet the user's multi-round conversation needs for navigation, entertainment, vehicle status query, etc. in real time.

[0194] In the field of smart homes: it supports multi-device control, multi-topic dialogue switching, and scenario-based operations.

[0195] In the field of online customer service: it provides efficient customer service and personalized solutions, improving service quality and efficiency.

[0196] In summary, by introducing knowledge graphs and semantic parsing technologies, this application significantly improves the context understanding ability, semantic parsing ability, and knowledge dynamic expansion ability of the voice multi-round interaction system. While solving the deficiencies of the existing technology, this application effectively improves the user experience and the intelligent level of voice interaction, and has broad commercial value and application prospects.

[0197] The embodiment of this application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a computer, it causes the computer to execute the method provided in the above embodiment.

[0198] The embodiment of this application also provides an electronic device, such as Figure 7As shown, the electronic device 700 includes a memory 701, a processor 702, and a computer program 703 stored in the memory 701 and executable on the processor 702. When the processor 702 executes the computer program 703, the electronic device 700 is caused to execute the method provided in the foregoing embodiment.

[0199] An embodiment of the present application further provides a computer program product, including a computer program which, when run, causes a computer to execute the method provided in the foregoing embodiment.

[0200] Those of ordinary skill in the art can understand that all or part of the steps of implementing the foregoing method embodiments can be completed by hardware related to program instructions. The foregoing program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps including those of the foregoing method embodiments; and the foregoing storage medium includes various media such as ROM, RAM, magnetic disk, or optical disc that can store program code.

[0201] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present application.

[0202] In the specification provided here, a large number of specific details are described. However, it can be understood that the embodiments of the present application can be practiced without these specific details. In some instances, well-known methods, structures, and technologies are not shown in detail so as not to obscure the understanding of this specification.

Claims

1. A conversation method, characterized in that: include: Receive a sentence to be replied currently input by the user, and extract first key information based on the sentence to be replied, where the first key information includes a first keyword and a first user intention; Determine whether there is a contextual association between the sentence to be replied and the historical conversation sentence, and if so, update the first keyword and the first user intention based on the historical conversation sentence to obtain second key information; Based on the second key information and the preset knowledge base, the preset first knowledge graph is extended and updated to obtain a second knowledge graph, where the first knowledge graph is a historical knowledge graph generated based on the historical conversation content; Output the conversation content according to the second key information and the second knowledge graph.

2. The dialogue method according to claim 1, characterized in that: After determining whether there is a contextual association between the sentence to be replied and the historical conversation sentence, the method further includes: If there is no contextual association between the sentence to be replied and the historical conversation sentence, generating a third knowledge graph based on the first key information and a preset knowledge base; Output the conversation content according to the first key information and the third knowledge graph.

3. The dialogue method according to claim 1, characterized in that: Outputting the conversation content according to the second key information and the second knowledge graph includes: According to the second key information, entity extraction and relationship construction are performed on the knowledge content of the second knowledge graph to obtain target entity information; Logical reasoning is performed based on the target entity information and preset reasoning rules to generate answers corresponding to the target entity information, and the answers are integrated and optimized to obtain and output the conversation content.

4. The dialogue method according to claim 1, characterized in that: The method of expanding and updating the preset first knowledge graph based on the second key information and the preset knowledge base to obtain the second knowledge graph includes: Based on the second key information, extract key entity information from the knowledge base; wherein the second key information includes a second keyword and a second user intent, and the key entity information includes entities, relationships and / or attributes related to the second keyword, and entities, relationships and / or attributes related to the second user intent; Based on the key entity information, entity information is added, modified or deleted to the first knowledge graph to obtain the second knowledge graph.

5. The dialogue method according to claim 1 or 2, characterized in that: The determining whether there is a contextual association between the statement to be replied and the historical conversation statement includes: The first key information is compared with the historical key information extracted from the historical conversation content. If the correlation is greater than or equal to a preset threshold, it is determined that there is a contextual association between the statement to be replied and the historical conversation statement; if the correlation is less than the preset threshold, it is determined that there is no contextual association between the statement to be replied and the historical conversation statement.

6. The dialogue method according to claim 1, characterized in that: The sentence to be replied currently input by the user is a voice sentence; and extracting the first key information based on the sentence to be replied includes: Converting the voice sentence into a text sentence; The first key information is obtained by performing semantic analysis based on the text sentence.

7. The dialogue method according to claim 1, characterized in that: The updating of the first keyword and the first user intention based on the historical dialogue sentence to obtain the second key information includes: Based on the historical keywords extracted from the historical dialogue sentences, the first keywords are updated to obtain second keywords; Based on the historical keywords and / or the association between the historical dialogue sentences and the sentences to be replied, the first user intention is updated to obtain a second user intention, and the second user intention and the second keyword constitute the second key information.

8. A dialogue system, characterized in that: include: An extraction unit, configured to receive a sentence to be replied currently input by a user, and extract first key information based on the sentence to be replied, wherein the first key information includes a first keyword and a first user intention; a judgment and updating unit, configured to judge whether there is a contextual association between the sentence to be replied and the historical conversation sentence, and if so, update the first keyword and the first user intention based on the historical conversation sentence to obtain second key information; An updating unit, configured to extend and update a preset first knowledge graph based on the second key information and a preset knowledge base to obtain a second knowledge graph, wherein the first knowledge graph is a historical knowledge graph generated based on the historical conversation content; A dialogue output unit is used to output the dialogue content according to the second key information and the second knowledge graph.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the electronic device executes the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is performed.

Citation Information

Cited By

  • Prompt word optimization method and device, electronic equipment, storage medium and product

    CN121029941A