Man-machine conversation processing method and device, equipment and medium

By constructing a dialogue state graph and using semantic recognition technology, the problems of phrasing jumps and repetitive questioning that occur in multi-turn dialogues of chatbots have been solved, achieving the coherence and fluency of dialogue and improving customer experience.

CN121029972APending Publication Date: 2025-11-28PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511203028.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-26
Publication Date
2025-11-28

AI Technical Summary

Technical Problem

Existing chatbots are prone to skipping lines or asking the same questions repeatedly as the number of conversation rounds increases, leading to conversation breaks and repetitive communication, resulting in a poor customer experience.

Method used

By constructing a dialogue state graph, recording the stages and content of dialogue, and using a graph network structure for semantic recognition and intent analysis, preliminary response content is generated. When there is inconsistency, the response content is regenerated according to the user's intent and context to maintain dialogue coherence.

Benefits of technology

It effectively solved the problems of dialogue breaks and repetitive communication, improved customer experience and dialogue fluency, and ensured that the response content conformed to the current dialogue flow.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121029972A_ABST
    Figure CN121029972A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence and natural language processing, can be applied to the field of intelligent medical treatment and financial science and technology, and discloses a man-machine conversation processing method, device, equipment and medium. Obtaining a dialogue state map associated with the dialogue verbal skill stage and the dialogue content thereof; obtaining dialogue voice information of a current round, and performing semantic recognition according to the dialogue voice information of the current round to obtain a user intention and a verbal skill stage label corresponding to the dialogue of the current round; generating initial response content according to the user intention, and detecting and judging whether the initial response content is consistent with the response trend of the dialogue content in the corresponding verbal skill stage of the dialogue state graph or not by utilizing the dialogue state graph according to the verbal skill stage label corresponding to the current round of dialogue; if not, the response content is regenerated according to the user intention and the context dialogue voice information and the response trend.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of artificial intelligence and natural language processing, and in particular to a human-computer dialogue processing method and device, equipment and medium. BACKGROUND

[0002] With the development of artificial intelligence technology, speech recognition technology and applications have gradually matured, such as some outbound dialogue robots, intelligent customer service, and intelligent assistants, etc. In some scenarios, through artificial intelligence technologies such as speech recognition and semantic understanding, the user's intention and questions can be understood, and services such as autonomous online question and answer interaction can be provided. For example, in the smart medical scenario, dialogue robots help guide hospital procedures to help patients quickly understand the corresponding medical procedures; in the financial scenario, dialogue robots help bank business recommendation, marketing, and collection outbound, which can better maintain customers and achieve rapid business promotion.

[0003] However, in actual question and answer interaction scenarios, dialogue robots are prone to dialogue tact jumps (such as suddenly jumping topics) or forgetting the content that the customer has said and repeating the questions, which causes dialogue breaks or repeated communication problems, resulting in a very poor customer experience. SUMMARY

[0004] The present application provides a human-computer dialogue processing method, device, equipment and medium to solve the technical problem of dialogue breaks or repeated communication caused by the increasing number of dialogue turns of dialogue robots.

[0005] In a first aspect, a human-computer dialogue processing method is provided, comprising:

[0006] According to the historical turn dialogue voice information, the dialogue voice information is recorded in a graph network structure to obtain a dialogue state graph associated with the dialogue tact stage and its dialogue content;

[0007] Obtain the dialogue voice information of the current turn, and perform semantic recognition according to the dialogue voice information of the current turn to obtain the user's intention and the tact stage label corresponding to the current turn dialogue;

[0008] Generate a preliminary answer content according to the user's intention, and use the dialogue state graph to detect whether the preliminary answer content is consistent with the response trend of the dialogue content in the corresponding tact stage of the dialogue state graph according to the tact stage label corresponding to the current turn dialogue;

[0009] If it is not consistent, the response content is regenerated according to the user's intention and the context dialogue voice information according to the response trend.

[0010] In a second aspect, a human-computer dialogue processing device is provided, comprising:

[0011] The atlas construction module is configured to record the dialogue voice information in a graph network structure manner according to the dialogue voice information of the historical turn, to obtain a dialogue state atlas associated with a dialogue dialogue stage and dialogue content thereof;

[0012] The dialogue stage recognition module is configured to acquire dialogue voice information of a current turn, and perform semantic recognition according to the dialogue voice information of the current turn to obtain a dialogue stage label corresponding to a user intention and the current turn dialogue;

[0013] The semantic tracking module is configured to generate a preliminary reply content according to the user intention, and detect and determine, by using the dialogue state atlas, whether the preliminary reply content is consistent with a reply trend of dialogue content in a corresponding dialogue stage of the dialogue state atlas according to the dialogue stage label corresponding to the current turn dialogue;

[0014] The output processing module is configured to re-generate a reply content according to the user intention and the context dialogue voice information according to the reply trend when the preliminary reply content is not consistent with the reply trend of the dialogue content in the corresponding dialogue stage of the dialogue state atlas.

[0015] In a third aspect, a computer device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor implements the steps of the above-mentioned human-computer dialogue processing method when executing the computer program.

[0016] In a fourth aspect, a computer readable storage medium is provided, which stores a computer program, and the computer program implements the steps of the above-mentioned human-computer dialogue processing method when executed by a processor.

[0017] The scheme implemented by the above-mentioned human-computer dialogue processing method, device, equipment and medium can obtain dialogue voice information of a historical round and dialogue voice information of a current round through a client, record the dialogue voice information in a graph network structure mode according to the dialogue voice information of the historical round, obtain a dialogue state graph atlas associated with a dialogue dialogue stage and dialogue content, perform semantic recognition according to the dialogue voice information of the current round, obtain a user intention and a dialogue stage label corresponding to the current round dialogue, generate a preliminary response content according to the user intention, and detect whether the preliminary response content is consistent with a response trend of dialogue content in a corresponding dialogue stage of the dialogue state graph atlas according to the dialogue stage label corresponding to the current round dialogue and the dialogue state graph atlas. If the preliminary response content is not consistent with the response trend, the response content is regenerated according to the user intention and context dialogue voice information according to the response trend, and the response content is transmitted to the client for output. In the present application, for online medical services of intelligent medical treatment and online services of financial and insurance business scenarios, when dialogue behaviors such as dialogue consultation and product marketing occur, the dialogue voice information of the historical round can be recorded in a graph network structure mode, and a dialogue state graph atlas associated with a dialogue dialogue stage and dialogue content can be constructed, that is, the dialogue history and context are effectively modeled by the dialogue state graph atlas. When the current round dialogue, the dialogue stage label corresponding to the current round dialogue is detected to determine whether the generated preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph atlas. If the preliminary response content is not consistent with the response trend, the response content is regenerated, and the response is ensured to be consistent with the current dialogue process, so as to maintain the context continuity in the dialogue, avoid the dialogue stage skipping or repeated questioning due to the increase of dialogue rounds, effectively solve the problem of dialogue break or repeated communication caused by too many dialogue rounds, and achieve the purpose of significantly improving the customer experience and dialogue fluency. BRIEF DESCRIPTION OF DRAWINGS

[0018] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following will briefly introduce the drawings needed to be used in the description of the embodiments of the present application. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0019] Figure 1 is an application environment schematic diagram of a human-computer dialogue processing method in an embodiment of the present application;

[0020] Figure 2 is a flowchart of a human-computer dialogue processing method in a first embodiment of the present application;

[0021] Figure 3 is Figure 2 is a specific implementation flowchart of step S101 in the embodiment;

[0022] Figure 4 is Figure 2 a flowchart of a specific embodiment of step S102 in the method;

[0023] Figure 5 is Figure 2 a flowchart of a specific embodiment of step S103 in the method;

[0024] Figure 6 is a flowchart of a human-computer dialogue processing method in a second embodiment of the present application;

[0025] Figure 7 is a structural diagram of a human-computer dialogue processing device in an embodiment of the present application;

[0026] Figure 8 is a structural diagram of a computer device in an embodiment of the present application;

[0027] Figure 9 is another structural diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION

[0028] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all the other embodiments obtained by those of ordinary skill in the art without creative work fall within the scope of the present application.

[0029] The human-computer dialogue processing method provided by the embodiments of the present application can be applied in, for example, Figure 1In an application environment of the application, the client communicates with the server through a network. The server can obtain dialogue voice information of a historical round and dialogue voice information of a current round through the client, record the dialogue voice information in a graph network structure according to the dialogue voice information of the historical round, obtain a dialogue state graph atlas associated with a dialogue dialogue stage and dialogue content, perform semantic recognition according to the dialogue voice information of the current round, obtain a user intention and a dialogue stage label corresponding to the current round, generate a preliminary response content according to the user intention, and detect whether the preliminary response content is consistent with a response trend of dialogue content in a corresponding dialogue stage of the dialogue state graph atlas according to the dialogue stage label corresponding to the current round. If the preliminary response content is not consistent with the response trend, the response content is regenerated according to the user intention and context dialogue voice information according to the response trend, and the response content is transmitted to the client for output. In the application, for online medical services of intelligent medical care and online services of financial and insurance business scenarios, when dialogue consultation and product marketing and other dialogue behaviors occur, the dialogue voice information of the historical round can be recorded in a graph network structure, and a dialogue state graph atlas associated with a dialogue dialogue stage and dialogue content can be constructed, that is, dialogue history and context are effectively modeled by the dialogue state graph atlas. When the current round of dialogue occurs, whether the currently generated preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph atlas can be detected according to the dialogue stage label corresponding to the current round. If the preliminary response content is not consistent with the response trend, the response content can be regenerated, and the response can be ensured to be consistent with the current dialogue process, so that the context is kept coherent in the dialogue, and the situation of dialogue skipping or repeated questioning due to an increase in dialogue rounds is avoided. The problems of dialogue breaking or repeated communication caused by too many dialogue rounds are effectively solved, and the purpose of significantly improving customer experience and dialogue fluency is achieved. The client can be, but is not limited to, various personal computers, notebook computers, smart phones, tablet computers and portable wearable devices. The server can be implemented by an independent server or a server cluster composed of multiple servers. The application will be described in detail below through specific embodiments.

[0030] Please refer to Figure 2 as shown, Figure 2 A flowchart of a human-computer dialogue processing method provided by the first embodiment of the application is shown in the figure, which includes the following steps:

[0031] S101: According to the dialogue voice information of the historical round, the dialogue voice information is recorded in a graph network structure, and a dialogue state graph atlas associated with a dialogue dialogue stage and dialogue content is obtained.

[0032] In this step, the historical round dialogue stage nodes and their transition paths are modeled in a graph network structure to dynamically represent the dialogue state and capture the association of dialogue stages and dialogue content trends. Among them, the dialogue stages can be stages corresponding to different dialogue topics such as greetings, introductions (such as product-related feature introductions when promoting medical / insurance products), objection handling (such as answering when users express objections to medical service process), and offers.

[0033] Specifically, as shown in FIG. 1, step S101 includes steps S1011-S1013 as follows: Figure 3

[0034] S1011, performing speech recognition on the dialogue voice information of the historical round to obtain a dialogue text sequence.

[0035] In this embodiment, the dialogue voice information is converted into text in real time by ASR (Automatic Speech Recognition) technology to obtain a dialogue text sequence.

[0036] S1012, performing semantic recognition on the dialogue text sequence to obtain corresponding dialogue content and dialogue dialogue stage labels.

[0037] In this step, the pre-trained BERT (Bidirectional Encoder Representations from Transformers) model based on the Transformer encoder can be used to perform semantic understanding and dialogue stage recognition on the dialogue text sequence and the context dialogue voice information of the historical round to obtain the dialogue stage label corresponding to the current round dialogue, and extract the key dialogue content such as the high-frequency phrase or embedding vector of the dialogue stage label in the historical dialogue.

[0038] S1013, taking the dialogue dialogue stage as a node, establishing edges between dialogue dialogue stage labels based on the order of the historical dialogue voice information round, obtaining a dialogue state graph, and associating the dialogue content to the node of the corresponding dialogue dialogue stage label.

[0039] In this embodiment, the TF-IDF or Sentence-BERT can be used to generate content embedding according to the dialogue content, and the content embedding is stored as a node attribute to associate the dialogue content to the node of the corresponding dialogue dialogue stage label.

[0040] ​Specifically, the edges between the dialogue script stage labels based on the turn order of the historical dialogue voice information can be: traversing the dialogue text sequence of the historical turn, creating edges of adjacent dialogue script stage labels based on the turn order and the dialogue script stage label, and taking a predefined key trigger word as a transfer condition. In this embodiment, a key trigger word can be predefined for each dialogue script, and the predefined key trigger word can be stored in the edge. For example, the key trigger word of the price objection handling stage can be defined as "expensive" and the like. If "expensive" is detected in the dialogue in the quotation stage of the medical product promotion, the response is transferred to the objection handling stage, and the stage response script generation and output are performed.

[0041] The human-computer dialogue processing method provided by the application can be applied to intelligent dialogue robots in various application scenarios. The intelligent dialogue robot is usually implemented through a server. The server can obtain dialogue voice information of historical turns, capture the association of the script stage and the dialogue content trend, and construct a dialogue state atlas. For example, in the medical service evaluation scenario in the field of intelligent medical care, the user can be guided to evaluate services such as registration, number taking, and medical treatment through an intelligent dialogue robot, and a dialogue state atlas is constructed according to dialogue voice information. In the product promotion or business marketing scenario in the field of financial technology, the product is often promoted to the user through an intelligent dialogue robot. A dialogue state atlas is constructed according to dialogue voice information of each turn to prevent repeated questioning (such as repeatedly asking the same information) and logical jumps (such as suddenly changing the topic) in the dialogue process, and the product or business promotion efficiency is improved.

[0042] S102: Obtain dialogue voice information of a current turn, and perform semantic recognition according to the dialogue voice information of the current turn to obtain a user intention and a script stage label corresponding to the dialogue of the current turn.

[0043] In this step, the ASR technology can be used to transcribe the dialogue voice information of the current turn into text in real time to obtain a dialogue text sequence of the current turn. Of course, other ways that can achieve the above function can be used instead, which will not be described here.

[0044] Specifically, as shown in Figure 4 The step S102 includes the following steps S1021-S1022:

[0045] S1021: Obtain dialogue voice information of a current turn, and perform voice recognition on the dialogue voice information of the current turn to obtain a dialogue text sequence of the current turn.

[0046] S1022: Use a pre-trained BERT model to perform intention recognition and script stage recognition on the dialogue text sequence of the current turn and the context dialogue voice information to obtain a user intention and a script stage label corresponding to the dialogue of the current turn.

[0047] S103: Generate preliminary response content based on the user intent, and use the dialogue state graph to detect and determine whether the preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph based on the dialogue stage label corresponding to the current round of dialogue.

[0048] In this step, the intelligent chatbot can generate preliminary response content based on the user's intent in the current round. To avoid incoherent dialogue context, a consistency detection mechanism is introduced in conjunction with the dialogue state graph to detect whether the preliminary response content conforms to the current dialogue flow.

[0049] Specifically, such as Figure 5 As shown, in this embodiment, step S103, based on the dialogue stage label corresponding to the current round of dialogue, uses the dialogue state graph to detect and determine whether the initial response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph, including the following steps S1031-S1033:

[0050] S1031. Query the dialogue content of the corresponding dialogue stage in the dialogue state graph according to the dialogue stage label corresponding to the current round of dialogue.

[0051] S1032. Detect the similarity between the preliminary response content and the corresponding response content in the queried dialogue content.

[0052] S1033. Based on the similarity, determine whether the initial response content is consistent with the response trend of the dialogue content in the corresponding speech stage of the dialogue state map.

[0053] For steps S1031-S1033, when detecting similarity, the similarity can be detected by calculating the average embedding vector (trend vector) of the corresponding response content in the queried dialogue content, converting the initial response content into an embedding, and calculating the cosine similarity between the generated initial response content embedding and the trend vector. If the cosine similarity is greater than a preset threshold (e.g., 0.7), it is determined to be consistent.

[0054] S104: If there is a discrepancy, the response content is regenerated according to the user intent and contextual dialogue voice information based on the response trend.

[0055] In this step, if the similarity is less than the preset threshold, it is determined that the preliminary response content is inconsistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph, at which time the regeneration or error correction mechanism is automatically triggered, the response content in the dialogue state graph is used to constrain the response generation of the intelligent robot, so as to maintain the dialogue coherence by forcing the trend consistency, that is, the response content in the dialogue state graph is taken as a condition to guide the generation of the expected response content. Preferably, in the present embodiment, a pre-trained LLM (Large Language Model) can be used to generate a response consistent with the trend in combination with the user's intention and context, and the context extraction can avoid asking for existing information.

[0056] For example, in a dialogue in which an intelligent dialogue robot introduces the price of a medical product, the user proposes that the price is high, and the dialogue theme needs to be converted from price introduction to price objection processing. When traversing the dialogue state graph, the node is transferred from the introduction stage node to the price objection processing node. At this time, the generated response content needs to conform to the response trend associated with the price objection processing node in the dialogue state graph. If the preliminary response content generated by the intelligent robot is a product feature introduction, dialogue rupture will occur. Using the human-computer dialogue processing method of the present application, a response related to price concessions can be regenerated according to the response trend associated with the price objection processing node in the dialogue state graph, such as generating a response of "This product has the following discounts...". The output can improve the natural fluency of multi-round dialogue, thereby improving the demand mining and conversion efficiency. It can be understood that the dialogue state graph can be incrementally updated after each new dialogue.

[0057] As can be seen, in the above scheme, for online services of intelligent medical care and online services of financial and insurance business scenarios, when dialogue consultation and product marketing promotion and other dialogue behaviors occur, the dialogue voice information of the historical rounds can be recorded in the form of a graph network structure, and a dialogue state graph associated with the dialogue stage and the dialogue content thereof can be constructed, that is, the dialogue history and context are effectively modeled by the dialogue state graph. In the current round of dialogue, whether the generated preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph can be determined according to the dialogue stage label corresponding to the current round of dialogue. If they are inconsistent, the response content can be regenerated to ensure that the response conforms to the current dialogue process, so as to maintain the context coherence in the dialogue and avoid the situation of dialogue stage skipping or repeated questioning due to the increase in dialogue rounds, so as to effectively solve the problems of dialogue rupture or repeated communication caused by too many dialogue rounds, and achieve the purpose of significantly improving the customer experience and dialogue fluency.

[0058] It should be understood that the size of the serial number of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0059] Referring to Figure 6 , Figure 6 A flowchart of a human-computer dialogue processing method provided by a second embodiment of the present application includes the following steps:

[0060] S201, before the dialogue is started, the customer portrait and the dialogue state corresponding to the dialogue object are retrieved and called.

[0061] In this step, the customer portrait and the dialogue state corresponding to the dialogue object can be automatically retrieved and called before the dialogue is started, wherein the customer portrait is a personalized label generated according to gender, age, shopping record, and communication style (such as language style, reaction speed, interruption frequency, and emotion during communication) of the customer in the historical call; the dialogue state includes the dialogue topic and the user feedback content, etc.

[0062] S202, determine the dialogue style according to the customer portrait, and generate a transition technique according to the last dialogue state.

[0063] In this step, the transition technique can be generated by using the context manager according to the last dialogue topic and the user feedback content, and the customer portrait can be used for dialogue style matching to adjust the dialogue strategy, adapt to different customer groups, ensure that the AI dialogue tone and recommendation rhythm are consistent with the customer communication style, and improve the customer satisfaction.

[0064] S203, start the dialogue after splicing the transition technique and the dialogue topic technique according to the dialogue style.

[0065] Specifically, the transition technique and the dialogue topic technique can be spliced according to the dialogue style by using the Prompt technology to generate the starting technique of this dialogue, for example, the last dialogue ended with the topic of recommending a financial-type insurance product, and the user did not purchase, the dialogue topic this time is to recommend a protection-type insurance product, and the starting technique such as "last time you said you were not interested in the financial-type insurance product, so I recommend a protection-type insurance product today, how about it?" can be generated, which realizes continuous communication, supports cross-day follow-up, and can enhance the stickiness.

[0066] S204, record the dialogue voice information in the form of a graph network structure according to the historical round dialogue voice information, and obtain a dialogue state atlas associated with the dialogue technique stage and the dialogue content.

[0067] This step is the same as or similar to step S101, and will not be described here.

[0068] S205, obtain the dialogue voice information of the current round, and perform semantic recognition according to the dialogue voice information of the current round to obtain the user intent and the technique stage label corresponding to the current round dialogue.

[0069] This step is the same as or similar to step S102, and will not be described here.

[0070] S206, generating preliminary response content according to the user intention, and determining whether the preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph according to the dialogue stage label of the current round of dialogue.

[0071] This step is the same as or similar to step S103, and will not be described here.

[0072] S207, if not consistent, re-generating response content according to the user intention and context dialogue voice information according to the response trend.

[0073] This step is the same as or similar to step S104, and will not be described here.

[0074] S208, detecting whether the current dialogue is interrupted, marking the interruption at the interruption dialogue stage when the current dialogue is interrupted, and saving the current dialogue state.

[0075] In the present application, when the call is unexpectedly interrupted (such as the user hanging up / muting), the interruption can be marked, and the current dialogue state can be automatically saved, wherein the current dialogue state includes the interruption dialogue stage position, the dialogue intention, the user feedback content before interruption, and the user emotion, etc., which facilitates the continuation of the next dialogue, that is, the intelligent robot will call the current dialogue state saved at the time of interruption to generate appropriate opening connection dialogue (such as: "We just talked about… Would you like to continue chatting?") to ensure the continuity of the context and improve the connection rate and completion rate.

[0076] As can be seen, in the above scheme, for the online service of intelligent medical treatment and the online service of business scenarios such as finance and insurance, when dialogue behaviors such as dialogue consultation and product marketing and promotion occur, the dialogue history and context can be effectively modeled by the dialogue state graph, and in the current round of dialogue, it can be determined whether the currently generated preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph according to the dialogue stage label corresponding to the current round of dialogue, and the response content can be re-generated when it is not consistent, ensuring that the response conforms to the current dialogue process, so as to maintain the context coherence in the dialogue, avoid the situation of dialogue stage jumping or repeated questioning due to the increase of dialogue rounds, significantly improve the customer experience and dialogue fluency, and also have an interruption recovery mechanism, which can mark the interruption and save the current dialogue state when interrupted, facilitating the smooth connection to the interruption place when the dialogue is resumed after interruption, and significantly improving the re-conversion potential of the unfinished call.

[0077] In one embodiment, a human-computer dialogue processing device is provided, which corresponds one-to-one with the human-computer dialogue processing method in the second embodiment described above. For example... Figure 7 As shown, the human-computer dialogue processing device includes a graph construction module 110, a speech recognition module 120, a semantic tracking module 130, and an output processing module 140. Detailed descriptions of each functional module are as follows:

[0078] The graph construction module 110 is used to record the dialogue voice information in a graph network structure according to the dialogue voice information of the historical rounds, and obtain a dialogue state graph that is associated with the dialogue speech stages and their dialogue content.

[0079] The speech recognition module 120 is used to acquire the dialogue voice information of the current round and perform semantic recognition based on the dialogue voice information of the current round to obtain the user intent and the speech stage label corresponding to the current round of dialogue.

[0080] The semantic tracking module 130 is used to generate preliminary response content according to the user intent, and to use the dialogue state graph to detect and determine whether the preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph according to the speech stage label corresponding to the current round of dialogue.

[0081] The output processing module 140 is used to regenerate the response content according to the user intent and contextual dialogue voice information when the initial response content is inconsistent with the response trend of the dialogue content in the corresponding speech stage of the dialogue state graph.

[0082] In one embodiment, the map construction module 110 is specifically used for:

[0083] Speech recognition is performed on the dialogue speech information from previous rounds to obtain the dialogue text sequence;

[0084] Semantic recognition is performed on the dialogue text sequence to obtain the corresponding dialogue content and dialogue stage labels;

[0085] Using the dialogue speech stage as a node, and based on the turn order of historical dialogue voice information, edges are established between dialogue speech stage labels to obtain a dialogue state graph, and the dialogue content is associated with the corresponding dialogue speech stage label node.

[0086] In one embodiment, the map construction module 110 is specifically used for:

[0087] Traverse the dialogue text sequence of historical rounds, and create edges between adjacent dialogue stage labels based on round order and dialogue stage labels, using predefined key trigger words as transition conditions.

[0088] In one embodiment, the script recognition module 120 is specifically used for:

[0089] Obtain the dialogue voice information of the current round, and perform speech recognition on the dialogue voice information of the current round to obtain the dialogue text sequence of the current round;

[0090] The pre-trained BERT model is used to perform intent recognition and speech phase recognition on the dialogue text sequence and contextual dialogue speech information of the current round, so as to obtain the user intent and the speech phase label corresponding to the current round of dialogue.

[0091] In one embodiment, the semantic tracking module 130 is specifically used for:

[0092] Based on the dialogue stage tag corresponding to the current round of dialogue, query the dialogue content of the corresponding dialogue stage in the dialogue state graph;

[0093] Detect the similarity between the initial response content and the corresponding response content in the queried dialogue content;

[0094] Based on the similarity, it is determined whether the initial response content is consistent with the response trend of the corresponding dialogue content in the dialogue state map during the corresponding speech stage.

[0095] In one embodiment, the output processing module 140 is further specifically used for:

[0096] Detect whether the current conversation is interrupted. If the current conversation is interrupted, mark the interruption at the interruption statement and save the current conversation state.

[0097] In one embodiment, the output processing module 140 is further configured to:

[0098] Before a conversation begins, retrieve and call the customer profile and conversation status corresponding to the conversation object;

[0099] The conversation style is determined based on the customer profile, and the connecting lines are generated based on the previous conversation status.

[0100] The dialogue is initiated by combining the connecting phrases and the topic phrases of the current conversation according to the stated dialogue style.

[0101] As can be seen, the present invention provides a human-computer dialogue processing device that dynamically represents the dialogue state using a graph network structure, captures the correlation between the dialogue stages and the trend of the dialogue content, and constructs a dialogue state graph. In the current round of dialogue, the device can detect and determine whether the initially generated response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph based on the dialogue stage label corresponding to the current round of dialogue. If they are inconsistent, the response content can be regenerated to ensure that the response conforms to the current dialogue flow, so as to maintain the contextual coherence in the dialogue and avoid the situation of dialogue jumps or repeated questions due to the increase of dialogue rounds. This effectively solves the problem of dialogue breakage or repeated communication caused by a large number of dialogue rounds, and significantly improves customer experience and dialogue fluency.

[0102] Specific limitations regarding the human-computer dialogue processing device can be found in the limitations of the human-computer dialogue processing method described above, and will not be repeated here. Each module in the aforementioned human-computer dialogue processing device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in or independent of the processor in the computer device in hardware form, or stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to each module.

[0103] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 8 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with external clients via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a human-computer interaction processing method on the server side.

[0104] In one embodiment, a computer device is provided, which may be a client, and its internal structure diagram may be as follows: Figure 9As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps on the client side of a human-computer interaction processing method.

[0105] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:

[0106] Based on the dialogue voice information of historical rounds, the dialogue voice information is recorded in a graph network structure to obtain a dialogue state graph that is associated with the dialogue speech stages and their dialogue content.

[0107] Acquire the dialogue voice information of the current round, and perform semantic recognition based on the dialogue voice information of the current round to obtain the user intent and the corresponding speech stage label of the current round of dialogue;

[0108] A preliminary response is generated based on the user's intent, and the dialogue state graph is used to detect and determine whether the preliminary response is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph based on the speech stage label corresponding to the current round of dialogue.

[0109] If there is a discrepancy, the response content is regenerated according to the user's intent and the contextual dialogue voice information, following the response trend.

[0110] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, performs the following steps:

[0111] Based on the dialogue voice information of historical rounds, the dialogue voice information is recorded in a graph network structure to obtain a dialogue state graph that is associated with the dialogue speech stages and their dialogue content.

[0112] Acquire the dialogue voice information of the current round, and perform semantic recognition based on the dialogue voice information of the current round to obtain the user intent and the corresponding speech stage label of the current round of dialogue;

[0113] A preliminary response is generated based on the user's intent, and the dialogue state graph is used to detect and determine whether the preliminary response is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph based on the speech stage label corresponding to the current round of dialogue.

[0114] If there is a discrepancy, the response content is regenerated according to the user's intent and the contextual dialogue voice information, following the response trend.

[0115] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and client side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.

[0116] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0117] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0118] The above-described embodiments are merely illustrative of the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. These modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention. Furthermore, any software tools or components not belonging to this company that appear in the embodiments of this application are merely illustrative examples and do not represent actual use.

Claims

1. A human-to-computer dialog processing method, the method comprising: Comprise: According to the historical round of dialogue voice information, record the dialogue voice information in the form of graph network structure, obtain the dialogue state graph associated with the dialogue dialogue stage and its dialogue content; Get the dialogue voice information of the current round, and perform semantic recognition according to the dialogue voice information of the current round to obtain the user intent and the dialogue corresponding to the current round of dialogue stage label; According to the user intent, generate a preliminary response content, and use the dialogue state graph to detect whether the preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph according to the dialogue corresponding to the current round of dialogue stage label; If it is not consistent, regenerate the response content according to the user intent and the context dialogue voice information according to the response trend.

2. The human-to-computer dialog processing method of claim 1, wherein, According to the historical round of dialogue voice information, record the dialogue voice information in the form of graph network structure, obtain the dialogue state graph associated with the dialogue dialogue stage and its dialogue content, comprising: Perform speech recognition on the dialogue voice information of the historical round to obtain a dialogue text sequence; According to the dialogue text sequence, perform semantic recognition to obtain the corresponding dialogue content and dialogue dialogue stage label; Take the dialogue dialogue stage as the node, establish the edge between the dialogue dialogue stage labels based on the round order of the historical dialogue voice information, obtain the dialogue state graph, and associate the dialogue content to the node of the corresponding dialogue dialogue stage label.

3. The human-to-computer dialog processing method of claim 2, wherein, The edge between the dialogue dialogue stage labels is established based on the round order of the historical dialogue voice information, comprising: Traverse the dialogue text sequence of the historical round, create the edge of the adjacent dialogue dialogue stage label based on the round order and the dialogue dialogue stage label, and take the predefined key trigger word as the transfer condition.

4. The human-to-computer dialog processing method of claim 1, wherein, The dialogue voice information of the current round is obtained, and the semantic recognition is performed according to the dialogue voice information of the current round to obtain the user intent and the dialogue corresponding to the current round of dialogue stage label, comprising: Get the dialogue voice information of the current round, and perform speech recognition on the dialogue voice information of the current round to obtain the dialogue text sequence of the current round; Use the pre-trained BERT model to identify the intent and dialogue stage of the dialogue text sequence of the current round and the context dialogue voice information to obtain the user intent and the dialogue corresponding to the current round of dialogue stage label.

5. The human-to-computer dialog processing method of claim 1, wherein, According to the dialogue corresponding to the current round of dialogue stage label, use the dialogue state graph to detect whether the preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph, comprising: According to the dialogue corresponding to the current round of dialogue stage label, query the dialogue content of the corresponding dialogue stage in the dialogue state graph; Detect the similarity between the preliminary response content and the corresponding response content in the dialogue content queried; According to the similarity, determine whether the preliminary response content is consistent with the response trend of the dialogue content in the corresponding dialogue stage of the dialogue state graph.

6. The human-to-computer dialog processing method of claim 1, wherein, The human-computer dialogue processing method further comprises: Detect whether the current dialogue is interrupted, mark the interruption at the interruption dialogue when the current dialogue is interrupted, and save the current dialogue state.

7. The human-to-computer dialog processing method of any one of claims 1-6, wherein, The human-computer dialogue processing method further comprises: Before the dialogue is started, a customer portrait and a dialogue state corresponding to a dialogue object are retrieved and called; A dialogue style is determined according to the customer portrait, and a connecting speech is generated according to the last dialogue state; The dialogue is started after the connecting speech and a dialogue topic speech are spliced according to the dialogue style.

8. A human-to-computer dialog processing apparatus, the apparatus comprising: The method comprises the following steps: A graph construction module is configured to record dialogue voice information in a graph network structure according to dialogue voice information of historical rounds, so as to obtain a dialogue state graph associated with a dialogue speech stage and dialogue content thereof; A speech recognition module is configured to obtain dialogue voice information of a current round, and perform semantic recognition according to the dialogue voice information of the current round, so as to obtain a user intention and a speech stage label corresponding to a dialogue of the current round; A semantic tracking module is configured to generate a preliminary response content according to the user intention, and determine whether the preliminary response content is consistent with a response trend of dialogue content in a corresponding speech stage of the dialogue state graph according to the speech stage label corresponding to the dialogue of the current round and the dialogue state graph; An output processing module is configured to regenerate a response content according to the user intention and context dialogue voice information according to a response trend when the preliminary response content is not consistent with the response trend of the dialogue content in the corresponding speech stage of the dialogue state graph.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, The processor executes the computer program to realize the steps of the human-computer dialogue processing method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1-9. The computer program is executed by the processor to realize the steps of the human-computer dialogue processing method according to any one of claims 1 to 7.