A method and device for intelligent interruption-based dialogue interaction decision-making in human-machine dialogue

By obtaining the user's first utterance in the intelligent dialogue robot and entering the interrupt model, determining the interrupt intention and executing the response strategy, the problem of processing when the user wants to interrupt the conversation is solved, and the user experience is improved.

CN114238565BActive Publication Date: 2025-07-08LINGXI TECHNOLOGY CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202111498428.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-09
Publication Date
2025-07-08
Estimated Expiration
2041-12-09

AI Technical Summary

Technical Problem

Existing smart conversation bots do not know what to do when users want to interrupt conversations, resulting in poor user experience.

Method used

By obtaining the user's first utterance, enter the interrupt model to determine whether there is an intention to interrupt, and obtain the response strategy based on the intention output by the model, and execute the corresponding response strategy.

Benefits of technology

It improves the ability of intelligent dialogue robots to cope with users when they have the intention to interrupt, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114238565B_ABST
    Figure CN114238565B_ABST
Patent Text Reader

Abstract

The present application provides a dialogue interaction decision-making method and device for intelligent interruption in human-machine dialogue. The method includes: during the communication with the user, obtaining the first utterance of the user; inputting the first utterance of the user into an interruption model to obtain an interruption judgment result and a model output intention output by the interruption model; determining whether the user has an interruption intention according to the interruption judgment result output by the interruption model; if the user has an interruption intention, taking the model output intention as the target interruption intention, and obtaining a coping strategy corresponding to the target interruption intention; and executing according to the coping strategy corresponding to the target interruption intention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent dialogue robots. Specifically, it relates to a dialogue interaction decision-making method and device for intelligent interruption in human-machine dialogue. Background Art

[0002] With the continuous development of robot technology, intelligent dialogue robots have emerged in the field of robots. For example, in the field of telephone communication, users have telephone conversations with intelligent dialogue robots, and in the field of e-commerce, users have voice or text conversations with intelligent dialogue robots to communicate about products. Existing intelligent dialogue robots often determine user intentions based on some keywords, and then give corresponding response strategies. For example, the keywords are singing, news broadcast, skin care products, hot pot, and Kobe Bryant. When users have negative emotions towards intelligent dialogue robots or when users think that intelligent dialogue robots do not understand what they want to express and want to interrupt the intelligent dialogue robots, the intelligent dialogue robots do not know how to handle it and still continue to talk to users. Summary of the Invention

[0003] The purpose of the embodiments of this application is to provide a dialogue interaction decision-making method and device for intelligent interruption in human-machine dialogue to solve the technical problem that existing intelligent dialogue robots do not know how to handle it when users want to interrupt the dialogue.

[0004] In a first aspect, a dialogue interaction decision-making method for intelligent interruption in human-machine dialogue is provided, including:

[0005] During the process of communicating with the user, obtain the first utterance of the user;

[0006] Input the first utterance of the user into an interruption model to obtain an interruption judgment result and a model output intention output by the interruption model;

[0007] Determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model;

[0008] If the user has an interruption intention, use the model output intention as the target interruption intention, and obtain a coping strategy corresponding to the target interruption intention;

[0009] Execute according to the coping strategy corresponding to the target interruption intention.

[0010] The above-mentioned dialogue interaction decision-making method for intelligent interruption in human-computer dialogue first obtains the first utterance of the user during the communication with the user; inputs the first utterance of the user into the interruption model to obtain the interruption judgment result and the model output intention output by the interruption model; determines whether the user has an interruption intention according to the interruption judgment result output by the interruption model; then if the user has an interruption intention, takes the model output intention as the target interruption intention and obtains the corresponding coping strategy for the target interruption intention; finally, executes according to the corresponding coping strategy for the target interruption intention. It can be seen that when it is determined according to the user's first utterance that the user has an interruption intention, obtaining the corresponding coping strategy for the target interruption intention to respond enables the intelligent dialogue robot to respond well when the user has an interruption intention, thus solving the technical problem in the prior art that the intelligent dialogue robot does not know how to handle when the user wants to interrupt the dialogue, and further improving the user experience to a certain extent.

[0011] In one embodiment, after the execution according to the corresponding coping strategy for the target interruption intention, it further includes:

[0012] Obtaining the second utterance of the user;

[0013] Determining the user intention of the user according to the first utterance and the second utterance of the user;

[0014] Obtaining the corresponding coping strategy for the user intention;

[0015] Executing according to the corresponding coping strategy for the user intention.

[0016] In one embodiment, the determining the user intention of the user according to the first utterance and the second utterance of the user includes:

[0017] Extracting target key phrases from the first utterance and the second utterance;

[0018] Inputting the target key phrases into the intention recognition model to obtain the user intention of the user output by the intention recognition model.

[0019] In one embodiment, before obtaining the first utterance of the user during the communication with the user, it further includes:

[0020] Training the intention recognition model.

[0021] In one embodiment, the inputting the target key phrases into the intention recognition model to obtain the user intention of the user output by the intention recognition model includes:

[0022] Inputting the target key phrases into the intention recognition model;

[0023] The intention recognition model combines the target keyword sentence with the target corpus in the corpus to obtain multiple combination results;

[0024] The intention recognition model calculates the similarity between the target keyword sentence and the target corpus in each combination result to obtain the target combination result with the largest similarity;

[0025] The corpus intention corresponding to the target corpus in the target combination result is used as the user intention of the user.

[0026] In one embodiment, before the intention recognition model combines the target keyword sentence with the target corpus in the corpus to obtain multiple combination results, it further includes:

[0027] Segment the target keyword sentence to obtain multiple segmentation results;

[0028] Obtain the occurrence probability of each segmentation result in the corpus;

[0029] Obtain the target segmentation result from the multiple segmentation results according to the occurrence probability of each segmentation result in the corpus;

[0030] In the corpus, obtain multiple target corpora containing the target segmentation result.

[0031] In a second aspect, a dialogue interaction decision-making device for intelligent interruption in human-computer dialogue is provided, including:

[0032] An acquisition module, configured to acquire the first utterance of the user during the process of communicating with the user;

[0033] A judgment module, configured to input the first utterance of the user into an interruption model to obtain an interruption judgment result and a model output intention output by the interruption model;

[0034] A determination module, configured to determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model;

[0035] A strategy module, configured to, if the user has an interruption intention, use the model output intention as the target interruption intention and obtain a coping strategy corresponding to the target interruption intention;

[0036] An execution module, configured to execute according to the coping strategy corresponding to the target interruption intention.

[0037] In one embodiment, the dialogue interaction decision-making device for intelligent interruption in human-computer dialogue further includes: a second module, configured to:

[0038] Obtain the second utterance of the user;

[0039] Determine the user intention of the user according to the first utterance and the second utterance of the user;

[0040] Obtain a coping strategy corresponding to the user intention;

[0041] Execute according to the coping strategy corresponding to the user intention.

[0042] In a third aspect, a computer device is provided, which is characterized by including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the dialogue interaction decision-making method for intelligent interruption in the above-mentioned human-computer dialogue are implemented.

[0043] In a fourth aspect, a computer-readable storage medium is provided. The computer-readable storage medium stores computer program instructions. When the computer program instructions are read and run by a processor, the steps of the dialogue interaction decision-making method for intelligent interruption in the above-mentioned human-computer dialogue are executed. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required to be used in the embodiments of the present application will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present application, and therefore should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0045] Figure 1 It is a schematic flowchart of the implementation process of the dialogue interaction decision-making method for intelligent interruption in the human-computer dialogue provided by the embodiment of the present application;

[0046] Figure 2 It is a schematic diagram for determining the interruption judgment result and the model output intention;

[0047] Figure 3 It is a schematic diagram of the composition structure of the dialogue interaction decision-making device for intelligent interruption in the human-computer dialogue provided by the embodiment of the present application;

[0048] Figure 4 It is a schematic diagram of the composition structure of the computer device provided by the embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0049] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0050] In one embodiment, a dialogue interaction decision method for intelligent interruption in human-machine dialogue is provided. The execution subject of the dialogue interaction decision method for intelligent interruption in human-machine dialogue described in the embodiments of the present invention is a computer device capable of implementing the dialogue interaction decision method for intelligent interruption in human-machine dialogue described in the embodiments of the present invention. An intelligent dialogue robot is set in the computer device, and the computer device may include, but is not limited to, a terminal and a server. Among them, the terminal includes a desktop terminal and a mobile terminal. The desktop terminal includes, but is not limited to, a desktop computer and a vehicle-mounted computer; the mobile terminal includes, but is not limited to, a mobile phone, a tablet, a laptop computer, and a smart watch. The server includes a high-performance computer and a high-performance computer cluster.

[0051] Figure 1 The dialogue interaction decision method for intelligent interruption in human-machine dialogue provided by the embodiments of the present application includes:

[0052] Step 100, during the process of communicating with the user, obtain the first utterance of the user.

[0053] During the process of communicating with the user, the intelligent dialogue robot obtains the first utterance of the user.

[0054] Among them, the first utterance refers to what the user says. The format of the first utterance obtained by the intelligent dialogue robot can be a voice format, a text format, or a video format. That is, the communication between the intelligent dialogue robot and the user can be a voice call communication (for example, making a phone call), a text communication (for example, sending a text message through an APP), or a video communication (for example, the user is deaf and mute, so the user communicates with the robot by signing). When the format of the first utterance is a voice format or a video format, it is necessary to convert the format of the first utterance to obtain the first utterance in text format for subsequent related processing.

[0055] The first utterance can be the sentence closest to the current moment, or multiple sentences replied by the user. During the process of communicating with the user, the intelligent dialogue robot may say multiple sentences. It can be understood that there will be a certain time interval between two sentences. For example, in an actual dialogue scenario, the intelligent dialogue robot has already said a sentence, and now it is the user's turn to reply to what the intelligent dialogue robot said. The user's reply contains multiple sentences. Then, obtain the sentence closest to the current moment from the multiple sentences, and then use the sentence closest to the current moment as the first utterance of the user. Of course, the sentences closest to the current moment among the multiple sentences can also be used as the first utterance of the user, or even the multiple sentences can be directly used as the first utterance of the user.

[0056] Step 200: input the user's first speech into an interruption model to obtain an interruption judgment result and a model output intention output by the interruption model.

[0057] Input the user's first speech into the interruption model; the interruption model divides the first speech into words to obtain multiple word segmentation results; the interruption model groups the first speech according to a preset length to obtain multiple grouping results; obtain the word vector of each word segmentation result in the multiple word segmentation results, and the word vector of each grouping result in the multiple grouping results; multiply the word vector of each word segmentation result in the multiple word segmentation results, and the word vector of each grouping result in the multiple grouping results with the first weight matrix respectively to obtain multiple multiplication result matrices; calculate the average of the multiple multiplication result matrices to obtain an average result matrix; multiply the average result matrix with the second weight matrix to obtain a target multiplication result; perform softmax processing on the target multiplication result to obtain the interruption judgment result and model output intention output by the interruption model.

[0058] For example Figure 2 As shown, assuming that the first utterance is "Don't say it", then the 4 word segmentation results are: "you", "don't", "say", "the", the preset length is 2, then the 5 grouping results are "<you", "don't you", "don't say", "said", "said>", where "<" indicates the beginning of the first utterance and ">" indicates the end of the first utterance, then the word vectors of "you", "don't", "say", "<you", "don't you", "don't say", "said", "said>" are obtained, the dimension of each word vector is assumed to be 1×N, the first weight matrix is ​​N×V, then each word vector is multiplied by the first weight matrix respectively, and 9 multiplication result matrices are obtained, each of which has a dimension of 1×V, and the 9 multiplication result matrices are added and averaged to obtain an average result matrix with a dimension of 1×V, the dimension of the second weight matrix is ​​V×N, so the average result matrix is ​​multiplied by the second weight matrix, and the dimension of the target multiplication result is 1×N. The target multiplication result is softmax processed to obtain the softmax result, which is also 1×N, that is, it corresponds to N probabilities, each probability corresponds to a model intent, and the maximum probability is obtained from the N probabilities. The model intent corresponding to the maximum probability is used as the model output intent. Assume that N=3, that is, there are 3 model intents, namely: understand moisturizing skin care products, contact tomorrow, and unwilling to continue communicating, and the probabilities are 0.01, 0.01, and 0.98 respectively. Therefore, the model output intent is: unwilling to continue communicating. Correspondingly, the interruption judgment result corresponding to the model output intent "unwilling to continue communicating" is: interrupt. If the model output intent is: understand moisturizing skin care products, then the interruption judgment result corresponding to the model output intent "understand moisturizing skin care products" is: do not interrupt. The corresponding relationship between the model output intent and the interruption judgment result is pre-set.

[0059] It should be noted that if the training process of the interruption model is interrupted, then after obtaining the softmax result, the cross-entropy function will be used to calculate the error between the softmax result and the annotation result, and then the parameters in the interruption model will be adjusted according to the calculated error until the error is less than the preset error, and the trained interruption model is obtained. Among them, the annotation result is the result corresponding to the first utterance in the training process. For example, the first utterance in the training process is also: Stop talking. So, the annotation result is [0 0 1], that is, the annotation result is: Not willing to continue the communication. The parameters in the interruption model may include, but are not limited to, the first weight matrix and the second weight matrix. For example, the parameters in the interruption model may also include: the word vectors of the word segmentation results and the grouping results, such as the word vectors of "you", "don't", "say", "already", "<you", "you don't", "don't say", "already said", "already>".

[0060] Step 300, determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model.

[0061] If the interruption judgment result is interruption, it is determined that the user has an interruption intention; if the interruption judgment result is non-interruption, it is determined that the user does not have an interruption intention.

[0062] Step 400, if the user has an interruption intention, use the intention output by the model as the target interruption intention, and obtain the corresponding coping strategy for the target interruption intention.

[0063] When the user has an interruption intention, use the intention output by the model, such as "Not willing to continue the communication", as the user's target interruption intention, and obtain the corresponding coping strategy for "Not willing to continue the communication". For example, the corresponding coping strategy for "Not willing to continue the communication" is: Pause communicating with the user.

[0064] Step 500, execute according to the coping strategy corresponding to the target interruption intention.

[0065] Continuing with the above example, since the corresponding coping strategy for "Not willing to continue the communication" is: Pause communicating with the user, so the intelligent dialogue robot will stop communicating with the user.

[0066] The above-mentioned dialogue interaction decision-making method for intelligent interruption in human-machine dialogue first obtains the first utterance of the user during the communication with the user; inputs the first utterance of the user into the interruption model to obtain the interruption judgment result and the model output intention output by the interruption model; determines whether the user has an interruption intention according to the interruption judgment result output by the interruption model; then, if the user has an interruption intention, takes the model output intention as the target interruption intention and obtains the corresponding coping strategy for the target interruption intention; finally, executes according to the corresponding coping strategy for the target interruption intention. It can be seen that when it is determined according to the user's first utterance that the user has an interruption intention, obtaining the corresponding coping strategy for the target interruption intention to respond enables the intelligent dialogue robot to respond well when the user has an interruption intention, thus solving the technical problem in the prior art that the intelligent dialogue robot does not know how to handle when the user wants to interrupt the dialogue, and further improving the user experience to a certain extent.

[0067] In one embodiment, after step 500 of executing according to the corresponding coping strategy for the target interruption intention, it further includes:

[0068] Step 600, obtaining the second utterance of the user.

[0069] After the intelligent dialogue robot executes according to the corresponding coping strategy for the target interruption intention, for example, the corresponding coping strategy for the target interruption intention is: the intelligent robot replies "Okay, please go ahead". After the intelligent robot replies "Okay, please go ahead", the user will then speak. For example, what the user says is, "I'm driving and it's not convenient. I'll contact you tomorrow." So, the second utterance of the user is: I'm driving and it's not convenient. I'll contact you tomorrow.

[0070] Step 700, determining the user intention of the user according to the first utterance and the second utterance of the user.

[0071] The first utterance and the second utterance of the user may both contain the user intention. Therefore, in order to more accurately determine the user intention, the user intention is determined according to the first utterance and the second utterance of the user. For example, the first utterance is: Stop stop stop, and the second utterance is: I'm driving and it's not convenient. I'll contact you tomorrow. So, the user intention is determined according to "Stop stop stop" and "I'm driving and it's not convenient. I'll contact you tomorrow", and the assumption is: Contact tomorrow; Another example is that the first utterance is: Shut up, and the second utterance is: I want to hang up the phone. So, the user intention is determined according to "Shut up" and "I want to hang up the phone", and the assumption is: Not convenient to contact; Another example is that the first utterance is: Stop and listen to me, and the second utterance is: I want to know which skin care products are moisturizing. So, the user intention is determined according to "Stop and listen to me" and "I want to know which skin care products are moisturizing", and the assumption is: Understand moisturizing skin care products.

[0072] Step 800: Obtain the corresponding coping strategy for the user intention.

[0073] Coping strategies corresponding to different user intentions are set in advance. For example, the coping strategy set for the user intention "Contact tomorrow" is: Reply: Okay, then I'll wait for you to contact me tomorrow; Another example, the coping strategy set for the user intention "Inconvenient to contact" is: Reply: Okay, then you're busy; Another example, the coping strategy set for the user intention "Learn about moisturizing skin care products" is: First reply: Please wait a moment, then obtain moisturizing skin care products, and finally provide feedback on the obtained moisturizing skin care products.

[0074] Step 900: Execute according to the coping strategy corresponding to the user intention.

[0075] Since the coping strategy corresponding to the user intention has been obtained, at this time, the intelligent dialogue robot can execute according to the coping strategy corresponding to the user intention.

[0076] In the above embodiment, after the intelligent dialogue robot executes according to the coping strategy corresponding to the target interruption intention, the intelligent dialogue robot also obtains the second utterance of the user, so as to determine the user intention of the user based on the first utterance and the second utterance, thereby improving the accuracy of user intention determination to a certain extent.

[0077] In one embodiment, step 700 of determining the user intention of the user according to the first utterance and the second utterance of the user includes:

[0078] Step 701: Extract target key phrases from the first utterance and the second utterance.

[0079] The target key phrase is a phrase extracted from the first utterance and the second utterance. For example, the first utterance is: Stop stop stop, and the second utterance is: I'm driving and it's inconvenient. I'll contact you tomorrow. So, the target key phrase extracted from the first utterance and the second utterance is: I'm driving; Another example, the first utterance is: Shut up, and the second utterance is: I'm going to hang up the phone. So, the target key phrase extracted from the first utterance and the second utterance is: I'm going to hang up the phone; Another example, the first utterance is: Stop and listen to me, and the second utterance is: I want to know which skin care products are moisturizing. So, the target key phrase extracted from the first utterance and the second utterance is: Which skin care products are moisturizing; Another example, the first utterance is: Stop talking and listen to me, and the second utterance is: I'm about to have a meeting soon and will contact you later. So, the target key phrase extracted from the first utterance and the second utterance is: I'm about to have a meeting soon.

[0080] Step 702: Input the target key phrase into the intention recognition model to obtain the user intention of the user output by the intention recognition model.

[0081] An intent recognition model is a model that can determine the user's intent based on the input target keyword sentence. Inputting the target keyword sentence into the intent recognition model can obtain the user's intent output by the intent recognition model.

[0082] In the above embodiment, the target keyword sentence is extracted from the first utterance and the second utterance, and inputting the target keyword sentence into the intent recognition model can obtain the user's intent.

[0083] In one embodiment, before step 100 of obtaining the first utterance of the user during the communication with the user, it further includes:

[0084] Training the intent recognition model.

[0085] Obtain training data. The training data includes multiple data groups. Each data group includes a keyword sentence and a user intent. Use the keyword sentence as the input of the intent recognition model, and use the user intent in the data group where the keyword sentence is located as the output of the intent recognition model to train the intent recognition model. For example, there are 3 data groups in the training data. Data group 1 is ["hang up the phone", "inconvenient to contact"], data group 2 is ["I'm driving", "inconvenient to contact"], and data group 3 is ["I'll contact you tomorrow", "contact tomorrow"]. Then, use "hang up the phone", "I'm driving", and "I'll contact you tomorrow" as the inputs of the intent recognition model respectively, and use "inconvenient to contact", "inconvenient to contact", and "contact tomorrow" as the corresponding outputs to train the intent recognition model.

[0086] The above embodiment illustrates the training process of the intent recognition model.

[0087] In one embodiment, step 702 of inputting the target keyword sentence into the intent recognition model to obtain the user's intent output by the intent recognition model includes:

[0088] Step 702A, input the target keyword sentence into the intent recognition model.

[0089] For example, the target keyword sentence is "I'm driving", so input "I'm driving" into the intent recognition model.

[0090] Step 702B, the intent recognition model combines the target keyword sentence with the target corpus in the corpus to obtain multiple combination results.

[0091] The target corpus can be all the corpora contained in the corpus, or it can be a part of the corpora. A corpus is a repository for storing corpora. There are multiple corpora stored in the corpus. For example, the corpora stored in the corpus are: I am driving, I am currently driving, I will drive later, I am going to eat, I want to eat, I am hungry, I am going to shop, I am going to work, I am inconvenient on the road, What are the moisturizing skin care products, I want to find a job. Then, the target keyword sentence is combined with the target corpus (assuming all the corpora) in the corpus to obtain multiple combination results, which are respectively: [I am driving, I am driving], [I am driving, I am currently driving], [I am driving, I will drive later], [I am driving, I am going to eat], [I am driving, I want to eat], [I am driving, I am hungry], [I am driving, I am going to shop], [I am driving, I am going to work], [I am driving, I am inconvenient on the road], [I am driving, What are the moisturizing skin care products], [I am driving, I want to find a job].

[0092] Step 702C, the intention recognition model calculates the similarity between the target keyword sentence and the target corpus in each of the combination results, and obtains the target combination result with the largest similarity.

[0093] A method for calculating similarity is provided, including: obtaining the word vectors of each character in the target keyword sentence; obtaining the word vector of the target keyword sentence according to the word vectors of each character in the target keyword sentence; obtaining the word vector of the target corpus; calculating the vector distance according to the word vector of the target keyword sentence and the word vector of the target corpus; performing a function transformation on the vector distance to obtain the similarity between the target keyword sentence and the target corpus. It can be understood that the larger the vector distance, the less similar the target keyword sentence is to the target corpus. On the contrary, the smaller the vector distance, the more similar the target keyword sentence is to the target corpus. Therefore, after obtaining the vector distance, a function transformation needs to be performed on the vector distance to obtain the similarity between the target keyword sentence and the target corpus.

[0094] Through the above method for calculating similarity, the similarity corresponding to each combination result can be obtained. Then, the maximum similarity is determined from the similarities corresponding to each combination result, and then the combination result corresponding to the maximum similarity is obtained, and the combination result corresponding to the maximum similarity is used as the target combination result.

[0095] Step 702D, taking the corpus intention corresponding to the target corpus in the target combination result as the user intention of the user.

[0096] The corpus intention corresponding to the target corpus in the target combination result is obtained, and then the obtained corpus intention is used as the user intention.

[0097] Set the corpus intention for each corpus in the corpus in advance. Continuing with the above example, the corpus intentions of the corpora "I am driving", "I am currently driving", and "I will drive later" are all: inconvenient to contact; the corpus intentions of the corpora "I am going to eat", "I want to eat", and "I am hungry" are all: want to eat; the corpus intention of the corpus "I am going shopping" is: the user is in a hurry to go to work; the corpus intention of the corpus "I am inconvenient on the road" is: inconvenient to contact; the corpus intention of the corpus "What are the moisturizing skin care products" is: understand moisturizing skin care products; the corpus intention of the corpus "I want to find a job" is: find a job.

[0098] The above embodiments illustrate how the intention recognition model obtains the user intention based on the target keyword sentence.

[0099] In one embodiment, before the intention recognition model combines the target keyword sentence with the target corpus in the corpus in step 702B to obtain multiple combination results, it further includes:

[0100] Step 702E, perform word segmentation on the target keyword sentence to obtain multiple word segmentation results.

[0101] For example, the target keyword sentence is "I am driving", so the multiple word segmentation results of this target keyword sentence are: "I", "am", "driving".

[0102] Step 702F, obtain the occurrence probability of each word segmentation result in the corpus.

[0103] Suppose the corpus stores the following corpora: "I like driving", "I am driving now", "I will go shopping", "I am looking for a job", "I am eating". Tokenize each corpus in the corpus to obtain the tokenization results of each corpus. The tokenization result of the corpus "I like driving" is: "I", "like", "driving". The tokenization result of the corpus "I am driving now" is: "I", "now", "am", "driving". The tokenization result of the corpus "I will go shopping" is "I", "will go", "shopping". The tokenization result of the corpus "I am looking for a job" is "I", "am", "looking for a job". The tokenization result of the corpus "I am eating" is: "I", "am", "eating". Thus, the occurrence probability of each tokenization result in the corpus is equal to the number of occurrences of each tokenization result divided by the total number of tokenization results corresponding to the corpus. For example, in this example, the number of occurrences of the tokenization result "I" is: 5. The determination of the total number of tokenization results corresponding to the corpus: Add up the number of tokenization results of each corpus to obtain the total number of tokenization results corresponding to the corpus. In this example, the total number of tokenization results corresponding to the corpus is: 3 + 4 + 3 + 3 + 3. Thus, the occurrence probability of the tokenization result "I" in the corpus is: 5 / (3 + 4 + 3 + 3 + 3), the occurrence probability of the tokenization result "am" is: 3 / (3 + 4 + 3 + 3 + 3), and the occurrence probability of the tokenization result "driving" is: 2 / (4 + 3 + 3 + 3). Since the tokenization results of the target keyword sentence are: "I", "am", "driving", thus, obtain the occurrence probabilities of the tokenization results "I", "am" and "driving" in the corpus, which are 5 / (3 + 4 + 3 + 3 + 3), 3 / (3 + 4 + 3 + 3 + 3) and 2 / (4 + 3 + 3 + 3) respectively.

[0104] Step 702G, obtain the target tokenization result from the multiple tokenization results according to the occurrence probability of each said tokenization result in the corpus.

[0105] Since the target keyword sentence has 3 tokenization results with occurrence probabilities of 5 / (3 + 4 + 3 + 3 + 3), 3 / (3 + 4 + 3 + 3 + 3) and 2 / (4 + 3 + 3 + 3) respectively, thus, obtain the minimum occurrence probability, that is, 2 / (4 + 3 + 3 + 3), and obtain the target tokenization result as: driving.

[0106] Step 702H, in the corpus, obtain multiple target corpora containing the target tokenization result.

[0107] The target tokenization result is: driving. Thus, obtain from the corpus the corpora containing the tokenization result "driving": "I like driving" and "I am driving now".

[0108] In the above embodiments, since there is a large amount of corpus in the corpus, if the target keyword sentences are combined with all the corpus and then the similarity is calculated, the processing efficiency of the intent recognition model will inevitably be low. Therefore, a part of the corpus is selected from the corpus as the target corpus, and then the target keyword sentences are combined with the target corpus, and then the similarity is calculated, which can improve the processing efficiency of the intent recognition model to a certain extent.

[0109] In one embodiment, as Figure 3 shown, a dialogue interaction decision device 300 for intelligent interruption in human-computer dialogue is provided, including:

[0110] An acquisition module 301, configured to acquire the first utterance of the user during the process of communicating with the user;

[0111] A judgment module 302, configured to input the first utterance of the user into the interruption model to obtain an interruption judgment result and a model output intent output by the interruption model;

[0112] A determination module 303, configured to determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model;

[0113] A policy module 304, configured to, if the user has an interruption intention, use the model output intent as the target interruption intent and obtain a coping strategy corresponding to the target interruption intent;

[0114] An execution module 305, configured to execute according to the coping strategy corresponding to the target interruption intent.

[0115] In one embodiment, the dialogue interaction decision device 300 for intelligent interruption in human-computer dialogue further includes: a second module, configured to:

[0116] Acquire the second utterance of the user;

[0117] Determine the user intent of the user according to the first utterance and the second utterance of the user;

[0118] Obtain a coping strategy corresponding to the user intent;

[0119] Execute according to the coping strategy corresponding to the user intent.

[0120] In one embodiment, the second module is specifically configured to:

[0121] Extract target keyword sentences from the first utterance and the second utterance;

[0122] Input the target keyword sentences into an intent recognition model to obtain the user intent of the user output by the intent recognition model.

[0123] In one embodiment, the dialogue interaction decision-making device 300 for intelligent interruption in human-computer dialogue further includes: a training module for:

[0124] Training the intention recognition model.

[0125] In one embodiment, the second module is specifically configured to:

[0126] Input the target keyword sentence into the intention recognition model;

[0127] The intention recognition model combines the target keyword sentence with the target corpus in the corpus to obtain multiple combination results;

[0128] The intention recognition model calculates the similarity between the target keyword sentence and the target corpus in each of the combination results to obtain the target combination result with the maximum similarity;

[0129] Take the corpus intention corresponding to the target corpus in the target combination result as the user intention of the user.

[0130] In one embodiment, the dialogue interaction decision-making device 300 for intelligent interruption in human-computer dialogue further includes: a target corpus module for:

[0131] Segment the target keyword sentence to obtain multiple segmentation results;

[0132] Obtain the occurrence probability of each segmentation result in the corpus;

[0133] Obtain the target segmentation result from the multiple segmentation results according to the occurrence probability of each segmentation result in the corpus;

[0134] In the corpus, obtain multiple target corpora containing the target segmentation result.

[0135] In one embodiment, such as Figure 4As shown, a computer device is provided, which may specifically be a terminal or a server. The computer device includes a processor, a memory, and a network interface connected via a system bus. The memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the dialogue interaction decision method for intelligent interruption in human-computer dialogue. The non-volatile memory may include a read-only memory (ROM), a programmable ROM (PROM), an electrically programmable ROM (EPROM), an electrically erasable programmable ROM (EEPROM), or a flash memory. The volatile memory may include a random access memory (RAM) or an external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can implement the dialogue interaction decision method for intelligent interruption in human-computer dialogue. Those skilled in the art can understand that Figure 4 The structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0136] The dialogue interaction decision method for intelligent interruption in human-computer dialogue provided by this application can be implemented in the form of a computer program, and the computer program can run on a computer device as shown in Figure 4 The memory of the computer device may store multiple program templates that make up the dialogue interaction decision device for intelligent interruption in human-computer dialogue. For example, an acquisition module 301, a judgment module 302, and a determination module 303.

[0137] A computer device includes a memory and a processor. The memory stores a computer program. When the computer program is executed by the processor, the processor performs the following steps:

[0138] During the process of communicating with the user, acquire the first utterance of the user;

[0139] Input the first utterance of the user into the interruption model to obtain the interruption judgment result and the model output intention output by the interruption model;

[0140] Determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model;

[0141] If the user has an interruption intention, use the model output intention as the target interruption intention and obtain the corresponding coping strategy for the target interruption intention;

[0142] Execute according to the coping strategy corresponding to the target interruption intention.

[0143] In one embodiment, a computer-readable storage medium is provided, storing a computer program, which when executed by a processor, causes the processor to perform the following steps: During the process of communicating with the user, obtain the first utterance of the user; Input the first utterance of the user into the interruption model to obtain the interruption judgment result and the model output intention output by the interruption model; Determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model; If the user has an interruption intention, use the model output intention as the target interruption intention and obtain the corresponding coping strategy for the target interruption intention;

[0144] Execute according to the coping strategy corresponding to the target interruption intention.

[0145] It should be noted that the above-mentioned intelligent interruption dialogue interaction decision-making method in human-computer dialogue, the intelligent interruption dialogue interaction decision-making device in human-computer dialogue, the computer device and the computer-readable storage medium belong to a general inventive concept, and the content in the embodiments of the intelligent interruption dialogue interaction decision-making method in human-computer dialogue, the intelligent interruption dialogue interaction decision-making device in human-computer dialogue, the computer device and the computer-readable storage medium can be mutually applicable.

[0146] In the embodiments provided in the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For another example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or units can be in electrical, mechanical or other forms. In addition, the units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Furthermore, in the various embodiments of the present application, the functional modules can be integrated together to form an independent part, or multiple modules can exist separately, or two or more modules can be integrated to form an independent part. In this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. The above are only the embodiments of the present application and are not used to limit the protection scope of the present application. For those skilled in the art, the present application can have various changes and modifications. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. A dialogue interaction decision-making method for intelligent interruption in human-computer dialogue, characterized in that, Including: During the process of communicating with the user, obtain the first utterance of the user; Input the first utterance of the user into the interruption model to obtain the interruption judgment result and the model output intention output by the interruption model; Determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model; If the user has an interruption intention, take the model output intention as the target interruption intention and obtain the corresponding coping strategy for the target interruption intention; Execute according to the coping strategy corresponding to the target interruption intention; After executing according to the coping strategy corresponding to the target interruption intention, it further includes: Obtain the second utterance of the user; Determine the user intention of the user according to the first utterance and the second utterance of the user; Obtain the coping strategy corresponding to the user intention; Execute according to the coping strategy corresponding to the user intention; Among them, inputting the first utterance of the user into the interruption model to obtain the interruption judgment result and the model output intention output by the interruption model includes: Input the first utterance of the user into the interruption model; The interruption model performs character segmentation on the first utterance to obtain multiple character segmentation results; The interruption model groups the first utterance according to a preset length to obtain multiple grouping results; Obtain the word vectors of each character segmentation result among the multiple character segmentation results, and the word vectors of each grouping result among the multiple grouping results; Multiply the word vectors of each character segmentation result among the multiple character segmentation results and the word vectors of each grouping result among the multiple grouping results by the first weight matrix respectively to obtain multiple multiplication result matrices; Calculate the average of the multiple multiplication result matrices to obtain an average result matrix; Multiply the average result matrix by the second weight matrix to obtain a target multiplication result; Perform softmax processing on the target multiplication result to obtain the interruption judgment result and the model output intention output by the interruption model.

2. The dialogue interaction decision-making method according to claim 1, wherein The determining the user intention of the user according to the first utterance and the second utterance of the user includes: Extract the target keyword sentences from the first utterance and the second utterance; Input the target keyword sentences into the intention recognition model to obtain the user intention of the user output by the intention recognition model.

3. The dialogue interaction decision-making method according to claim 2, wherein Before obtaining the first utterance of the user during the process of communicating with the user, it further includes: Train the intention recognition model.

4. The dialogue interaction decision-making method according to claim 2, wherein, Inputting the target keyword sentences into the intention recognition model to obtain the user intention of the user output by the intention recognition model includes: Input the target keyword sentences into the intention recognition model; The intention recognition model combines the target keyword sentences with the target corpus in the corpus to obtain multiple combination results; The intention recognition model calculates the similarity between the target keyword sentences and the target corpus in each combination result to obtain the target combination result with the largest similarity; Take the corpus intention corresponding to the target corpus in the target combination result as the user intention of the user.

5. The dialogue interaction decision-making method according to claim 4, wherein Before the intention recognition model combines the target keyword sentences with the target corpus in the corpus to obtain multiple combination results, it further includes: Segment the target keyword sentence to obtain multiple segmentation results; Obtain the occurrence probability of each of the segmentation results in the corpus; Obtain the target segmentation result from the multiple segmentation results according to the occurrence probability of each of the segmentation results in the corpus; In the corpus, obtain multiple target corpora containing the target segmentation result.

6. A dialogue interaction decision-making device for intelligent interruption in human-machine dialogue, characterized in that, It includes: An acquisition module, configured to acquire the first utterance of the user during the process of communicating with the user; A judgment module, configured to input the first utterance of the user into an interruption model to obtain an interruption judgment result and a model output intention output by the interruption model; A determination module, configured to determine whether the user has an interruption intention according to the interruption judgment result output by the interruption model; A strategy module, configured to, if the user has an interruption intention, use the model output intention as the target interruption intention, and obtain a corresponding coping strategy for the target interruption intention; An execution module, configured to execute according to the coping strategy corresponding to the target interruption intention; A second module, configured to: acquire the second utterance of the user; determine the user intention of the user according to the first utterance and the second utterance of the user; acquire the coping strategy corresponding to the user intention; execute according to the coping strategy corresponding to the user intention; Wherein, the judgment module is specifically configured to: input the first utterance of the user into the interruption model; the interruption model performs character segmentation on the first utterance to obtain multiple character segmentation results; the interruption model groups the first utterance according to a preset length to obtain multiple grouping results; obtain the word vector of each character segmentation result among the multiple character segmentation results, and the word vector of each grouping result among the multiple grouping results; multiply the word vector of each character segmentation result among the multiple character segmentation results, and the word vector of each grouping result among the multiple grouping results, respectively, with a first weight matrix to obtain multiple multiplication result matrices; calculate the average of the multiple multiplication result matrices to obtain an average result matrix; multiply the average result matrix with a second weight matrix to obtain a target multiplication result; perform softmax processing on the target multiplication result to obtain the interruption judgment result and the model output intention output by the interruption model.

7. A computer device, characterized in that, It includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of the intelligent interruption dialogue interaction decision method in the human-computer dialogue according to any one of claims 1 to 5 are implemented.

8. A computer-readable storage medium, characterized in that, Computer program instructions are stored in the computer-readable storage medium. When the computer program instructions are read and run by the processor, the steps of the intelligent interruption dialogue interaction decision method in the human-computer dialogue according to any one of claims 1 to 5 are executed.

Citation Information

Patent Citations

  • Voice processing method and device based on human-computer interaction, equipment and storage medium

    CN111970409A

  • Voice interrupt processing method and device, computer equipment and storage medium

    CN112037799A

  • Man-machine dialogue control method and device, computer equipment and storage medium

    CN112669842A