Human-computer interaction method, intelligent robot and storage medium

By obtaining the user's voice signal and text information, using multimodal data to generate fusion feature vectors to identify whether the semantics are complete, the problem of low semantic recognition accuracy of intelligent robots is solved and the fluency of human-computer interaction is achieved.

CN114936560BActive Publication Date: 2025-08-29ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210375013.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2025-08-29
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

The existing intelligent robots have low semantic recognition accuracy, which leads to failure of human-computer interaction.

Method used

By obtaining the user's voice signal and text information, using multimodal data for feature extraction, generating a fused feature vector, identifying whether the semantics are complete, and responding according to the classification results.

Benefits of technology

It improves the accuracy of semantic recognition, reduces sentence breaking errors, and ensures the fluency of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114936560B_ABST
    Figure CN114936560B_ABST
Patent Text Reader

Abstract

An embodiment of the present invention provides a human-computer interaction method, an intelligent robot, and a storage medium. The method includes: obtaining a first voice signal generated by a user and a first text message corresponding to the first voice signal. Then, a fused feature vector is obtained based on the respective feature vectors of the first voice signal and the first text message. A classification result reflecting whether the semantics of the first voice signal are complete is determined based on the fused feature vector, and the first voice signal is responded to based on the classification result. Among them, the feature vector of the first voice signal reflects the user's speaking state; the feature vector of the first text message reflects the user's semantics, and the fused feature vector will contain the above-mentioned speaking state and semantics at the same time. Therefore, it can improve the accuracy of identifying whether the semantics are complete, that is, improve the accuracy of the intelligent robot's sentence segmentation, reduce the occurrence of failure to respond to the first voice signal due to sentence segmentation errors, and ensure the fluency of human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a human-computer interaction method, an intelligent robot, and a storage medium. Background Art

[0002] With the development of artificial intelligence technology, various intelligent robots are increasingly entering people's lives, such as service robots, cleaning robots, and self-propelled vending robots. In addition to the aforementioned robots, intelligent voice robots for customer service scenarios have emerged in recent years, such as intelligent outbound call robots and intelligent customer service robots.

[0003] Each of the aforementioned intelligent robots with voice interaction capabilities can achieve interaction by collecting user-generated voice signals, performing semantic recognition on them, and then outputting responses based on the recognition results. Specifically, after collecting user-generated voice signals, the intelligent robots can determine whether they are semantically complete and perform sentence segmentation on semantically complete voice signals, treating the collected voice signals as semantically complete, performing semantic recognition on them, and ultimately outputting responses that align with the semantic recognition results.

[0004] However, in practice, intelligent robots' accuracy in identifying semantic completeness is not high. As a result, they may output incorrect responses or even fail to output responses, leading to failed human-computer interactions. Based on the above description, how to improve the accuracy of semantic completeness recognition to ensure smooth human-computer interactions has become an urgent issue. Summary of the Invention

[0005] In view of this, embodiments of the present invention provide a human-computer interaction method, an intelligent robot, and a storage medium for accurately identifying whether semantics are complete, thereby ensuring the fluency of human-computer interaction.

[0006] In a first aspect, an embodiment of the present invention provides a human-computer interaction method, comprising:

[0007] Acquire a first voice signal generated by a user and first text information corresponding to the first voice signal;

[0008] Determining a fusion feature vector based on the feature vectors of the first speech signal and the first text information;

[0009] determining, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete;

[0010] Respond to the first speech signal according to the classification result.

[0011] In a second aspect, an embodiment of the present invention provides an intelligent robot comprising a processor and a memory, wherein the memory is configured to store one or more computer instructions, wherein the one or more computer instructions, when executed by the processor, implement the human-computer interaction method described in the first aspect. The electronic device may further include a communication interface for communicating with other devices or a communication network.

[0012] In a third aspect, an embodiment of the present invention provides a non-temporary machine-readable storage medium, on which executable code is stored. When the executable code is executed by a processor of an electronic device, the processor can at least implement the human-computer interaction method described in the first aspect.

[0013] The human-computer interaction method provided by an embodiment of the present invention obtains a first voice signal generated by a user and first text information corresponding to the first voice signal. Feature extraction is then performed on the first voice signal and the first text information, and a fused feature vector is generated from the feature vectors of the two features. A classification result is determined based on the fused feature vector, reflecting whether the semantics of the first voice signal are complete, and a response is provided to the first voice signal based on the classification result.

[0014] It can be seen that in the above process, voice signals and text information are used simultaneously to identify whether the semantics of the voice signal are complete, that is, multimodal data is used to identify whether the semantics are complete. And because the feature vector of the first voice signal can reflect the speaking state of the user who generates the first voice signal, such as speaking speed and intonation, etc.; the feature vector of the first text information can reflect the semantics of the first voice signal, therefore, the fused feature vector obtained in the above manner also includes the user's speaking state and semantics. The intelligent robot can use multimodal data to identify whether the semantics are complete from multiple angles (i.e., speaking state and voice), thereby improving the accuracy of recognition, that is, improving the accuracy of the intelligent robot's punctuation, reducing the occurrence of failure to respond to the first voice signal due to punctuation errors, and ensuring the fluency of human-computer interaction. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0016] Figure 1 A flowchart of a human-computer interaction method provided by an embodiment of the present invention;

[0017] Figure 2A flowchart of another human-computer interaction method provided by an embodiment of the present invention;

[0018] Figure 3 A flowchart of another human-computer interaction method provided by an embodiment of the present invention;

[0019] Figure 4 A flowchart of another human-computer interaction method provided by an embodiment of the present invention;

[0020] Figure 5 A schematic diagram of the human-computer interaction method provided by an embodiment of the present invention applied in a customer service scenario;

[0021] Figure 6 A schematic diagram of the human-computer interaction method provided by an embodiment of the present invention applied in an outbound call scenario;

[0022] Figure 7 A schematic structural diagram of an intelligent robot provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more clear, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.

[0024] The terms used in the embodiments of the present invention are for the purpose of describing specific embodiments only and are not intended to limit the present invention. The singular forms "a," "the," and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms, unless the context clearly indicates otherwise. "A plurality" generally includes at least two, but does not exclude the inclusion of at least one.

[0025] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0026] As used herein, the words “if” and “if” may be interpreted as “at the time of” or “when” or “in response to determining” or “in response to identifying,” depending on the context. Similarly, the phrases “if it is determined” or “if (stated condition or event) is identified” may be interpreted as “when it is determined” or “in response to determining” or “when identifying (stated condition or event)” or “in response to identifying (stated condition or event),” depending on the context.

[0027] It should also be noted that the terms "include," "comprises," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a product or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such product or system. In the absence of further limitations, an element defined by the phrase "comprises a..." does not exclude the presence of other identical elements in the product or system comprising the element.

[0028] Before describing the human-computer interaction methods provided by various embodiments of the present invention, some possible human-computer interaction scenarios may be briefly described:

[0029] As mentioned in the background technology, common intelligent robots may include service robots, cleaning robots, self-moving vending robots, intelligent voice robots, etc. These robots collect voice signals generated by users and output corresponding response voice signals to users based on the semantics of the voice signals.

[0030] For example, a service robot may be a guide robot in a shopping mall lobby, which may proactively say to the user in front of the robot: "How may I help you?" The user may then say: "Which floor is store A on?" The guide robot, after determining that the voice signal generated by the user is semantically complete, recognizes the semantics of the voice signal and ultimately outputs a response voice signal: "Store A is on the 2nd floor."

[0031] Another example is an intelligent outbound call robot, capable of proactively making calls to users, such as for typical service follow-up calls, payment collection calls, and schedule reminders. The intelligent outbound call robot can respond to voice signals it collects from users. For example, when a user answers a service follow-up call, the intelligent outbound call robot might ask, "Are you satisfied with Product A you purchased?" The user can respond, "I'm very satisfied with Product A." The intelligent outbound call robot then recognizes the semantic integrity of the user's voice signal and outputs a response voice signal to the user, "OK, thank you for your support."

[0032] Another example is an intelligent customer service robot, capable of answering user-initiated calls, such as inquiries about related matters. The intelligent customer service robot can also perform semantic recognition of user-generated voice signals to answer questions. For example, a user can initiate a government consultation call, and the intelligent customer service robot answers the call, saying, "This is the Provident Fund Service Platform. How may I help you?" The user then generates a voice signal, "I would like to inquire about that Provident Fund." After recognizing the complete semantics of the user's voice signal, the intelligent customer service robot can then respond with a voice message, "Please enter your ID number."

[0033] As can be seen, in each of the above scenarios, after collecting the audio signal generated by the user, the intelligent robot needs to determine whether its semantics are complete and respond to the voice signal only after confirming the semantics are complete. In this case, the human-computer interaction method provided by the various embodiments of the present invention can be used to more accurately identify whether the semantics are complete, thereby ensuring the smooth flow of human-computer interaction.

[0034] In addition, the usage scenarios of the embodiments provided by the present invention are not limited to the above scenarios, and any scenario that requires judgment on whether semantics is complete can use the human-computer interaction method provided by the embodiments of the present invention.

[0035] Based on the above description, some embodiments of the present invention are described in detail below with reference to the accompanying drawings. In the absence of conflicts between the embodiments, the following embodiments and the features in the embodiments may be combined with each other. In addition, the step timings in the following method embodiments are only examples and are not strictly limiting.

[0036] Figure 1 This is a flow chart of a human-computer interaction method provided by an embodiment of the present invention. The human-computer interaction method provided by an embodiment of the present invention can be executed by an intelligent robot with voice interaction function. Figure 1 As shown, the method may include the following steps:

[0037] S101: Acquire a first voice signal generated by a user and first text information corresponding to the first voice signal.

[0038] A user can generate a first voice signal to the intelligent robot. Optionally, the intelligent robot can use Voice Activity Detection (VAD) technology to determine whether the user is speaking, thereby collecting the first voice signal generated by the user. The received first voice signal is then converted into text information, thereby obtaining a first text message corresponding to the first voice signal.

[0039] S102: Determine a fusion feature vector based on the feature vectors of the first speech signal and the first text information.

[0040] Next, the intelligent robot extracts features from the first voice signal and the first text information to obtain their respective feature vectors. The feature vector of the first voice signal can be expressed as a u , the feature vector of the first text information can be expressed as t u According to the aforementioned eigenvector a u and the eigenvector t u A fused feature vector can be obtained, and the fused feature vector can be expressed as v. Optionally, the fused feature vector v can be determined by directly concatenating the feature vectors or by linear fusion, and the linear fusion can be directly adding or subtracting corresponding elements in the feature vectors.

[0041] Among them, the feature vector a of the first speech signal u It can reflect the user's speaking speed and stress when generating the voice signal, as well as the user's mood, age, etc., that is, the feature vector a of the first voice signal u It can reflect the user's paralinguistic information. The feature vector t of the first text information u It can reflect the semantics of the first speech signal, that is, it can reflect the user's linguistic information. Therefore, the fused feature vector v obtained above also contains the user's linguistic information and paralinguistic information. The fused feature vector v contains a variety of information for identifying whether the semantics are complete.

[0042] S103: Determine, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete.

[0043] Furthermore, the intelligent robot can determine a classification result reflecting whether the semantics of the first speech signal are complete based on the fused feature vector.

[0044] When the feature vector a of the first speech signal u When the user's current speaking speed is slow and the user is older, the user may not be sure about the semantics he wants to express, which greatly reduces the possibility of the user expressing complete semantics. When the fused feature vector v containing the feature vector of the first text information is used for classification, the possibility of the classification result being incomplete semantics will greatly increase; on the contrary, when the user's speaking speed is fast and the user is younger, the user is very clear about the semantics he wants to express. When the fused feature vector v containing the feature vector corresponding to the first text information is used for classification, the possibility of the classification result being complete semantics will greatly increase.

[0045] It should be noted that, in the process of determining the classification result based on the fused feature vector v, the intelligent robot determines whether the semantics of the first voice signal are complete, but does not perform semantic recognition on the first voice signal, that is, it does not know the specific semantics of the first voice signal.

[0046] S104: Respond to the first voice signal according to the classification result.

[0047] Ultimately, the intelligent robot can further adopt different response methods to the first voice signal based on the classification results.

[0048] Specifically, if the semantics of the first voice signal are complete, the intelligent robot can perform semantic recognition on the first voice signal and output a successful response voice signal corresponding to the semantic recognition result. Alternatively, after obtaining the semantic recognition result, the intelligent robot can determine the response text information corresponding to the recognition result from a locally stored preset question and answer set, and broadcast this response text information to form a response voice signal corresponding to the first voice signal. Alternatively, the intelligent robot can generate the response text information in real time based on the recognition result using its own configured sentence generation model to form a response voice signal.

[0049] If the semantics of the first voice signal are incomplete, the intelligent robot may output a preset answer failure voice signal, which may be, for example: "Sorry, I didn't hear clearly, please repeat it."

[0050] In this embodiment, a first voice signal generated by a user and first text information corresponding to the first voice signal are obtained. Feature extraction is then performed on each of the first voice signal and the first text information, and a fused feature vector is generated from the feature vectors of the two features. A classification result is determined based on the fused feature vector, reflecting whether the semantics of the first voice signal are complete, and a response is made to the first voice signal based on the classification result.

[0051] It can be seen that in the above process, voice signals and text information are used simultaneously to identify whether the semantics of the voice signal are complete, that is, multimodal data is used to identify whether the semantics are complete. And because the feature vector of the first voice signal can reflect the speaking state of the user who generates the first voice signal, such as speaking speed and intonation, etc.; the feature vector of the first text information can reflect the semantics of the first voice signal, therefore, the fused feature vector obtained in the above manner also includes the user's speaking state and semantics. The intelligent robot can use multimodal data to identify whether the semantics are complete from multiple angles, thereby improving the accuracy of recognition, that is, improving the accuracy of the intelligent robot's punctuation, reducing the occurrence of failure to respond to the first voice signal due to punctuation errors, and ensuring the fluency of human-computer interaction.

[0052] It should be noted that the above steps S102 to S103 can be specifically performed by a classification model configured in the intelligent robot, that is, the first voice signal and the first text information are used as inputs of the classification model, so that the classification model extracts and fuses the feature vectors of the two, and classifies them according to the fused feature vector v. The feature vectors of the first voice signal and the first text signal can be linearly fused according to the following formula: v = a u W0t u +b0. Where W0 and b0 are model parameters in the classification model.

[0053] Then in Figure 1 In the human-computer interaction method shown, optionally, the feature vector a of the first speech signal u It can be extracted by the first sub-model in the classification model, which can be any one of the Convolutional Neural Networks (CNN) model, the Bidirectional Encoder Representation from Transformers (BERT) model, and the Recurrent Neural Network (RNN) model. The feature vector t of the first text information u It can be extracted by the second sub-model in the classification model, and the second sub-model can be any one of the RNN model based on the gated recurrent unit (GRU) and the long short-term memory neural network (LSTM) model.

[0054] The training process for the classification model deployed in an intelligent robot can be described as follows: voice signals generated by the user or the intelligent robot are collected and converted into text. The voice signals are also annotated to determine whether the voice signals are semantically complete. The collected voice signals and text information serve as training samples, and the annotated voice signals serve as supervisory information to train the classification model. During training, backpropagation and gradient descent algorithms can be used to adjust model parameters until the classification model converges.

[0055] Regarding the speech signals used in training classification models, in practice, it is less likely that intelligent robots will generate semantically incomplete speech signals. Therefore, the speech signals generated by intelligent robots are usually used as positive samples for training models; semantically incomplete speech signals are often generated by users and can be used as negative samples for training models.

[0056] According to the description in the above embodiment, the intelligent robot can use VAD technology to collect the voice signal generated by the user. In practice, a common situation is that due to personal habits or uncertainty about the semantics of what the user wants to express, the user may pause for a long time in the process of generating a voice signal with complete semantics. When the pause time reaches the preset silence duration, the intelligent voice robot will mistakenly believe that the user has stopped speaking and will perform sentence segmentation on the voice signal collected before the preset silence duration is reached. That is, the voice signal collected before the pause is treated as a semantically complete voice signal for semantic recognition. However, since the semantics of this voice signal are incomplete, the response voice signal output by the intelligent robot may not be the answer the user wants, and may even output a failed response voice signal, resulting in failure of human-computer interaction.

[0057] To improve this situation, if the classification result of the first voice signal is determined to be semantically incomplete, the preset silence period can be appropriately extended, that is, the intelligent robot waits for a preset period of time, and further determines the response voice signal corresponding to the first voice signal based on whether the user generates a new voice signal within this preset period of time. For convenience of description, the new voice signal generated by the user can be referred to as the third voice signal.

[0058] In one case, if the classification result of the first voice signal is semantically incomplete and the user generates a third voice signal within a preset time, indicating that the user has generated a new voice signal, namely the third voice signal, after a long pause, and the first voice signal and the third voice signal are highly semantically related, then the intelligent robot can splice the first voice signal and the third voice signal to obtain a spliced ​​voice signal. At this time, the preset silence time is restored. When the user does not generate a new voice signal within the restored preset silence time, the intelligent robot determines whether the semantics of the spliced ​​voice signal are complete. If the classification result of the spliced ​​voice signal is semantically complete, the intelligent robot performs semantic recognition on the spliced ​​voice signal and outputs the corresponding response content, namely the successful response voice signal, based on the semantic recognition result. If the classification result of the spliced ​​voice signal is semantically incomplete, the intelligent robot outputs a failed response statement.

[0059] In the above scenario, if the user pauses once while expressing a complete meaning, the intelligent robot will need to perform a splicing operation. In reality, users may also pause multiple times while expressing a complete meaning. In this case, the intelligent robot will splice the voice signals generated by these multiple pauses and respond to the resulting spliced ​​voice signal.

[0060] In another case, if the classification result of the first voice signal is semantically incomplete and the user does not generate a third voice signal within the preset time length, that is, the user only generates the first voice signal within the preset silence time length + the preset time length, and the intelligent robot determines that the semantics of this first voice signal is incomplete, then the intelligent robot outputs a failed response voice signal corresponding to the first voice signal.

[0061] Optionally, when a long pause occurs after the user generates a semantically incomplete first voice signal, the intelligent robot may output a guiding audio signal to the user, such as "Yes, OK, please continue," for a period of time equal to the preset silence duration plus the preset duration, to guide the user to complete the semantics, i.e., to guide the user to generate a third voice signal. Optionally, the guiding audio signal may be pre-set.

[0062] Optionally, in Figure 1 In the illustrated embodiment, the intelligent robot can determine whether the semantics of the first voice signal generated by the user are complete based on the fusion feature vector v. Optionally, on this basis, the first text information can also be used to determine whether the semantics of the first voice signal are complete. Specifically, if the intelligent robot recognizes that the word at the preset position in the first text information is a preset word, it can be considered that the semantics of the first voice signal are incomplete. For example, assuming that the last word in the first text information is a preset word such as "that", "and", "I also want", etc., which obviously indicates incomplete semantics, the intelligent robot can determine that the semantics of the first voice signal are incomplete.

[0063] When the intelligent robot determines that the semantics are incomplete based on the fused feature vector v or the first text information, it can continue to determine how to respond to the first voice signal based on whether the user generates a third voice signal in the above manner.

[0064] exist Figure 1 On the basis of the embodiment shown, in order to further improve the accuracy of identifying whether the semantics of the first speech signal is complete, Figure 2 Flowchart of another human-computer interaction method provided by an embodiment of the present invention. Figure 2 As shown, the method may include the following steps:

[0065] S201: Acquire a first voice signal generated by a user and first text information corresponding to the first voice signal.

[0066] The execution process of step S201 can be found in Figure 1 The relevant descriptions in the embodiments are not repeated here.

[0067] S202: Acquire second text information corresponding to a second voice signal generated by the intelligent robot, where the second voice signal is generated before the first voice signal.

[0068] As can be seen from the above description, a user can respond to a voice signal generated by an intelligent robot to generate a first voice signal. For simplicity, the voice signal generated by the intelligent robot before the first voice signal can be referred to as a second voice signal, and the two voice signals are semantically related. The intelligent robot can also convert the second voice signal generated by the intelligent robot into text information to obtain second text information corresponding to the second voice signal.

[0069] In practice, before the first voice signal is generated, the intelligent robot may generate multiple voice signals. In order to ensure the closest semantic correlation with the first voice signal, among the multiple voice signals generated by the intelligent robot, the voice signal whose generation time is closest to that of the first voice signal may be determined as the second voice signal.

[0070] S203: Determine a fused text feature vector based on the feature vectors of the first text information and the second text information.

[0071] Next, the feature vectors of the first and second text information can be obtained separately and a fused text feature vector can be determined based on the feature vectors. Furthermore, since the first and second text information are semantically related, the fused text feature vector contains the semantics of the first speech signal as well as the contextual information between the first and second text information.

[0072] For determining the fusion text feature vector, it is optional to directly concatenate the feature vectors or to perform linear fusion, such as directly adding or subtracting corresponding elements in the feature vectors. The feature vector of the first text information can be expressed as t u , the feature vector of the second text information can be expressed as t h , then the eigenvector t u and the eigenvector t h Fusion is performed to obtain the fused text feature vector t uh .

[0073] S204: Determine a fused feature vector according to the fused text feature vector and the feature vector of the first speech signal.

[0074] Furthermore, according to the fusion text feature vector t uh and the feature vector a of the first speech signal u , determine the fusion feature vector v. Since the feature vector a u Can reflect the user's speaking status and integrate text feature vector t uhIt can reflect the user's semantic and contextual information. Therefore, the fused feature vector v also includes the above-mentioned speech state, semantics, and contextual information. Optionally, the fused feature vector v can be determined by directly concatenating the feature vectors or linearly fusion, such as directly adding or subtracting corresponding elements in the feature vectors.

[0075] S205 : Determine, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete.

[0076] The execution process of step S205 can be found in Figure 1 The relevant descriptions in the embodiment are not repeated here. However, it should be noted that since the fused feature vector v contains the context information between the first text information and the second text information, when there is an omission of a component in the first speech signal generated by the user, the intelligent robot can also combine the fused text feature vector t uh The context information in the first speech signal is used to determine whether the semantics are complete, that is, the omission of components in the first speech signal does not affect the determination of whether the semantics are complete.

[0077] S206: Respond to the first voice signal according to the classification result.

[0078] The execution process of step S206 can be found in Figure 1 The relevant descriptions in the embodiments are not repeated here.

[0079] In this embodiment, a second speech signal that is semantically closely related to the first speech signal is first obtained. Then, a fused text feature vector t is determined based on the feature vectors of the text information corresponding to the two speech signals. uh Then, according to the fusion text feature vector t uh and the feature vector a of the first speech signal u The final fused feature vector v is obtained. This fused feature vector v not only includes the contextual information between the two voice signals, but also includes the user's speaking state and semantics. Therefore, it can more accurately determine the semantic completeness of the first voice signal, thereby improving the intelligent robot's sentence segmentation accuracy and ensuring smooth human-computer interaction. Furthermore, thanks to the contextual information contained in the fused feature vector, omitted components in the first voice signal do not affect the determination of semantic completeness, thus ensuring smooth human-computer interaction.

[0080] It should be noted that the above steps S203 to S205 can be specifically performed by a classification model configured in the intelligent robot, that is, the first voice signal, the first text information, and the second text information are used as inputs of the classification model, so that the classification model extracts and fuses feature vectors and performs classification based on the feature vectors. Optionally, the feature vectors of the first text information and the second text information can also be linearly fused according to the following formula: Fusion text feature vector t uh and the feature vector a of the first speech signal u Linear fusion can also be performed according to the following formula: v = a u W0t uh +b0. Where W1, b1, W0, and b0 are model parameters in the classification model.

[0081] Then in Figure 2 In the illustrated human-computer interaction method, the training process for the classification model configured in the intelligent robot can optionally be described as follows: Voice signals generated by humans or machines with question-and-answer relationships are pre-collected and converted into text. The user's voice signals are then annotated to determine whether the voice signals are semantically complete. The collected voice signals and text information with question-and-answer relationships are then used as training samples, and the annotated content of the voice signals is used as supervision information to train the classification model until the classification model converges.

[0082] exist Figure 2 On the basis of the embodiment shown, in order to further improve the accuracy of identifying whether the semantics are complete, optionally, Figure 3 Flowchart of another human-computer interaction method provided by an embodiment of the present invention. Figure 3 As shown, the method may include the following steps:

[0083] S301: Acquire a first voice signal generated by a user and first text information corresponding to the first voice signal.

[0084] S302, obtaining second text information corresponding to a second voice signal generated by the intelligent robot, where the second voice signal is generated before the first voice signal.

[0085] S303: Determine a fused text feature vector based on the feature vectors of the first text information and the second text information.

[0086] The execution process of steps S301 to S303 can refer to the relevant description in the above embodiment and will not be repeated here.

[0087] S304: Adjust the information amount of the feature vector of the first speech signal according to the fused text feature vector to obtain a first adjustment result.

[0088] According to the fusion text feature vector t uh , adjust the feature vector a of the first speech signal u The amount of information to obtain the first adjustment result. The above adjustment process also realizes the interaction between the data of the voice signal mode and the data of the text information mode. The feature vector a of the first voice signal can be filtered out through the modal interaction. uThe less important information in the feature vector a is reduced u While increasing the amount of information, it will not cause the loss of important information.

[0089] Optionally, the amount of information can be adjusted in the following manner: a' u =a u σ(t uh ).

[0090] Among them, a' u is the first adjustment result, σ(t uh ) is the preset Sigmoid function, which is used to transform the feature vector t uh Normalize the element values ​​in .

[0091] S305: Adjust the information amount of the fused text feature vector according to the feature vector of the first speech signal to obtain a second adjustment result.

[0092] Similarly, the feature vector a of the first speech signal can be used to u , adjust the fusion text feature vector t uh Similarly, the fusion text feature vector t can be filtered out through modal interaction. uh The less important information in the whole fusion text feature vector t uh While increasing the amount of information, it will not cause the loss of important information.

[0093] Optionally, the amount of information can be adjusted as follows: uh =t uh σ(a u ).

[0094] Among them, t' uh is the second adjustment result, σ(a u ) is the preset Sigmoid function, which is used to transform the feature vector a u Normalize the element values ​​in .

[0095] S306: Determine a fusion feature vector according to the first adjustment result and the second adjustment result.

[0096] Furthermore, according to the first adjustment result t' obtained above uh and the second adjustment result a' u Determine the fused feature vector v. Optionally, the determination of the fused feature vector v may specifically be the first adjustment result t' uh and the second adjustment result a' u Direct concatenation can also be linear fusion, such as direct addition or subtraction of corresponding elements in the feature vector.

[0097] S307: Determine, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete.

[0098] S308: Respond to the first voice signal according to the classification result.

[0099] The execution process of steps S307 to S308 can refer to the relevant description in the above embodiment and will not be repeated here.

[0100] In this embodiment, Figure 2 Based on the embodiment shown, the intelligent robot first obtains the fusion text feature vector t uh Then the fusion text feature vector t uh and the feature vector a of the first speech signal u Modal interaction is performed to adjust the information content of the two feature vectors. The information adjustment results are then fused to obtain the final fused feature vector v. The fused feature vector v obtained after fusion and modal interaction includes not only the contextual information between the two voice signals, but also the user's speaking status and semantics. While reducing the information content of the feature vector, it will not cause the loss of important information. Therefore, the use of the fused feature vector v with the above characteristics can more accurately identify whether the semantics of the first voice signal are complete, that is, to improve the accuracy of the intelligent robot's punctuation, reduce the occurrence of failure to respond to the first voice signal due to punctuation errors, and ensure the fluency of human-computer interaction.

[0101] It should be noted that the above steps S303 to S307 can also be executed by the classification model configured in the intelligent robot. Figure 2 The embodiment shown is the same, that is, the first speech signal, the first text information, and the second text information are used as inputs of the classification model, so that the classification model extracts and fuses feature vectors, and performs classification based on the feature vectors. Optionally, the feature vectors of the first text information and the second text information can also be linearly fused according to the following formula: The first adjustment result a' u and the second adjustment result t' uh Linear fusion can also be performed according to the following formula: v = a' u W0t' uh +b0. Where W1, b1, W0, and b0 are model parameters in the classification model.

[0102] exist Figure 1 On the basis of the embodiment shown, in order to further improve the accuracy of identifying whether the semantics are complete, Figure 4 Flowchart of another human-computer interaction method provided by an embodiment of the present invention. Figure 4 As shown, the method may include the following steps:

[0103] S401: Acquire a first voice signal generated by a user and first text information corresponding to the first voice signal.

[0104] The execution process of step S401 can refer to the relevant description in the above embodiment and will not be repeated here.

[0105] S402: Adjust the information amount of the feature vector of the first speech signal according to the feature vector of the first text information to obtain a third adjustment result.

[0106] The intelligent robot can perform modal interaction on the obtained multimodal data, namely the first voice signal and the first text information. The feature vector a of the first voice signal can be filtered out through modal interaction. u The less important information in the feature vector a is reduced u While increasing the amount of information, it will not cause the loss of important information.

[0107] Specifically, the feature vector t of the first text information can be u , adjust the feature vector a of the first speech signal u The amount of information is adjusted to obtain the third adjustment result. Similarly, the amount of information can be adjusted in the following manner: a” u =a u σ(t u ).

[0108] Among them, a” u is the third adjustment result, σ(t u ) is the preset Sigmoid function, which is used to transform the feature vector t u Normalize the element values ​​in .

[0109] S403: Adjust the information amount of the feature vector of the first text information according to the feature vector of the first speech signal to obtain a fourth adjustment result.

[0110] Similar to step S402, the feature vector a of the first speech signal may be used to generate the first speech signal. u , adjust the feature vector t of the first text information u The amount of information is used to obtain the fourth adjustment result. The eigenvector t can be filtered out through modal interaction. u The less important information in the whole feature vector t u While increasing the amount of information, it will not cause the loss of important information.

[0111] Optionally, the amount of information can be adjusted as follows: u =t u σ(a u ).

[0112] Among them, t” u is the fourth adjustment result, σ(a u ) is the preset Sigmoid function, which is used to transform the feature vector a u Normalize the element values ​​in .

[0113] S404: Determine a fusion feature vector according to the third adjustment result and the fourth adjustment result.

[0114] Further, according to the third adjustment result a" u and the fourth adjustment result t" u Determine the fused feature vector v. Optionally, the fused feature vector v may be specifically determined as the third adjustment result a” u and the fourth adjustment result t" u Direct concatenation can also be linear fusion, such as direct addition or subtraction of corresponding elements in the feature vector.

[0115] S405 : Determine, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete.

[0116] S406: Respond to the first voice signal according to the classification result.

[0117] The execution process of steps S405 to S406 can refer to the relevant description in the above embodiment and will not be repeated here.

[0118] In this embodiment, after obtaining multimodal data, namely the first voice signal and the first text information, the intelligent robot can also perform modal fusion on the feature vectors of each of the multimodal data to obtain an adjustment result, and generate a fused feature vector v based on the adjustment result. The fused feature vector v includes the user's speaking state and semantics, and through modal interaction, while reducing the amount of feature vector information, it will not cause the loss of important information. Therefore, using the fused feature vector with the above characteristics can more quickly and accurately identify whether the semantics of the first voice signal are complete, that is, improve the accuracy of the intelligent robot's punctuation, reduce the occurrence of failures to respond to the first voice signal due to punctuation errors, and ensure the fluency of human-computer interaction.

[0119] It should be noted that the above steps S402 to S405 can be specifically performed by a classification model configured in the intelligent robot, that is, the first voice signal and the first text information are used as inputs of the classification model, and the classification model extracts feature vectors, performs modal interaction and fusion processing on the two, and performs classification based on the fused feature vectors. Among them, the third adjustment result a" u and the fourth adjustment result t" u You can also perform linear fusion according to the following formula: v = a" u W0t” u+b0. Where W0 and b0 are model parameters in the classification model. Figure 4 In the human-computer interaction method shown in the figure, the training process of the classification model in the intelligent robot is the same as Figure 1 The training method of the classification model in the illustrated embodiments is the same, and reference may be made to the above-mentioned related description, which will not be repeated here.

[0120] In summary, the above embodiments can improve the accuracy of semantic integrity to varying degrees and ensure the procedurality of human-computer interaction.

[0121] in, Figure 1 In the illustrated embodiment, the feature vectors of the first speech signal and the first text information are used to identify whether the semantics of the first speech signal are complete.

[0122] Figure 2 In the embodiment shown, Figure 1 On the basis of the illustrated embodiment, a second voice signal and its corresponding second text information are added. The intelligent robot uses the feature vectors of multiple voice signals and multiple text information with semantic association to identify whether the semantics of the first voice signal are complete, thereby further improving the recognition accuracy.

[0123] Figure 3 In the embodiment shown, Figure 2 On the basis of the illustrated embodiment, a modal interaction process is added, which further improves the accuracy of semantic integrity recognition while reducing the amount of information in the feature vector without losing important information.

[0124] Figure 4 In the embodiment shown, Figure 1 On the basis of the illustrated embodiment, a modal interaction process of the first voice signal and the first text information is newly added, thereby further improving the recognition accuracy.

[0125] For ease of understanding, the specific implementation process of the human-computer interaction method provided above is illustrated in combination with the customer service scenario. Figure 5 understand.

[0126] Suppose a user actively calls the customer service number of the Provident Fund Service Hall. Upon receiving the call, the intelligent customer service robot generates voice signal 1 to the user: "This is the Provident Fund Service Platform. How can I help you?" The user responds to voice signal 1 with voice signal 2: "I would like to inquire about that." After a preset silence duration, such as 3 seconds, the intelligent customer service robot segments voice signal 1 and begins to determine whether voice signal 2 is semantically complete.

[0127] The intelligent robot can first convert voice signal 1 into text information 1, and voice signal 2 into text information 2. Based on this, one judgment process can be: the intelligent robot extracts the feature vectors of voice signal 1 and text information 1 respectively, obtains a fused feature vector by direct concatenation or linear fusion, and judges that the semantics of voice signal 1 is incomplete based on this feature vector.

[0128] Another judgment process can be: after the intelligent robot extracts the feature vectors of speech signal 1 and text information 1 respectively, it can perform modal fusion on the extracted feature vectors, and then directly concatenate or linearly fuse the fusion results to obtain a fused feature vector. Based on this feature vector, it is determined that the semantics of speech signal 2 is incomplete.

[0129] Another judgment process can be: the intelligent robot first extracts feature vectors from text information 1 and text information 2, then directly concatenates or linearly fuses them to obtain a fused text feature vector. The fused text feature vector is then directly concatenated or linearly fused with the feature vector of speech signal 2 to obtain a fused feature vector. Based on this feature vector, the robot determines that speech signal 2 is semantically incomplete.

[0130] Another judgment process may be: the intelligent robot first obtains a fused text feature vector based on the feature vectors of text information 1 and text information 2, then performs modal fusion on the fused text feature vector and the feature vector of the first speech signal, and obtains a fused feature vector based on the modal fusion result. Based on this feature vector, the intelligent robot determines that the semantics of speech signal 2 are incomplete.

[0131] The various judgment processes mentioned above can all be executed by the classification model configured in the intelligent robot.

[0132] Another judgment process: Since the last word in the text information 1 is "that", a preset word indicating incomplete semantics, the intelligent robot can directly determine that the semantics of the voice signal 2 is incomplete.

[0133] After recognizing that Voice Signal 2 is semantically incomplete, the intelligent robot can appropriately extend the preset silence duration, even if the intelligent robot waits for a preset duration, such as 3 seconds, and further determine whether the user has generated Voice Signal 3 within this preset duration. If the user generates Voice Signal 3 ("Provident Fund") within the 6-second period (preset silence duration + preset duration), the intelligent robot can concatenate Voice Signal 1 and Voice Signal 3 to produce a concatenated voice signal: "I want to inquire about that Provident Fund." At this point, the preset silence duration is restored. If the user does not generate any new voice signals within the restored 3-second silence duration, the intelligent robot punctuates the concatenated voice signal and begins to determine whether the concatenated voice signal is semantically complete.

[0134] At this point, the intelligent robot can determine that the spliced ​​voice signal is semantically complete and can then output response voice 4 to the user: "Please enter your ID number," thereby ensuring smooth human-computer interaction. Furthermore, even if the user pauses for a long time while generating the voice signal, the intelligent robot will not make punctuation errors, which would cause the human-computer interaction to fail.

[0135] Optionally, within the 6 seconds of the preset silence duration + the preset duration, the intelligent robot may also output a guiding audio signal to the user: "Well, OK, please continue speaking" to guide the user to generate voice signal 3.

[0136] In practice, in addition to the above-mentioned customer service scenarios, intelligent robots can also be outbound call robots. In order to facilitate understanding, the specific implementation process of the human-computer interaction method provided above is exemplified in combination with the outbound call scenario. The following process can be combined with Figure 6 understand.

[0137] For example, in a customer service call, the outbound call robot can generate voice signal 1: "Are you satisfied with product A you purchased?" After the user answers the call, they can respond with voice signal 2: "I am very satisfied with product A." The intelligent outbound call robot can then perform sentence segmentation and determine the semantic integrity of voice signal 2 using the methods described in the previous embodiments. The intelligent robot can then respond with voice signal 3: "OK, thank you for your support," thus achieving smooth human-computer interaction.

[0138] In a possible design, the human-computer interaction method provided in the above embodiments can be applied to an intelligent robot, such as Figure 7 As shown, the intelligent robot may include: a processor 21 and a memory 22. The memory 22 is used to store the electronic device to perform the above Figures 1 to 4 The program of the human-computer interaction method provided in the illustrated embodiment, the processor 21 is configured to execute the program stored in the memory 22 .

[0139] The program includes one or more computer instructions, wherein the one or more computer instructions, when executed by the processor 21, can implement the following steps:

[0140] Acquire a first voice signal generated by a user and first text information corresponding to the first voice signal;

[0141] Determining a fusion feature vector based on the feature vectors of the first speech signal and the first text information;

[0142] determining, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete;

[0143] Respond to the first speech signal according to the classification result.

[0144] Optionally, the processor 21 is further configured to execute the aforementioned Figures 1 to 4 All or part of the steps in the illustrated embodiments.

[0145] The structure of the electronic device may further include a communication interface 23 for the electronic device to communicate with other devices or a communication network.

[0146] In addition, an embodiment of the present invention provides a computer storage medium for storing computer software instructions used by the above electronic device, which includes instructions for executing the above Figures 1 to 4 The procedures involved in the human-computer interaction method in the method embodiment shown.

[0147] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.

Claims

1. A human-computer interaction method, characterized in that: Applied to intelligent robots, including: Obtaining a first voice signal generated by a user and first text information corresponding to the first voice signal, and second text information corresponding to a second voice signal generated by the intelligent robot, where the second voice signal is generated before the first voice signal, and the first voice signal and the second voice signal are semantically associated; Determining a fusion feature vector based on the feature vectors of the first speech signal and the first text information; determining, based on the fused feature vector, a classification result reflecting whether the first speech signal is semantically complete; responding to the first voice signal according to the classification result; The step of determining a fusion feature vector based on the feature vectors of the first speech signal and the first text information includes: Determining a fused text feature vector based on the feature vectors of the first text information and the second text information; reducing the information content of the feature vector of the first speech signal according to the fused text feature vector to obtain a first adjustment result; reducing the information content of the fused text feature vector according to the feature vector of the first speech signal to obtain a second adjustment result; The fused feature vector is determined according to the first adjustment result and the second adjustment result, where the fused feature vector includes context information, user speaking state, and semantics between the first voice signal and the second voice signal.

2. The method according to claim 1, characterized in that The responding to the first voice signal according to the classification result includes: If the classification result is semantically complete, performing semantic recognition on the first speech signal; According to the recognition result, a successful answer voice signal corresponding to the first voice signal is output.

3. The method according to claim 1, characterized in that The responding to the first voice signal according to the classification result includes: If the classification result is semantically incomplete, the response result of the first voice signal is determined according to whether the user generates a third voice signal within a preset time period.

4. The method according to claim 3, characterized in that The determining a response result of the first voice signal according to whether the user generates a third voice signal within a preset time period includes: If the user fails to generate the third voice signal within the preset time period, an answer failure voice signal corresponding to the first voice signal is output.

5. The method according to claim 3, characterized in that The determining a response result of the first voice signal according to whether the user generates a third voice signal within a preset time period includes: If the user generates the third voice signal within the preset time period, splicing the first voice signal and the third voice signal to obtain a spliced ​​voice signal; If the classification result of the spliced ​​speech signal is semantically complete, performing semantic recognition on the spliced ​​speech signal; According to the recognition result, a successful answer voice signal corresponding to the first voice signal is output.

6. The method according to claim 1, characterized in that The method further comprises: If the word at the preset position in the first text information is a preset word, it is determined that the classification result is semantically incomplete.

7. An intelligent robot, characterized in that: include: A memory and a processor; wherein the memory stores executable code, and when the executable code is executed by the processor, the processor executes the human-computer interaction method according to any one of claims 1 to 6.

8. A non-transitory machine-readable storage medium, characterized in that The non-transitory machine-readable storage medium stores executable code, and when the executable code is executed by a processor of an electronic device, the processor is caused to execute the human-computer interaction method according to any one of claims 1 to 6.

Citation Information

Patent Citations

  • Multi-modal semantic integrity recognition method and device and electronic equipment

    CN112101045A

  • Data modeling method and system based on intelligent marketing scene

    CN114022192A