Voice interaction method and device, electronic equipment and computer program product

By allowing electronic devices to act as the initiator of voice interaction, actively output interactive voice and analyze user response voice, the problem of poor user experience in the existing technology is solved, and a wider range of voice interaction application scenarios and better user experience is achieved.

CN119943043APending Publication Date: 2025-05-06SHENZHEN YOUBIXUAN MEDICAL ROBOT CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411985684.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-06

Smart Images

  • Figure CN119943043A_ABST
    Figure CN119943043A_ABST
Patent Text Reader

Abstract

The invention is suitable for the technical field of computers, and particularly relates to a voice interaction method and device, electronic equipment and a computer program product. According to the method, in a target application scene, when voice interaction with a user needs to be carried out, the electronic equipment can serve as an initiator of voice interaction to actively output interaction voice to carry out voice interaction with the user, and the user can serve as a responder of voice interaction to respond to the interaction voice output by the electronic equipment; the electronic equipment can obtain the response voice of the user, and can obtain the interaction result based on the response voice of the user, so that the application scene of voice interaction can be expanded, the user requirements in many application scenes can be met, and the user experience can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application belongs to the field of computer technology, and in particular, relates to a voice interaction method, device, electronic device, and computer program product. Background Art

[0002] At present, voice interaction generally involves the user as the initiator of voice interaction, and the electronic device as the responder of voice interaction responds to the voice initiated by the user. For example, after the user initiates a voice question, the electronic device obtains the user's voice. After obtaining the user's voice, the electronic device can convert the voice into text through automatic speech recognition (automatic speech recognition, ASR) technology, and then perform semantic analysis of the text through natural language processing (natural language processing, NLP) technology to determine the response content based on the analyzed semantics, and output the response content. This voice interaction method cannot meet the user needs in many application scenarios, resulting in a poor user experience. Summary of the invention

[0003] The embodiments of the present application provide a voice interaction method, device, electronic device and computer program product, which can realize voice interaction with the electronic device as the initiator of voice interaction and the user as the responder of voice interaction, meet user needs in multiple application scenarios and improve user experience.

[0004] In a first aspect, an embodiment of the present application provides a voice interaction method, which is applied to an electronic device, and the method includes:

[0005] The electronic device determines a target application scenario for voice interaction;

[0006] The electronic device outputs a first interactive voice according to the target application scenario;

[0007] The electronic device acquires a first response voice, where the first response voice is a voice in which the user gives feedback on the first interactive voice;

[0008] The electronic device analyzes the first response voice to obtain a first interaction result.

[0009] In the voice interaction method provided above, in the target application scenario, when voice interaction with the user is required, the electronic device can actively output interactive voice as the initiator of voice interaction to interact with the user, and the user can respond to the interactive voice output by the electronic device as the responder of voice interaction. The electronic device can obtain the user's response voice, and can obtain the interaction result based on the user's response voice, which can expand the application scenario of voice interaction, meet the user needs in many application scenarios, and improve user experience.

[0010] In some embodiments, after the electronic device acquires the first response voice, the method further includes:

[0011] The electronic device determines the emotional state of the user according to the first response voice;

[0012] The electronic device determines an emotional feedback voice according to the emotional state, and outputs the emotional feedback voice.

[0013] In one embodiment, the electronic device determines the emotional state of the user according to the first response voice, including:

[0014] The electronic device acquires the physiological sign data of the user, and / or acquires image data containing the facial information of the user;

[0015] The electronic device determines the emotional state of the user based on the first response voice, the physiological sign data and / or the image data.

[0016] In some embodiments, the electronic device outputs a first interactive voice according to the target application scenario, including:

[0017] The electronic device determines a text interaction question corresponding to the target application scenario and a text answer corresponding to the text interaction question;

[0018] The electronic device converts the text interaction question into the first interaction voice, and outputs the first interaction voice;

[0019] The electronic device analyzes the first response voice to obtain a first interaction result, including:

[0020] The electronic device converts the first response voice into text to obtain interactive content corresponding to the first response voice;

[0021] The electronic device obtains the first interaction result according to the text answer and the interaction content.

[0022] In some embodiments, after the electronic device analyzes the first response voice and obtains a first interaction result, the method further includes:

[0023] When the first interaction result is a first preset result, the electronic device outputs a second interaction voice; the first preset result is used to indicate that the interaction content corresponding to the first response voice cannot be determined; the second interaction voice is associated with the first interaction voice;

[0024] The electronic device acquires a second response voice, where the second response voice is a voice in which the user gives feedback on the second interactive voice;

[0025] The electronic device analyzes the second response voice to obtain a second interaction result.

[0026] In some other embodiments, after the electronic device analyzes the first response voice to obtain a first interaction result, the method further includes:

[0027] When the first interaction result is the second preset result, the electronic device outputs a prompt voice; the prompt voice is used to prompt the user to provide correct feedback on the first interaction voice, and the second preset result is used to indicate that the degree of match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is less than or equal to a preset threshold.

[0028] In some other embodiments, after the electronic device analyzes the first response voice to obtain a first interaction result, the method further includes:

[0029] When the first interaction result is the third preset result, the electronic device outputs a third interaction voice according to the target application scenario; the third preset result is used to indicate that the degree of match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is greater than a preset threshold.

[0030] In a second aspect, an embodiment of the present application provides a voice interaction device, which is applied to an electronic device, and the device includes:

[0031] A scenario determination module is used to determine the target application scenario of voice interaction;

[0032] An interactive voice output module, used to output a first interactive voice according to the target application scenario;

[0033] A response voice acquisition module, used to acquire a first response voice, where the first response voice is a voice in which the user gives feedback on the first interactive voice;

[0034] The voice analysis module is used to analyze the first response voice to obtain a first interaction result.

[0035] In a third aspect, an embodiment of the present application provides an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein when the processor executes the computer program, the electronic device implements the voice interaction method described in any one of the first aspects above.

[0036] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a computer, the computer implements the voice interaction method described in any one of the first aspects above.

[0037] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a computer, the computer implements the voice interaction method described in any one of the first aspects above.

[0038] It can be understood that the beneficial effects of the second to fifth aspects mentioned above can be found in the relevant description of the first aspect mentioned above, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0040] Figure 1 It is a flowchart of the voice interaction method provided in the embodiment of the present application;

[0041] Figure 2 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 1 ;

[0042] Figure 3 is an example diagram of a state machine provided in an embodiment of the present application;

[0043] Figure 4 is a structural diagram of a voice interaction device provided in an embodiment of the present application;

[0044] Figure 5 This is a schematic diagram of the structure of the electronic device provided in the embodiment of the present application. Figure 2 . DETAILED DESCRIPTION

[0045] In the following description, specific details such as specific system structures, technologies, etc. are provided for the purpose of illustration rather than limitation, so as to provide a thorough understanding of the embodiments of the present application. However, it should be clear to those skilled in the art that the present application may also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to prevent unnecessary details from obstructing the description of the present application.

[0046] It should be understood that when used in the present specification and the appended claims, the term "comprising" indicates the presence of described features, integers, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integers, steps, operations, elements, components and / or combinations thereof.

[0047] It should also be understood that the term “and / or” used in the specification and appended claims refers to any and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0048] As used in the specification and appended claims of this application, the term "if" can be interpreted as "when" or "uponce" or "in response to determining" or "in response to detecting", depending on the context. Similarly, the phrase "if it is determined" or "if [described condition or event] is detected" can be interpreted as meaning "uponce it is determined" or "in response to determining" or "uponce [described condition or event] is detected" or "in response to detecting [described condition or event]", depending on the context.

[0049] In addition, in the description of the present application specification and the appended claims, the terms "first", "second", "third", etc. are only used to distinguish the descriptions and cannot be understood as indicating or implying relative importance.

[0050] References to "one embodiment" or "some embodiments" etc. described in the specification of this application mean that one or more embodiments of the present application include specific features, structures or characteristics described in conjunction with the embodiment. Therefore, the statements "in one embodiment", "in some embodiments", "in some other embodiments", "in some other embodiments", etc. that appear in different places in this specification do not necessarily refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized in other ways. The terms "including", "comprising", "having" and their variations all mean "including but not limited to", unless otherwise specifically emphasized in other ways.

[0051] Voice interaction generally refers to the user as the initiator of voice interaction, and the electronic device as the responder of voice interaction responds to the voice initiated by the user. For example, after the user initiates a voice question, the electronic device can obtain the user's voice. After obtaining the user's voice, the electronic device can first convert the voice into text through ASR technology, and then perform semantic analysis on the text through NLP technology to determine the response content based on the analyzed semantics, and output the response content. Among them, in application scenarios such as education, medical consultation, cognitive assessment, psychological inquiry or companion communication, it is generally required that the electronic device initiates voice interaction as the active party, and the user responds as the responder. However, this voice interaction method in which the user is the initiator and the electronic device is the responder cannot meet the user needs in many application scenarios, resulting in a poor user experience.

[0052] To solve the above problems, the embodiments of the present application provide a voice interaction method, device, electronic device and computer program product. In the method, after determining the target application scenario of the voice interaction, the electronic device can output the interactive voice according to the target application scenario to perform voice interaction with the user. Subsequently, the electronic device can obtain the response voice of the user to the interactive voice, and can analyze the response voice to obtain the interaction result. That is, in the embodiment of the present application, in the target application scenario, the electronic device can actively output the interactive voice as the initiator of the voice interaction to perform voice interaction with the user. The user can respond to the interactive voice output by the electronic device as the responder of the voice interaction. The electronic device can obtain the user's response voice, and can obtain the interaction result based on the user's response voice, which can expand the application scenario of the voice interaction, meet the user's needs in many application scenarios, improve the user's physical examination, and has strong ease of use and practicality.

[0053] The voice interaction method provided in the embodiments of the present application can be applied to electronic devices with voice interaction functions, such as robots, mobile phones, tablet computers, vehicle-mounted devices, wearable devices, smart large screens, augmented reality (AR) / virtual reality (VR) devices, laptop computers, ultra-mobile personal computers (UMPC), netbooks, personal digital assistants (PDA), etc. The embodiments of the present application do not impose any restrictions on the specific types of electronic devices.

[0054] The voice interaction method provided in the embodiment of the present application is described in detail below with reference to the accompanying drawings and specific application scenarios.

[0055] See also Figure 1 , Figure 1The schematic flow chart of the voice interaction method provided in the embodiment of the present application is shown. The method can be applied to electronic devices with voice interaction functions such as robots, mobile phones, laptops, tablet computers or smart large screens. Figure 1 As shown, the method may include:

[0056] S101. The electronic device determines a target application scenario for voice interaction.

[0057] In an embodiment of the present application, when voice interaction is required, for example, when an instruction for voice interaction is obtained, the electronic device can determine the target application scenario of the voice interaction. Among them, the target application scenario can be any one of the application scenarios such as education, medical consultation, cognitive assessment, psychological inquiry or companion communication. The embodiment of the present application does not specifically limit the target application scenario and can be determined according to the actual scenario. It should be understood that the instruction for voice interaction can be triggered by the user, or can be automatically triggered by the electronic device. For example, when the electronic device detects the presence of a user facing the electronic device, it can automatically trigger the instruction for voice interaction to instruct the electronic device to perform voice interaction with the user.

[0058] It should be noted that the target application scenario of voice interaction can be set by default by the electronic device, or can be customized by the user, or can be determined by the electronic device according to the current location of the electronic device. For example, the electronic device can set the target application scenario of voice interaction to a cognitive assessment scenario by default. For example, the user can customize the target application scenario of voice interaction to a companion communication scenario. For example, when it is determined that the electronic device is currently located in a school, the electronic device can determine the target application scenario to be an education scenario. For example, when it is determined that the electronic device is currently located in a hospital, the electronic device can determine the target application scenario to be a medical consultation scenario.

[0059] S102: The electronic device outputs a first interactive voice according to a target application scenario.

[0060] In some embodiments, the interactive questions and answers corresponding to the target application scenario can be set in advance according to the information required to be collected for the target application scenario. Exemplarily, the interactive questions and answers in text form corresponding to the target application scenario can be set in advance. It should be understood that the interactive questions can also be interactive questions in voice form, and the answers can also be answers in voice form, which is not limited in the embodiments of the present application.

[0061] For example, when the target application scenario requires the collection of personal information such as age and gender, interactive questions can be set in advance, including text-based question A1 and voice-based question A2, and the answers include text-based answer B1 corresponding to question A1 and voice-based answer B2 corresponding to question A2. Question A1 can be "May I ask what is your gender?", and the answer B1 corresponding to question A1 can be male or female. Question A2 can be "May I ask how old you are?", and the answer B2 corresponding to question A2 can be any value between 5 and 100 years old.

[0062] For example, a question library and a standard answer library corresponding to the target application scenario may be stored in an electronic device or other device that is communicatively connected to the electronic device (the electronic device will be used as an example for exemplary explanation below). The question library may include interactive questions corresponding to the target application scenario. The standard answer library may include answers corresponding to each interactive question. Among them, when setting interactive questions and answers for the target application scenario, for any candidate question set, the electronic device may match the candidate question with the question library to determine whether there are interactive questions similar to the candidate question in the question library. When there are no interactive questions similar to the candidate question in the question library, the electronic device may add the candidate question to the question library, and may add the answer corresponding to the candidate question to the standard answer library. Otherwise, the electronic device may not add the candidate question and the answer.

[0063] It should be understood that the embodiments of the present application do not limit the specific method of determining whether there are already interactive questions similar to the candidate question in the question library, and can be determined according to the actual scenario. For example, the text similarity between the candidate question and each interactive question in the question library can be determined to determine whether there are already interactive questions similar to the candidate question in the question library based on the text similarity. For example, keyword extraction can be performed on the candidate question and each interactive question in the question library respectively, and the similarity between the keywords can be determined to determine whether there are already interactive questions similar to the candidate question in the question library based on the similarity between the keywords.

[0064] In one embodiment, after determining each interactive question corresponding to the target application scenario, for each interactive question in text form, the electronic device can convert each interactive question in text form into a corresponding interactive voice. For example, the electronic device can convert each interactive question in text form into a corresponding interactive voice through ASR technology. When performing voice interaction, the electronic device can select one or more interactive voices from the interactive voices corresponding to the target application scenario (i.e., interactive questions in voice form) as the first interactive voice, and can output the first interactive voice to the user to perform voice interaction with the user through the first interactive voice.

[0065] In another embodiment, when performing voice interaction, for each interactive question in text form corresponding to the target application scenario, the electronic device can select one or more interactive questions (for example, can be called text interactive questions) from the interactive questions in text form. Subsequently, the electronic device can convert the selected text interactive question into a first interactive voice, and can output the first interactive voice to the user to perform voice interaction with the user through the first interactive voice. It should be understood that when the electronic device selects the text interactive question, it can also select one or more interactive questions in voice form as the first interactive voice, that is, the first interactive voice can include the interactive voice converted from the interactive question in text form, and / or can include the interactive question in voice form.

[0066] It should be noted that the embodiments of the present application do not limit the specific manner in which the electronic device selects the first interactive voice from the interactive voices corresponding to the target application scenario, and the specific manner in which the electronic device selects the text interactive question from the interactive questions in text form corresponding to the target application scenario, which can be determined according to the actual scenario. For example, the electronic device can select the first interactive voice or text interactive question by random selection. For example, the importance corresponding to the interactive question can be set in advance in the electronic device, and the electronic device can select the first interactive voice or text interactive question according to the importance corresponding to the interactive question, and so on.

[0067] S103: The electronic device obtains a first response voice, where the first response voice is a voice in which the user provides feedback on the first interactive voice.

[0068] In the embodiment of the present application, after the electronic device outputs the first interactive voice to interact with the user by voice, the user can reply to the first interactive voice. For example, the user can reply to the first interactive voice by voice, and the electronic device can obtain the voice of the user's reply (i.e., the first response voice).

[0069] In some embodiments, after acquiring the first response voice, the electronic device can determine the user's emotional state based on the first response voice, and can perform corresponding processing based on the user's emotional state. For example, when the user's emotional state is a negative emotional state such as tension, anxiety, sadness, anger, fear or irritability, the electronic device can determine the emotional feedback voice based on the emotional state, and can output the emotional feedback voice to soothe the user's emotions. Among them, the emotional feedback voice can be a comforting voice, a motivational voice, or an encouraging voice, etc.

[0070] In other embodiments, the electronic device may also obtain the user's physiological sign data, and / or obtain image data containing the user's facial information, and may determine the user's emotional state based on the first response voice, and the physiological sign data and / or image data. That is, the electronic device may determine the user's emotional state based on the first response voice and the physiological sign data; or, the electronic device may determine the user's emotional state based on the first response voice and the image data; or, the electronic device may determine the user's emotional state based on the first response voice, the physiological sign data, and the image data. The physiological sign data may include heart rate, respiratory rate, or blood pressure, etc.

[0071] It should be noted that the embodiments of the present application do not limit the specific manner in which the electronic device determines the user's emotional state based on the first response voice, and the specific manner in which the electronic device determines the user's emotional state based on the first response voice and physiological sign data and / or image data, which can be determined according to the actual scenario.

[0072] In a possible implementation, an emotional speech library may be provided in the electronic device or other device that is communicatively connected to the electronic device (the electronic device will be used as an example for illustrative explanation below). The emotional speech library may include one or more emotional feedback voices corresponding to each emotional state. Therefore, after determining the emotional state of the user, the electronic device may determine one or more emotional feedback voices corresponding to the emotional state of the user from the emotional speech library, and may output the determined emotional feedback voices.

[0073] It should be understood that the embodiments of the present application do not limit the specific manner in which the electronic device determines one or more emotional feedback voices corresponding to the user's emotional state from the emotional voice library, and can be determined according to the actual scenario. For example, the electronic device can determine one or more emotional feedback voices corresponding to the user's emotional state from the emotional voice library in a random manner. For example, the electronic device can determine the historical soothing effect of the emotional feedback voice, and can determine one or more emotional feedback voices corresponding to the user's emotional state from the emotional voice library based on the historical soothing effect.

[0074] It should be noted that the electronic device described above outputs an emotional feedback voice according to the user's emotional state to soothe the user's emotions for only exemplary purposes and should not be construed as limiting the embodiments of the present application. In the embodiments of the present application, the electronic device may also output other types of information to soothe the user's emotions. For example, the electronic device may soothe the user's emotions by playing music or displaying related images.

[0075] S104: The electronic device analyzes the first response voice to obtain a first interaction result.

[0076] In an embodiment of the present application, after obtaining the user's first response voice, the electronic device can perform text conversion on the first response voice to obtain the interactive content corresponding to the first response voice, and can obtain the interactive result (i.e., the first interactive result) based on the text answer corresponding to the first interactive voice and the interactive content corresponding to the first response voice. For example, the electronic device can perform text conversion on the first response voice through ASR to obtain the interactive content corresponding to the first response voice. After obtaining the interactive content corresponding to the first response voice, the electronic device can perform semantic analysis on the interactive content corresponding to the first response voice through NLP to obtain the semantic analysis result, so as to determine the first interactive result based on the semantic analysis result.

[0077] In some embodiments, the first interaction result may be one of a first preset result, a second preset result, and a third preset result. The first preset result is used to indicate that the interaction content corresponding to the first response voice cannot be determined. The second preset result is used to indicate that the degree of match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is less than or equal to a preset threshold. The third preset result is used to indicate that the degree of match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is greater than a preset threshold.

[0078] It should be noted that the preset threshold can be specifically determined according to the actual scenario, and the embodiments of the present application are not limited to this.

[0079] In one embodiment, when the first interaction result is the first preset result, the electronic device can output a second interaction voice, and the second interaction voice can be associated with the first interaction voice. That is, when the electronic device cannot determine the interaction content corresponding to the first response voice, the electronic device can re-output the second interaction voice to re-interact with the user. Among them, the core of the re-output second interaction voice and the first interaction voice can be the same, that is, the information to be obtained by the second interaction voice and the first interaction voice can be the same or similar, but the text content corresponding to the second interaction voice may be different from the text content corresponding to the first interaction voice. For example, when the first interaction voice output by the electronic device is "Have you seen a cat?", when the electronic device cannot determine the interaction content corresponding to the user's first response voice, the electronic device can re-output the second interaction voice, for example, the second interaction voice can be "Which animal's call are you most interested in?".

[0080] It should be noted that the electronic device being unable to determine the interactive content corresponding to the first response voice may include the electronic device being unable to perform text conversion and / or semantic analysis on the first response voice, or may include the electronic device performing text conversion and semantic analysis on the first response voice but being unable to determine the meaning expressed by the analyzed content, etc.

[0081] After outputting the second interactive voice, the electronic device may obtain the user's response voice to the second interactive voice (i.e., the second response voice). After obtaining the second response voice, the electronic device may analyze the second response voice to obtain an interactive result (i.e., the second interactive result). It should be understood that after obtaining the second interactive result, the electronic device may perform subsequent processing with reference to the processing method of the first interactive result.

[0082] For example, the second interaction result may be one of the first preset result, the second preset result, and the third preset result. The first preset result is used to indicate that the interaction content corresponding to the second response voice cannot be determined. The second preset result is used to indicate that the degree of match between the interaction content corresponding to the second response voice and the text answer corresponding to the second interaction voice is less than or equal to a preset threshold. The third preset result is used to indicate that the degree of match between the interaction content corresponding to the second response voice and the text answer corresponding to the second interaction voice is greater than a preset threshold.

[0083] In another embodiment, when the first interaction result is the second preset result, the electronic device can output a prompt voice. The prompt voice can be used to prompt the user to give correct feedback to the first interaction voice. That is, when the electronic device determines that the match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is less than or equal to the preset threshold, the electronic device can determine that the first response voice is an incorrect feedback. At this time, the electronic device can output a prompt voice to prompt the user to make a correct reply. For example, when the first interaction voice is "Do you know how old you are?", if the electronic device performs text conversion and semantic analysis on the user's first response voice, and determines that the interaction content corresponding to the first response voice is "I am eating", the electronic device can determine that the user's reply is an incorrect feedback. At this time, the electronic device can output a prompt voice to prompt the user to make a correct reply.

[0084] In another embodiment, when the first interaction result is the third preset result, the electronic device can output a third interaction voice according to the target application scenario. That is, when the electronic device determines that the match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is greater than the preset threshold, the electronic device can determine that the first response voice is correct feedback. At this time, the electronic device can determine that the interaction of the first interaction voice is completed, can determine the next interaction voice (i.e., the third interaction voice), and can output the third interaction voice to interact with the user. It should be understood that when the first interaction voice is the last interaction voice required for this voice interaction, when it is determined that the first interaction result is the third preset result, the electronic device can determine that this voice interaction is completed, can stop outputting the interaction voice, that is, can no longer output the third interaction voice, and can save the interaction data of this voice interaction (such as all response voices, or all interaction voices and all response voices) to the relevant database.

[0085] It should be noted that, in order to reduce the impact of voice deviation on voice analysis, when determining the degree of match between the interactive content corresponding to the first response voice and the text answer corresponding to the first interactive voice, the embodiment of the present application can use a fuzzy matching method to perform matching. The preset threshold can be determined based on the fuzzy matching. Among them, the specific method of fuzzy matching can be determined according to the actual scenario, and the embodiment of the present application does not limit this.

[0086] In some embodiments, in order to reduce the impact of possible environmental noise on the accuracy of voice analysis, after obtaining the user's first response voice, the electronic device may first perform environmental sound filtering on the first response voice to remove the environmental noise in the first response voice. It should be understood that the embodiments of the present application do not limit the specific manner in which the electronic device performs environmental sound filtering on the first response voice, which can be determined specifically according to the actual scenario, for example, by performing environmental sound filtering on the first response voice in any existing manner.

[0087] In other embodiments, in order to improve the efficiency and accuracy of voice analysis and improve the efficiency of voice interaction, after obtaining the user's first response voice, the electronic device may first identify and compare the voice features of the first response voice to determine whether the first response voice contains valid response content. When it is determined that the first response voice contains valid response content, the electronic device may analyze the first response voice to obtain a first interaction result, for example, the first response voice may be converted into text and semantically parsed to determine the first interaction result based on the semantic parsing result. Among them, when it is determined that the first response voice does not contain valid response content, the electronic device may determine that the interaction content corresponding to the first response voice cannot be identified, that is, the first interaction result may be directly determined to be the first preset result (that is, the interaction result of the interaction content corresponding to the first response voice cannot be determined).

[0088] It should be noted that the valid response content may refer to the first response voice containing human speech. That is, when the electronic device recognizes and compares the voice features of the first response voice, if it is determined that the first response voice does not contain human speech, the electronic device can determine that the first response voice definitely does not contain content that replies to the first interactive voice. At this time, the electronic device may not perform voice conversion and semantic analysis on the first response voice, and may directly determine that the first interactive result is the first preset result, thereby reducing the voice conversion and semantic analysis process of the electronic device and improving the efficiency of voice interaction. When it is determined that the first response voice contains human speech, the electronic device can determine that the first response voice may contain content that replies to the first interactive voice, and can then perform subsequent voice analysis and other processing.

[0089] In some embodiments, after completing this voice interaction, the electronic device can save the interaction data of this voice interaction (e.g., all response voices, or all interaction voices and all response voices) to a relevant database. In addition, the electronic device can also input the user's response voice into a big data comparison model for comparison, so as to compare the user's response with the responses of other users, obtain a comparison result, and display the comparison result to the user, so that the user can understand the comparison between his own response and the responses of other users.

[0090] Among them, the big data comparison model may include the response voices corresponding to multiple users, and the response scores corresponding to each response voice, etc. After the user's response voice is input into the big data comparison model, the big data comparison model can extract the features corresponding to the user's response voice (such as timbre, audio, semantic keywords, volume or sound sequence, etc.), and can compare the features corresponding to the user's response voice with the features corresponding to the response voices of other users to determine the comparison result between the user's response and the responses of other users.

[0091] For example, see Figure 2 , Figure 2 The structure of the electronic device provided by the embodiment of the present application is shown in FIG. Figure 1 .

[0092] like Figure 2 As shown, the electronic device may include an instruction module 210 , a state machine 220 , an error or fault handling module 230 , and a big data comparison module 240 .

[0093] The instruction module 210 can be used to set interactive questions and answers corresponding to the target application scenario, and can be used to convert interactive questions in text form into corresponding interactive voices, etc.

[0094] The state machine 220 can be used to obtain an interactive voice (e.g., a first interactive voice) from the instruction module 210, and can output the first interactive voice to the user, and can obtain the user's response voice to the first interactive voice (e.g., a first response voice). In addition, the state machine 220 can also be used to filter the ambient sound of the first response voice, and to identify and compare the voice features of the first response voice to determine whether the first response voice contains valid response content. When it is determined that the first response voice contains valid response content, the state machine 220 can perform voice conversion and semantic analysis on the first response voice to obtain a semantic analysis result corresponding to the first response voice, so as to determine the interaction result corresponding to the first interactive voice (e.g., the first interaction result) based on the semantic analysis result corresponding to the first response voice, and can perform relevant processing based on the first interaction result.

[0095] For example, when it is determined that the first interaction result is the third preset result, the state machine 220 can obtain a new interactive voice according to the target application scenario to perform voice interaction with the user. When it is determined that the first interaction result is the first preset result, the state machine 220 can obtain an interactive voice associated with the first interactive voice (for example, a second interactive voice) to re-interact with the user. Among them, when performing voice interaction based on the second interactive voice, if the state machine 220 determines that the interaction result corresponding to the second interactive voice is still the first preset result, the state machine 220 can notify the error or error handling module 230 to process the first response voice. In addition, when it is determined that the first interaction result is the second preset result, the state machine 220 can also notify the error or error handling module 230 to process the first response voice.

[0096] In addition, after obtaining the first response voice, the state machine 220 can also determine the user's emotional state according to the first response voice. After determining the user's emotional state, the state machine 220 can determine the corresponding emotional feedback voice according to the user's emotional state, and can output the corresponding emotional feedback voice to the user.

[0097] The error or error handling module 230 can be used to process the response voice whose interactive content cannot be determined or the response voice of the error type, that is, to process the scenario of the first preset result or the second preset result. For example, when the first interactive result is the first preset result, the error or error handling module 230 can output the second interactive voice associated with the first interactive voice to interact with the user again. For example, when the first interactive result is the second preset result, the error or error handling module 230 can output a prompt voice to prompt the user to give correct feedback on the first interactive voice through the prompt voice.

[0098] The big data comparison module 240 can be used to compare the user's current answer with the answers of other users to obtain a comparison result, so that the user can understand the comparison between his own answer and the answers of other users. For example, the big data comparison module 240 can include a big data comparison model. After the user completes the current answer, the big data comparison module 240 can compare the user's current answer with the answers of other users through the big data comparison model to obtain a comparison result.

[0099] For example, see Figure 3 , Figure 3 An example diagram of a state machine provided by an embodiment of the present application is shown.

[0100] like Figure 3As shown, the state machine 220 may include an ambient sound filtering module 310 , a speech feature recognition and comparison module 320 , an ARS speech conversion module 330 , an NLP semantic analysis module 340 , a speech loop analysis module 350 , a speech output module 360 ​​and an emotional speech library 370 .

[0101] After obtaining the user's response voice (e.g., the first response voice), the state machine 220 can perform environmental sound filtering on the first response voice through the environmental sound filtering module 310 to remove environmental noise in the first response voice. In addition, the state machine 220 can also perform voice feature recognition and comparison on the first response voice through the voice feature recognition and comparison module 320 to determine whether the first response voice contains valid response content.

[0102] When it is determined that the first response voice does not contain valid response content, the voice feature recognition and comparison module 320 can directly determine that the interaction result corresponding to the first interactive voice is the first preset result. At this time, the voice feature recognition and comparison module 320 can input the first response voice into the error and error analysis module to process the first response voice through the error and error analysis module.

[0103] When it is determined that the first response voice contains valid response content, the state machine 220 can perform text conversion on the first response voice through the ARS voice conversion module 330 to convert the first response voice into text content. After converting the first response voice into text content, the state machine 220 can perform semantic analysis on the text content corresponding to the first response voice through the NLP semantic analysis module 340 to obtain the semantic analysis result corresponding to the first response voice.

[0104] After obtaining the semantic analysis result corresponding to the first response voice, the state machine 220 can analyze the first interactive voice and the semantic analysis result through the voice loop analysis module 350 to obtain the first interactive result, and can perform relevant processing based on the first interactive result.

[0105] For example, when the first interaction result is the third preset result, the third interaction voice can be output according to the target application scenario through the voice output module 360 ​​to output the next interaction voice to perform voice interaction with the user. For example, when the first interaction result is the first preset result, the second interaction voice associated with the first interaction voice can be output through the voice output module 360 ​​to re-interact with the user through the second interaction voice. Among them, when re-interacting, if the voice loop analysis module 350 still cannot determine the interaction content corresponding to the response voice (such as the second response voice) obtained by the re-interaction, that is, when the interaction result corresponding to the voice interaction re-performed by the voice loop analysis module 350 is still the first preset result, the voice loop analysis module 350 can input the first response voice into the error and error analysis module to process the first response voice through the error and error analysis module.

[0106] Exemplarily, after obtaining the user's first response voice, the state machine 220 may also perform emotion analysis on the first response voice through the voice loop analysis module 350 to determine the user's corresponding emotional state, and may select one or more emotional feedback voices from the emotional voice library 370 according to the user's corresponding emotional state. After determining the emotional feedback voice, the emotional feedback voice may be output to the user through the voice output module 360.

[0107] In the embodiment of the present application, the electronic device can actively output interactive voice as the initiator of voice interaction to perform voice interaction with the user. The user can respond to the interactive voice output by the electronic device as the responder of voice interaction. The electronic device can obtain the user's response voice, and can obtain the interaction result based on the user's response voice, which can expand the application scenarios of voice interaction, meet the user needs in many application scenarios, and improve user experience.

[0108] It should be understood that the size of the serial numbers of the steps in the above embodiments does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0109] Corresponding to the voice interaction method described in the above embodiment, Figure 4 The structural block diagram of the voice interaction device provided in an embodiment of the present application is shown, and the device can be applied to electronic devices. Figure 4 For ease of explanation, only the parts related to the embodiments of the present application are shown.

[0110] Reference Figure 4 , the device may include:

[0111] A scenario determination module 401 is used to determine a target application scenario for voice interaction;

[0112] An interactive voice output module 402, configured to output a first interactive voice according to the target application scenario;

[0113] The response voice acquisition module 403 is used to acquire a first response voice, where the first response voice is a voice in which the user gives feedback on the first interactive voice;

[0114] The voice analysis module 404 is used to analyze the first response voice to obtain a first interaction result.

[0115] In some embodiments, the apparatus may further include:

[0116] The emotional feedback module is used to determine the emotional state of the user according to the first response voice; determine the emotional feedback voice according to the emotional state, and output the emotional feedback voice.

[0117] In one embodiment, the emotional feedback module is also used to obtain physiological sign data of the user, and / or obtain image data containing facial information of the user; and determine the emotional state of the user based on the first response voice, and the physiological sign data and / or the image data.

[0118] In some embodiments, the interactive speech output module 402 is specifically used to determine the text interactive question corresponding to the target application scenario and the text answer corresponding to the text interactive question; convert the text interactive question into the first interactive speech, and output the first interactive speech;

[0119] The speech analysis module 404 is specifically used to convert the first response speech into text to obtain the interaction content corresponding to the first response speech; and obtain the first interaction result according to the text answer and the interaction content.

[0120] In some embodiments, the interactive voice output module 402 is further used to output a second interactive voice when the first interactive result is a first preset result; the first preset result is used to indicate that the interactive content corresponding to the first response voice cannot be determined; the second interactive voice is associated with the first interactive voice;

[0121] The response voice acquisition module 403 is further used to acquire a second response voice, where the second response voice is a voice in which the user gives feedback on the second interactive voice;

[0122] The voice analysis module 404 is further configured to analyze the second response voice to obtain a second interaction result.

[0123] In other embodiments, the interactive voice output module 402 is also used to output a prompt voice when the first interactive result is a second preset result; the prompt voice is used to prompt the user to provide correct feedback on the first interactive voice, and the second preset result is used to indicate that the degree of match between the interactive content corresponding to the first response voice and the text answer corresponding to the first interactive voice is less than or equal to a preset threshold.

[0124] In other embodiments, the interactive voice output module 402 is also used to output a third interactive voice according to the target application scenario when the first interactive result is a third preset result; the third preset result is used to indicate that the degree of match between the interactive content corresponding to the first response voice and the text answer corresponding to the first interactive voice is greater than a preset threshold.

[0125] It should be noted that the information interaction, execution process, etc. between the above-mentioned devices / units are based on the same concept as the method embodiment of the present application. Their specific functions and technical effects can be found in the method embodiment part and will not be repeated here.

[0126] The technicians in the relevant field can clearly understand that for the convenience and simplicity of description, only the division of the above-mentioned functional units and modules is used as an example for illustration. In practical applications, the above-mentioned function allocation can be completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiment can be integrated in a processing unit, or each unit can exist physically separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of software functional units. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of this application. The specific working process of the units and modules in the above-mentioned system can refer to the corresponding process in the aforementioned method embodiment, which will not be repeated here.

[0127] Figure 5 A schematic diagram of the structure of the electronic device provided in the embodiment of the present application Figure 2 .like Figure 5 As shown, the electronic device 5 of this embodiment includes: at least one processor 50 ( Figure 5 Only one is shown), a memory 51, and a computer program 52 stored in the memory 51 and executable on the at least one processor 50, wherein the processor 50 implements the steps in any of the above-mentioned embodiments of the voice interaction method when executing the computer program 52.

[0128] The electronic device 5 may be a robot, a mobile phone, a wearable device, a desktop computer, a notebook, a PDA, a smart screen, or other electronic device with a voice interaction function. The electronic device may include, but is not limited to, a processor 50 and a memory 51. Those skilled in the art will appreciate that Figure 5 It is only an example of the electronic device 5 and does not constitute a limitation on the electronic device 5. It may include more or fewer components than shown in the figure, or a combination of certain components, or different components. For example, it may also include input and output devices, network access devices, etc.

[0129] The processor 50 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor, etc.

[0130] In some embodiments, the memory 51 may be an internal storage unit of the electronic device 5, such as a hard disk or memory of the electronic device 5. In other embodiments, the memory 51 may also be an external storage device of the electronic device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the electronic device 5. Further, the memory 51 may also include both an internal storage unit and an external storage device of the electronic device 5. The memory 51 is used to store an operating system, an application program, a boot loader (BootLoader), data, and other programs, such as the program code of the computer program. The memory 51 may also be used to temporarily store data that has been output or is to be output.

[0131] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by an electronic device, the electronic device implements the steps in the above-mentioned method embodiments.

[0132] An embodiment of the present application provides a computer program product, which includes a computer program. When the computer program is executed by a computer, the computer is enabled to implement the steps in the above-mentioned method embodiments.

[0133] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present application implements all or part of the processes in the above-mentioned embodiment method, which can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Among them, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form. The computer-readable storage medium may at least include: any entity or device that can carry the computer program code to the device / electronic device, a recording medium, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electric carrier signal, a telecommunication signal, and a software distribution medium. For example, a USB flash drive, a mobile hard disk, a magnetic disk or an optical disk. In some jurisdictions, according to legislation and patent practice, a computer-readable storage medium cannot be an electric carrier signal and a telecommunication signal.

[0134] In the above embodiments, the description of each embodiment has its own emphasis. For parts that are not described or recorded in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0135] Those of ordinary skill in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.

[0136] In the embodiments provided in the present application, it should be understood that the disclosed devices / electronic devices and methods can be implemented in other ways. For example, the device / electronic device embodiments described above are merely schematic. For example, the division of the modules or units is only a logical function division. There may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0137] The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0138] The embodiments described above are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, a person skilled in the art should understand that the technical solutions described in the aforementioned embodiments may still be modified, or some of the technical features may be replaced by equivalents. Such modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application, and should all be included in the protection scope of the present application.

Claims

1. A voice interaction method, characterized in that: Applied to electronic equipment, the method comprises: The electronic device determines a target application scenario for voice interaction; The electronic device outputs a first interactive voice according to the target application scenario; The electronic device acquires a first response voice, where the first response voice is a voice in which the user gives feedback on the first interactive voice; The electronic device analyzes the first response voice to obtain a first interaction result.

2. The method according to claim 1, characterized in that After the electronic device acquires the first response voice, the method further includes: The electronic device determines the emotional state of the user according to the first response voice; The electronic device determines an emotional feedback voice according to the emotional state, and outputs the emotional feedback voice.

3. The method according to claim 2, characterized in that The electronic device determines the emotional state of the user according to the first response voice, including: The electronic device acquires the physiological sign data of the user, and / or acquires image data containing the facial information of the user; The electronic device determines the emotional state of the user based on the first response voice, the physiological sign data and / or the image data.

4. The method according to claim 1, characterized in that: The electronic device outputting a first interactive voice according to the target application scenario includes: The electronic device determines a text interaction question corresponding to the target application scenario and a text answer corresponding to the text interaction question; The electronic device converts the text interaction question into the first interaction voice, and outputs the first interaction voice; The electronic device analyzes the first response voice to obtain a first interaction result, including: The electronic device converts the first response voice into text to obtain interactive content corresponding to the first response voice; The electronic device obtains the first interaction result according to the text answer and the interaction content.

5. The method according to any one of claims 1 to 4, characterized in that After the electronic device analyzes the first response voice and obtains a first interaction result, the method further includes: When the first interaction result is a first preset result, the electronic device outputs a second interaction voice; the first preset result is used to indicate that the interaction content corresponding to the first response voice cannot be determined; the second interaction voice is associated with the first interaction voice; The electronic device acquires a second response voice, where the second response voice is a voice in which the user gives feedback on the second interactive voice; The electronic device analyzes the second response voice to obtain a second interaction result.

6. The method according to any one of claims 1 to 4, characterized in that After the electronic device analyzes the first response voice and obtains a first interaction result, the method further includes: When the first interaction result is the second preset result, the electronic device outputs a prompt voice; the prompt voice is used to prompt the user to provide correct feedback on the first interaction voice, and the second preset result is used to indicate that the degree of match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is less than or equal to a preset threshold.

7. The method according to any one of claims 1 to 4, characterized in that After the electronic device analyzes the first response voice and obtains a first interaction result, the method further includes: When the first interaction result is the third preset result, the electronic device outputs a third interaction voice according to the target application scenario; the third preset result is used to indicate that the degree of match between the interaction content corresponding to the first response voice and the text answer corresponding to the first interaction voice is greater than a preset threshold.

8. A voice interaction device, characterized in that: Applied to electronic equipment, the device comprises: A scenario determination module is used to determine the target application scenario of voice interaction; An interactive voice output module, used to output a first interactive voice according to the target application scenario; A response voice acquisition module, used to acquire a first response voice, where the first response voice is a voice in which the user gives feedback on the first interactive voice; The voice analysis module is used to analyze the first response voice to obtain a first interaction result.

9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the electronic device implements the voice interaction method according to any one of claims 1 to 7.

10. A computer program product, comprising a computer program, characterized in that: When the computer program is executed by an electronic device, the electronic device implements the voice interaction method as described in any one of claims 1 to 7.