Voice interaction method, device, equipment and storage medium

By recognizing keyword phrases in vehicle user voice commands and using an associated database for guidance, the problem of user intent being difficult to understand accurately during vehicle operation is solved, thus improving voice interaction efficiency and user experience.

CN117524223BActive Publication Date: 2026-08-04CHERY AUTOMOBILE CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHERY AUTOMOBILE CO LTD
Filing Date
2023-11-29
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

While the vehicle is in motion, the user's attention is focused on driving, making it difficult to organize accurate voice commands. This results in the vehicle being unable to understand the user's intentions, leading to low voice interaction efficiency and a poor user experience.

Method used

If the user's intent cannot be determined by recognizing keyword groups in the voice command, voice guidance is provided to supplement the intent. Guidance information is generated using a keyword association database to guide the user to supplement the voice intent until the user's intent is determined and the corresponding function is activated.

Benefits of technology

It improves the efficiency and user experience of voice interaction. When a single voice command cannot determine the intent, it guides the user to supplement the command, ensuring that the function is executed accurately.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117524223B_ABST
    Figure CN117524223B_ABST
Patent Text Reader

Abstract

This application discloses a voice interaction method, apparatus, device, and storage medium, belonging to the field of intelligent interaction. The method includes: receiving a voice command from a target user; identifying keywords in the voice command to obtain a keyword group; if the target user's voice intent cannot be determined through the keyword group, then providing voice guidance to the target user based on the keyword group to allow the target user to supplement their voice intent; if the target user's voice intent can be determined through the supplemented voice input, then activating the target function corresponding to the voice intent. This application improves the efficiency of voice interaction and enhances the user experience by guiding the target user to supplement their voice input based on voice guidance interaction when a single voice command cannot determine their voice intent, thereby determining the target user's voice intent and activating the target function corresponding to the voice intent.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent interaction, and in particular to a voice interaction method, device, equipment and storage medium. Background Technology

[0002] With the rapid development of vehicle technology and the gradual improvement of vehicle intelligence, voice interaction technology is becoming increasingly common in vehicle usage scenarios. However, due to the complexity of voice scenarios, when users interact with vehicles via voice, their voice commands need to be highly accurate so that the vehicle can accurately understand the user's intent and thus achieve the corresponding functions.

[0003] While the vehicle is in motion, the user's attention needs to be focused on judging the driving environment and controlling the vehicle. They may not be able to form accurate voice commands, causing the vehicle to be unable to understand the user's intentions and perform the corresponding functions. Therefore, the current voice interaction method is inefficient and the user experience is poor. Summary of the Invention

[0004] This application provides a voice interaction method, apparatus, device, and storage medium, which can improve the efficiency of voice interaction and enhance the user experience. The technical solution is as follows:

[0005] On the one hand, a voice interaction method is provided, the method comprising:

[0006] Receive voice commands from the target user;

[0007] Identify the keywords in the voice command to obtain the keyword group;

[0008] If the target user's voice intent cannot be determined through the keyword group, then the target user will be given voice guidance based on the keyword group to allow the target user to supplement the voice intent;

[0009] If the target user's voice intent can be determined through the supplementary voice information, then the target function corresponding to that voice intent will be activated.

[0010] Optionally, the voice guidance provided to the target user based on the keyword group includes:

[0011] Determine multiple target related terms corresponding to the keyword group from the keyword association database, as well as the relevance between the multiple target related terms and the keyword group;

[0012] Let i = 1. According to the ranking result of the related words, select the i-th related word from the multiple related words, generate and play voice guidance information based on the i-th related word to guide the target user to supplement the voice intent. The ranking result of the related words is obtained by sorting the multiple target related words in descending order of relevance between the multiple target related words and the keyword group.

[0013] If a voice supplement command is received from the target user, the voice intent of the target user cannot be determined through the voice supplement command, and the number of times the voice supplement command is issued does not exceed the number threshold, then the keywords in the voice supplement command are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database.

[0014] If no voice supplementation command is received from the target user, and the multiple target related words have not been traversed or the number of times the voice supplementation command is issued is not greater than the number threshold, then let i = i + 1, and return the step of selecting the i-th related word from the multiple related words according to the sorting result of the related words.

[0015] Optionally, determining multiple target related terms corresponding to the keyword group from the keyword association database, and the relevance between the multiple target related terms and the keyword group, includes:

[0016] From the keyword association database, determine multiple candidate related words and the relevance between these multiple candidate related words and the corresponding keywords in the keyword group. These multiple candidate related words refer to words that are related to the keywords in the keyword group.

[0017] Determine the logical relationships between the keywords in this keyword group;

[0018] Based on this logical relationship and the multiple candidate related words, the multiple target related words are determined;

[0019] Based on the relevance between the multiple candidate related terms and the corresponding keywords in the keyword group, the relevance between the multiple target related terms and the keyword group is determined.

[0020] Optionally, before the step of selecting the i-th related word from the plurality of related words, where i = i + 1 is set and the sorting result according to the related words is returned, the method further includes:

[0021] Supplementary related words are determined from the keyword association database. These supplementary related words refer to words that are related to the i-th related word.

[0022] Based on the supplementary related words, audio recommendation information is generated and played.

[0023] If no voice recommendation response is received from the target user, then execute the step of setting i = i + 1, returning the sorted results of the associated words, and selecting the i-th associated word from the multiple associated words.

[0024] Optionally, after generating and playing the audio recommendation information based on the supplementary related words, the method further includes:

[0025] If a voice recommendation response from the target user is received, the keywords in the voice recommendation response are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database.

[0026] Optionally, after generating and playing the voice guidance information based on the i-th associated word, the method further includes:

[0027] Based on the target user's feedback on the voice guidance information, update the relevance between the i-th related word in the keyword association database and the keywords in the keyword group.

[0028] On the other hand, a voice interaction device is provided, the device comprising:

[0029] The voice receiving module is used to receive voice commands issued by the target user.

[0030] The speech recognition module is used to identify keywords in the speech command and obtain keyword groups;

[0031] The voice guidance module is used to provide voice guidance to the target user based on the keyword group if the target user's voice intent cannot be determined through the keyword group, so that the target user can supplement the voice intent.

[0032] The function activation module is used to activate the target function corresponding to the voice intent if the voice intent of the target user can be determined through the voice supplementation results of the target user.

[0033] Optionally, the voice guidance module includes:

[0034] The related term determination submodule is used to determine multiple target related terms corresponding to the keyword group from the keyword association database, as well as the relevance between the multiple target related terms and the keyword group;

[0035] The speech generation submodule is used to set i=1, select the i-th related word from the multiple related words according to the related word sorting result, generate and play voice guidance information based on the i-th related word to guide the target user to supplement the voice intent. The related word sorting result is obtained by sorting the multiple target related words in descending order of relevance between the multiple target related words and the keyword group.

[0036] The associated word determination submodule is also used to add the keywords in the voice supplement instruction to the keyword group if the target user sends a voice supplement instruction, the voice intention of the target user cannot be determined by the voice supplement instruction, and the number of times the voice supplement instruction is sent is not greater than the number of times threshold, and then return to the step of determining multiple target associated words corresponding to the keyword group from the keyword association database.

[0037] The speech generation submodule is also used to select the i-th associated word from the multiple associated words if no speech supplementation instruction is received from the target user, and the multiple target associated words have not been traversed or the number of times the speech supplementation instruction is issued is not greater than the number threshold. In this case, if i = i + 1 is not received, the process of selecting the i-th associated word from the multiple associated words is returned according to the sorting result of the associated words.

[0038] Optionally, the associated term determines the submodule, specifically for:

[0039] From the keyword association database, determine multiple candidate related words and the relevance between these multiple candidate related words and the corresponding keywords in the keyword group. These multiple candidate related words refer to words that are related to the keywords in the keyword group.

[0040] Determine the logical relationships between the keywords in this keyword group;

[0041] Based on this logical relationship and the multiple candidate related words, the multiple target related words are determined;

[0042] Based on the relevance between the multiple candidate related terms and the corresponding keywords in the keyword group, the relevance between the multiple target related terms and the keyword group is determined.

[0043] Optionally, the device further includes a voice recommendation module, which is used for:

[0044] Supplementary related words are determined from the keyword association database. These supplementary related words refer to words that are related to the i-th related word.

[0045] Based on the supplementary related words, audio recommendation information is generated and played.

[0046] If no voice recommendation response is received from the target user, then execute the step of setting i = i + 1, returning the sorted results of the associated words, and selecting the i-th associated word from the multiple associated words.

[0047] Optionally, the associated term, which identifies the submodule, is also used for:

[0048] If a voice recommendation response from the target user is received, the keywords in the voice recommendation response are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database.

[0049] Optionally, the associated term, which identifies the submodule, is also used for:

[0050] Based on the target user's feedback on the voice guidance information, update the relevance between the i-th related word in the keyword association database and the keywords in the keyword group.

[0051] On the other hand, a computer device is provided, the computer device including a memory and a processor, the memory for storing computer programs, and the processor for executing the computer programs stored in the memory to implement the steps of the voice interaction method described above.

[0052] On the other hand, a computer-readable storage medium is provided, wherein a computer program is stored therein, and when the computer program is executed by a processor, it implements the steps of the above-described voice interaction method.

[0053] On the other hand, a computer program product containing instructions is provided, which, when executed on a computer, cause the computer to perform the steps of the voice interaction method described above.

[0054] The technical solution provided in this application can bring at least the following beneficial effects:

[0055] Upon receiving a voice command from a target user, keywords within the command are identified to obtain keyword groups, which are then used to determine the user's voice intent. Considering that a single voice command may not include all keywords for activating the target function, if the user's intent cannot be determined from the keyword group alone, voice guidance is provided based on the keywords included in the command. This guided interaction allows for the acquisition of additional voice input from the user, revealing the remaining keywords for activating the target function. If the user's intent can be determined from these additional inputs, the corresponding target function is activated. Therefore, even when a single voice command cannot determine the user's intent, guided interaction guides the user to provide additional voice input to confirm their intent and activate the corresponding target function. This improves the efficiency of voice interaction and enhances the user experience. Attached Figure Description

[0056] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0057] Figure 1 This is a schematic diagram of an implementation environment provided in an embodiment of this application;

[0058] Figure 2 This is a flowchart of a voice interaction method provided in an embodiment of this application;

[0059] Figure 3 This is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this application;

[0060] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0061] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the implementation methods of this application will be further described in detail below with reference to the accompanying drawings.

[0062] Before providing a detailed explanation of the voice interaction method provided in the embodiments of this application, the implementation environment involved in the embodiments of this application will be introduced first.

[0063] Please refer to Figure 1 , Figure 1 This is a schematic diagram illustrating an implementation environment according to an exemplary embodiment. The implementation environment includes a voice interaction terminal 101, a processor 102, and at least one functional module 103. The processor 102 can communicate with both the voice interaction terminal 101 and the functional module 103. This communication connection can be wired or wireless; this embodiment does not limit the specific connection.

[0064] The voice interaction terminal 101 is used to enable voice interaction with a user. For example, the voice interaction terminal 101 may include a microphone and a speaker. The microphone receives voice commands and supplementary voice intentions issued by the user, and the speaker provides voice guidance to the user.

[0065] The processor 102 is used to determine the user's voice intent based on the user's voice commands and voice supplementation results. For example, the processor 102 can determine the user's voice intent through semantic analysis based on the user's voice commands and voice supplementation results.

[0066] The processor 102 is also used to enable the function corresponding to the user's voice intent through the corresponding function module 103.

[0067] The processor 102 can be a general-purpose CPU (Central Processing Unit), NP (Network Processor), microprocessor, or one or more integrated circuits for implementing the scheme of this application, such as ASIC (Application-Specific Integrated Circuit), PLD (Programmable Logic Device), or a combination thereof. The aforementioned PLD can be CPLD (Complex Programmable Logic Device), FPGA (Field-Programmable Gate Array), GAL (Generic Array Logic), or any combination thereof.

[0068] Functional module 103 is used to implement corresponding functions. Depending on the application scenario of this application embodiment, the functional module is used to implement different functions. For example, when this application embodiment is applied to vehicle voice interaction, the functional module 103 can be an execution module for vehicle in-vehicle functions, so that different in-vehicle functions can be implemented through different execution modules. For instance, the functional module 103 may include the vehicle's air conditioning fan, used to turn the in-vehicle air conditioning on or off.

[0069] Those skilled in the art should understand that the above-described voice interaction terminal 101, processor 102, and functional module 103 are merely examples. Other existing or future voice interaction terminals, processors, or functional modules that are applicable to the embodiments of this application should also be included within the scope of protection of the embodiments of this application, and are hereby incorporated by reference.

[0070] It should be noted that the implementation environment described in the embodiments of this application is for the purpose of more clearly illustrating the technical solutions of the embodiments of this application, and does not constitute a limitation on the technical solutions provided in the embodiments of this application. As those skilled in the art will know, as the implementation environment evolves, the technical solutions provided in the embodiments of this application are also applicable to similar technical problems.

[0071] The voice interaction method provided in the embodiments of this application will now be explained in detail.

[0072] Figure 2This is a flowchart of a voice interaction method provided in an embodiment of this application, which is applied to the processor 102 described above. Please refer to... Figure 2 The method includes the following steps.

[0073] Step 201: Receive voice commands from the target user.

[0074] In some embodiments, the processor can be in a normally-on state, in which case the processor can receive voice commands issued by the target user in real time. In other embodiments, in order to reduce the power consumption of the processor and considering that the target user only has the intention to activate a specific function in a specific scenario, the working scenario for receiving voice commands issued by the target user can be limited so that the processor is in a sleep state in non-interactive mode, that is, it does not accept voice commands issued by the target user, and only accepts voice commands issued by the target user in interactive mode.

[0075] For example, when the processor is in sleep mode, it can monitor only the instruction to enter the interactive mode. Upon receiving the activation instruction from the target user (i.e., the instruction to enter the interactive mode), it enters the voice interactive mode and then begins to receive voice commands issued by the target user.

[0076] The activation command can be a button command or a voice command. For example, pressing a specific button can activate the voice interaction mode, or uttering a voice containing specific keywords (such as "Xiao A Xiao A", "Start voice interaction", etc.) can activate the voice interaction mode. The specific activation method of the voice interaction mode can be selected based on actual usage needs, and this application embodiment does not limit it.

[0077] Step 202: Identify the keywords in the voice command to obtain the keyword group.

[0078] In some embodiments, voice commands can be converted into text, and then semantic segmentation and feature extraction can be performed on the text based on text analysis to determine the keywords in the voice commands; alternatively, the keywords in the voice commands can be determined directly based on semantic analysis and feature extraction of the voice.

[0079] It should be noted that a voice command can contain multiple keywords, and then a keyword group can be formed based on the multiple keywords contained in the voice command.

[0080] For example, if the target user issues the voice command "open the car window", based on keyword recognition, we can obtain the two keywords "open" and "car window", resulting in the keyword phrase "open, car window".

[0081] It is understandable that in some scenarios, the voice commands issued by the target user may include multiple virtual words (i.e., words unrelated to the intended function). When identifying keywords in voice commands, these virtual words can be filtered out through semantic analysis. For example, the voice commands issued by the target user are "please open the car window" or "open the car window briefly." After keyword extraction, we can obtain words such as "open," "car window," "please," and "briefly." However, after semantic analysis, we can determine that words such as "please" and "briefly" are unrelated to the intended function in the user's voice commands. Therefore, we can identify the words "open" and "car window" as keywords in the voice commands, resulting in the keyword combination "open, car window."

[0082] Step 203: If the target user's voice intent cannot be determined through the keyword group, then the target user is given voice guidance based on the keyword group to allow the target user to supplement the voice intent.

[0083] It should be noted that the target user's voice intent can be understood as the intention to perform a function that points to a unique and specific function. For example, in semantic analysis, determining the target user's voice intent requires ensuring that the keyword group includes keywords of the categories of "action," "object," and "attribute." This can be understood as the keyword group needing to fill the action slots, object slots, and attribute slots. Action keywords are used to indicate the execution method of the schematic diagram (function), object keywords are used to indicate the execution object of the schematic diagram (function), and attribute keywords are used to limit the "action" and "object" to ensure that the intent (function) is a unique and accurate intent.

[0084] In some embodiments, if the keyword group in the voice command can fill all the action slots, object slots, and attribute slots, it is determined that the target user's voice intent can be determined through the keyword group; if the keyword group in the voice command cannot fill any of the action slots, object slots, and attribute slots, it is determined that the target user's voice intent cannot be determined through the keyword group. Here, the ability of the keyword group to fill all the action slots, object slots, and attribute slots can be understood as the attribute slots in the keyword group constraining the action slots and object slots, ensuring that the constrained actions and objects point to a unique function.

[0085] It should be noted that in some voice commands, the keyword phrase can be made to point to a unique function simply by limiting the action or object. For example, in voice commands such as turning on the headlights or opening the sunroof, the keyword phrase can be made to point to a unique function simply by limiting the position of the headlights (such as high beams, low beams, interior lights, etc.) and the degree of opening of the sunroof (such as fully open, half open, etc.).

[0086] In some embodiments, when the object is unique and the degree of execution of the object's action is also unique, it may not be necessary to include attribute words in the keyword group to determine the target user's voice intent, such as turning on the windshield wipers.

[0087] In some embodiments, the type of keyword can be determined based on the position of each keyword in the keyword group in the voice command and its part of speech.

[0088] For example, the keyword phrase "open, car window" indicates that the target user's initial intention is to open the car window. However, since the vehicle has four windows and the degree of opening varies, and the instruction does not specify the location or degree of opening of the window, the keyword phrase "open, car window" cannot point to a single specific function. In other words, the target user's voice intent cannot be determined by the keyword phrase "open, car window".

[0089] For example, the keyword phrase is "open, driver's seat, half, car window". Based on this keyword phrase, we can obtain the target user's initial intention as to open the driver's seat car window and open the window halfway. Since the keyword phrase "open, driver's seat, half, car window" can point to a single specific function, we can assume that the target user's voice intention can be determined through this keyword phrase.

[0090] In some embodiments, such as the example above, if the target user's voice intent can be determined through keyword groups, the target function corresponding to that voice intent is directly activated.

[0091] In some embodiments, when the voice intent of the target user cannot be determined by keyword groups, voice guidance for the target user based on keyword groups can be achieved through the following steps (1)-(4).

[0092] (1) Determine the multiple target related words corresponding to the keyword group from the keyword association database, and the relevance between the multiple target related words and the keyword group.

[0093] A keyword association database can be understood as a word database (lexicon). This database contains multiple words and the relationships between them. These relationships can be understood as the connections between words within a given voice command. For example, taking "play" as a verb, the related words could include object words such as "music," "opera," "crosstalk," "talk show," and "news"—the objects that implement the action of "playing"—as well as attribute words such as "volume level" and "time"—the attribute qualifiers for the action of "playing."

[0094] In some embodiments, the keyword association databases corresponding to different application scenarios may differ, and the keyword association databases corresponding to different scenarios can be determined based on different data platforms. For example, for in-vehicle voice interaction devices, the keyword association database during the in-vehicle voice interaction process can be obtained based on the in-vehicle voice interaction data platform, which may include voice interaction data from multiple in-vehicle devices.

[0095] In some embodiments, the keyword association databases corresponding to different users in the same application scenario may also be different. An initial keyword association database can be obtained based on the data platform, and then as the target user uses the voice interaction function, the words in the keyword association database and the relationships between the words can be modified to obtain the keyword association database corresponding to the target user.

[0096] The target related words corresponding to the keyword group can be understood as the words corresponding to the missing slots when the target user's voice intent cannot be determined through the keyword group, that is, when the keywords in the keyword group cannot fill all the slots.

[0097] For example, if the voice command issued by the target user is "open the car window", based on semantic analysis, it can be seen that the voice command lacks attribute qualifiers for "open" and "car window". Therefore, based on the keyword association database, multiple attribute qualifiers corresponding to the keyword group "open" and "car window" are determined, and these multiple attribute qualifiers are used as target related words (such as location, degree).

[0098] For example, if the voice command issued by the target user only includes "open," semantic analysis reveals that the voice command lacks the attribute and object qualifiers for "open." Therefore, based on a keyword association database, multiple attribute and object qualifiers corresponding to "open" are determined. These multiple attribute and object qualifiers are then used as target related terms (such as object, degree).

[0099] In some embodiments, multiple candidate related words and their relevance to the corresponding keywords in the keyword group can be determined from the keyword association database. The multiple candidate related words refer to words that are associated with the keywords in the keyword group. The logical relationship between the keywords in the keyword group can be determined. Based on the logical relationship and the multiple candidate related words, multiple target related words can be determined. Based on the relevance between the multiple candidate related words and the corresponding keywords in the keyword group, the relevance between the multiple target related words and the keyword group can be determined.

[0100] It should be noted that since the voice commands issued by the target user usually include multiple keywords, it is necessary to determine the logical relationship between the keywords. Different logical relationships result in completely different outcomes. In order to more accurately understand the target user's voice intent, it is necessary to analyze the logical relationship between the keywords in order to determine the target related words corresponding to the keyword group.

[0101] Logical relationships typically include "AND", "OR", and "NOT". For example, a target user's voice command may include the keyword "A" and the keyword "B", but "A AND B", "A OR B", and "A NOT B" express completely different meanings.

[0102] In some embodiments, if there is no "OR" / "NOT" logical connector between keyword "A" and keyword "B" in the voice command, the logical relationship between keyword "A" and keyword "B" can generally be considered as an "AND" relationship. In this case, the intersection of multiple candidate related words corresponding to keyword "A" and multiple candidate related words corresponding to keyword "B" can be obtained to obtain the target related word corresponding to the keyword group "A, B". That is, when the logical relationship between keyword "A" and keyword "B" is an "AND" relationship, the intersection of multiple candidate related words corresponding to keyword "A" and multiple candidate related words corresponding to keyword "B" can be taken as the target related word corresponding to the keyword group "A, B".

[0103] For example, if the keyword group obtained from the voice command issued by the target user is "open, car window", then it is necessary to perform an intersection operation on the candidate related words corresponding to "open" and "car window" to obtain words that are related to both "open" and "car window" as target related words (such as: location, degree).

[0104] Because the voice command "open the car window" lacks attribute limitations for both "open" and "car window," it's necessary to determine the attribute limitations related to "car window" within the attribute limitations of "open," and the attribute limitations related to "open" within the attribute limitations of "car window," thus obtaining target related words that are associated with both "open" and "car window." Taking "open" as an example, the keyword "open" can have multiple attribute limitation types in the keyword association database, such as "direction," "degree," and "frequency." However, since only "degree" is related to "car window," "degree" is determined as the target related word.

[0105] In some embodiments, if there is a logical connector with an "or" meaning between keyword "A" and keyword "B" (such as words like "or" or "or"), the logical relationship between keyword "A" and keyword "B" can generally be considered as an "or" relationship. In this case, the union of multiple candidate related words corresponding to keyword "A" and multiple candidate related words corresponding to keyword "B" can be obtained to obtain the target related words corresponding to the keyword group "A, B".

[0106] It should be noted that the "or" relationship mainly appears in scenarios where the target user may have a relatively casual attitude towards the implementation of some functions.

[0107] When the logical relationship between keyword "A" and keyword "B" is "OR", multiple candidate related words corresponding to keyword "A" and multiple candidate related words corresponding to keyword "B" can be used as the target related words corresponding to the keyword group "A, B". In other words, the union of the candidate related words corresponding to keyword "A" and the candidate related words corresponding to keyword "B" is taken as the target related words corresponding to the keyword group "A, B".

[0108] For example, in a scenario where a target user needs to navigate to a restaurant, the target user may not have a specific preference for the location. For instance, if the target user's voice command is "go eat Sichuan cuisine or Hunan cuisine," it can be determined that the target user's initial intention is either "go eat Sichuan cuisine" or "go eat Hunan cuisine." Therefore, based on logical AND operations, multiple first-level related words corresponding to "go eat Sichuan cuisine" and multiple second-level related words corresponding to "go eat Hunan cuisine" can be determined. Then, based on logical OR operations, the first-level related words and the second-level related words can be combined to obtain the target related words corresponding to "go, eat, Sichuan cuisine or Hunan cuisine."

[0109] Semantic analysis reveals that the action in the voice command is "go," the object is "restaurant (i.e., eat)," and the attribute of "eat" is limited to "Sichuan cuisine or Hunan cuisine." Therefore, a logical OR operation can be performed on the results of the limitation of "eating" by "Sichuan cuisine" (multiple Sichuan restaurants) and the results of the limitation of "eating" by "Hunan cuisine" (multiple Hunan restaurants) to obtain multiple restaurants with the cuisine attribute of Sichuan or Hunan. However, due to the lack of attribute limitation for "go," and the fact that the attribute limitation for "eating" is only cuisine-based, multiple limitation results are obtained. Further limitation of "eating" is needed through other attribute limitations. Therefore, it is necessary to determine the attribute limitations related to "eating" in the attribute limitations of "go" (such as including "method"), and the attribute limitation types related to "go" in the attribute limitation types of "eating" (such as including "average consumption per person" and "type"), thereby obtaining target related words that are associated with "go," "eating," and "Sichuan cuisine or Hunan cuisine."

[0110] For example, the target related terms obtained are "method", "average consumption per person", and "type", where "method" is an attribute limitation for "go", and "average consumption per person" and "type" are attribute limitations for "eat".

[0111] In some embodiments, specific attributes can be set to default values ​​based on the voice interaction scenario. For example, in the above example, when the voice interaction in this application embodiment is applied to the scenario of in-vehicle voice interaction, the default "method" attribute of "go" is limited to "driving".

[0112] In some embodiments, to ensure the accuracy of voice interaction results, when the voice command issued by the target user contains an "or" logical relation word, it is necessary to ensure that the keywords before and after the logical relation word are similar (such as keywords of the same type), such as "open the rear window or the front window" or "play music or opera", that is, the initial intent of the voice command is basically clear (open the window, play entertainment programs, etc.); if the keywords before and after the logical relation word are unrelated keywords, such as "open the window or the shopping mall" or "play music or the restaurant", since the keywords before and after the logical relation word are unrelated keywords, such as the keyword group composed of "open" and "shopping mall" and the keyword group composed of "play" and "restaurant", it is impossible to determine the user's initial intent corresponding to the voice command. At this time, a confirmation voice can be generated and played to the target user, such as "I didn't hear you clearly, please say it again" or "Cannot be executed at present, please resend the voice".

[0113] In some embodiments, if the logical relationship between the keyword "A" or keyword "B" and the logical connectors containing the meaning of "NOT" (such as "except A", "except B", "don't A", etc.) is "NOT", then it is necessary to calculate the difference between the multiple candidate related words corresponding to the keyword "A" and the multiple candidate related words corresponding to the keyword "B" to obtain the target related words corresponding to the keyword group "A, B". The specific difference operation method needs to be determined in combination with the position of the logical connector.

[0114] When the logical relationship between keyword "A" and keyword "B" is "NOT" and the logical meaning is "A is not B", the candidate related words that do not correspond to keyword "B" among the multiple candidate related words corresponding to keyword "A" can be taken as the target keywords corresponding to the keyword group "A, B". In other words, the candidate related words that only correspond to keyword "A" and not to keyword "B" are taken as the target keywords corresponding to the keyword group "A, B". Or, the target keywords corresponding to the keyword group "A, B" include the candidate related words corresponding to keyword "A" but do not include the candidate related words corresponding to keyword "B".

[0115] When the logical relationship between keyword "A" and keyword "B" is "NOT" and the logical meaning is "B is not A", the candidate related words that do not correspond to keyword "A" among the multiple candidate related words corresponding to keyword "B" can be taken as the target keywords corresponding to the keyword group "A, B". In other words, the candidate related words that only correspond to keyword "B" and not keyword "A" are taken as the target keywords corresponding to the keyword group "A, B". Or, the target keywords corresponding to the keyword group "A, B" include the candidate related words corresponding to keyword "B" but do not include the candidate related words corresponding to keyword "A".

[0116] It should be noted that logical "NOT" relationships are typically used to describe how attribute keywords limit action keywords or object keywords. Therefore, the specific logical meaning can be determined based on the position of the "NOT" logical connector and the contextual information of the keywords. For example, when the logical connector "besides" appears in a voice command, the keywords preceding the logical connector need to be combined to limit the attribute of the action keyword or object keyword. Furthermore, for ease of understanding, the above explanation uses keywords "A" and "B" as examples. In some embodiments, keyword "A" can also be called the first keyword, and keyword "B" can also be called the second keyword.

[0117] For example, taking "go to a restaurant" as an example, if the user's voice command is "go to a restaurant other than Sichuan cuisine", then based on semantic analysis, we can obtain that the target user's initial intention is "go to a restaurant" but not "Sichuan cuisine restaurant". Therefore, based on logical AND operation, we can determine multiple first-related words corresponding to "go to a restaurant" and multiple second-related words corresponding to "Sichuan cuisine restaurant". Then, based on logical NOT operation, we can perform a difference operation on the first-related words and the second-related words to obtain the target related words corresponding to "go, other than Sichuan cuisine, restaurant".

[0118] Semantic analysis reveals that the action in the voice command is "go," the object is "restaurant," and the attribute of "restaurant" is limited to "excluding Sichuan cuisine." Therefore, a logical NOT operation can be performed between the object result of "restaurant" (all restaurants) and the result of the limitation of "restaurant" by "Sichuan cuisine" (multiple Sichuan restaurants) to obtain restaurants with the cuisine attribute of "non-Sichuan cuisine." However, due to the lack of attribute limitation for "go," and the fact that the attribute limitation for "restaurant" is only cuisine limitation, multiple objects are obtained. Further limitation of "restaurant" is needed through other attribute limitations. Therefore, it is necessary to determine the attribute limitations of "go" related to "restaurant" (such as including "method"), and the attribute limitation types of "restaurant" related to "go" (such as including "average consumption per person" and "type"), thereby obtaining target related words that are associated with "go," "excluding Sichuan cuisine," and "restaurant."

[0119] For example, the target related terms obtained are "method", "average consumption per person", and "size", where "method" is an attribute limitation for "go", and "average consumption per person" and "type" are attribute limitations for "restaurant".

[0120] (2) Let i = 1. According to the ranking result of the related words, select the i-th related word from the multiple related words, generate and play voice guidance information based on the i-th related word to guide the target user to supplement the voice intent. The ranking result of the related words is obtained by ranking the multiple target related words in descending order of the relevance between the multiple target related words and the keyword group.

[0121] For example, for the voice command "open the car window", the first associated word can be "location". Then, based on "location", the voice guidance information "which location" is generated and played to guide the target user to supplement the voice intent (such as "all locations", "rear left", etc.).

[0122] In some embodiments, in order to make the voice guidance information more intelligent, the voice guidance information can be generated based on related words and keyword phrases. For example, in the above example of "open the car window", the voice guidance information can also be "Which window do you want to open?" etc.

[0123] Relevance can be understood as a correlation parameter between words. For any given word, since there are usually multiple words that are related to it, the degree of association between words can be described based on the relevance parameter in order to select and determine the target related words. For example, taking "restaurant" as an example, its corresponding related words can include "cuisine", "average consumption per person", "type", "distance", etc. The degree of association between the above words and "restaurant" can be described based on the relevance parameter. For example, the relevance parameter between "restaurant" and "cuisine" can be 80, the relevance parameter between "restaurant" and "average consumption per person" can be "85", the relevance parameter between "restaurant" and "type" can be 90, and so on.

[0124] Furthermore, based on the correlation between multiple keywords in the keyword group and a certain related word, the correlation between the certain related word and the keyword group is obtained. The calculation method of the correlation between the related word and the keyword group can be determined based on actual usage needs, and this application embodiment does not limit it.

[0125] For example, if the keyword phrase includes "go" and "restaurant", then the correlation between "cuisine" and the keyword phrase can be obtained (e.g., 45) based on the correlation between "cuisine" and "go" (e.g., 10) and the correlation between "cuisine" and "restaurant" (e.g., 80).

[0126] In some embodiments, in order to improve the accuracy of voice guidance information, related words can be filtered by the relevance between related words and keyword groups. For example, when the relevance between a related word and a keyword group is lower than the weak relevance threshold (e.g., lower than 10), the related word is considered to have a weak relevance to the keyword group, and the related word is rejected as the target related word corresponding to the keyword group.

[0127] In some embodiments, in order to improve the efficiency of voice interaction with the target user, voice guidance information can be generated based on multiple related words. For example, taking the voice command "go to a restaurant" as an example, if the determined target related words are ordered from most to least related in order of relevance as method, type, cuisine, average cost per person, etc., then the voice guidance information "How do you want to go, what type of restaurant do you want to go to?" can be generated based on a fixed number of related words, such as the two related words "distance" and "type". The number of related words can be determined based on specific usage needs.

[0128] In some embodiments, the relevance between the i-th associated word and the keywords in the keyword group in the keyword association database can be updated based on the target user's feedback on the voice guidance information.

[0129] It should be noted that after generating and playing voice guidance information based on related words, if a voice supplement instruction is received from the target user, it indicates that the related word (attribute) is a related word that the target user believes needs to be expressed, that is, an attribute limitation that the target user is more concerned about; if no voice supplement instruction is received from the target user, it indicates that the related word is a related word that the target user does not care about.

[0130] For example, the keyword phrase is "go, eat", the corresponding related word is "cuisine", and the generated and played voice guidance information is "what cuisine do you want to eat?" If the target user does not issue a voice supplementary instruction based on the voice guidance information, it indicates that the target user is relatively casual about the choice of cuisine and does not think that there is a need to restrict the cuisine. If the target user issues a voice supplementary instruction based on the voice guidance information, such as issuing the instruction "I want to eat Sichuan cuisine", it indicates that the target user has certain requirements for the choice of cuisine when eating.

[0131] Therefore, in order to make the generated voice guidance information more closely match the needs of the target users, the relevance between related words and keywords can be updated based on the target users' feedback on the voice guidance information. For example, in the above example, if the target user does not issue a voice supplementary command based on the voice guidance information after generating and playing "What cuisine do you want to eat?", the relevance between "cuisine" and the keyword group "go, eat" is reduced; if the target user issues a voice supplementary command based on the voice guidance information, "I want to eat Sichuan cuisine", the relevance between "cuisine, Sichuan cuisine" and the keyword group "go, eat" is increased.

[0132] In some embodiments, in order to further improve the efficiency of determining the voice intent of the target user and the intelligence of voice interaction, as the target user uses the product, when the correlation between the related word and the keyword group is high, the related word can be directly identified as the implicit keyword of the keyword group.

[0133] Taking the voice supplementary command "I want to eat Sichuan cuisine" as an example, it can be understood that as the target user's usage time increases, if the target user's voice supplementary command is "I want to eat Sichuan cuisine" after generating and playing "What cuisine do you want to eat?" multiple times, the correlation between "Sichuan cuisine" and the keyword group "go, eat" will continue to increase. If this correlation is greater than the strong correlation threshold (e.g., greater than 80), and the target user's voice command is "go to eat", then the keyword group corresponding to the voice command is determined to be "go, eat, Sichuan cuisine".

[0134] (3) If a voice supplement instruction is received from the target user, the voice intention of the target user cannot be determined through the voice supplement instruction, and the number of times the voice supplement instruction is issued is not greater than the number threshold, then the keywords in the voice supplement instruction are added to the keyword group, and the step of determining the multiple target related words corresponding to the keyword group from the keyword association database is returned.

[0135] It should be noted that the voice interaction process in this application embodiment can be understood as follows: when the functional intent composed of keywords in the target user's voice command is not unique, the target user is guided by voice guidance information, and then the keywords in the initial voice command are further limited based on the keywords in the supplementary voice command issued by the target user, until all the keywords in the voice issued by the target user in the voice interaction can form a unique functional intent (i.e., voice intent), thereby accurately realizing the functional intent.

[0136] Therefore, when a voice supplement instruction is received from the target user, and the target user's accurate voice intent cannot be determined through the voice supplement instruction, the keywords in the voice supplement instruction need to be added to the keyword group (equivalent to treating both the voice supplement instruction and the voice instruction issued by the target user as the target user's expression of voice intent), and the process returns to the step of determining multiple target related words corresponding to the keyword group based on the keyword association database.

[0137] It should be noted that, in the embodiments of this application, the above examples are mainly used to illustrate relatively complex usage scenarios with many selection conditions (such as "going to eat" mentioned above). Under normal circumstances, the voice intent of the target user is usually a relatively simple function implementation intent. For example, in the scenario of in-vehicle voice interaction, the user's voice intent is usually to turn on or off a certain in-vehicle function, such as turning on the windows, air conditioning, and windshield wipers, playing music, or making a phone call. The attribute limitations of such in-vehicle functions are usually relatively simple, and a unique and accurate function intent can be obtained through a few voice interactions.

[0138] Therefore, to avoid reducing the user experience of the target user due to excessive voice interaction in some scenarios, a threshold can be set for the number of times voice supplementary commands can be issued. If the number of times the specified voice supplementary commands are issued exceeds the threshold, although it may still be impossible to determine the unique functional intent based on the voice supplementary commands and the voice commands, in order to ensure the user experience of the target user, no more voice guidance information will be generated. Instead, the most matching voice intent will be determined directly based on all the voice supplementary commands issued by the target user and the keywords in the voice commands, and the target function corresponding to the voice intent will be activated.

[0139] For example, taking the voice command "Go eat" as an example, assuming the threshold for the number of times is 3, after issuing 3 supplementary voice commands, we get "Sichuan cuisine", "shopping mall" and "average cost per person is 100 yuan". At this point, although we still cannot get a unique restaurant, in order to ensure the user experience, we directly determine a target restaurant based on the attribute constraints of "Sichuan cuisine", "shopping mall" and "average cost per person is 100 yuan", and generate and display navigation information to the target restaurant. The target restaurant can be randomly generated within the current constraints, or it can be generated based on specific conditions (such as selecting the nearest restaurant, the highest-rated restaurant, etc.). Specifically, it can be selected in combination with the application scenario and conditions such as the next related words.

[0140] (4) If no voice supplementation instruction is received from the target user, and the multiple target related words have not been traversed or the number of times the voice supplementation instruction is issued is not greater than the number threshold, then let i = i + 1, return the step of selecting the i-th related word from the multiple related words according to the sorting result of the related words.

[0141] Based on the above description, if no supplementary voice command is received from the target user after the voice guidance information is played, it usually indicates that the content of the associated word is something that the target user does not care about. At this time, the next associated word can be selected based on the associated word sorting result (i.e., let i = i + 1), and the voice guidance information can be generated and played based on the next associated word.

[0142] In some embodiments, considering that there is a certain response time for the target user's voice supplementary command, it is usually necessary to wait for a certain period of time to determine whether the target user has issued a voice supplementary command. In order to avoid waiting for a long time for the target user to issue a voice supplementary command, a timer can be started after playing the voice guidance information. If the duration of not receiving the voice supplementary command issued by the target user is greater than a first duration threshold, it is considered that the voice supplementary command issued by the target user has not been received. The first duration threshold can be determined in combination with the specific use case, such as 3 seconds, 5 seconds, etc.

[0143] In some embodiments, considering that the target user's voice intent may change, such as the target user no longer wishing to enable the target function after issuing a voice command due to changes in time or scenario, the voice interaction can be terminated after generating and playing the voice guidance information, and the voice interaction can be terminated when the accumulated duration is greater than or equal to the second duration threshold.

[0144] For example, taking a first duration threshold of 5 seconds and a second duration threshold of 15 seconds as an example, after playing the voice guidance information, a timer is started. If no supplementary voice instruction from the target user is received within five seconds, the voice guidance information is regenerated and played based on the next related word, and a timer is started again. If no supplementary voice instruction from the target user is received within five seconds, the cumulative waiting time is counted as 10 seconds. Then, the voice guidance information is regenerated and played again based on the next related word, and a timer is started again. If no supplementary voice instruction from the target user is received within 5 seconds, the cumulative waiting time is counted as 15 seconds, and the voice interaction ends directly.

[0145] In some embodiments, before returning the sorted results of the related words and selecting the i-th related word from the plurality of related words, supplementary related words can be determined from the keyword association database. The supplementary related words refer to words that are related to the i-th related word. Based on the supplementary related words, voice recommendation information is generated and played. If no voice recommendation response is received from the target user, the step of setting i = i + 1 is executed, and the sorted results of the related words are returned to select the i-th related word from the plurality of related words.

[0146] In some scenarios, if no supplementary voice command is received from the target user after the voice guidance information is played, it may be because the target user does not know how to select or what content is included. To avoid misjudging the target user's intent, if no command is received from the target user, supplementary related words corresponding to the identified related words can be generated and played first. If no voice recommendation response is received from the target user after playing the voice recommendation information, since the target user has not issued a supplementary voice command for the related word and has not issued a voice recommendation response for the supplementary related word, it can be assumed that the content of the related word is content that the target user does not care about.

[0147] Among them, supplementary related words can be understood as the related words in the keyword association database that are most relevant to the related word, or as the subordinate supplementary words of the related word.

[0148] For example, the i-th related word can be a cuisine, and the generated voice guidance information is "What cuisine do you want to eat?". If no supplementary voice instruction is received from the target user, supplementary related words (such as Sichuan cuisine, Shaanxi cuisine, Hunan cuisine, Shandong cuisine, etc.) can be determined from the keyword association database based on the related word (cuisine). Based on the supplementary related words, the voice recommendation information "Do you want to eat Sichuan cuisine?" is generated and played.

[0149] In some embodiments, voice recommendation information can also be generated and played based on multiple supplementary keywords. For example, in the above example, the voice recommendation information could also be "Do you want to eat Sichuan cuisine, Shaanxi cuisine, or Hunan cuisine?".

[0150] In some embodiments, if a voice recommendation response from the target user is received, the keywords in the voice recommendation response are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database.

[0151] In some embodiments, the voice recommendation response can include three types: an affirmative response, a negative response, and an update response. Using the example above, if the voice recommendation information is "Do you want to eat Sichuan cuisine, Shaanxi cuisine, or Hunan cuisine?", and the target user's voice recommendation response is "I want to eat Sichuan cuisine," which is an affirmative response, then "Sichuan cuisine" is added to the keyword group, and the logical relationship between "restaurant" and this keyword is an "AND" relationship. The process then returns to the step of determining multiple target related words corresponding to this keyword group from the keyword association database. If the target user's voice recommendation response is "except Sichuan cuisine" or "I don't want to eat Sichuan cuisine," which is a negative response, then "Sichuan cuisine" is added to the keyword group, and the logical relationship between "restaurant" and this keyword is a "NOT" relationship. If the target user's voice recommendation response is "Change to another one" or "Are there any others?", it is an update response. In this case, based on the related word "cuisine", supplementary related words that have not been used to generate voice recommendation information (such as "Shandong cuisine", "Cantonese cuisine", etc.) are determined from the keyword association database. Based on these supplementary related words, voice recommendation information is regenerated and played, such as "Do you want to eat Shandong cuisine?".

[0152] It should be noted that, since the voice commands, supplementary voice commands, and recommended voice responses issued by the target user during voice interaction mainly rely on the target user's subjective awareness and are highly random, in some embodiments, it is sufficient as long as the keywords in the target user's supplementary voice commands and recommended voice commands are related to the keywords in the keyword group.

[0153] For example, if the target user's voice command is "go eat", the voice guidance information generated based on the keyword group's related words is "what kind of cuisine do you want to eat", and the target user's supplementary voice command is "I want to go to the mall to eat", or "anything is fine, I want to go to the mall to eat", then "mall" can be added to the keyword group, and the steps of determining the multiple target related words corresponding to the keyword group from the keyword association database can be returned.

[0154] For example, if the voice recommendation message is "Do you want to eat Sichuan cuisine?", and the target user's voice recommendation response is "I want to eat Cantonese cuisine", then "Cantonese cuisine" can be added to the keyword group, and the steps to determine the multiple target related words corresponding to the keyword group from the keyword association database can be returned.

[0155] Step 204: If the target user's voice intent can be determined through the voice supplementation results, then activate the target function corresponding to the voice intent.

[0156] The voice supplement result can be understood as the result of the voice command issued by the target user and the voice intent supplement issued based on the voice guidance of the target user (such as the voice supplement command mentioned above).

[0157] In some embodiments, it can be determined whether the target user's voice intent can be determined based on the keywords in the keyword group and the keywords in the voice intent supplement. The specific determination of whether the target user's voice intent can be combined with the relevant description at step 203 above in the embodiments of this application, which will not be repeated here.

[0158] This application provides a voice interaction method. When the target user's voice intent cannot be determined by the keyword group obtained from the voice command, a keyword association database is used to determine the target related words corresponding to the keyword group. Based on the relevance between the target related words and the keyword group, voice guidance messages are generated and played sequentially. These messages guide the target user to provide supplementary voice input. Different guidance processes are executed based on the target user's different responses to the voice guidance messages, until the target user's voice intent can be determined based on the supplementary voice input, and the target function corresponding to the voice intent is activated. Furthermore, to avoid frequent voice interaction affecting the user experience, the number of times the target user issues supplementary voice commands is limited. Voice interaction is performed in a loop as long as the number of supplementary voice commands does not exceed a threshold, thus improving the efficiency of voice interaction while ensuring a good user experience. Moreover, the relevance between related words and keyword groups is updated based on the target user's feedback on the voice guidance information. As the number of times the target user uses the information increases, the keyword association database is updated, improving its timeliness and enhancing the intelligence of the voice interaction.

[0159] Figure 3 This is a schematic diagram of the structure of a voice interaction device provided in an embodiment of this application. The voice interaction device can be implemented as part or all of a voice interaction equipment by software, hardware, or a combination of both. The voice interaction equipment can be... Figure 1 The processor shown. Please refer to... Figure 3 The device includes: a voice receiving module 301, a voice recognition module 302, a voice guidance module 303, a function activation module 304, and a voice recommendation module 305.

[0160] The voice receiving module 301 is used to receive voice commands issued by the target user;

[0161] The speech recognition module 302 is used to recognize keywords in the speech command and obtain keyword groups;

[0162] The voice guidance module 303 is used to provide voice guidance to the target user based on the keyword group if the target user's voice intent cannot be determined through the keyword group, so that the target user can supplement the voice intent.

[0163] The function activation module 304 is used to activate the target function corresponding to the voice intent if the voice intent of the target user can be determined through the voice supplementation results of the target user.

[0164] Optionally, the voice guidance module 303 includes:

[0165] The related term determination submodule is used to determine multiple target related terms corresponding to the keyword group from the keyword association database, as well as the relevance between the multiple target related terms and the keyword group;

[0166] The speech generation submodule is used to set i=1, select the i-th related word from the multiple related words according to the related word sorting result, generate and play voice guidance information based on the i-th related word to guide the target user to supplement the voice intent. The related word sorting result is obtained by sorting the multiple target related words in descending order of relevance between the multiple target related words and the keyword group.

[0167] The associated word determination submodule is also used to add the keywords in the voice supplement instruction to the keyword group if the target user sends a voice supplement instruction, the voice intention of the target user cannot be determined by the voice supplement instruction, and the number of times the voice supplement instruction is sent is not greater than the number of times threshold, and then return to the step of determining multiple target associated words corresponding to the keyword group from the keyword association database.

[0168] The speech generation submodule is also used to select the i-th associated word from the multiple associated words if no speech supplementation instruction is received from the target user, and the multiple target associated words have not been traversed or the number of times the speech supplementation instruction is issued is not greater than the number threshold. In this case, if i = i + 1 is not received, the process of selecting the i-th associated word from the multiple associated words is returned according to the sorting result of the associated words.

[0169] Optionally, the related term determination submodule is specifically used for:

[0170] From the keyword association database, determine multiple candidate related words and the relevance between these multiple candidate related words and the corresponding keywords in the keyword group. These multiple candidate related words refer to words that are related to the keywords in the keyword group.

[0171] Determine the logical relationships between the keywords in this keyword group;

[0172] Based on this logical relationship and the multiple candidate related words, the multiple target related words are determined;

[0173] Based on the relevance between the multiple candidate related terms and the corresponding keywords in the keyword group, the relevance between the multiple target related terms and the keyword group is determined.

[0174] Optionally, the voice recommendation module 305 is used to determine supplementary related words from the keyword association database, where the supplementary related words refer to words that are related to the i-th related word;

[0175] Based on the supplementary related words, audio recommendation information is generated and played.

[0176] If no voice recommendation response is received from the target user, then execute the step of setting i = i + 1, returning the sorted results of the associated words, and selecting the i-th associated word from the multiple associated words.

[0177] Optionally, the associated term, which identifies the submodule, is also used for:

[0178] If a voice recommendation response from the target user is received, the keywords in the voice recommendation response are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database.

[0179] Optionally, the associated term, which identifies the submodule, is also used for:

[0180] Based on the target user's feedback on the voice guidance information, update the relevance between the i-th related word in the keyword association database and the keywords in the keyword group.

[0181] In this embodiment, when the target user's voice intent cannot be determined by the keyword group obtained from the voice command, a keyword association database is used to determine the target related words corresponding to the keyword group. Based on the relevance between the target related words and the keyword group, voice guidance messages are generated and played sequentially. These messages guide the target user to provide additional voice information. Different guidance processes are executed based on the target user's feedback to these messages, until the target user's voice intent can be determined from their additional voice information, and the corresponding target function is activated. Furthermore, to avoid frequent voice interaction affecting the user experience, the number of times the target user issues additional voice commands is limited. Voice interaction is performed in a loop as long as the number of additional voice commands does not exceed a threshold, improving efficiency while maintaining a positive user experience. Additionally, the relevance between related words and keyword groups is updated based on the target user's feedback on the voice guidance information. This updates the keyword association database as the user's usage increases, improving its timeliness and enhancing the intelligence of the voice interaction.

[0182] It should be noted that the voice interaction device provided in the above embodiments is only illustrated by the division of the above functional modules when implementing voice interaction. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the voice interaction device and the voice interaction method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.

[0183] Figure 4 This is a structural block diagram of an electronic device 400 provided in an embodiment of this application. The electronic device 400 can be a portable mobile electronic device, such as a smartphone, tablet computer, MP3 player (Moving Picture Experts Group Audio Layer III), MP4 player (Moving Picture Experts Group Audio Layer IV), laptop computer, or desktop computer. The electronic device 400 may also be referred to as a user device, portable electronic device, laptop electronic device, desktop electronic device, or other names.

[0184] Typically, electronic device 400 includes a processor 401 and a memory 402.

[0185] Processor 401 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. Processor 401 may be implemented using at least one hardware form selected from DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), and PLA (Programmable Logic Array). Processor 401 may also include a main processor and a coprocessor. The main processor, also known as a CPU (Central Processing Unit), is used to process data in the wake-up state; the coprocessor is a low-power processor used to process data in the standby state. In some embodiments, processor 401 may integrate a GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the screen. In some embodiments, processor 401 may also include an AI (Artificial Intelligence) processor, which is used to handle computational operations related to machine learning.

[0186] The memory 402 may include one or more computer-readable storage media, which may be non-transitory. The memory 402 may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory 402 are used to store at least one instruction, which is executed by the processor 401 to implement the voice interaction method provided in the method embodiments of this application.

[0187] In some embodiments, the electronic device 400 may optionally include a peripheral device interface 403 and at least one peripheral device. The processor 401, memory 402, and peripheral device interface 403 can be connected via a bus or signal line. Each peripheral device can be connected to the peripheral device interface 403 via a bus, signal line, or circuit board. Specifically, the peripheral device includes at least one of the following: a radio frequency circuit 404, a display screen 405, a camera assembly 406, an audio circuit 407, a positioning assembly 408, and a power supply 409.

[0188] Peripheral device interface 403 can be used to connect at least one I / O (Input / Output) related peripheral device to processor 401 and memory 402. In some embodiments, processor 401, memory 402 and peripheral device interface 403 are integrated on the same chip or circuit board; in some other embodiments, any one or two of processor 401, memory 402 and peripheral device interface 403 can be implemented on separate chips or circuit boards, which is not limited in this embodiment.

[0189] The radio frequency (RF) circuit 404 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The RF circuit 404 communicates with communication networks and other communication devices via electromagnetic signals. The RF circuit 404 converts electrical signals into electromagnetic signals for transmission, or converts received electromagnetic signals back into electrical signals. Optionally, the RF circuit 404 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a user identity module card, etc. The RF circuit 404 can communicate with other electronic devices through at least one wireless communication protocol. This wireless communication protocol includes, but is not limited to: the World Wide Web, metropolitan area networks, intranets, various generations of mobile communication networks (2G, 3G, 4G, and 5G), wireless local area networks, and / or WiFi (Wireless Fidelity) networks. In some embodiments, the RF circuit 404 may also include circuitry related to NFC (Near Field Communication), which is not limited in this application embodiment.

[0190] Display screen 405 is used to display a UI (User Interface). This UI may include graphics, text, icons, videos, and any combination thereof. When display screen 405 is a touch display screen, it also has the ability to collect touch signals on or above its surface. These touch signals can be input as control signals to processor 401 for processing. In this case, display screen 405 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 405, which serves as the front panel of electronic device 400; in other embodiments, there may be at least two display screens, respectively disposed on different surfaces of electronic device 400 or in a folded design; in still other embodiments, display screen 405 may be a flexible display screen, disposed on a curved or folded surface of electronic device 400. Furthermore, display screen 405 may be configured as a non-rectangular irregular shape, i.e., a non-rectangular screen. Display screen 405 may be made of materials such as LCD (Liquid Crystal Display) or OLED (Organic Light-Emitting Diode).

[0191] Camera assembly 406 is used to acquire images or videos. Optionally, camera assembly 406 includes a front-facing camera and a rear-facing camera. Typically, the front-facing camera is located on the front panel of the electronic device, and the rear-facing camera is located on the back of the electronic device. In some embodiments, there are at least two rear-facing cameras, which are any one of a main camera, a depth-sensing camera, a wide-angle camera, and a telephoto camera, to achieve background blurring by fusion of the main camera and the depth-sensing camera, panoramic shooting by fusion of the main camera and the wide-angle camera, VR (Virtual Reality) shooting, or other fusion shooting functions. In some embodiments, camera assembly 406 may also include a flash. The flash can be a single-color temperature flash or a dual-color temperature flash. A dual-color temperature flash refers to a combination of a warm-light flash and a cool-light flash, which can be used for light compensation at different color temperatures.

[0192] The audio circuit 407 may include a microphone and a speaker. The microphone is used to collect sound waves from the user and the environment, converting the sound waves into electrical signals that are input to the processor 401 for processing, or input to the radio frequency circuit 404 for voice communication. For stereo sound acquisition or noise reduction purposes, multiple microphones may be used, each located in a different part of the electronic device 400. The microphone may also be an array microphone or an omnidirectional microphone. The speaker is used to convert the electrical signals from the processor 401 or the radio frequency circuit 404 into sound waves. The speaker may be a conventional diaphragm speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can convert electrical signals not only into audible sound waves but also into inaudible sound waves for purposes such as distance measurement. In some embodiments, the audio circuit 407 may also include a headphone jack.

[0193] Those skilled in the art will understand that Figure 4 The structure shown does not constitute a limitation on the electronic device 400, and may include more or fewer components than shown, or combine certain components, or use different component arrangements.

[0194] In some embodiments, a computer-readable storage medium is also provided, which stores a computer program that, when executed by a processor, implements the steps of the voice interaction method described in the above embodiments. For example, the computer-readable storage medium may be a ROM, RAM, CD-ROM, magnetic tape, floppy disk, or optical data storage device.

[0195] It is worth noting that the computer-readable storage medium mentioned in the embodiments of this application can be a non-volatile storage medium, in other words, it can be a non-transient storage medium.

[0196] It should be understood that all or part of the steps of the above embodiments can be implemented by software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented wholly or partially in the form of a computer program product. The computer program product includes one or more computer instructions. The computer instructions can be stored in the above-described computer-readable storage medium.

[0197] That is, in some embodiments, a computer program product containing instructions is also provided, which, when run on a computer, causes the computer to perform the steps of the voice interaction method described above.

[0198] It should be understood that "at least one" as mentioned herein refers to one or more, and "multiple" refers to two or more. In the description of the embodiments of this application, unless otherwise stated, " / " means "or," for example, A / B can mean A or B; "and / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, and B existing alone. In addition, in order to clearly describe the technical solutions of the embodiments of this application, the terms "first," "second," etc., are used in the embodiments of this application to distinguish identical or similar items with substantially the same function and effect. Those skilled in the art will understand that the terms "first," "second," etc., do not limit the quantity or execution order, and the terms "first," "second," etc., are not necessarily different.

[0199] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, data stored, data displayed, etc.) and signals involved in the embodiments of this application are all authorized by the user or fully authorized by all parties, and the collection, use and processing of related data must comply with the relevant laws, regulations and standards of the relevant countries and regions.

[0200] The above descriptions are embodiments provided in this application and are not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.

Claims

1. A voice interaction method, characterized in that, The method includes: Receive voice commands from the target user; Identify keywords in the voice commands to obtain keyword groups; If the target user's voice intent cannot be determined through the keyword group, then the target user is given voice guidance based on the keyword group to allow the target user to supplement the voice intent; If the target user's voice intent can be determined through the voice supplementation results, then the target function corresponding to the voice intent is activated; The voice guidance provided to the target user based on the keyword group includes: Determine multiple target related words corresponding to the keyword group from the keyword association database, and the relevance between the multiple target related words and the keyword group respectively; Let i=1, select the i-th target related word from the plurality of target related words according to the sorting result of related words, generate and play voice guidance information based on the i-th target related word to guide the target user to supplement the voice intent. The sorting result of related words is obtained by sorting the plurality of target related words in descending order of relevance between the plurality of target related words and the keyword group. If a voice supplement instruction is received from the target user, and the voice intention of the target user cannot be determined through the voice supplement instruction, and the number of times the voice supplement instruction is issued is not greater than the number threshold, then the keywords in the voice supplement instruction are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database. If no voice supplementation command is received from the target user, and the multiple target related words have not been traversed or the number of times the voice supplementation command is issued is not greater than the number threshold, then let i = i + 1, and return to the step of selecting the i-th target related word from the multiple target related words according to the sorting result of the related words.

2. The method as described in claim 1, characterized in that, The step of determining multiple target related words corresponding to the keyword group from the keyword association database, and the relevance between the multiple target related words and the keyword group, includes: From the keyword association database, multiple candidate related words and the relevance between the multiple candidate related words and the corresponding keywords in the keyword group are determined. The multiple candidate related words refer to words that are related to the keywords in the keyword group. Determine the logical relationships between the keywords in the keyword group; Based on the logical relationship and the multiple candidate related words, the multiple target related words are determined; Based on the relevance between the multiple candidate related words and the corresponding keywords in the keyword group, the relevance between the multiple target related words and the keyword group is determined.

3. The method as described in claim 1, characterized in that, Before the step of setting i=i+1 and returning the sorted results according to the associated words, and selecting the i-th target associated word from the plurality of target associated words, the method further includes: Supplementary related words are determined from the keyword association database, where the supplementary related words refer to words that are associated with the i-th target related word; Based on the supplementary related words, generate and play audio recommendation information; If no voice recommendation response is received from the target user, then the step of setting i=i+1 and returning to the sorting result of the associated words, and selecting the i-th target associated word from the plurality of target associated words is executed.

4. The method as described in claim 3, characterized in that, After generating and playing the audio recommendation information based on the supplementary related words, the method further includes: If a voice recommendation response from the target user is received, the keywords in the voice recommendation response are added to the keyword group, and the process returns to the step of determining multiple target related words corresponding to the keyword group from the keyword association database.

5. The method as described in claim 1, characterized in that, After generating and playing the voice guidance information based on the i-th target related word, the method further includes: Based on the target user's feedback on the voice guidance information, update the relevance between the i-th target related word in the keyword association database and the keywords in the keyword group.

6. A voice interaction device, characterized in that, The device includes: The voice receiving module is used to receive voice commands issued by the target user. The speech recognition module is used to recognize keywords in the speech commands and obtain keyword groups; The voice guidance module is used to provide voice guidance to the target user based on the keyword group if the target user's voice intent cannot be determined through the keyword group, so that the target user can supplement the voice intent. The function activation module is used to activate the target function corresponding to the voice intent if the voice intent of the target user can be determined through the voice supplementation results of the target user. The voice guidance module includes: The related word determination submodule is used to determine multiple target related words corresponding to the keyword group from the keyword association database, and the relevance between the multiple target related words and the keyword group respectively; The speech generation submodule is used to set i=1, select the i-th target related word from the plurality of target related words according to the related word sorting result, generate and play voice guidance information based on the i-th target related word to guide the target user to supplement the voice intent. The related word sorting result is obtained by sorting the plurality of target related words in descending order of relevance between the plurality of target related words and the keyword group. The associated word determination submodule is further configured to, if it receives a voice supplementation instruction from the target user, and the voice intention of the target user cannot be determined through the voice supplementation instruction, and the number of times the voice supplementation instruction is issued is not greater than the number of times threshold, add the keywords in the voice supplementation instruction to the keyword group, and return to the step of determining multiple target associated words corresponding to the keyword group from the keyword association database; The speech generation submodule is further configured to, if no speech supplementation instruction is received from the target user, and the plurality of target related words have not been traversed or the number of times the speech supplementation instruction is issued is not greater than the number threshold, then set i=i+1 and return the step of selecting the i-th target related word from the plurality of target related words according to the sorting result of the related words.

7. A computer device, characterized in that, The computer device comprises a memory for storing a computer program and a processor for executing the computer program stored in the memory to implement the steps of the method according to any one of claims 1-5.

8. A computer-readable storage medium, characterized in that, The storage medium stores a computer program, and the computer program is executed by the processor to implement the steps of the method according to any one of claims 1-5.