Voice interaction method and device
By setting preset conditions and multiple functions, intents, and parameters in electronic devices, and combining local and cloud-based semantic recognition results, the accuracy and latency issues of electronic devices in speech recognition are solved, achieving efficient response in different scenarios.
Patent Information
- Application Number
- CN202511693019.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2021-04-17
- Publication Date
- 2026-02-10
AI Technical Summary
Electronic devices suffer from insufficient accuracy and long response times in voice recognition, especially when network conditions are poor, making it difficult to provide users with timely and appropriate responses.
By setting preset conditions and multiple functions, intents, and parameters in electronic devices, and combining local and cloud-based semantic recognition results, processing methods can be flexibly selected to improve the balance between response latency and accuracy.
Reduce response latency and improve response efficiency in scenarios where electronic devices excel; improve response accuracy in scenarios where they are not good at it, balancing the accuracy and efficiency of speech recognition.
Smart Images

Figure CN121506121A_ABST
Abstract
Description
[0001] This application is a divisional application of the original application with the application number 202180005755.9 and the original filing date of April 17, 2021, and the entire contents of the original application are incorporated herein by reference. TECHNICAL FIELD
[0002] The present application relates to the field of electronic devices, and more particularly, to a method and apparatus for voice interaction. BACKGROUND
[0003] A user and an electronic device can interact with each other through voice. The user can speak a voice instruction to the electronic device. The electronic device can acquire the voice instruction and perform an operation indicated by the voice instruction.
[0004] The electronic device itself can recognize the voice instruction. The capability of the electronic device to recognize the voice instruction can be relatively limited. If the voice instruction is recognized only by the electronic device, the recognition result can be inaccurate, and thus the electronic device can not give a proper response to the user.
[0005] The electronic device can also upload voice information related to the voice instruction to a cloud. The cloud can recognize the voice information and feed back the recognition result to the electronic device. The cloud can have relatively strong capability of voice recognition and natural language understanding. However, the interaction between the electronic device and the cloud can depend on the current network status of the electronic device. That is, the interaction between the electronic device and the cloud can have a relatively long time delay. If the voice instruction is recognized by the cloud, the electronic device can not acquire the recognition result from the cloud in time, and thus can not quickly give a response to the user. SUMMARY
[0006] The present application provides a method and apparatus for voice interaction, aiming to balance the accuracy and efficiency of voice recognition, and thus to facilitate giving a proper and quick response to the user.
[0007] In a first aspect, a method for voice interaction is provided, applied to a first apparatus, and the method comprises: acquiring first voice information from a voice sensor; determining a first semantic recognition result according to the first voice information; determining to perform a first operation determined by the first apparatus according to the first semantic recognition result, or determining to perform a second operation indicated by a second apparatus, according to the first semantic recognition result and a first preset condition.
[0008] In one possible example, the second apparatus can indicate the second operation to the first apparatus by sending one or more of a semantic recognition result and operation information to the first apparatus.
[0009] In the voice interaction scene in which the first device is relatively good at, the first device can determine the operation corresponding to the voice instruction of the user without the information provided by the second device. In this way, it is beneficial to reduce the response delay of the first device in executing the voice instruction of the user and improve the response efficiency. In the voice interaction scene in which the first device is relatively not good at, the first device can determine the operation corresponding to the voice instruction of the user according to the information provided by the second device. In this way, it is beneficial to improve the accuracy of the first device in responding to the voice instruction of the user. Through the above scheme, the processing manner of the voice instruction can be flexibly selected according to the voice interaction scene in which the first device is good at, and the response delay and the response accuracy are balanced.
[0010] Optionally, the method further includes: the first device sending the first voice information to the second device.
[0011] The first device sends the first voice information to the second device, and the second device can send feedback to the first device for the first voice information. If the first device executes the first operation, the first device can adjust the semantic recognition model, the voice control model and the like of the first device according to the feedback of the second device, so as to improve the accuracy of the semantic recognition result output by the first device, optimize the applicability of the operation of responding to the voice instruction of the user, or ignore the feedback of the second device, for example, when the feedback of the second device has a larger time delay than the voice recognition result of the first device, the feedback of the second device is ignored. If the first device executes the second operation, the first device can relatively quickly obtain the feedback of the second device for the first voice information. Therefore, it is beneficial to shorten the time of the first device in responding to the voice instruction of the user.
[0012] With reference to the first aspect, in some implementations of the first aspect, the determining, according to the first semantic recognition result and a first preset condition, to execute a first operation determined by the first device according to the first semantic recognition result, includes: In a case where the first semantic recognition result meets the first preset condition, it is determined to execute the first operation.
[0013] Optionally, in a case where the first semantic recognition result does not meet the first preset condition, it is determined to execute a second operation indicated by the second device.
[0014] The first preset condition is beneficial to the first device in judging whether the first device can accurately recognize the current voice information, and is further beneficial to balancing the accuracy and efficiency of voice recognition.
[0015] With reference to the first aspect, in some implementations of the first aspect, the first device is preset with a plurality of functions, and the first semantic recognition result meets a first preset condition, including: The first semantic recognition result indicates a first function, and the first function belongs to the plurality of functions.
[0016] Optionally, the first device has multiple preset functions, and the first semantic recognition result does not meet the first preset condition, including: the first semantic recognition result indicates a first function, and the first function does not belong to the multiple functions.
[0017] Optionally, the plurality of functions may include one or more of the following functions: vehicle control function, navigation function, audio function, and video function.
[0018] The first device may have multiple preset functions, such as those supported by the first device. The first device may have a relatively higher semantic recognition capability for these preset functions. Conversely, the first device may have a relatively lower semantic recognition capability for other functions not preset by the first device. Based on these preset functions and the first function, the first device can determine whether it can relatively accurately recognize the current speech information, thereby balancing the accuracy and efficiency of speech recognition.
[0019] In conjunction with the first aspect, in some implementations of the first aspect, the first device has multiple preset intentions, and the first semantic recognition result satisfies a first preset condition, including: The first semantic recognition result indicates a first intent, which belongs to the plurality of intents.
[0020] Optionally, the method further includes: the first device has multiple preset intentions, and the first semantic recognition result does not meet the first preset condition, including: the first semantic recognition result indicates a first intention, and the first intention does not belong to the multiple intentions.
[0021] Optionally, the plurality of intents includes one or more of the following intents: hardware activation intent, path planning intent, audio playback intent, and video playback intent.
[0022] The first device can preset multiple intents, such as multiple intents supported by the first device. The first device can have a relatively higher semantic recognition capability for the preset multiple intents. The first device can have a relatively lower semantic recognition capability for other intents not preset by the first device. The first device can determine whether it can relatively accurately recognize the current speech information based on the preset multiple intents and the first intent, thereby helping to balance the accuracy and efficiency of speech recognition.
[0023] In conjunction with the first aspect, in some implementations of the first aspect, the first device has multiple preset parameters, and the first semantic recognition result satisfies a first preset condition, including: The first semantic recognition result includes a first parameter, which belongs to the plurality of parameters.
[0024] Optionally, the first device has multiple preset parameters, and the first semantic recognition result does not meet the first preset condition, including: the first semantic recognition result indicates the first parameter, and the first parameter does not belong to the multiple parameters.
[0025] The first device may have multiple preset parameters, such as those supported by the first device. The first device may have a relatively higher semantic recognition capability for the preset parameters. The first device may have a relatively lower semantic recognition capability for other parameters not preset by the first device. Based on the preset parameters and the first parameter, the first device can determine whether it can relatively accurately recognize the current speech information, thus balancing the accuracy and efficiency of speech recognition.
[0026] In conjunction with the first aspect, in some implementations of the first aspect, the first device is pre-set with multiple intents corresponding to the first function, and the first semantic recognition result satisfies the first preset condition, further including: The first semantic recognition result also indicates a first intent, which belongs to the plurality of intents.
[0027] Optionally, the first device has multiple preset intentions corresponding to the first function. If the first semantic recognition result does not meet the first preset condition, it further includes: the first semantic recognition result also indicates a first intention, and the first intention does not belong to the multiple intentions.
[0028] The intent corresponding to the first function is usually not unlimited. Establishing a correspondence between multiple functions and multiple intents helps the first device to more accurately determine whether it can recognize the current voice information, thus helping to balance the accuracy and efficiency of voice recognition.
[0029] In conjunction with the first aspect, in some implementations of the first aspect, the first device is preset with multiple parameters corresponding to the first function, and the first semantic recognition result satisfies the first preset condition, further including: The first semantic recognition result also indicates a first parameter, which belongs to the plurality of parameters.
[0030] Optionally, the first device has multiple parameters preset corresponding to the first function. If the first semantic recognition result does not meet the first preset condition, it further includes: the first semantic recognition result also indicates a first parameter, and the first parameter does not belong to the multiple parameters.
[0031] Optionally, the multiple parameter types include one or more of the following parameter types: hardware identifier, time, temperature, location, artist, song, playlist, audio playback mode, movie, TV series, actor, and video playback mode.
[0032] The parameters corresponding to the first function are usually not unlimited. Establishing a correspondence between multiple functions and multiple parameters helps the first device to more accurately determine whether it can recognize the current voice information, thus helping to balance the accuracy and efficiency of voice recognition.
[0033] In conjunction with the first aspect, in some implementations of the first aspect, the first device is preset with multiple parameters corresponding to the first intent, and the first semantic recognition result satisfies the first preset condition, further including: The first semantic recognition result also indicates a first parameter, which belongs to the plurality of parameters.
[0034] Optionally, the first device has multiple parameters preset to correspond to the first intent. If the first semantic recognition result does not meet the first preset condition, it further includes: the first semantic recognition result also indicates a first parameter, and the first parameter does not belong to the multiple parameters.
[0035] The parameters corresponding to the first intent are usually not unlimited. Establishing a correspondence between multiple intents and multiple parameters helps the first device to more accurately determine whether it can recognize the current speech information relatively accurately, thus helping to balance the accuracy and efficiency of speech recognition.
[0036] In conjunction with the first aspect, in certain implementations of the first aspect, the first semantic recognition result indicates a first function and indicates a first parameter, and the first semantic recognition result satisfies a first preset condition, further including: The first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
[0037] Optionally, the first semantic recognition result indicates a first function and indicates a first parameter. If the first semantic recognition result does not meet the first preset condition, it further includes: the first function indicated by the first semantic recognition result and the first parameter indicated by the first semantic recognition result correspond to different parameter types.
[0038] The following examples illustrate the parameter types that correspond to the functions.
[0039] For example, the parameter types corresponding to vehicle control functions can be time, temperature, hardware identifiers, etc.
[0040] For example, the parameter type corresponding to the temperature control function can be temperature, etc.
[0041] For example, the parameter types for navigation functions can be location, time, etc.
[0042] For example, the parameter types corresponding to audio functions can be artist, song, playlist, time, audio playback mode, etc.
[0043] For example, the parameter types corresponding to the video function can be movies, TV series, actors, time, video playback mode, etc.
[0044] The following examples illustrate the parameter types that parameters can correspond to.
[0045] For example, the parameter types corresponding to air conditioning, cameras, seats, and windows can be hardware identifiers.
[0046] For example, the parameter type corresponding to 5℃, 28℃, etc. can be temperature.
[0047] For example, the parameter type corresponding to 1 hour, 1 minute, etc. can be time.
[0048] For example, the parameter type corresponding to position A, position B, etc. can be position.
[0049] For example, the parameter type corresponding to singer A, singer B, etc. can be singer.
[0050] For example, the parameter type corresponding to song A, song B, etc. can be "song".
[0051] For example, the parameter type corresponding to playlist A, playlist B, etc. can be playlist.
[0052] For example, the parameter types corresponding to standard playback, high-quality playback, and lossless playback can be audio playback modes.
[0053] For example, the parameter type corresponding to Movie A, Movie B, etc. can be Movie.
[0054] For example, the parameter type corresponding to TV series A and TV series B can be TV series.
[0055] For example, the parameter type corresponding to actor A, actor B, etc. can be actor.
[0056] For example, the parameter types corresponding to standard definition playback, high definition playback, extended playback, and Blu-ray playback can be video playback modes.
[0057] If the first device obtains in advance that the first parameter corresponds to the first parameter type and the first function corresponds to the first parameter type, then the accuracy of the first device in recognizing the first speech information may be relatively high. If the first device obtains in advance that the first parameter corresponds to the first parameter type, but the first function does not correspond to the first parameter type, then the accuracy of the first device in analyzing the first speech information may be relatively low. Establishing the relationship between multiple parameters and multiple functions through parameter types helps the first device to more accurately determine whether it can relatively accurately recognize the current speech information, thus helping to balance the accuracy and efficiency of speech recognition.
[0058] Furthermore, if multiple functions correspond to the same type of parameter, then that type of parameter can be derived across multiple functions. This helps reduce the complexity of establishing relationships between parameters and functions.
[0059] In conjunction with the first aspect, in some implementations of the first aspect, determining the first semantic recognition result based on the first speech information includes: Based on the first voice information, a second semantic recognition result is determined, wherein the second semantic recognition result indicates a second function and indicates the first parameter; If the second function is not included in the preset functions of the first device, and the first parameter is included in the preset parameters of the first device, the second function in the second semantic recognition result is corrected to the first function to obtain the first semantic recognition result. The first function and the second function are two different functions of the same type.
[0060] For example, the first function could be a local translation function, and the second function could be a cloud translation function. Both the first and second functions can be translation-related features.
[0061] For example, the first function could be local navigation, and the second function could be cloud navigation. Both the first and second functions can be navigation-type functions.
[0062] For example, the first function could be local audio, and the second function could be cloud audio. Both the first and second functions can be audio playback functions.
[0063] For example, the first function could be a local video function, and the second function could be a cloud video function. Both the first and second functions can be video playback functions.
[0064] The second function is not among the multiple functions preset by the first device, meaning that the first device may have a relatively weak semantic recognition ability for the second function. For example, the first device can learn the voice commands related to the second function multiple times, thereby gradually improving its semantic recognition ability for the second function. In other words, by modifying the second function to the first function, the first device can apply its learned skills in areas where it is relatively less proficient, which helps increase the applicable scenarios for edge-side decision-making and thus improves the efficiency of speech recognition.
[0065] In conjunction with the first aspect, in certain implementations of the first aspect, the first semantic recognition result indicates a first intent and indicates a first parameter, and the first semantic recognition result satisfies a first preset condition, further including: The first intent indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
[0066] Optionally, the first semantic recognition result indicates a first intent and indicates a first parameter. If the first semantic recognition result does not meet the first preset condition, it further includes: the first intent indicated by the first semantic recognition result and the first parameter indicated by the first semantic recognition result correspond to different parameter types.
[0067] The following examples illustrate the parameter types that correspond to intents.
[0068] For example, the parameter type corresponding to the hardware activation intent can be time, temperature, hardware identifier, etc.
[0069] For example, the parameter type corresponding to the path planning intent can be location, time, etc.
[0070] For example, the parameter types corresponding to the intention to play audio can be artist, song, playlist, time, audio playback mode, etc.
[0071] For example, the parameter type corresponding to the intention to play a video can be movie, TV series, actor, time, video playback mode, etc.
[0072] The following examples illustrate the parameter types that parameters can correspond to.
[0073] For example, the parameter types corresponding to air conditioning, cameras, seats, and windows can be hardware identifiers.
[0074] For example, the parameter type corresponding to 5℃, 28℃, etc. can be temperature.
[0075] For example, the parameter type corresponding to 1 hour, 1 minute, etc. can be time.
[0076] For example, the parameter type corresponding to position A, position B, etc. can be position.
[0077] For example, the parameter type corresponding to singer A, singer B, etc. can be singer.
[0078] For example, the parameter type corresponding to song A, song B, etc. can be "song".
[0079] For example, the parameter type corresponding to playlist A, playlist B, etc. can be playlist.
[0080] For example, the parameter types corresponding to standard playback, high-quality playback, and lossless playback can be audio playback modes.
[0081] For example, the parameter type corresponding to Movie A, Movie B, etc. can be Movie.
[0082] For example, the parameter type corresponding to TV series A and TV series B can be TV series.
[0083] For example, the parameter type corresponding to actor A, actor B, etc. can be actor.
[0084] For example, the parameter types corresponding to standard definition playback, high definition playback, extended playback, and Blu-ray playback can be video playback modes.
[0085] If the first device obtains in advance that the first parameter corresponds to the first parameter type and the first intent corresponds to the first parameter type, then the accuracy of the first device's analysis of the first speech information may be relatively high. If the first device obtains in advance that the first parameter corresponds to the first parameter type, but the first intent does not correspond to the first parameter type, then the accuracy of the first device's analysis of the first speech information may be relatively low. Establishing the relationship between multiple parameters and multiple intents through parameter types helps the first device to more accurately determine whether it can relatively accurately recognize the current speech information, thus helping to balance the accuracy and efficiency of speech recognition.
[0086] Furthermore, if multiple intents correspond to the same type of parameters, then those parameters can be extrapolated across multiple intents. This helps reduce the complexity of establishing relationships between parameters and intents.
[0087] In conjunction with the first aspect, in some implementations of the first aspect, determining the first semantic recognition result based on the first speech information includes: Based on the first voice information, a third semantic recognition result is determined, wherein the third semantic recognition result indicates a second intention and indicates the first parameter; If the second intention is not included in the multiple intentions preset by the first device, and the first parameter is included in the multiple parameters preset by the first device, the second intention in the third semantic recognition result is corrected to the first intention to obtain the first semantic recognition result. The first intention and the second intention are two different intentions of the same type.
[0088] For example, the first intention could be local translation of English, while the second intention could be cloud translation of English. Both the first and second intentions can fall under the category of translating English.
[0089] For example, the first intent could be a local route planning intent, while the second intent could be a cloud route planning intent. Both the first and second intents can belong to the route planning type of intent.
[0090] For example, the first intent could be to play local audio, and the second intent could be to play cloud audio. Both the first and second intents can belong to the audio playback type of intent.
[0091] For example, the first intent could be to play a local video, while the second intent could be to play a cloud video. Both the first and second intents can fall under the category of video playback intents.
[0092] The second intent does not belong to the multiple intents preset by the first device, meaning that the first device may have a relatively weak semantic recognition ability for the second intent. For example, the first device can learn voice commands related to the second intent multiple times, thereby gradually improving its semantic recognition ability for the second intent. In other words, by modifying the second intent to the first intent, the first device can apply its learned skills in areas where it is relatively less proficient, which helps increase the applicable scenarios for edge-side decision-making and thus improves the efficiency of speech recognition.
[0093] In conjunction with the first aspect, in certain implementations of the first aspect, the first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result includes a first indicator bit, which indicates that the first semantic recognition result meets the first preset condition.
[0094] Optionally, if the first semantic recognition result does not meet the first preset condition, it includes: the first semantic recognition result includes a second indicator bit, the second indicator bit indicating that the first semantic recognition result does not meet the first preset condition.
[0095] In conjunction with the first aspect, in some implementations of the first aspect, The step of determining the first semantic recognition result based on the first voice information includes: Based on the first voice information, a fourth semantic recognition result is determined, wherein the fourth semantic recognition result includes a first function and a first parameter; When the first function belongs to multiple functions preset by the first device, and the first parameter belongs to multiple parameters preset by the first device, and the first function and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
[0096] The first device may have a relatively weak semantic recognition capability for the first function. For example, the first device can learn voice commands related to the first function multiple times, thereby gradually improving its semantic recognition capability for the first function. In other words, by carrying a first indicator bit in the semantic recognition result, the first device can apply its learned skills in areas where it is relatively less proficient, which helps increase the applicable scenarios for edge-side decision-making and thus improves the efficiency of speech recognition.
[0097] In conjunction with the first aspect, in some implementations of the first aspect, The step of determining the first semantic recognition result based on the first voice information includes: Based on the first voice information, a fifth semantic recognition result is determined, wherein the fifth semantic recognition result includes a first intent and a first parameter; When the first intent belongs to a plurality of intents preset by the first device, and the first parameter belongs to a plurality of parameters preset by the first device, and the first intent and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
[0098] The first intent may have a relatively weak semantic recognition capability. The first device can, for example, learn voice commands related to the first intent multiple times, thereby gradually improving its semantic recognition capability for the first intent. In other words, by carrying a first indicator bit in the semantic recognition result, the first device can apply the learned skills in areas where it is relatively less proficient, which helps to increase the applicable scenarios for edge decision-making and thus improve the efficiency of speech recognition.
[0099] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: Send the first voice information to the second device; The sixth semantic recognition result from the second device is discarded.
[0100] In conjunction with the first aspect, in certain implementations of the first aspect, determining to execute the second operation instructed by the second device based on the first semantic recognition result and the first preset condition includes: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed.
[0101] The first preset condition helps the first device determine whether it can relatively accurately recognize the current voice information, thereby helping to balance the accuracy and efficiency of voice recognition.
[0102] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: Based on the sixth semantic recognition result, determine the second parameter and the type of the second parameter; Save the association between the second parameter and the second parameter type.
[0103] The first device can learn new parameters. This helps increase the applicable scenarios for edge-side decision-making, thereby improving the efficiency of speech recognition.
[0104] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: Acquire second voice information from the voice sensor; Based on the second speech information, determine the seventh semantic recognition result; Based on the seventh semantic recognition result and the second preset condition, it is determined to execute the operation indicated by the first voice information, or to execute the third operation determined by the first device based on the seventh semantic recognition result, or to execute the fourth operation indicated by the second device.
[0105] In multi-turn voice interaction scenarios, the user and the first device can engage in voice dialogue for a specific scenario or domain. In one possible scenario, the user may not be able to achieve the desired voice control with a single voice command. Adjacent voice interactions in multi-turn interactions are usually related. However, the user's responses may have a degree of randomness. The user's response may be unrelated to the voice information requested or sought by the first device. If the first device completely follows the user's response, the content of previous voice interactions may be rendered invalid, potentially increasing the number of times the user needs to control the first device. If the first device completely ignores the user's response, it may be unable to respond to the user's instructions in certain specific scenarios, rendering the user's voice commands ineffective. A second preset condition is used to indicate whether the first device should end the multi-turn voice interaction, allowing the first device to appropriately choose whether to exit the multi-turn voice interaction.
[0106] In conjunction with the first aspect, in certain implementations of the first aspect, determining to execute the operation indicated by the first voice information based on the seventh semantic recognition result and the second preset condition, or determining to execute a third operation determined by the first device based on the seventh semantic recognition result, or determining to execute a fourth operation indicated by the second device, includes: If the seventh semantic recognition result satisfies both the first preset condition and the second preset condition, the third operation is to be performed. If the seventh semantic recognition result does not meet the first preset condition but meets the second preset condition, the fourth operation is to be performed. If the seventh semantic recognition result does not meet the second preset condition, it is determined to execute the operation corresponding to the first semantic recognition result.
[0107] In one example, if the first operation is performed and the first device determines to perform the third operation, the new end-side voice interaction can end the previous round of end-side voice interaction.
[0108] In one example, if the second operation is performed and the first device determines to perform the third operation, the new end-side voice interaction can end the previous round of cloud-side voice interaction.
[0109] In one example, if the first operation is performed and the first device determines to perform the fourth operation, the new cloud-side voice interaction can end the previous round of end-side voice interaction.
[0110] In one example, if the second operation is performed and the first device determines to perform the fourth operation, the new cloud-side voice interaction can end the previous round of end-to-cloud voice interaction.
[0111] The first device can comprehensively judge the first preset condition and the second preset condition, which is conducive to the first device appropriately choosing whether to exit the multi-round voice interaction, and is also conducive to balancing the accuracy and efficiency of voice recognition.
[0112] In conjunction with the first aspect, in certain implementations of the first aspect, the seventh semantic recognition result satisfies the second preset condition, including: The priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result.
[0113] Users can end the current multi-round voice interaction by using high-priority voice commands.
[0114] In conjunction with the first aspect, in certain implementations of the first aspect, the priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: The function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; The intent indicated by the seventh semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; The parameters indicated by the seventh semantic recognition result have a higher priority than the parameters indicated by the first semantic recognition result.
[0115] Functions, intentions, and parameters are better able to reflect the current voice interaction scenario. The priority of functions, intentions, and parameters helps a device to relatively accurately determine whether to exit the current multi-turn voice interaction.
[0116] In conjunction with the first aspect, in certain implementations of the first aspect, determining to execute the second operation instructed by the second device based on the first semantic recognition result and the first preset condition includes: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed; The seventh semantic recognition result satisfies the second preset condition, including: The priority of the seventh semantic recognition result is higher than the priority of the sixth semantic recognition result.
[0117] The sixth semantic recognition result is indicated by the second device, and this result can have relatively higher accuracy. By comparing the priority of the seventh semantic recognition result with that of the sixth semantic recognition result, the first device can more appropriately choose whether to exit the multi-round voice interaction.
[0118] In conjunction with the first aspect, in certain implementations of the first aspect, the priority of the seventh semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: The function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; The intent indicated by the seventh semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; The parameter indicated by the seventh semantic recognition result has a higher priority than the parameter indicated by the sixth semantic recognition result.
[0119] Functions, intentions, and parameters are better able to reflect the current voice interaction scenario. The priority of functions, intentions, and parameters helps a device to relatively accurately determine whether to exit the current multi-turn voice interaction.
[0120] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: Send a second voice message to the second device; The step of determining to execute the operation indicated by the first voice information, or to determine to execute the fourth operation indicated by the second device, based on the seventh semantic recognition result and the second preset condition, includes: If the seventh semantic recognition result does not meet the first preset condition, or if the seventh semantic recognition result does not meet the second preset condition and the first semantic recognition result does not meet the first preset condition, an eighth semantic recognition result is obtained from the second device. Based on the eighth semantic recognition result and the second preset condition, it is determined to execute the operation indicated by the first voice information, or to execute the fourth operation.
[0121] If the seventh semantic recognition result does not meet the first preset condition, the seventh semantic recognition result obtained by the first device may be relatively inaccurate, while the eighth semantic recognition result obtained by the second device may be relatively accurate. The first device can determine whether the eighth semantic recognition result meets the second preset condition, which helps to determine relatively accurately whether to end the current multi-round voice interaction.
[0122] If the seventh semantic recognition result does not meet the second preset condition, it means that the first device can determine the operation to be performed without relying on the seventh semantic recognition result. If the first semantic recognition result does not meet the first preset condition, it may mean that the current multi-turn voice interaction belongs to cloud-side voice interaction. In this case, the first device can obtain the eighth semantic recognition result from the second device to continue the current cloud-side multi-turn voice interaction. This is beneficial for maintaining cloud-side multi-turn voice interaction.
[0123] In conjunction with the first aspect, in some implementations of the first aspect, determining to execute the operation indicated by the first voice information, or determining to execute the fourth operation, based on the eighth semantic recognition result and the second preset condition, includes: If the eighth semantic recognition result satisfies the second preset condition, the fourth operation is to be performed. If the eighth semantic recognition result does not meet the second preset condition, it is determined to execute the operation indicated by the first voice information.
[0124] The eighth semantic recognition result is indicated by the second device, and this result can have relatively higher accuracy. Based on the priority of the eighth semantic recognition result, the first device can more appropriately choose whether to exit the multi-round voice interaction.
[0125] In conjunction with the first aspect, in certain implementations of the first aspect, the eighth semantic recognition result satisfies the second preset condition, including: The priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result.
[0126] Users can end the current multi-round voice interaction by using high-priority voice commands.
[0127] In conjunction with the first aspect, in certain implementations of the first aspect, the priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: The function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; The intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; The fifth parameter indicated by the eighth semantic recognition result has a higher priority than the parameter indicated by the first semantic recognition result.
[0128] Functions, intentions, and parameters are better able to reflect the current voice interaction scenario. The priority of functions, intentions, and parameters helps a device to relatively accurately determine whether to exit the current multi-turn voice interaction.
[0129] In conjunction with the first aspect, in certain implementations of the first aspect, determining to execute the second operation instructed by the second device based on the first semantic recognition result and the first preset condition includes: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed; The eighth semantic recognition result satisfies the second preset condition, including: The priority of the eighth semantic recognition result is higher than the priority of the sixth semantic recognition result.
[0130] The sixth and eighth semantic recognition results are both indicated by the second device, and both can have relatively higher accuracy. By comparing the priority of the eighth semantic recognition result and the priority of the sixth semantic recognition result, the first device can more appropriately choose whether to exit the multi-round voice interaction.
[0131] In conjunction with the first aspect, in certain implementations of the first aspect, the priority of the eighth semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: The function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; The intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; The parameter indicated by the eighth semantic recognition result has a higher priority than the parameter indicated by the sixth semantic recognition result.
[0132] Functions, intentions, and parameters are better able to reflect the current voice interaction scenario. The priority of functions, intentions, and parameters helps a device to relatively accurately determine whether to exit the current multi-turn voice interaction.
[0133] In conjunction with the first aspect, in some implementations of the first aspect, the second voice information is unrelated to the operation indicated by the first voice information.
[0134] For example, the correlation between the second voice information and the operation indicated by the first voice information is lower than a second preset threshold. Alternatively, the first voice information and the second voice information are unrelated. Or, the correlation between the first voice information and the second voice information is lower than the second preset threshold. Or, one or more of the functions, intentions, or parameters indicated by the first and second voice information are different.
[0135] In conjunction with the first aspect, in some implementations of the first aspect, the method further includes: Respond to user input and execute voice wake-up.
[0136] Secondly, a voice interaction device is provided, comprising: An acquisition unit is used to acquire first speech information from a speech sensor; The processing unit is configured to determine a first semantic recognition result based on the first speech information; The processing unit is further configured to determine, based on the first semantic recognition result and the first preset condition, to execute a first operation determined by the first device based on the first semantic recognition result, or to determine to execute a second operation indicated by the second device.
[0137] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: If the first semantic recognition result meets the first preset condition, the first operation is executed.
[0138] In conjunction with the second aspect, in some implementations of the second aspect, the first device has multiple preset functions, and the first semantic recognition result satisfies a first preset condition, including: The first semantic recognition result indicates a first function, which belongs to the plurality of functions.
[0139] In conjunction with the second aspect, in some implementations of the second aspect, the first device has multiple preset intentions, and the first semantic recognition result satisfies a first preset condition, including: The first semantic recognition result indicates a first intent, which belongs to the plurality of intents.
[0140] In conjunction with the second aspect, in some implementations of the second aspect, the first device has multiple preset parameters, and the first semantic recognition result satisfies a first preset condition, including: The first semantic recognition result indicates the first parameter, which belongs to the plurality of parameters.
[0141] In conjunction with the second aspect, in some implementations of the second aspect, the first semantic recognition result indicates a first function and indicates a first parameter, and the first semantic recognition result satisfies a first preset condition, further including: The first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
[0142] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: Based on the first voice information, a second semantic recognition result is determined, wherein the second semantic recognition result indicates a second function and indicates the first parameter; If the second function is not included in the preset functions of the first device, and the first parameter is included in the preset parameters of the first device, the second function in the second semantic recognition result is corrected to the first function to obtain the first semantic recognition result. The first function and the second function are two different functions of the same type.
[0143] In conjunction with the second aspect, in some implementations of the second aspect, the first semantic recognition result indicates a first intent and indicates a first parameter, and the first semantic recognition result satisfies a first preset condition, further including: The first intent indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
[0144] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: Based on the first voice information, a third semantic recognition result is determined, wherein the third semantic recognition result indicates a second intention and indicates a first parameter; If the second intention is not included in the multiple intentions preset by the first device, and the first parameter is included in the multiple parameters preset by the first device, the second intention in the third semantic recognition result is corrected to the first intention to obtain the first semantic recognition result. The first intention and the second intention are two different intentions of the same type.
[0145] In conjunction with the second aspect, in certain implementations of the second aspect, the first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result includes a first indicator bit, which indicates that the first semantic recognition result meets the first preset condition.
[0146] In conjunction with the second aspect, in some implementations of the second aspect, The processing unit is specifically used for: Based on the first voice information, a fourth semantic recognition result is determined, wherein the fourth semantic recognition result includes a first function and a first parameter; When the first function belongs to multiple functions preset by the first device, and the first parameter belongs to multiple parameters preset by the first device, and the first function and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
[0147] In conjunction with the second aspect, in some implementations of the second aspect, The processing unit is specifically used for: Based on the first voice information, a fifth semantic recognition result is determined, wherein the fifth semantic recognition result includes a first intent and a first parameter; When the first intention belongs to a plurality of intentions preset by the first device, and the first parameter belongs to a plurality of parameters preset by the first device, and the first intention and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
[0148] In conjunction with the second aspect, in some implementations of the second aspect, the apparatus further includes: A sending unit is configured to send the first voice information to the second device; The processing unit is also used to discard the sixth semantic recognition result from the second device.
[0149] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed.
[0150] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is further configured to: Based on the sixth semantic recognition result, determine the second parameter and the type of the second parameter; The device further includes a storage unit for storing the association between the second parameter and the second parameter type.
[0151] In conjunction with the second aspect, in some implementations of the second aspect, The acquisition unit is further configured to acquire second voice information from the voice sensor; The processing unit is further configured to determine the seventh semantic recognition result based on the second speech information; The processing unit is further configured to, based on the seventh semantic recognition result and the second preset conditions, determine to execute the operation indicated by the first voice information, or determine to execute the third operation determined by the first device based on the seventh semantic recognition result, or determine to execute the fourth operation indicated by the second device.
[0152] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for, If the seventh semantic recognition result satisfies both the first preset condition and the second preset condition, the third operation is to be performed. If the seventh semantic recognition result does not meet the first preset condition but meets the second preset condition, the fourth operation is to be performed. If the seventh semantic recognition result does not meet the second preset condition, it is determined to execute the operation corresponding to the first semantic recognition result.
[0153] In conjunction with the second aspect, in certain implementations of the second aspect, the seventh semantic recognition result satisfies the second preset condition, including: The priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result.
[0154] In conjunction with the second aspect, in some implementations of the second aspect, the priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: The function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; The intent indicated by the seventh semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; The parameters indicated by the seventh semantic recognition result have a higher priority than the parameters indicated by the first semantic recognition result.
[0155] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed; The seventh semantic recognition result satisfies the second preset condition, including: The priority of the seventh semantic recognition result is higher than the priority of the sixth semantic recognition result.
[0156] In conjunction with the second aspect, in some implementations of the second aspect, the priority of the seventh semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: The function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; The intent indicated by the seventh semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; The parameter indicated by the seventh semantic recognition result has a higher priority than the parameter indicated by the sixth semantic recognition result.
[0157] In conjunction with the second aspect, in some implementations of the second aspect, the apparatus further includes: The sending unit is used to send second voice information to the second device; The processing unit is specifically used for: If the seventh semantic recognition result does not meet the first preset condition, or if the seventh semantic recognition result does not meet the second preset condition and the first semantic recognition result does not meet the first preset condition, an eighth semantic recognition result is obtained from the second device. Based on the eighth semantic recognition result and the second preset condition, it is determined to execute the operation indicated by the first voice information, or to execute the fourth operation.
[0158] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: If the eighth semantic recognition result satisfies the second preset condition, the fourth operation is to be performed. If the eighth semantic recognition result does not meet the second preset condition, it is determined to execute the operation indicated by the first voice information.
[0159] In conjunction with the second aspect, in certain implementations of the second aspect, the eighth semantic recognition result satisfies the second preset condition, including: The priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result.
[0160] In conjunction with the second aspect, in some implementations of the second aspect, the priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: The function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; The intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; The parameters indicated by the eighth semantic recognition result have a higher priority than the parameters indicated by the first semantic recognition result.
[0161] In conjunction with the second aspect, in some implementations of the second aspect, the processing unit is specifically used for: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed; The eighth semantic recognition result satisfies the second preset condition, including: The priority of the eighth semantic recognition result is higher than the priority of the sixth semantic recognition result.
[0162] In conjunction with the second aspect, in some implementations of the second aspect, the priority of the eighth semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: The function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; The intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; The parameter indicated by the eighth semantic recognition result has a higher priority than the parameter indicated by the sixth semantic recognition result.
[0163] In conjunction with the second aspect, in some implementations of the second aspect, the second voice information is unrelated to the operation indicated by the first voice information.
[0164] In conjunction with the second aspect, in some implementations of the second aspect, the apparatus further includes: The wake-up module is used to respond to user input and perform voice wake-up operations.
[0165] Thirdly, a voice interaction device is provided, the device comprising: a processor and a memory, the processor being coupled to the memory for storing a computer program, and the processor for executing the computer program stored in the memory such that the device performs the method described in any possible implementation of the first aspect above.
[0166] Fourthly, a computer-readable medium is provided that stores program code for execution by a device, the program code including methods for performing any of the implementations of the first aspect described above.
[0167] Fifthly, a computer program product containing instructions is provided, which, when run on a computer, causes the computer to perform the method described in any of the implementations of the first aspect above.
[0168] In a sixth aspect, a chip is provided, the chip including a processor and a data interface, wherein the processor reads instructions stored in a memory through the data interface and executes the method described in any of the implementations of the first aspect above.
[0169] Optionally, as one implementation, the chip may further include a memory storing instructions, and the processor is configured to execute the instructions stored in the memory. When the instructions are executed, the processor is configured to perform the method described in any of the implementations of the first aspect.
[0170] In a seventh aspect, a voice interaction system is provided, the voice interaction system comprising a first device as described in any possible implementation of the first aspect above, and a second device as described in any possible implementation of the first aspect above, wherein the first device is configured to execute the method described in any possible implementation of the first aspect above.
[0171] The solution provided in this application allows the first device to determine its ability to independently recognize user voice commands. In voice interaction scenarios where the first device is relatively proficient, it can independently determine the operation corresponding to the user's voice command, thereby reducing the response latency and improving response efficiency. In voice interaction scenarios where the first device is relatively less proficient, it can choose to execute operations instructed by other devices, which helps improve the accuracy of its response to user voice commands. Furthermore, by processing the voice commands collected by the sensors both locally and via the cloud, and adaptively selecting to execute operations from either the local processor or the cloud, a balance between response efficiency and accuracy can be struck. The solution provided in this application allows the first device to continuously learn new voice commands, thus expanding its proficiency in various voice interaction scenarios. The solution also allows the first device to appropriately choose whether to exit multi-round voice interactions, thereby improving the voice interaction effect. Attached Figure Description
[0172] Figure 1 This is a schematic diagram of a voice interaction system.
[0173] Figure 2 This is a schematic diagram of a system architecture.
[0174] Figure 3 This is a schematic diagram of a voice interaction system.
[0175] Figure 4 This is a schematic diagram of a voice interaction system.
[0176] Figure 5 This is a voice interaction system provided in the embodiments of this application.
[0177] Figure 6 This is a schematic flowchart illustrating a voice interaction method provided in an embodiment of this application.
[0178] Figure 7 This is a schematic flowchart illustrating a voice interaction method provided in an embodiment of this application.
[0179] Figure 8 This is a schematic flowchart illustrating a voice interaction method provided in an embodiment of this application.
[0180] Figure 9 This is a schematic flowchart illustrating a voice interaction method provided in an embodiment of this application.
[0181] Figure 10 is a schematic flowchart of a voice interaction method provided in an embodiment of this application.
[0182] Figure 11This is a schematic structural diagram of a voice interaction device provided in an embodiment of this application.
[0183] Figure 12 This is a schematic structural diagram of a voice interaction device provided in an embodiment of this application. Detailed Implementation
[0184] The technical solutions of the embodiments of this application will be described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0185] The following section introduces several possible application scenarios for speech recognition.
[0186] Application Scenario 1: Application Scenario of Intelligent Driving In intelligent driving applications, users can control intelligent driving devices via voice commands. For example, users can issue voice commands to the in-vehicle voice assistant to control these devices. In some possible examples, users can use voice commands to adjust the seat back tilt, adjust the in-vehicle air conditioning temperature, turn the seat heaters on or off, turn the headlights on or off, open or close the windows, open or close the trunk, plan navigation routes, and play personalized playlists. In intelligent driving applications, voice interaction helps provide users with a convenient driving environment.
[0187] Application Scenario 2: Application Scenarios of Smart Home In smart home applications, users can control smart home devices via voice commands. For example, users can issue voice commands to IoT devices (e.g., smart home devices) or IoT control devices (such as mobile phones) to control them. In some possible examples, users can use voice commands to control the temperature of a smart air conditioner, control a smart TV to play a user-specified TV series, control smart cooking appliances to start at a user-specified time, control smart curtains to open or close, and control smart lights to adjust their color temperature. In smart home applications, voice interaction helps provide users with a comfortable home environment.
[0188] Figure 1 This is a schematic diagram of a voice interaction system 100.
[0189] The execution device 110 may be a device with speech recognition capabilities, natural language understanding capabilities, etc. For example, the execution device 110 may be a server. Optionally, the execution device 110 may also work in conjunction with other computing devices, such as data storage, routers, load balancers, etc. The execution device 110 may be deployed on a single physical site or distributed across multiple physical sites. The execution device 110 may use data from the data storage system 150 or call program code from the data storage system 150 to implement at least one of the following functions: speech recognition, machine learning, deep learning, model training, etc. Figure 1 The data storage system 150 can be integrated on the execution device 110, or it can be set up in the cloud or on other network servers.
[0190] Users can interact with execution device 110 by operating their respective local devices (such as local device 101 and local device 102). Figure 1 The local devices shown can represent, for example, various voice interaction terminals.
[0191] The user's local device can interact with the execution device 110 through a wired or wireless communication network. The type or standard of the communication network is not limited and can be a wide area network, a local area network, a point-to-point connection, or any combination thereof.
[0192] In one implementation, local device 101 can provide local data or feedback calculation results to execution device 110.
[0193] In another implementation, all or part of the functions of execution device 110 may be implemented by a local device. For example, local device 101 implements the functions of execution device 110 and provides services to its own users, or provides services to users of local device 102.
[0194] Figure 2 This is a schematic diagram of a system architecture 200.
[0195] The data acquisition device 260 can be used to collect training data. The data acquisition device 260 can also be used to store the training data in the database 230. The training device 220 can train the target model / rule 201 based on the training data maintained in the database 230. The trained target model / rule 201 can be used to execute the voice interaction method of this application embodiment. The training device 220 does not necessarily have to train the target model / rule 201 entirely based on the training data maintained in the database 230; it can also obtain training data from the cloud or other sources for model training. The above description should not be construed as limiting the embodiments of this application.
[0196] The training data maintained in the database 230 may not all originate from the data acquisition device 260; it may also be received from other devices. In one example, the training data in the database 230 can be obtained through the client device 240 or through the execution device 210. The client device 240 may include, for example, various voice interaction terminals. The execution device 210 may be a device with speech recognition capabilities, natural language understanding capabilities, etc. For example, by obtaining speech information through the data acquisition device 260 and performing related processing, training data such as text features of the input text and phonetic features of the target speech can be obtained; the data acquisition device 260 can also obtain text features of the input text and phonetic features of the target speech. Alternatively, speech information can be directly used as training data. In another example, the same account can log in to multiple client devices 240, and the data collected by these multiple client devices 240 can all be maintained in the database 230.
[0197] Optionally, the training data mentioned above may include one or more of the following: speech, corpus, and hot words. Speech refers to sounds that carry certain linguistic meaning. Corpus refers to language materials, which can refer to texts and their contextual relationships that describe language in the real world. Hot words are popular vocabulary. Hot words can be a lexical phenomenon that reflects the issues, topics, or things that people are relatively concerned about during a certain period.
[0198] In one possible example, the training data mentioned above may include input speech (which may come from the user or be speech acquired by other devices).
[0199] In another possible example, the training data described above could include feature vectors of the input speech (such as phonetic features, which could reflect the phonetic transcription of the input speech). The feature vectors of the input speech can be obtained by feature extraction from the input speech.
[0200] In another possible example, the training data mentioned above could include, for example, the target text corresponding to the input speech.
[0201] In another possible example, the training data described above could include, for instance, the text features of the target text corresponding to the input speech. The target text can be obtained by preprocessing the input speech for features. The text features of the target text can be obtained by feature extraction from the target text.
[0202] For example, assume that the pronunciation of the input voice is "nǐhǎo", and the corresponding target text can be "你好". Feature extraction of "nǐhǎo" can obtain the phonetic features of the input voice. After feature preprocessing and feature extraction of "nǐhǎo", the text features of the target text "你好" can be obtained.
[0203] It should be understood that the input voice can be sent by the client device 240 to the data acquisition device 260, can also be obtained by the data acquisition device 260 by reading from the storage device, or can be obtained through real-time acquisition.
[0204] Optionally, the data acquisition device 260 can determine the training data from the above phonetic features and / or text features.
[0205] The above feature preprocessing of the input voice can include processing such as normalization, Chinese character conversion of phonetic sounds, and prosodic pause prediction. Normalization can refer to converting non-Chinese characters such as numbers and symbols in the text into Chinese characters according to semantics. Chinese character conversion of phonetic sounds can refer to predicting the corresponding pinyin for each voice, and then generating a Chinese character text sequence for each voice. Prosodic pause prediction can refer to predicting stress marks, prosodic phrases, intonation phrase marks, etc.
[0206] Feature preprocessing can be performed by the data acquisition device 260, or can be performed by the client device 240 or other devices. When the data acquisition device 260 obtains the input voice, the data acquisition device 260 can perform feature preprocessing and feature extraction on the input voice to obtain the text features of the target text. Or, when the client device 240 obtains the input voice, the client device 240 can perform feature preprocessing on the input voice to obtain the target text; the data acquisition device 260 can perform feature extraction on the target text.
[0207] Taking Chinese as an example below, the feature extraction of the input voice will be introduced.
[0208] For example, if the pronunciation of the input voice is "nǐmenhǎo", the following phonetic features can be generated: S_n_i_3_SP0_m_en_0_SP1_h_ao_3_E In the phonetic features of "nǐmenhǎo", "S" can be a sentence start marker, or can be understood as a start marker; "E" can be a sentence end marker, or can be understood as an end marker; the numbers "0", "1", "2", "3", "4" can be tone markers; "SP0", "SP"1 can be different pause level markers; the initials and finals of Chinese pinyin can be used as phonemes; different phonemes / markers can be separated by a space "_". In this example, there can be 13 phonetic feature elements in the phonetic features.
[0209] For example, if the target text corresponding to another input speech is "Hello everyone", then the following text features can be generated: S_d_a_4_SP0_j_ia_1_SP1_h_ao_3_E In the text features of "Hello everyone", "S" can be a sentence beginning marker, or a start marker; "E" can be a sentence ending marker, or an end marker; the numbers "0", "1", "3", and "4" can be tone markers; "SP0" and "SP1" can be markers for different pause levels; the initials and finals of Pinyin can be used as phonemes; different phonemes / markers can be separated by spaces "_". In this example, there can be 13 text feature elements.
[0210] It should be noted that there is no limitation on the language used in this embodiment; in addition to the Chinese example mentioned above, other languages such as English, German, and Japanese can also be used. This embodiment mainly uses Chinese as an example for illustration.
[0211] The following describes the process by which training device 220 trains the target model / rule 201 based on language training data.
[0212] The training device 220 can input the acquired training data into the target model / rule 201. For example, it can compare the phonetic feature results output by the target model / rule 201 with the phonetic features corresponding to the current input speech, or it can compare the text feature results output by the target model / rule 201 with the text features corresponding to the current input speech, thereby completing the training of the target model / rule 201.
[0213] The target model / rule 201, trained using training device 220, can be a model built on a neural network. This neural network can be a convolutional neural network (CNN), a recurrent neural network (RNN), a long-short-term memory (LSTM) neural network, a bidirectional long-short-term memory (BLSTM) neural network, a deep convolutional neural network (DCNN), etc. Furthermore, the target model / rule 201 can be implemented based on a self-attention neural network. The type of target model / rule 201 can be, for example, an automatic speech recognition (ASR) model, a natural language processing (NLP) model, etc.
[0214] The target model / rule 201 obtained from the training device 220 described above can be applied to different systems or devices. Figure 2 In the system architecture 200 shown, the execution device 210 can be configured with an input / output (I / O) interface 212. Through this I / O interface 212, the execution device 210 can interact with external devices. For example... Figure 2 As shown, a "user" can input data to the I / O interface 212 through the client device 240. For example, the user can input intermediate prediction results to the I / O interface 212 through the client device 240, and then the client device 240 will send the intermediate prediction results, after certain processing, to the execution device 210 through the I / O interface 212. The intermediate prediction results may be, for example, the target text corresponding to the input speech.
[0215] Optionally, the training device 220 can generate corresponding target models / rules 201 based on different training data for different objectives or tasks. The corresponding target models / rules 201 can be used to achieve the above objectives or complete the above tasks, thereby providing the user with the required results.
[0216] The execution device 210 can call data, code, etc. in the data storage system 250, and can also store data, instructions, etc. in the data storage system 250.
[0217] Optionally, the execution device 210 can further segment the target model / rule 201 obtained from the training device 220 to obtain its sub-models / sub-rules, and deploy the obtained sub-models / sub-rules on the client device 240 and the execution device 210 respectively. In one example, the execution device 210 can send the personalized sub-model of the target model / rule 201 to the client device 240, which will then deploy it within its device. Optionally, the general sub-model of the target model / rule 201 does not update its parameters during training and therefore remains unchanged.
[0218] For example, training device 220 can obtain training data from database 230. Training device 220 can train a speech model from the training data. Training device 220 can send the trained speech model to execution device 210, which can then partition the speech model to obtain personalized speech sub-models and general speech sub-models. Alternatively, training device 220 can first partition the trained speech model to obtain personalized speech sub-models and general speech sub-models, and then send the personalized speech sub-models and general speech sub-models to execution device 210.
[0219] Optionally, the target model / rule 201 can be trained based on a base speech model. During training, a portion of the target model / rule 201 can be updated, while another portion may remain unchanged. The updated portion of the target model / rule 201 can correspond to a personalized speech sub-model. The unupdated portion of the target model / rule 201 can correspond to a general speech sub-model. The base speech model can be pre-trained by the training device 220 using speech data from multiple users, corpora, etc., or it can be an existing speech model.
[0220] Client device 240 and computing module 211 can work together. Client device 240 and computing module 211 can process data input to client device 240 and / or data input to execution device 210 (e.g., intermediate prediction results from client device 240) according to the aforementioned personalized and general speech sub-models. In one example, client device 240 can process the input user speech to obtain phonetic or textual features corresponding to the user speech; then, client device 240 can input these phonetic or textual features to computing module 211. In other examples, preprocessing module 213 of execution device 210 can receive input speech from I / O interface 112 and perform feature preprocessing and feature extraction on the input speech to obtain textual features of the target text. Preprocessing module 213 can input the textual features of the target text to computing module 211. Computing module 211 can input these phonetic or textual features into target model / rule 201 to obtain the output result of speech recognition (e.g., semantic recognition result, operation corresponding to voice command, etc.). The calculation module 211 can input the output result to the client device 240, so that the client device 240 can perform corresponding operations to respond to the user's voice commands.
[0221] I / O interface 212 can send input data to the corresponding module of execution device 210, or return the output result to client device 240 for the user. For example, I / O interface 212 can send the intermediate prediction result corresponding to the input speech to calculation module 211, or return the result obtained after speech recognition to client device 240.
[0222] exist Figure 2 In the system architecture 200 shown, users can input voice, speech data, and other data into the client device 240, and view the results output by the execution device 210 on the client device 240. The specific presentation format can be sound or a combination of sound and display. The client device 240 can also act as a data acquisition terminal, storing the acquired voice, speech data, and other data into the database 230. Alternatively, data can be collected without going through the client device 240; other devices can collect the user's voice, speech data, and the output results of the I / O interface 212 as new sample data and store them into the database 230.
[0223] exist Figure 2In the system architecture 200 shown, the execution device 210 and the data storage system 250 can be integrated into different devices depending on the data processing capabilities of the client device 240. For example, when the client device 240 has strong data processing capabilities, the execution device 210 and the data storage system 250 can be integrated into the client device 240; while when the client device 240 has weak data processing capabilities, the execution device 210 and the data storage system 250 can be integrated into a dedicated data processing device. Figure 2 The database 230, training device 220, and data acquisition device 260 can be integrated into a dedicated data processing device, or set up on other servers in the cloud or on the network, or set up separately in the client device 240 and the data processing device.
[0224] It is worth noting that, Figure 2 This is merely a schematic diagram of a system architecture provided in an embodiment of this application. Figure 2 The positional relationships between the devices, components, modules, etc., shown do not constitute any limitation. For example, in Figure 2 In this context, the data storage system 250 is an external memory relative to the execution device 210. However, in other cases, the data storage system 250 may also be located within the execution device 210. For example, in some possible examples, the execution device 210 may be located within the client device 240. The general speech sub-model of the target model / rule 201 may be the factory-installed speech model of the client device 240. After the client device 240 leaves the factory, the personalized speech sub-model of the target model / rule 201 can be updated based on the data collected by the client device 240.
[0225] To better understand the solutions of the embodiments of this application, the following will first combine... Figure 3 , Figure 4 This section introduces some voice interaction systems. Among them, Figure 3 , Figure 4 This could be a diagram of two voice interaction systems.
[0226] Figure 3 The voice interaction system 300 shown may include at least one first device. The first device may be various types of voice interaction terminals.
[0227] exist Figure 3 In the voice interaction system 300 shown, the first device can acquire the user's voice commands through an interactive interface. The first device can recognize the voice commands and obtain the recognition results. The first device can perform corresponding operations based on the recognition results to respond to the user's voice commands. Optionally, the first device can also perform related processing such as machine learning, deep learning, and model training through a memory for storing data and a processor for data processing.
[0228] Figure 3 The voice interaction system 300 shown is Figure 1 The voice interaction system 100 shown can have certain corresponding relationships. Figure 3 In the voice interaction system 300 shown, the first device can be, for example, equivalent to Figure 1 The local device 101 or local device 102 shown.
[0229] Figure 3 The voice interaction system 300 shown is Figure 2 The system architecture 200 shown can have certain corresponding relationships. Figure 3 In the voice interaction system 300 shown, the first device can be, for example, equivalent to Figure 2 The client equipment 240 shown is optional. Figure 3 The first device shown may have Figure 2 The execution device 210 shown has some or all of its functions.
[0230] exist Figure 3 In the voice interaction system 300 shown, voice recognition can primarily rely on the first device. If the first device has strong voice recognition and natural language understanding capabilities, it can typically respond quickly to the user's voice commands. This means... Figure 3 The voice interaction system 300 shown places relatively high demands on the processing power of the first device. If the first device's speech recognition and natural language understanding capabilities are relatively weak, the first device will be unable to accurately respond to the user's voice commands, which may reduce the user's voice interaction experience.
[0231] Figure 4 The voice interaction system 400 shown may include at least one first device and at least one second device. The first device may be various types of voice interaction terminals. The second device may be a cloud device, such as a server, that has speech recognition capabilities, natural language understanding capabilities, etc.
[0232] exist Figure 4 In the voice interaction system 400 shown, a first device can receive or acquire a user's voice, which may contain the user's voice commands. The first device can forward the user's voice to a second device. The first device can obtain the result of recognizing the voice from the second device. The first device can perform operations based on the result of recognizing the voice to respond to the user's voice commands.
[0233] exist Figure 4 In the voice interaction system 400 shown, the second device can acquire voice from the first device through an interactive interface and perform voice recognition on the voice. The second device can also forward the result of the recognized voice to the first device. Figure 4The term "memory" can be a general term that includes local storage as well as databases that store historical data. Figure 4 The database can be on a second device or on other devices.
[0234] Optionally, both the first and second devices can perform machine learning, deep learning, model training, speech recognition, and other related processing through a memory for storing data and a processor for processing data.
[0235] Figure 4 The voice interaction system 400 shown is Figure 1 The voice interaction system 100 shown can have certain corresponding relationships. Figure 4 In the voice interaction system 400 shown, the first device can be, for example, equivalent to Figure 1 The local device 101 or local device 102 shown; the second device may be equivalent to, for example, local device 101 or local device 102. Figure 1 The execution device 110 shown.
[0236] Figure 4 The voice interaction system 400 shown is Figure 2 The system architecture 200 shown can have certain corresponding relationships. Figure 4 In the voice interaction system 400 shown, the first device can be, for example, equivalent to Figure 2 The client equipment 240 shown; the second device may be equivalent to, for example, the client equipment 240. Figure 2 The execution device 210 shown.
[0237] exist Figure 4 In the voice interaction system 400 shown, voice recognition can primarily rely on the second device. The first device may not process the voice, or may only perform simple preprocessing. Figure 4 The voice interaction system 400 shown reduces the processing power requirements of the first device. However, the interaction between the second device and the first device may introduce latency. This could hinder the first device from responding quickly to the user's voice commands, potentially reducing the user's voice interaction experience.
[0238] Figure 5 This application provides a voice interaction system 500. The voice interaction system 500 may include a first device and a second device.
[0239] The first device can be various voice interaction devices, such as cars, in-vehicle systems, on-board computers (or on-board PCs), chips (such as automotive chips, voice processing chips, etc.), processors, mobile phones, personal computers, smart bracelets, tablets, smart cameras, set-top boxes, game consoles, voice-enabled in-vehicle devices, smart cars, media consumption devices, smart home devices, smart voice assistants on wearable devices that can produce sound, smart speakers, or various machines or devices that can converse with people. The first device can be, for example, […]. Figure 1 The local device 101 or local device 102 shown, or a unit or module of local device 101 or local device 102. The first device may also be, for example, a... Figure 2 The client equipment 240 shown is either a unit or module of the client equipment 240.
[0240] Optionally, the first device may include, for example, a semantic recognition module, an operation decision module, and a transceiver module.
[0241] The semantic recognition module can be used to recognize speech information from a speech sensor. Speech information can be, for example, audio directly acquired by the speech sensor, or a processed signal carrying the content of that audio. The semantic recognition module can output the semantic recognition result. The semantic recognition result can be structured information that reflects the user's speech content. The semantic recognition module can, for example, store speech interaction models, such as ASR models or NLP models. The semantic recognition module can perform semantic recognition on speech information using these speech interaction models.
[0242] The first transceiver module can be used to forward voice information from the voice sensor to the second device and obtain voice analysis results indicated by the second device from the second device. The voice analysis results may include, for example, semantic recognition results and / or operation information.
[0243] The operation decision module can be used to determine the response to a user's operation. For example, the operation decision module can obtain the semantic recognition result from the semantic recognition module; based on the semantic recognition result, the operation decision module can determine to execute the corresponding operation. As another example, the operation decision module can obtain the voice analysis result indicated by the second device from the transceiver module, and then determine to execute the operation indicated by the second device. Furthermore, the operation decision module can determine whether the operation responding to voice information is indicated by the semantic recognition module or by the second device.
[0244] The second device can be independent of the first device. The second device may be, for example, a remote service platform or server. The second device can be implemented by one or more servers. Servers may include, for example, one or more of the following: web servers, application servers, management servers, cloud servers, edge servers, virtual servers (servers virtualized using multiple physical resources). The second device can also be other devices with speech recognition capabilities, natural language understanding capabilities, etc., such as in-vehicle infotainment systems, in-vehicle computers, chips (e.g., in-vehicle chips, voice processing chips), processors, mobile phones, personal computers, or tablet computers, etc., devices with data processing functions. The second device may be, for example, a remote service platform or server. Figure 1 The illustrated execution device 110 may be a unit or module of the execution device 110. The second device may also be... Figure 2 The execution device 210 shown may be a unit or module of the execution device 210.
[0245] Optionally, the second device may include, for example, a second transceiver module and a voice analysis module.
[0246] The second transceiver module can be used to acquire voice information from the first device. The second transceiver module can also be used to send the voice analysis results output by the voice analysis module to the first device.
[0247] The speech analysis module can be used to acquire speech information from the first device via the second transceiver module and perform speech analysis on that speech information. For example, the speech analysis module can store speech interaction models, AS (Automatic Speech), NLP (Natural Language Processing) models, etc. The speech analysis module can then process the speech information using these speech interaction models.
[0248] In one example, the speech analysis module can output the semantic recognition results obtained after semantic recognition.
[0249] In another example, the speech analysis module can analyze the semantic recognition results to obtain operation information corresponding to the speech information. This operation information can be used to instruct the operation to be performed by the first device. The operation information can be represented in the form of, for example, a command.
[0250] Figure 6 This is a schematic flowchart of a voice interaction method 600 provided in an embodiment of this application. Figure 6 The method 600 shown can be applied, for example, to... Figure 5 The voice interaction system 500 shown is shown.
[0251] 601, The first device responds to the user's input operation and performs a voice wake-up operation.
[0252] For example, a user can speak a wake word to the first device. The first device can detect the wake word input by the user and thus activate the voice interaction function of the first device.
[0253] For example, a user can press a wake-up button to input an operation into the first device. The first device can detect the user's input to the wake-up button and thus activate the voice interaction function of the first device.
[0254] Step 601 can be an optional step. Optionally, the first device may include a wake-up module that can be used to detect user input.
[0255] 602, The first device acquires first voice information from the voice sensor.
[0256] A voice sensor can be used to record a user's voice. The voice sensor can be a device with recording capabilities. For example, a voice sensor may include a microphone.
[0257] Optionally, the voice sensor can also be a device with recording and data processing capabilities. The voice sensor can also perform related voice processing on the user's voice, such as noise reduction, amplification, and code modulation.
[0258] In other examples, the above-mentioned speech processing can also be performed by other modules.
[0259] The first voice information can be, for example, audio directly acquired by a voice sensor, or a processed signal carrying the audio content.
[0260] In one example, the user utters a first voice command, which instructs the first device to perform a target operation. The first voice command can be converted into first voice information using a voice sensor, and this first voice information can then be used to instruct the target operation.
[0261] 603a, the first device determines, based on the first voice information, to execute the target operation indicated by the first voice information.
[0262] By analyzing the first voice information, the first device can perceive or understand the specific meaning of the first voice command, and thus determine the target operation indicated by the first voice information. Optionally, the first device can execute the target operation in response to the user's first voice command.
[0263] Optionally, the first device determines the target operation indicated by the first voice information based on the first voice information, including: the first device determining a first semantic recognition result based on the first voice information; and the first device determining to execute a first operation based on the first semantic recognition result. The first operation may be an operation determined by the first device based on the first voice information. The first operation may correspond to the aforementioned target operation.
[0264] The process of converting first speech information into a first semantic recognition result may include, for example, converting first speech information containing audio content into first text information; performing semantic extraction on the first text information to obtain a structured first semantic recognition result. The first semantic recognition result may include one or more of the following structured information: function, intent, and parameters. The function of the first semantic recognition result can be a specific value of a functional result. The function of the first semantic recognition result can represent a class of strategies. The function of the first semantic recognition result can also be referred to as a domain. The intent can be a specific value of an intent result. The parameters can be specific values of slot results.
[0265] The following example illustrates one possible method by which the first device determines the first operation based on first semantic information.
[0266] The user can speak a first voice command to the first device. The first voice command can be converted into first text information, such as "Turn on device A." Device A can be, for example, an air conditioner. The natural meaning of the first text information can indicate the first operation. However, based solely on the first text information, the first device usually cannot directly determine the natural meaning of the first voice command, nor can it determine the operation indicated by the first voice command.
[0267] The first textual information can be converted into a first semantic recognition result, and the first semantic recognition result can indicate a first operation.
[0268] In one example, the first semantic recognition result may include a first function, a first intent, and a first parameter; the first function may be, for example, "vehicle control function"; the first intent may be "open"; and the first parameter may be "device A". Based on the first semantic recognition result, the first device may determine that: the first operation may be an operation within the vehicle control function; the intent of the first operation may be "open" or "enable"; and the object being opened or enabled may be device A.
[0269] In one example, the first semantic recognition result may include a first function, a first intent, and a first parameter; the first function may be, for example, "vehicle control function"; the first intent may be "turn on device A"; and the first parameter may be "empty". The first device may determine, based on the first semantic recognition result, that: the first operation may be an operation within the vehicle control function; the intent of the first operation may be to turn on device A or to put device A in an on state; and the first semantic recognition result may not include specific parameters of the on state of device A (such as temperature parameters, timing parameters, mode parameters, etc.).
[0270] In one example, the first semantic recognition result may include a first function and a first parameter; the first function may be, for example, "vehicle control function"; the first parameter may be "device A". The first device may determine, based on the first semantic recognition result, that: the first operation may be an operation within the vehicle control function; and the object of the first operation may be device A.
[0271] In one example, the first semantic recognition result may include a first intent and a first parameter; the first function may be, for example, "open"; the first intent may be "device A". The first device may determine, based on the first semantic recognition result, that: the intent of the first operation may be "open" or "enable"; and the object to be opened or enabled may be device A.
[0272] In one example, the first semantic recognition result may include a first intent; the first intent may be "turn on device A". The first device may determine, based on the first semantic recognition result, that the intent of the first operation may be to turn on device A, or to put device A in an on state; the first semantic recognition result may not include specific parameters of the on state of device A (such as temperature parameters, timing parameters, mode parameters, etc.).
[0273] In one example, the first semantic recognition result may include a first parameter; the first parameter may be "device A". The first device may determine, based on the first semantic recognition result, that the object of the first operation may be device A.
[0274] Optionally, the first device determines the target operation indicated by the first voice information based on the first voice information, including: the first device determines a first semantic recognition result based on the first voice information; the first device determines the first operation to be performed based on the first semantic recognition result and a first preset condition.
[0275] In other words, the first device can determine whether a first preset condition is met, and then, based on the determination result, determine whether to execute the first operation determined by the first device based on the first semantic recognition result. For example, if the first semantic recognition result meets the first preset condition, the first device can execute 603a.
[0276] 603b, the first device sends the first voice information to the second device.
[0277] For example, the first device is a mobile phone, and the second device is a server. That is to say, the mobile phone can send the first voice information to the server.
[0278] For example, the first device is a mobile phone, and the second device is a mobile phone. That is to say, the mobile phone can send the first voice message to the mobile phone.
[0279] One possible approach is for the first device to send the first voice information to the second device via another device. This other device could be, for example, a mobile phone.
[0280] In one example, 603b and 603a can be executed synchronously or sequentially within a certain time period. The execution order of 603b and 603a is not limited in this embodiment. Optionally, if the first semantic recognition result satisfies the first preset condition, the first device can also execute 603b.
[0281] Both the first device and the second device can perform semantic recognition on the first voice information.
[0282] Optionally, the second device may send a sixth semantic recognition result to the first device based on the first voice information.
[0283] In other words, the second device can provide the first device with the recognition result based on the first voice information.
[0284] Optionally, based on the first voice information, the first device and the second device may obtain the same or different semantic recognition results. For example, the first device may determine a first semantic recognition result based on the first voice information, and the first semantic recognition result may be used to indicate a first operation; the second device may determine a sixth semantic recognition result based on the first voice information, and the sixth semantic recognition result may be used to indicate a second operation; at least one of the first operation and the second operation may correspond to the aforementioned target operation.
[0285] The content of 603a describes some possible examples of the first device determining the first semantic recognition result based on the first voice information. The method by which the second device determines the sixth semantic recognition result based on the first voice information can refer to 603a, and will not be described in detail here.
[0286] Optionally, in addition to semantic recognition results, the second device can also instruct the first device on a target operation through target operation information. Target operation information can, for example, be represented by operation signaling.
[0287] In one possible scenario, the first device can determine a first operation in response to a first voice command based on the first semantic recognition result it analyzes. Optionally, the first device can obtain a sixth semantic recognition result and / or target operation information from a second device. The sixth semantic recognition result and / or target operation information can be determined by the second device based on the first voice information. The sixth semantic recognition result can be used to reflect the meaning of the first voice information and to indirectly instruct the first device to perform a second operation; the target operation information can be used to directly instruct the first device to perform the second operation. Since the sixth semantic recognition result and / or target operation information fed back by the second device may have a time delay, the first device can obtain the sixth semantic recognition result and / or target operation information after performing the first operation. The first device can adjust its own semantic recognition model, voice control model, etc., based on the sixth semantic recognition result and / or target operation information to improve the accuracy of the semantic recognition result output by the first device and optimize the applicability of the operation in response to user voice commands.
[0288] Optionally, the first device may discard the sixth semantic recognition result from the second device.
[0289] The first device may receive the sixth semantic recognition result and then discard it. Alternatively, the first device may choose not to receive the sixth semantic recognition result. In other words, if the first device can determine the operation to be performed based on the first voice information, the second device provides feedback on the first voice information.
[0290] In another possible scenario, before the first device determines the target operation to be performed based on the first voice information, the first device may be unsure whether it has the ability to recognize voice information. In other words, the first device may not be able to determine the target operation indicated by the first voice information. The first device may send the first voice information to the second device in advance. The second device may recognize and process the first voice information to obtain a sixth semantic recognition result and / or target operation information; the sixth semantic recognition result can be used to reflect the meaning of the first voice information and to indirectly instruct the first device to perform the target operation; the target operation information can be used to directly instruct the first device to perform the target operation. If the first device can determine the target operation based on the first voice information, the first device may skip or discard the result fed back by the second device based on the first voice information. If the first device cannot determine the target operation based on the first voice information, since the first device has sent the first voice information to the second device in advance, the first device can obtain the sixth semantic recognition result and / or target operation information from the second device relatively faster. Therefore, this helps to shorten the time for the first device to respond to user commands.
[0291] In another example, only one of 603b and 603a is executed by the first device.
[0292] In other words, the first device can determine that only one of the following will be executed: the first device determines the first operation to be executed based on the first voice information; the first device determines the second operation to be executed based on the result fed back by the second device.
[0293] For example, in voice interaction scenarios where the first device excels, it can determine the operation corresponding to the user's voice command without relying on information provided by the second device. This improves the efficiency of the first device in responding to user voice commands. Furthermore, the first device can choose not to send the first voice information to the second device, thus reducing the amount of signaling transmission. In voice interaction scenarios where the first device is less adept, it can determine the operation corresponding to the user's voice command based on information provided by the second device. This improves the accuracy of the first device in responding to user voice commands.
[0294] Optionally, the first device determines a first semantic recognition result based on the first voice information; the first device determines to execute a second operation instructed by the second device based on the first semantic recognition result and a first preset condition.
[0295] The first device can determine whether the first preset condition is met, and then determine whether to execute the second operation instructed by the second device based on the determination result. For example, if the first semantic recognition result does not meet the first preset condition, the first device can execute 603b and not execute 603a; if the first semantic recognition result meets the first preset condition, the first device can execute 603a and not execute 603b.
[0296] The following examples illustrate how the first device determines, based on the first semantic recognition result and the first preset conditions, whether to execute the first operation determined by the first device based on the first semantic recognition result or to execute the second operation instructed by the second device.
[0297] Optionally, if the first semantic recognition result satisfies the first preset condition, the first device may determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0298] If the first semantic recognition result meets the first preset condition, the first device can determine to perform the corresponding operation based on the first semantic recognition result, so that the first device can respond to the user's voice command.
[0299] Optionally, if the first semantic recognition result does not meet the first preset condition, a second operation instructed by the second device is determined to be performed.
[0300] If the first semantic recognition result does not meet the first preset condition, the first device can determine the operation to be performed according to the instruction of the second device, so that the first device can respond to the user's voice command. In one possible example, the second device can instruct the first device to perform a second operation through the semantic recognition result and / or operation information.
[0301] Optionally, determining to execute the second operation indicated by the second device based on the first semantic recognition result and the first preset condition includes: if the first semantic recognition result does not meet the first preset condition, obtaining a sixth semantic recognition result from the second device, the sixth semantic recognition result being used to indicate the second operation; and determining to execute the second operation based on the sixth semantic recognition result.
[0302] The second device can recognize the first voice information to obtain a sixth semantic recognition result. The second device can use the sixth semantic recognition result to indirectly instruct the first device to perform a second operation, enabling the first device to respond to the user's voice command. In other words, the first device can determine the operation to be performed as the second operation based on the sixth semantic recognition result.
[0303] Optionally, the first semantic recognition result satisfying the first preset condition means that the function included (corresponding to or indicating) in the first semantic recognition result is a function with a preset priority level by the first device. The first device may preset multiple functions. When the first function included (corresponding to or indicating) in the first semantic recognition result belongs to the multiple functions, the first semantic recognition result satisfies the first preset condition. Taking the preset multiple functions as a first semantic recognition list as an example, other preset multiple functions are similar. The method further includes: the first device obtaining a first semantic recognition list, the first semantic recognition list including multiple functions; the first semantic recognition result satisfying the first preset condition includes: the first semantic recognition result includes a first function, and the first function belongs to the multiple functions.
[0304] Optionally, the first semantic recognition result not satisfying the first preset condition means that the function included (corresponding to or indicating) in the first semantic recognition result is not a function with a preset priority level by the first device. The first device may preset multiple functions. When the first function included (corresponding to or indicating) in the first semantic recognition result does not belong to the multiple functions, the first semantic recognition result does not satisfy the first preset condition. Taking the preset multiple functions as a first semantic recognition list as an example, other preset multiple functions are similar. The method further includes: the first device obtaining a first semantic recognition list, the first semantic recognition list including multiple functions; the first semantic recognition result not satisfying the first preset condition includes: the first semantic recognition result includes a first function, and the first function does not belong to the multiple functions.
[0305] In other words, the first device can perform semantic recognition on the first voice information to obtain a first semantic recognition result, which may include a first function; the first device can search for the first function in the first semantic recognition list. In one example, if the first semantic recognition list includes the first function, the first device can determine that the first semantic recognition result meets a first preset condition, and then the first device can determine to execute a first operation determined by the first device based on the first semantic recognition result. In another example, if the first semantic recognition list does not include the first function, the first device can determine that the first semantic recognition result does not meet the first preset condition, and then the first device can determine to execute a second operation indicated by the second device.
[0306] The first semantic recognition list may be a list pre-stored by the first device. The first semantic recognition list may include multiple functions supported by the first device. The first device may have relatively higher semantic recognition capabilities for multiple functions in the first semantic recognition list. The first device may have relatively lower semantic recognition capabilities for other functions outside the first semantic recognition list. If a first function is included in the first semantic recognition list, it may mean that the accuracy of the first semantic recognition result is relatively high. If the first function is not included in the first semantic recognition list, it may mean that the accuracy of the first semantic recognition result is relatively low.
[0307] For example, multiple functions in the first semantic recognition list may include "vehicle control function". That is, the semantic recognition capability of the first device can be relatively good for "vehicle control function". The first text information may be, for example, "turn on device A". The first text information may indicate "vehicle control function". The first semantic recognition result may include a first function, which may be "vehicle control function". The first device can determine, based on the first semantic recognition list and the first semantic recognition result, that the first semantic recognition result meets a first preset condition, and determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0308] For example, the first semantic recognition list may include multiple functions such as "local translation function" but not "cloud translation function". That is, the semantic recognition capability of the first device can be relatively good for the "local translation function"; however, its semantic recognition capability can be relatively poor for the "cloud translation function" (e.g., the first device may not be able to promptly identify current hot words or translate all foreign languages). The first text information may be, for example, "Translate the following content using foreign language A", where "foreign language A" may not be a foreign language that the first device can translate. The first semantic recognition result may include the first function, which may be the "cloud translation function". The first device can determine, based on the first semantic recognition list and the first semantic recognition result, that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0309] Optionally, the first semantic recognition result satisfying the first preset condition means that the intent included (corresponding to or indicating) in the first semantic recognition result is an intent with a preset priority level by the first device. The first device may preset multiple intents. When the first intent included (corresponding to or indicating) in the first semantic recognition result belongs to the multiple intents, the first semantic recognition result satisfies the first preset condition. Taking the preset multiple intents as a second semantic recognition list as an example, other forms of preset multiple intents are similar. The method further includes: the first device obtaining a second semantic recognition list, the second semantic recognition list including multiple intents; the first semantic recognition result satisfying the first preset condition includes: the first semantic recognition result includes a first intent, and the first intent belongs to the multiple intents.
[0310] Optionally, the first semantic recognition result not satisfying the first preset condition means that the intent included (corresponding to or indicating) in the first semantic recognition result is not an intent with a preset priority level by the first device. The first device may preset multiple intents. When the first intent included (corresponding to or indicating) in the first semantic recognition result does not belong to the multiple intents, the first semantic recognition result does not satisfy the first preset condition. Taking the preset multiple intents as a second semantic recognition list as an example, other preset multiple intents are similar. The method further includes: the first device obtaining a second semantic recognition list, the second semantic recognition list including multiple intents; the first semantic recognition result not satisfying the first preset condition includes: the first semantic recognition result includes a first intent, and the first intent does not belong to the multiple intents.
[0311] The first device can perform semantic recognition on the first voice information to obtain a first semantic recognition result, which may include a first intent. The first device can search for the first intent in a second semantic recognition list. In one example, if the second semantic recognition list includes the first intent, the first device can determine that the first semantic recognition result meets a first preset condition, and then the first device can determine to execute a first operation determined by the first device based on the first semantic recognition result. In another example, if the second semantic recognition list does not include the first intent, the first device can determine that the first semantic recognition result does not meet the first preset condition, and then the first device can determine to execute a second operation indicated by the second device.
[0312] The second semantic recognition list may be, for example, a list pre-stored by the first device. The second semantic recognition list may include, for example, multiple intents supported by the first device. The first device may have relatively higher semantic recognition capabilities for multiple intents in the second semantic recognition list. The first device may have relatively lower semantic recognition capabilities for other intents outside the second semantic recognition list. If the first intent is included in the second semantic recognition list, it may mean that the accuracy of the first semantic recognition result is relatively high. If the first intent is not included in the second semantic recognition list, it may mean that the accuracy of the first semantic recognition result is relatively low.
[0313] For example, multiple intents in the second semantic recognition list may include "open device A". That is, the semantic recognition capability of the first device for the intent "open device A" can be relatively good. The first text information may, for example, be "open device A". The first text information can indicate "open device A". The first semantic recognition result may include the first intent, which can be "open device A". The first device can determine, based on the second semantic recognition list and the first semantic recognition result, that the first semantic recognition result meets a first preset condition, and determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0314] For example, the multiple intents in the second semantic recognition list may include "local audio playback intent" but not "cloud audio playback intent". That is, for "local audio playback intent", the semantic recognition capability of the first device can be relatively good (e.g., the first device can identify the audio resource indicated in the voice command based on locally stored audio data); for "cloud audio playback intent", the semantic recognition capability of the first device can be relatively poor (e.g., the first device may not support the ability to recognize cloud audio data). The first text information may be, for example, "start playing from minute 1 of song A", where "song A" may belong to cloud audio data. The first semantic recognition result may include the first intent, which may be "play song A". The first device can determine, based on the second semantic recognition list and the first semantic recognition result, that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0315] Optionally, the first semantic recognition result satisfying the first preset condition means that the parameters included (corresponding to or indicating) in the first semantic recognition result indicate that the first device presets information such as functions, intentions, scenarios, devices, or locations with priority processing levels. The first device can preset multiple parameters. When the first parameter included in the first semantic recognition result belongs to the multiple parameters, the first semantic recognition result satisfies the first preset condition. Taking the preset multiple parameters as a third semantic recognition list as an example, other preset multiple parameter forms are similar. The method further includes: the first device obtaining a third semantic recognition list, the third semantic recognition list including multiple parameters; the first semantic recognition result satisfying the first preset condition includes: the first semantic recognition result includes a first parameter, and the first parameter belongs to the multiple parameters.
[0316] Optionally, the first semantic recognition result not satisfying the first preset condition means that the parameters included (corresponding to or indicating) in the first semantic recognition result indicate that the first device does not have priority processing level information such as functions, intentions, scenarios, devices, or locations. The first device can preset multiple parameters. When the first parameter included in the first semantic recognition result does not belong to the multiple parameters, the first semantic recognition result does not satisfy the first preset condition. Taking the preset multiple parameters as a third semantic recognition list as an example, other preset multiple parameter forms are similar. The method further includes: the first device obtaining a third semantic recognition list, the third semantic recognition list including multiple parameters; the first semantic recognition result not satisfying the first preset condition includes: the first semantic recognition result includes a first parameter, and the first parameter does not belong to the multiple parameters.
[0317] The first device can perform semantic recognition on the first speech information to obtain a first semantic recognition result, which may include a first parameter. The first device can search for the first parameter in a third semantic recognition list. In one example, if the third semantic recognition list includes the first parameter, the first device can determine that the first semantic recognition result meets a first preset condition, and then the first device can determine to execute a first operation determined by the first device based on the first semantic recognition result. In another example, if the third semantic recognition list does not include the first parameter, the first device can determine that the first semantic recognition result does not meet the first preset condition, and then the first device can determine to execute a second operation indicated by the second device.
[0318] The third semantic recognition list may be a list pre-stored by the first device. The third semantic recognition list may include multiple parameters supported by the first device. The first device may have relatively higher semantic recognition capabilities for multiple parameters in the third semantic recognition list. The first device may have relatively lower semantic recognition capabilities for other parameters outside the third semantic recognition list. If the first parameter is included in the third semantic recognition list, it may mean that the accuracy of the first semantic recognition result is relatively high. If the first parameter is not included in the third semantic recognition list, it may mean that the accuracy of the first semantic recognition result is relatively low.
[0319] For example, multiple parameters in the third semantic recognition list may include "Device A". That is, the semantic recognition capability of the first device for the parameter "Device A" can be relatively good. The first text information may be, for example, "Open Device A". The first text information may indicate "Device A". The first semantic recognition result may include a first parameter, which may be "Device A". The first device can determine, based on the third semantic recognition list and the first semantic recognition result, that the first semantic recognition result meets a first preset condition, and determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0320] For example, the parameters in the third semantic recognition list may not include "Location A". That is, the semantic recognition capability of the first device may be relatively poor for the parameter "Location A". The first text information may be, for example, "Navigate to Location B, via Location A". The first semantic recognition result may include the first parameter, which may be "Location A". The first device can determine, based on the third semantic recognition list and the first semantic recognition result, that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0321] The various implementation methods described above for determining whether the first semantic recognition result meets the first preset condition can be used independently or in combination. For example, optionally, the first semantic recognition result meeting the first preset condition may include at least two of the following: the first semantic recognition result includes a first function, and the first function belongs to the first semantic recognition list; the first semantic recognition result includes a first intent, and the first intent belongs to the second semantic recognition list; the first semantic recognition result includes a first parameter, and the first parameter belongs to the third semantic recognition list.
[0322] For example, the first function of the first semantic recognition result is "vehicle control function", the first intent of the first semantic recognition result is "open", and the first parameter of the first semantic recognition result is "device A". If "vehicle control function" belongs to the first semantic recognition list, "open" belongs to the second semantic recognition list, and "device A" belongs to the third semantic recognition list, then the first semantic recognition result can satisfy the first preset condition.
[0323] In one possible example, the first semantic recognition list, the second semantic recognition list, and the third semantic recognition list could be three independent lists.
[0324] In another possible example, at least two of the first semantic identification list, the second semantic identification list, and the third semantic identification list may come from the same list. For example, the first semantic identification list, the second semantic identification list, and the third semantic identification list may belong to a master table. This master table may, for example, include multiple sublists, and the first semantic identification list, the second semantic identification list, and the third semantic identification list may, for example, be three sublists of this master table.
[0325] Optionally, the first semantic recognition list may further include multiple intents corresponding to the first function. If the first semantic recognition result satisfies a first preset condition, it may also include: the first semantic recognition result may further include a first intent, and the first intent belongs to the multiple intents.
[0326] Optionally, the first semantic recognition list may also include multiple intents corresponding to the first function. If the first semantic recognition result does not meet the first preset condition, it may also include: the first semantic recognition result may also include a first intent, and the first intent does not belong to the multiple intents.
[0327] The intents corresponding to the first function are usually not unlimited. For example, a navigation function may correspond to intents such as route planning or voice packs; however, a navigation function typically does not correspond to intents such as hardware activation or deactivation. Similarly, an audio function may correspond to intents such as audio playback or lyrics; however, an audio function typically does not correspond to intents such as route planning. The first device can pre-record the correspondence between multiple functions and multiple intents in a first semantic recognition list.
[0328] If the first semantic recognition list indicates that the first function and the first intent of the first semantic recognition result correspond, then the first semantic recognition list can satisfy the first preset condition.
[0329] If the first function and the first intent of the first semantic recognition result both belong to the first semantic recognition list, but the first intent and the first function do not correspond in the first semantic recognition list, then it can be determined that the first semantic recognition result does not meet the first preset condition.
[0330] If one or more of the first function and first intent of the first semantic recognition result do not belong to the first semantic recognition list, then it can be determined that the first semantic recognition result does not meet the first preset condition.
[0331] For example, the first function of the first semantic recognition result is "vehicle control function", and the first intent of the first semantic recognition result is "open device A"; when both "vehicle control function" and "open device A" belong to the first semantic recognition list, and "open device A" corresponds to "vehicle control function" in the first semantic recognition list, the first semantic recognition result can satisfy the first preset condition. The first device can determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0332] For example, the first function of the first semantic recognition result is "navigation function," and the first intent of the first semantic recognition result is "play song A." The title of "song A" and the place name of "location A" can be the same in words. That is, the same words can represent different meanings. Multiple functions of the first semantic recognition list can include "navigation function"; multiple intents of the first semantic recognition list can include "play song A." However, in the first semantic recognition list, "navigation function" and "play song A" do not correspond. In the first semantic recognition list, "navigation function" can correspond to "passing through location A," "destination location A," "originating location A," etc.; in the first semantic recognition list, "audio function" can correspond to "play song A," etc. Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device can determine to execute the second operation instructed by the second device.
[0333] For example, the first function of the first semantic recognition result is "vehicle control function," and the first intent of the first semantic recognition result is "open device B." Multiple functions in the first semantic recognition list may include "vehicle control function"; multiple intents in the first semantic recognition list may not include "open device B" (e.g., device B is not installed in the vehicle). Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0334] Optionally, the first semantic recognition list may further include multiple parameters corresponding to the first function, and the first semantic recognition result may satisfy a first preset condition, and may also include: the first semantic recognition result may further include a first parameter, and the first parameter belongs to the multiple parameters.
[0335] Optionally, the first semantic recognition list may further include multiple parameters corresponding to the first function. If the first semantic recognition result does not meet the first preset condition, it may also include: the first semantic recognition result may further include a first parameter, and the first parameter does not belong to the multiple parameters.
[0336] The parameters corresponding to the first function are usually not unlimited. For example, a navigation function may correspond to parameters such as location; however, a navigation function typically does not correspond to parameters such as audio playback mode. Similarly, an audio function may correspond to parameters such as artist and song; however, an audio function typically does not correspond to parameters such as temperature. The first device can pre-record the correspondence between multiple functions and multiple parameters in a first semantic recognition list.
[0337] If the first semantic recognition list indicates that the first function and the first parameter of the first semantic recognition result correspond, then the first semantic recognition list can satisfy the first preset condition.
[0338] If the first function and the first parameter of the first semantic recognition result both belong to the first semantic recognition list, but the first parameter and the first function do not correspond in the first semantic recognition list, then it can be determined that the first semantic recognition result does not meet the first preset condition.
[0339] If one or more of the first function or first parameter of the first semantic recognition result do not belong to the first semantic recognition list, it can be determined that the first semantic recognition result does not meet the first preset condition.
[0340] For example, the first function of the first semantic recognition result is "temperature control function", and the first parameter of the first semantic recognition result is "28℃". If both "temperature control function" and "28℃" belong to the first semantic recognition list, and "temperature control function" corresponds to "28℃" in the first semantic recognition list, then the first semantic recognition result can satisfy the first preset condition. The first device can then determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0341] For example, the first function of the first semantic recognition result is "audio function," and the first parameter of the first semantic recognition result is "high-definition playback." Multiple functions in the first semantic recognition list may include "audio function"; multiple parameters in the first semantic recognition list may include "high-definition playback." However, in the first semantic recognition list, "audio function" does not correspond to "high-definition playback." In the first semantic recognition list, "audio function" may correspond to "standard playback," "high-quality playback," or "lossless playback"; in the first semantic recognition list, "audio function" may correspond to "standard definition playback," "high-definition playback," "ultra-high definition playback," "Blu-ray playback," etc. Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0342] For example, the first function of the first semantic recognition result is "play function", and the first parameter of the first semantic recognition result is "singer A". Multiple functions of the first semantic recognition list may include "play function"; multiple parameters of the first semantic recognition list may not include "singer A" (e.g., the first device has not previously played songs by singer A). Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0343] Optionally, the first semantic recognition result satisfying the first preset condition further includes: in the first semantic recognition list, the first parameter corresponds to the first intent.
[0344] Optionally, if the first semantic recognition result does not meet the first preset condition, it further includes: in the first semantic recognition list, the first parameter does not correspond to the first intent.
[0345] In addition to the first function and the first intent having a corresponding relationship, the first function and the first parameter can also have a corresponding relationship, and the first intent and the first parameter can also have a corresponding relationship.
[0346] For example, the first function of the first semantic recognition result is "vehicle control function", the first intent of the first semantic recognition result is "turn on the air conditioner", and the first parameter of the first semantic recognition result is "28℃". If "vehicle control function", "turn on the air conditioner", and "28℃" all belong to the first semantic recognition list, and "vehicle control function" corresponds to "turn on the air conditioner" and "turn on the air conditioner" corresponds to "28℃" in the first semantic recognition list, then the first semantic recognition result can satisfy the first preset condition. The first device can then determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0347] For example, the first function of the first semantic recognition result is "vehicle control function," the first intent of the first semantic recognition result is "turn on the air conditioner," and the first parameter of the first semantic recognition result is "5℃." Multiple functions in the first semantic recognition list may include "vehicle control function." Multiple parameters in the first semantic recognition list may include "turn on the air conditioner." Multiple parameters in the first semantic recognition list may include "5℃." In the first semantic recognition list, "vehicle control function" can correspond to "turn on the air conditioner"; "vehicle control function" can correspond to "5℃." However, in the first semantic recognition list, "turn on the air conditioner" and "5℃" may not correspond. In the first semantic recognition list, "turn on the air conditioner" may, for example, correspond to a temperature value of 17~30℃, while "5℃" may exceed the adjustable temperature range of the air conditioner; in the first semantic recognition list, "turn on the car refrigerator" may, for example, correspond to "5℃." Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0348] Optionally, the second semantic recognition list further includes multiple parameters corresponding to the first intent, and the first semantic recognition result satisfies a first preset condition, further including: the first semantic recognition result also includes a first parameter, and the first parameter belongs to the multiple parameters.
[0349] Optionally, the second semantic recognition list may further include multiple parameters corresponding to the first intent. If the first semantic recognition result does not meet the first preset condition, it may also include: the first semantic recognition result may further include a first parameter, and the first parameter does not belong to the multiple parameters.
[0350] The parameters corresponding to the first intent are usually not unlimited. For example, the intent to enable hardware may correspond to parameters such as hardware identifier; however, the intent to enable hardware usually does not correspond to parameters such as location. Similarly, the intent to plan a route may correspond to parameters such as location; however, the intent to plan a route usually does not correspond to parameters such as song. The first device can pre-record the correspondence between multiple intents and multiple parameters in a second semantic recognition list.
[0351] If the second semantic recognition list indicates that the first intent and the first parameter of the first semantic recognition result correspond, then the second semantic recognition list can satisfy the first preset condition.
[0352] If the first intention and the first parameter of the first semantic recognition result both belong to the second semantic recognition list, but the first parameter and the first intention do not correspond in the second semantic recognition list, then it can be determined that the first semantic recognition result does not meet the first preset condition.
[0353] If one or more of the first intent or first parameters of the first semantic recognition result do not belong to the second semantic recognition list, then it can be determined that the first semantic recognition result does not meet the first preset condition.
[0354] For example, the first intent of the first semantic recognition result is "turn on device A", and the first parameter of the first semantic recognition result is "1 hour"; if both "turn on device A" and "1 hour" belong to the second semantic recognition list, and "turn on device A" corresponds to "1 hour" in the second semantic recognition list, then the first semantic recognition result can satisfy the first preset condition. The first device can determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0355] For example, the first intent of the first semantic recognition result is "play audio", and the first parameter of the first semantic recognition result is "photo A". Multiple intents in the second semantic recognition list may include "play audio". Multiple parameters in the second semantic recognition list may include "photo A". However, in the second semantic recognition list, "play audio" and "photo A" may not correspond. In the second semantic recognition list, "play audio" may correspond to parameters such as "singer", "song", or "playlist"; in the second semantic recognition list, "photo A" may correspond to intents such as "share" or "upload". Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0356] For example, the first intent of the first semantic recognition result is "play video", and the first parameter of the first semantic recognition result is "actor A". Multiple intents in the second semantic recognition list may include "play video"; multiple parameters in the second semantic recognition list may not include "actor A" (e.g., the first device has not previously played any film or television work featuring actor A). Therefore, it can be determined that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0357] Optionally, the first semantic recognition result may indicate at least two of a first function, a first intent, and a first parameter. For example, if the first semantic recognition result can indicate both a first function and a first parameter, then the first semantic recognition result satisfies a first preset condition when the first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result; otherwise, the first semantic recognition result does not satisfy the first preset condition. As another example, if the first semantic recognition result can indicate both a first intent and a first parameter, then the first semantic recognition result satisfies the first preset condition when the first intent indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result; otherwise, the first semantic recognition result does not satisfy the first preset condition.
[0358] When the first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result, the first semantic recognition result satisfies the first preset condition. In one implementation, the third semantic recognition list is further used to indicate that the first parameter corresponds to the first parameter type, and the method further includes: obtaining a first semantic recognition list, the first semantic recognition list including multiple functions and multiple parameter types corresponding to the multiple functions; the first semantic recognition result satisfying the first preset condition further includes: the first semantic recognition result also includes a first function, the first function belongs to the multiple functions, and in the first semantic recognition list, the first function corresponds to the first parameter type.
[0359] When the first function indicated by the first semantic recognition result corresponds to a different parameter type than the first parameter indicated by the first semantic recognition result, the first semantic recognition result does not meet the first preset condition. In one implementation, the third semantic recognition list is further used to indicate that the first parameter corresponds to a first parameter type, and the method further includes: obtaining a first semantic recognition list, the first semantic recognition list including multiple functions and multiple parameter types corresponding to the multiple functions; if the first semantic recognition result does not meet the first preset condition, it further includes: the first semantic recognition result also includes a first function, the first function belongs to the multiple functions, and in the first semantic recognition list, the first function does not correspond to the first parameter type.
[0360] The first device can pre-record functions and their corresponding one or more parameter types. The correspondence (or association) between multiple functions and multiple parameter types can be stored in the first semantic recognition list.
[0361] For example, the parameter types corresponding to vehicle control functions can be time, temperature, hardware identifiers, etc.
[0362] For example, the parameter type corresponding to the temperature control function can be temperature, etc.
[0363] For example, the parameter types for navigation functions can be location, time, etc.
[0364] For example, the parameter types corresponding to audio functions can be artist, song, playlist, time, audio playback mode, etc.
[0365] For example, the parameter types corresponding to the video function can be movies, TV series, actors, time, video playback mode, etc.
[0366] The first device can pre-store parameters and their corresponding one or more parameter types. The correspondence or association between multiple parameters and multiple parameter types can be stored in a third semantic recognition list.
[0367] For example, the parameter types corresponding to air conditioning, cameras, seats, and windows can be hardware identifiers.
[0368] For example, the parameter type corresponding to 5℃, 28℃, etc. can be temperature.
[0369] For example, the parameter type corresponding to 1 hour, 1 minute, etc. can be time.
[0370] For example, the parameter type corresponding to position A, position B, etc. can be position.
[0371] For example, the parameter type corresponding to singer A, singer B, etc. can be singer.
[0372] For example, the parameter type corresponding to song A, song B, etc. can be "song".
[0373] For example, the parameter type corresponding to playlist A, playlist B, etc. can be playlist.
[0374] For example, the parameter types corresponding to standard playback, high-quality playback, and lossless playback can be audio playback modes.
[0375] For example, the parameter type corresponding to Movie A, Movie B, etc. can be Movie.
[0376] For example, the parameter type corresponding to TV series A and TV series B can be TV series.
[0377] For example, the parameter type corresponding to actor A, actor B, etc. can be actor.
[0378] For example, the parameter types corresponding to standard definition playback, high definition playback, extended playback, and Blu-ray playback can be video playback modes.
[0379] Users can use a variety of voice messages to achieve a specific function of the first device. For this function, the slots in the voice messages are typically filled with parameters of a limited number of types. The number of parameters corresponding to a single parameter type can be, for example, any number.
[0380] For example, a user can instruct the first device to perform navigation-related operations via voice commands. These voice commands can be converted into semantic recognition results related to the navigation function. The parameter type corresponding to the slot in the semantic recognition result can be, for example, location, time, etc. The parameter type corresponding to the slot in the semantic recognition result does not necessarily have to be temperature, song, playlist, audio playback mode, movie, TV series, video playback mode, etc.
[0381] For example, a user can instruct the first device to perform vehicle control functions via voice commands. These voice commands can be converted into semantic recognition results related to the vehicle control functions. The parameter type corresponding to the slot in the semantic recognition result can be, for example, time, temperature, hardware identifier, etc. The parameter type corresponding to the slot in the semantic recognition result does not have to be location, artist, song, playlist, audio playback mode, movie, TV series, actor, video playback mode, etc.
[0382] The first semantic recognition result may include the first function and the first parameter.
[0383] If the first device obtains in advance that the first parameter corresponds to the first parameter type and the first function corresponds to the first parameter type, then the accuracy of the first device's analysis of the first speech information may be relatively high. In this case, the first device can determine the corresponding operation based on its own obtained first semantic recognition result.
[0384] For example, suppose the first semantic recognition list may include navigation functionality, and the parameter type corresponding to the navigation functionality in the first semantic recognition list may include location; suppose the third semantic recognition list may include location A, and the parameter type corresponding to location A in the first semantic recognition list may be location. In one possible scenario, a user can instruct the first device to perform an operation related to navigation functionality via the voice command "Navigate to location A". This voice command can be converted into a first semantic recognition result, the first function of which may be navigation functionality; the first parameter of which may be location A. The first device can determine whether the first semantic recognition result meets a first preset condition based on the first semantic recognition list and the third semantic recognition list. The first device can determine the first operation to be performed based on the first semantic recognition result.
[0385] If the first device obtains a first parameter and a first parameter type that correspond beforehand, but the first function does not correspond to the first parameter type, the accuracy of the first device's analysis of the first voice information may be relatively low; for example, the specific meaning of the first parameter may be incorrect. In this case, if the first device determines the corresponding operation solely based on its own derived first semantic recognition result, a response error may occur. The first device may choose to execute the operation indicated by the second device to respond to the user's voice command.
[0386] For example, suppose the first semantic recognition list may include navigation functionality, and the parameter type corresponding to the navigation functionality in the first semantic recognition list may include location; suppose the third semantic recognition list may include song A, and the parameter type corresponding to song A in the first semantic recognition list may be song. Song A may have the same name as location A, but location A is not included in the third semantic recognition list. In one possible scenario, a user can instruct the first device to perform an operation related to navigation functionality via the voice command "Navigate to location A". However, because location A and song A have the same name, the first device may identify location A as song A. For example, the first function of the first semantic recognition result may be navigation functionality; the first parameter of the first semantic recognition result may be song A. The first device can determine, based on the first semantic recognition list and the third semantic recognition list, that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0387] Optionally, the first device can correct the semantic recognition result. The first device determines the semantic recognition result (referred to as the second semantic recognition result for distinction) based on the first voice information. The second semantic recognition result indicates the second function and the second parameter (which may be the same parameter as the first parameter mentioned above). If the preset multiple functions do not include the second function and the preset multiple parameters include the second parameter, the second function in the second semantic recognition result is corrected to the first function to obtain the first semantic recognition result. The first function and the second function are two different functions of the same type.
[0388] In one implementation, determining the first semantic recognition result based on the first voice information includes: determining a second semantic recognition result based on the first voice information, the second semantic recognition result including a second function and the first parameter; if the second function is not included in the first semantic recognition list, correcting the second function in the second semantic recognition result to the first function based on the third semantic recognition list to obtain the first semantic recognition result, wherein the first function and the second function are two different functions of the same type.
[0389] The first device can recognize the first speech information to obtain a second semantic recognition result. The second semantic recognition result may be, for example, the initial result obtained after semantic recognition. The second function of the second semantic recognition result does not belong to the first semantic recognition list, which means that the first device may have a relatively weak semantic recognition capability for the second function.
[0390] However, the first device may, for example, have the ability to learn and train. The first device may, for example, learn voice commands related to the second function multiple times, thereby gradually improving its semantic recognition ability for the second function. Some voice commands related to the second function may initially be unrecognizable by the first device, but as the first device learns and trains on the semantic commands related to the second function, the voice commands related to the second function can gradually be accurately recognized by the first device, and the number of voice commands related to the second function accurately recognized by the first device can gradually increase.
[0391] In one possible scenario, the parameter types corresponding to the second function may be relatively diverse, and the first device may not be able to fully store all the parameters corresponding to the second function. That is, the probability that the first voice information indicating the second function can be accurately recognized by the first device may be relatively low. However, if the first parameter belongs to the third semantic recognition list, it means that the first device may have learned the first parameter beforehand. Then, the first device can modify the second function in the second semantic recognition result to the first function, obtaining the first semantic recognition result. The first semantic recognition result can be a modified semantic recognition result.
[0392] It should be noted that the first function and the second function should be two different functions of the same type.
[0393] For example, the first function could be a local translation function, and the second function could be a cloud translation function. Both the first and second functions can be translation-related features.
[0394] For example, the first function could be local navigation, and the second function could be cloud navigation. Both the first and second functions can be navigation-type functions.
[0395] For example, the first function could be local audio, and the second function could be cloud audio. Both the first and second functions can be audio playback functions.
[0396] For example, the first function could be a local video function, and the second function could be a cloud video function. Both the first and second functions can be video playback functions.
[0397] For example, suppose the first semantic recognition list may include local navigation functions but not cloud navigation functions, and the parameter type corresponding to the local navigation function in the first semantic recognition list may include location; suppose the third semantic recognition list may include location A, and the parameter type corresponding to location A in the third semantic recognition list may be location. In one possible scenario, a user can instruct the first device to perform an operation related to navigation by using the voice command "Navigate to location A". This voice command can be converted into a second semantic recognition result, the second function of which may be cloud navigation; the first parameter of the second semantic recognition result may be location A. Since the first parameter belongs to the third semantic recognition list (e.g., the first device has previously navigated to location A), the first device can modify the second function of the second semantic recognition result to the local navigation function to obtain the first semantic recognition result, where the first function of the first semantic recognition result is the local navigation function. Both local navigation and cloud navigation functions belong to the navigation type, and local navigation and cloud navigation functions may be two different functions. Afterwards, the first device can determine whether the first semantic recognition result meets the first preset condition based on the first semantic recognition list and the third semantic recognition list. Since the local navigation function belongs to the first semantic recognition list, and the parameter type corresponding to the local navigation function in the first semantic recognition list includes location, and the parameter type corresponding to position A in the third semantic recognition list is also location, the first device can determine that the first semantic recognition result satisfies the first preset condition. The first device can determine the first operation to be performed based on the first semantic recognition result.
[0398] When the first semantic recognition result indicates a first intent and the first parameter indicated by the first semantic recognition result correspond to the same parameter type, the first semantic recognition result satisfies a first preset condition. In one implementation, the third semantic recognition list is further used to indicate that the first parameter corresponds to a first parameter type, and the method further includes: obtaining a second semantic recognition list, the second semantic recognition list including multiple intents and multiple parameter types corresponding to the multiple intents; the first semantic recognition result satisfying the first preset condition further includes: the first semantic recognition result also includes a first intent, the first intent belongs to the multiple intents, and in the second semantic recognition list, the first intent corresponds to the first parameter type.
[0399] When the first semantic recognition result indicates a first intent and the first parameter indicated by the first semantic recognition result correspond to different parameter types, the first semantic recognition result does not satisfy the first preset condition. In one implementation, the third semantic recognition list is further used to indicate that the first parameter corresponds to a first parameter type, and the method further includes: obtaining a second semantic recognition list, the second semantic recognition list including multiple intents and multiple parameter types corresponding to the multiple intents; if the first semantic recognition result does not satisfy the first preset condition, it further includes: the first semantic recognition result also includes a first intent, the first intent belongs to the multiple intents, and in the second semantic recognition list, the first intent does not correspond to the first parameter type.
[0400] The first device can store intents and their corresponding parameter types in advance. The correspondence (or association) between multiple intents and multiple parameter types can be stored in the second semantic recognition list.
[0401] For example, the parameter type corresponding to the hardware activation intent can be time, temperature, hardware identifier, etc.
[0402] For example, the parameter type corresponding to the path planning intent can be location, time, etc.
[0403] For example, the parameter types corresponding to the intention to play audio can be artist, song, playlist, time, audio playback mode, etc.
[0404] For example, the parameter type corresponding to the intention to play a video can be movie, TV series, actor, time, video playback mode, etc.
[0405] The first device can pre-record parameters and their corresponding parameter types. The correspondence or association between multiple parameters and multiple parameter types can be stored in a third semantic recognition list. The correspondence between multiple parameters and multiple parameter types has already been explained above, and will not be elaborated further here.
[0406] Users can achieve a certain type of intent from the first device through a variety of voice information. For this type of intent, the slots in the voice information are typically filled with parameters of a finite number of types. The number of parameters corresponding to a parameter type can be, for example, any number.
[0407] For example, a user instructs a first device to perform an operation related to route planning via voice command. This voice command can be converted into a semantic recognition result related to the route planning intention. The parameter type corresponding to the slot in this semantic recognition result can be, for example, a location or time parameter. The parameter type corresponding to the slot in the semantic recognition result does not necessarily have to be temperature, artist, song, playlist, audio playback mode, movie, TV series, actor, or video playback mode.
[0408] For example, a user instructs a first device to perform an operation related to hardware activation via voice command. This voice command can be converted into a semantic recognition result related to the hardware activation intention. The parameter type corresponding to the slot of this semantic recognition result can be, for example, time, temperature, hardware identifier, etc. The parameter type corresponding to the slot of the semantic recognition result does not have to be location, artist, song, playlist, audio playback mode, movie, TV series, actor, video playback mode, etc.
[0409] The first semantic recognition result may include the first intent and the first parameters.
[0410] If the first device obtains in advance that the first parameter corresponds to the first parameter type, and the first intent corresponds to the first parameter type, then the accuracy of the first device's analysis of the first speech information may be relatively high. In this case, the first device can determine the corresponding operation based on its own obtained first semantic recognition result.
[0411] For example, suppose the second semantic recognition list may include a route planning intent, and the parameter type corresponding to the route planning intent in the second semantic recognition list may include location; suppose the third semantic recognition list may include location A and location B, and the parameter type corresponding to location A and location B in the second semantic recognition list may both be location. In one possible scenario, a user can instruct the first device to perform an operation related to the route planning intent via the voice command "Navigate to location A, via location B". This voice command can be converted into a first semantic recognition result, the first intent of the first semantic recognition result may be a route planning intent; the first parameter of the first semantic recognition result may include location A and location B. The first device can determine whether the first semantic recognition result meets a first preset condition based on the second and third semantic recognition lists. The first device can determine the first operation to be performed based on the first semantic recognition result.
[0412] If the first device obtains a first parameter and a first parameter type beforehand, but the first intent does not correspond to the first parameter type, the accuracy of the first device's analysis of the first voice information may be relatively low; for example, the specific meaning of the first parameter may be incorrect. In this case, if the first device determines the corresponding operation solely based on its own derived first semantic recognition result, voice response errors may occur. The first device may choose to execute the operation indicated by the second device to respond to the user's voice command.
[0413] For example, suppose the second semantic recognition list may include the intent to play audio, and the parameter type corresponding to the navigation intent in the second semantic recognition list may include a singer; suppose the third semantic recognition list may include actor A, and the parameter type corresponding to actor A in the second semantic recognition list may be actor. Actor A may also be a singer; actor A is not only an actor but also a singer. However, in the third semantic recognition list, the parameter type corresponding to actor A does not include singer. In one possible scenario, the user can instruct the first device to perform an operation related to the intent to play audio via voice command "Play actor A's song." However, since the first device may not be able to recognize actor A as a singer, it may recognize actor A as an actor. For example, the first intent of the first semantic recognition result may be the intent to play audio; the first parameter of the first semantic recognition result may be actor A. The first device can determine, based on the second and third semantic recognition lists, that the first semantic recognition result does not meet the first preset condition. Optionally, the first device may determine to execute the second operation instructed by the second device.
[0414] The first device can correct the semantic recognition result. The first device determines the semantic recognition result (referred to as the third semantic recognition result for distinction) based on the first voice information. The third semantic recognition result indicates the second intention and the third parameter (which may be the same parameter as the first parameter mentioned above). If the second intention is not included in the preset multiple intentions and the third parameter is included in the preset multiple parameters, the second intention in the second semantic recognition result is corrected to the first intention to obtain the first semantic recognition result. The first intention and the second intention are two different intentions of the same type.
[0415] In one implementation, determining the first semantic recognition result based on the first voice information includes: determining a third semantic recognition result based on the first voice information, the third semantic recognition result including a second intent and the first parameter; if the second semantic recognition list does not include the second intent, correcting the second intent in the third semantic recognition result to the first intent based on the third semantic recognition list to obtain the first semantic recognition result, wherein the first intent and the second intent are two different intents of the same type.
[0416] The first device can recognize the first speech information to obtain a third semantic recognition result. The third semantic recognition result may be, for example, the initial result obtained after semantic recognition. The second intent of the third semantic recognition result does not belong to the second semantic recognition list, which means that the first device may have a relatively weak semantic recognition ability for the second intent.
[0417] However, the first device may, for example, have the ability to learn and train. The first device may, for example, learn the voice commands related to the second intention multiple times, thereby gradually improving its semantic recognition ability for the second intention. Some voice commands related to the second intention may initially be unrecognizable by the first device, but as the first device learns and trains on the semantic commands related to the second intention, the voice commands related to the second intention can gradually be accurately recognized by the first device, and the number of voice commands related to the second intention accurately recognized by the first device can gradually increase.
[0418] In one possible scenario, the parameter types corresponding to the second intent may be relatively diverse, and the first device may not be able to fully store all the parameters corresponding to the second intent. That is, the probability that the first speech information indicating the second intent can be accurately recognized by the first device may be relatively low. However, if the first parameter belongs to the third semantic recognition list, it means that the first device may have learned the first parameter beforehand. Then, the first device can modify the second intent in the third semantic recognition result to the first intent, obtaining the first semantic recognition result. The first semantic recognition result can be a modified semantic recognition result.
[0419] It is important to note that the first intention and the second intention should be two different intentions of the same type.
[0420] For example, the first intention could be local translation of English, while the second intention could be cloud translation of English. Both the first and second intentions can fall under the category of translating English.
[0421] For example, the first intent could be a local route planning intent, while the second intent could be a cloud route planning intent. Both the first and second intents can belong to the route planning type of intent.
[0422] For example, the first intent could be to play local audio, and the second intent could be to play cloud audio. Both the first and second intents can belong to the audio playback type of intent.
[0423] For example, the first intent could be to play a local video, while the second intent could be to play a cloud video. Both the first and second intents can fall under the category of video playback intents.
[0424] For example, suppose the second semantic recognition list may include the intent to play local audio but not the intent to play cloud audio, and the parameter type corresponding to the intent to play local audio in the second semantic recognition list may include a singer; suppose the third semantic recognition list may include singer A, and the parameter type corresponding to singer A in the second semantic recognition list may be a singer. In one possible scenario, a user can instruct the first device to perform an operation related to the intent to play audio via the voice command "Play a song by singer A". This voice command can be converted into a third semantic recognition result, the second intent of which may be the intent to play cloud audio; the first parameter of which may be singer A. Since the first parameter belongs to the third semantic recognition list (e.g., the first device has previously played a song by singer A), the first device can modify the second intent of the third semantic recognition result to the intent to play local audio, obtaining the first semantic recognition result, where the first intent of the first semantic recognition result is the intent to play local audio. Both the intent to play local audio and the intent to play cloud audio belong to the audio playback type, and the local audio intent and the intent to play cloud audio can be two different intents. Afterwards, the first device can determine whether the first semantic recognition result meets the first preset condition based on the second and third semantic recognition lists. Since the intent to play local audio belongs to the second semantic recognition list, and the parameter type corresponding to the intent to play local audio in the second semantic recognition list includes singer, and the parameter type corresponding to singer A in the third semantic recognition list is also singer, the first device can determine that the first semantic recognition result satisfies the first preset condition. The first device can determine the first operation to be performed based on the first semantic recognition result.
[0425] Optionally, the first semantic recognition result satisfies the first preset condition, including: the first semantic recognition result includes a first indicator bit, the first indicator bit indicating that the first semantic recognition result satisfies the first preset condition.
[0426] Optionally, if the first semantic recognition result does not meet the first preset condition, it includes: the first semantic recognition result includes a second indicator bit, the second indicator bit indicating that the first semantic recognition result does not meet the first preset condition.
[0427] If the first semantic recognition result carries a first indication bit, the first device can determine to execute the first operation determined by the first device based on the first semantic recognition result. The specific content of the first indication bit may include information such as "local" or "end-side" to instruct the first device itself to perform the operation.
[0428] If the first semantic recognition result carries a second indication bit, the first device can determine to perform the second operation indicated by the second device. The specific content of the second indication bit may include information such as "cloud" or "cloud side" to instruct the first device to perform the operation according to the instruction of the second device.
[0429] Optionally, if the first semantic recognition result does not carry the first indication bit, and other contents of the first semantic recognition result do not meet the first preset condition, the first device may determine to execute the second operation indicated by the second device.
[0430] Optionally, if the first semantic recognition result carries a second indicator bit, and other contents of the first semantic recognition result do not meet the first preset condition, then the first device may determine to execute the first operation determined by the first device based on the first semantic recognition result.
[0431] The fourth semantic recognition result can indicate at least two of the first function, first intention, and first parameter. For example, if the fourth semantic recognition result can indicate both the first function and the first parameter, then when the first function indicated by the fourth semantic recognition result corresponds to the same parameter type as the first parameter indicated by the fourth semantic recognition result, a first semantic recognition result carrying a first indicator bit can be determined; when the first function indicated by the fourth semantic recognition result corresponds to a different parameter type than the first parameter indicated by the fourth semantic recognition result, a first semantic recognition result carrying a second indicator bit can be determined. Similarly, if the fourth semantic recognition result can indicate both the first intention and the first parameter, then when the first intention indicated by the fourth semantic recognition result corresponds to the same parameter type as the first parameter indicated by the fourth semantic recognition result, a first semantic recognition result carrying a first indicator bit can be determined; when the first intention indicated by the fourth semantic recognition result corresponds to a different parameter type than the first parameter indicated by the fourth semantic recognition result, a first semantic recognition result carrying a second indicator bit can be determined.
[0432] When the first function indicated by the fourth semantic recognition result corresponds to the same parameter type as the first parameter indicated by the fourth semantic recognition result, it can be determined that the first semantic recognition result carries a first indicator bit. In one implementation, determining the first semantic recognition result based on the first voice information includes: determining a fourth semantic recognition result based on the first voice information, the fourth semantic recognition result including a first function and a first parameter; determining the first semantic recognition result when the fourth semantic recognition list includes the first function, the fourth semantic recognition list indicates that the first function corresponds to the first parameter type, the third semantic recognition list includes the first parameter, and the third semantic recognition list indicates that the first parameter corresponds to the first parameter type, the semantic recognition result including the first indicator bit.
[0433] The first device can recognize the first speech information to obtain a fourth semantic recognition result. The fourth semantic recognition result can be, for example, the initial result obtained after semantic recognition. The first function of the fourth semantic recognition result does not belong to the first semantic recognition list, meaning that the first device may have a relatively weak semantic recognition capability for the first function.
[0434] However, the first device may, for example, possess learning and training capabilities. The first device may, for example, learn voice commands related to the first function multiple times, thereby gradually improving its semantic recognition capability for the first function. Some voice commands related to the first function may initially be unrecognizable by the first device, but as the first device learns and trains on the semantic commands related to the first function, the voice commands related to the first function can gradually be accurately recognized by the first device, and the number of voice commands related to the first function accurately recognized by the first device can gradually increase. Optionally, the fourth semantic recognition list may be used, for example, to record functions newly learned (or trained, updated, etc.) by the first device.
[0435] In one possible scenario, the parameter types corresponding to the first function may be relatively diverse, and the first device may not be able to fully store all the parameters corresponding to the first function. That is, the probability that the first voice information indicating the first function can be accurately recognized by the first device may be relatively low. However, if the first parameter belongs to the third semantic recognition list, it means that the first device may have learned the first parameter beforehand. In the fourth semantic recognition list, the first function corresponds to the first parameter type, and in the third semantic recognition list, the first parameter corresponds to the first parameter type; therefore, the probability that the first device accurately recognizes the first voice information is relatively high. The first device can determine the first semantic recognition result and indicate that the first semantic recognition result meets the first preset condition through the first indicator bit in the first semantic recognition result. The first semantic recognition result can be a modified semantic recognition result.
[0436] For example, suppose the first semantic recognition list may not include navigation functionality; the fourth semantic recognition list may include navigation functionality, and the parameter type corresponding to navigation functionality in the fourth semantic recognition list may include location; suppose the third semantic recognition list may include location A, and the parameter type corresponding to location A in the third semantic recognition list may be location. In one possible scenario, a user can instruct the first device to perform an operation related to navigation functionality via the voice command "Navigate to location A". This voice command can be converted into a fourth semantic recognition result, the first function of which may be navigation functionality; the first parameter of which may be location A. Since location A belongs to the third semantic recognition list (e.g., the first device has previously navigated to location A), the parameter type corresponding to location A in the third semantic recognition list is location, and the parameter type corresponding to navigation functionality in the fourth semantic recognition list includes location, therefore, the first device can determine a first semantic recognition result based on the fourth semantic recognition result, the first semantic recognition list, the third semantic recognition list, and the fourth semantic recognition list. This first semantic recognition result may include a first indicator bit. The first device can determine that the first semantic recognition result satisfies a first preset condition. The first device can determine the first operation to be performed based on the first semantic recognition result.
[0437] When the first intent indicated by the fourth semantic recognition result corresponds to the same parameter type as the first parameter indicated by the fourth semantic recognition result, the first semantic recognition result carrying the first indicator bit can be determined. In one implementation, the method further includes: obtaining a fifth semantic recognition list, the fifth semantic recognition list including multiple intents and multiple parameter types corresponding to the multiple intents; obtaining a third semantic recognition list, the third semantic recognition list including multiple parameters and multiple parameter types corresponding to the multiple intents; determining the first semantic recognition result based on the first voice information includes: determining a fifth semantic recognition result based on the first voice information, the fifth semantic recognition result including a first intent and a first parameter; when the fifth semantic recognition list includes the first intent, the fifth semantic recognition list indicates that the first intent corresponds to the first parameter type, the third semantic recognition list includes the first parameter, the third semantic recognition list indicates that the first parameter corresponds to the first parameter type, determining the first semantic recognition result, the semantic recognition result including the first indicator bit.
[0438] The first device can recognize the first speech information to obtain a fifth semantic recognition result. The fifth semantic recognition result can be, for example, the initial result obtained after semantic recognition. If the first intent of the fifth semantic recognition result does not belong to the second semantic recognition list, it means that the first device may have a relatively weak semantic recognition capability for the first intent.
[0439] However, the first device may, for example, possess the ability to learn and train. The first device may, for example, learn voice commands related to the first intent multiple times, thereby gradually improving its semantic recognition ability for the first intent. Some voice commands related to the first intent may initially be unrecognizable by the first device, but as the first device learns and trains on the semantic commands related to the first intent, the voice commands related to the first intent can gradually be accurately recognized by the first device, and the number of voice commands related to the first intent accurately recognized by the first device can gradually increase. Optionally, the fifth semantic recognition list may be used, for example, to record newly learned (or trained, updated, etc.) intents by the first device.
[0440] In one possible scenario, the parameter types corresponding to the first intent may be relatively diverse, and the first device may not be able to fully store all the parameters corresponding to the first intent. That is, the probability that the first speech information indicating the first intent can be accurately recognized by the first device may be relatively low. However, if the first parameter belongs to the third semantic recognition list, it means that the first device may have learned the first parameter beforehand. In the fifth semantic recognition list, the first intent corresponds to the first parameter type, and in the third semantic recognition list, the first parameter corresponds to the first parameter type; therefore, the probability that the first device accurately recognizes the first speech information is relatively high. The first device can determine the first semantic recognition result and indicate that the first semantic recognition result meets the first preset condition through the first indicator bit in the first semantic recognition result. The first semantic recognition result can be a modified semantic recognition result.
[0441] For example, suppose the second semantic recognition list may not include the intent to play audio; the fifth semantic recognition list may include the intent to play audio, and the parameter type corresponding to the intent to play audio may include a singer; suppose the third semantic recognition list may include singer A, and the parameter type corresponding to singer A may be a singer. In one possible scenario, a user can instruct the first device to perform an operation related to the intent to play audio via the voice command "navigate to singer A". This voice command can be converted into a fifth semantic recognition result, where the first intent of the fifth semantic recognition result can be the intent to play audio; the first parameter of the fifth semantic recognition result can be singer A. Since singer A belongs to the third semantic recognition list (e.g., the first device previously navigated to singer A), the parameter type corresponding to singer A in the third semantic recognition list is singer, and the parameter type corresponding to the intent to play audio includes singer in the fifth semantic recognition list, the first device can determine a first semantic recognition result based on the fifth semantic recognition result, the second semantic recognition list, the third semantic recognition list, and the fifth semantic recognition list. This first semantic recognition result may include a first indicator bit. The first device can determine that the first semantic recognition result satisfies a first preset condition. The first device can determine the first operation to be performed based on the first semantic recognition result.
[0442] In the above example, the third semantic recognition list may include the first parameter, and the first parameter in the third semantic recognition list may correspond to a first parameter type. In other possible cases, if the third semantic recognition list does not include the first parameter, or if the first parameter in the third semantic recognition list may correspond to a parameter type other than the first parameter type, the first semantic recognition result obtained by the first device from the first speech information may be relatively inaccurate. The first device may obtain a sixth semantic recognition result from the second device, which may be determined by the second device based on the first speech information.
[0443] Optionally, the first device can determine the second parameter and the second parameter type based on the sixth semantic recognition result indicated by the second device, and save the association between the second parameter and the second parameter type. In one implementation, the method further includes: determining the second parameter and the second parameter type based on the sixth semantic recognition result; and recording the association between the second parameter and the second parameter type in a third semantic recognition list.
[0444] The sixth semantic recognition result may include the second parameter and the second parameter type. The first device can record the association between the second parameter and the second parameter type in the third semantic recognition list. In subsequent voice interactions, if the first device encounters a voice command related to the second parameter and the second parameter type again, it can more easily and accurately recognize the user's voice command because the third semantic recognition list includes the association between the second parameter and the second parameter type.
[0445] Optionally, the method further includes: obtaining N semantic recognition results from the second device; recording the association between the second parameter and the second parameter type in the third semantic recognition list includes: each of the N semantic recognition results includes the second parameter, the second parameter of each semantic recognition result corresponds to the second parameter type, and N is greater than a first preset threshold, then recording the association between the second parameter and the second parameter type in the third semantic recognition list.
[0446] After the association between the second parameter and the second parameter type appears multiple times, the first device can record the association between the second parameter and the second parameter type in the third semantic recognition list. In other words, the first device can learn the association between the second parameter and the second parameter type a relatively large number of times. However, the representativeness of a single speech recognition session regarding the association between parameters and parameter types may be relatively low. The first device's repeated learning of the association between the second parameter and the second parameter type makes the association between the second parameter and the second parameter type relatively accurate or relatively important, thereby improving the accuracy of the first device in recognizing speech commands.
[0447] Figure 7 This is a schematic flowchart of a voice interaction method 700 provided in an embodiment of this application. Figure 7 The method 700 shown can be applied, for example, to... Figure 5 The voice interaction system 500 shown is shown.
[0448] 701, The first device responds to the user's input operation and performs a voice wake-up operation.
[0449] 702, The first device acquires first voice information from the voice sensor.
[0450] In one possible example, the first voice information may be indicated by the user in the first round of voice interaction. The first round of voice interaction may be the first round of voice interaction in a multi-round voice interaction or a non-first round of voice interaction, and the first round of voice interaction may not be the last round of voice interaction in a multi-round voice interaction.
[0451] 703, the first device determines the first semantic recognition result based on the first voice information.
[0452] 704, the first device determines to execute a first operation determined by the first device based on the first semantic recognition result and the first preset condition, or determines to execute a second operation indicated by the second device.
[0453] The specific implementation methods of 701 to 704 can be referred to Figure 6 The numbers 601, 602, 603a, and 603b shown will not be described in detail here.
[0454] Optionally, before or after the first round of voice interaction ends, the first device may perform a first operation or a second operation. That is, the first device performing the first operation or the second operation may signify the end of the first round of voice interaction or the start of the next round of voice interaction.
[0455] 705, The first device acquires second voice information from the voice sensor.
[0456] The second voice message can be indicated by the user during the second round of voice interaction. Optionally, the second round of voice interaction can be the next round of voice interaction following the first round of voice interaction.
[0457] 706, the first device determines the seventh semantic recognition result based on the second voice information.
[0458] The specific implementation methods of 705 to 706 can be referred to 702 to 703, and will not be described in detail here.
[0459] Multi-turn voice interaction can be applied to scenarios with a large amount of voice interaction information. In multi-turn voice interaction scenarios, the user and the first device can conduct voice dialogues targeting a specific scenario or domain. In one possible scenario, the user may not be able to achieve the desired voice control purpose with a single voice command. For example, through voice interaction, the user can control the first device to purchase a plane ticket from city A to city B. However, there may be a large number of flights from city A to city B. The user may need to interact with the first device multiple times to finally determine the flight information they need to purchase.
[0460] In multi-turn voice interaction, adjacent turns are usually related. For example, the first device might ask the user which flight time slot is the best to book, and the user could reply with a specific time slot. Similarly, the first device might ask the user which artist's work they prefer to listen to, and the user could reply with the artist's name. In other words, the second voice message can be related to the action indicated by the first voice message. In this case, the first device can determine to execute the action indicated by the second voice message to continue the current multi-turn voice interaction.
[0461] However, user responses may have a degree of randomness. A user's response may be unrelated to the voice information requested or to be obtained by the first device. For example, the first device might ask the user which time slot is the best for booking a flight, but the user might reply with the name of a particular singer. Or, the first device might ask the user which time slot is the best for booking a flight, but the user might reply that a car accident has occurred.
[0462] If the primary device completely follows the user's responses, previous voice interactions may be rendered invalid, potentially increasing the number of times the user needs to manipulate the device. For example, through multiple rounds of voice interaction, the user might inadvertently provide the primary device with information such as the flight they intend to purchase and their personal details. However, before making the final payment, the user might unintentionally provide voice information unrelated to the flight or ticket. If the primary device then ends the multi-round voice interaction regarding purchasing the ticket and responds to the unintentional instruction, this not only fails to meet the user's expectations but also invalidates the previously provided flight and personal information. If the user wants to purchase the ticket again, they will need to provide the flight and personal information to the primary device once more.
[0463] If the first device completely ignores the user's response, it may be unable to respond to the user's instructions in certain special scenarios, rendering the user's voice commands ineffective.
[0464] During two adjacent rounds of voice interaction, if the two voice information pieces acquired by the first device are unrelated or weakly related, or if the voice information in the later round of voice interaction does not correspond to or is unrelated to the operation in the previous round, the first device may choose whether to end the multi-round voice interaction based on the situation. For example, if the first and second voice information are unrelated or their correlation is less than a second preset threshold, or if the second voice information is not a feedback to the first or second operation, in order to satisfy the user's voice control experience, the first device may determine whether to end the multi-round voice interaction based on a second preset condition. The second preset condition may be a preset condition for whether to end the multi-round voice interaction. For example, if the second preset condition is met, the first device may end the multi-round voice interaction and start a new voice interaction; if the second preset condition is not met, the first device may continue the current multi-round voice interaction.
[0465] Optionally, the first device can determine whether the second voice information is related to the operation indicated by the first voice information. For example, the first operation is used to ask the user question A, and the second voice information is the answer to question A.
[0466] The second voice information is related to the operation indicated by the first voice information. For example, it may include a relationship between the first and second voice information, or the correlation between the first and second voice information may be higher than a second preset threshold. For example, the second voice information is related to the operation indicated by the first voice information when both the first and second voice information indicate one or more of the following: the same function, the same intention, or the same parameters.
[0467] The following scenarios further illustrate this: the second voice information is unrelated to the operation indicated by the first voice information; or, the correlation between the second voice information and the operation indicated by the first voice information is lower than a second preset threshold; or, the first voice information is unrelated to the second voice information; or, the correlation between the first voice information and the second voice information is lower than the second preset threshold. In this scenario, the first device may need to end the current multi-round voice interaction, or it may need to repeat the operation corresponding to the previous voice interaction to continue the current multi-round voice interaction.
[0468] 707a, the first device determines to execute the operation indicated by the first voice information based on the seventh semantic recognition result and the second preset condition.
[0469] 707c, the first device determines to execute a third operation based on the seventh semantic recognition result and the second preset condition.
[0470] 707d, the first device determines to perform the fourth operation instructed by the second device based on the seventh semantic recognition result and the second preset conditions.
[0471] Any of 707a, 707b, or 707c can be executed.
[0472] In 707a, even if the first device acquires the second voice information, the first device can still perform the operation from the previous round of voice interaction.
[0473] The operation indicated by the first voice information can be either the first operation or the second operation. In one example, if the operation performed by the first device in 704 is the first operation, then in 707a the first device can still determine to perform the first operation. In this case, the current multi-turn voice interaction can be an edge-side multi-turn voice interaction. If the operation performed by the first device in 704 is the second operation, then in 707a the first device can still determine to perform the second operation. In this case, the current multi-turn voice interaction can be a cloud-side multi-turn voice interaction.
[0474] In 707b and 707c, the first device acquires the second voice information and can instruct the operation indicated by the second voice information. Specifically, in 707b, the first device can determine to execute a third operation determined by the first device based on the seventh semantic recognition result; in 707c, the first device can determine to execute a fourth operation instructed by the second device. A detailed implementation of the first device determining the third operation can be found in [reference needed]. Figure 6 The 603a shown will not be described in detail here. For a detailed description of the implementation method of the first device determining the fourth operation, please refer to [reference needed]. Figure 6 The 603b shown will not be described in detail here.
[0475] In one example, if the operation performed by the first device in 704 is a first operation, the first device determines in 707b to perform a third operation. In this case, the new end-side voice interaction can end the previous round of end-side voice interaction.
[0476] In one example, if the operation performed by the first device in 704 is the first operation, the first device determines in 707c to perform the fourth operation. In this case, the new cloud-side voice interaction can end the previous round of end-side voice interaction.
[0477] In one example, if the operation performed by the first device in 704 is the second operation, the first device determines in 707b to perform the third operation. In this case, the new edge-side voice interaction can end the previous round of cloud-side voice interaction.
[0478] In one example, if the operation performed by the first device in 704 is the second operation, the first device determines in 707c to perform the fourth operation. In this case, the new cloud-side voice interaction can end the previous round of cloud-side voice interaction.
[0479] As can be seen from the above, the first preset condition can be used to instruct the first device whether to determine the corresponding operation based on its own determined semantic recognition result, and the second preset condition can be used to instruct the first device whether to end the multi-round voice interaction. The first device can determine the operation to be performed on the second voice information based on the seventh semantic recognition result, the first preset condition, and the second preset condition.
[0480] For example, the seventh semantic recognition result may not meet the second preset condition, meaning that the first device may not end the multi-round voice interaction and may determine to execute the operation indicated by the first voice information, i.e., execute the above-mentioned 707a. The first device may determine whether the seventh semantic recognition result meets the first preset condition, or it may not determine whether the seventh semantic recognition result meets the first preset condition.
[0481] For example, the seventh semantic recognition result can satisfy both the second preset condition and the first preset condition. The seventh semantic recognition result satisfying the second preset condition means that the first device can end the multi-round voice interaction and determine to execute the operation indicated by the second voice information. The seventh semantic recognition result satisfying the first preset condition means that the first device can determine to execute the operation determined by the first device based on the seventh semantic recognition result, i.e., execute 707b as described above.
[0482] For example, the seventh semantic recognition result may satisfy the second preset condition, but may not satisfy the first preset condition. If the seventh semantic recognition result satisfies the second preset condition, it means that the first device can end the multi-round voice interaction and determine to execute the operation indicated by the second voice information. If the seventh semantic recognition result does not satisfy the first preset condition, it means that the first device can determine to execute the operation indicated by the second device, that is, to execute 707c as described above.
[0483] Optionally, the method further includes: sending second voice information to the second device; determining to execute the fourth operation indicated by the second device based on the seventh semantic recognition result and the second preset condition includes: obtaining an eighth semantic recognition result from the second device when the seventh semantic recognition result, or when the seventh semantic recognition result does not meet the second preset condition and the first semantic recognition result does not meet the first preset condition; determining to execute the operation indicated by the first voice information, or determining to execute the fourth operation based on the eighth semantic recognition result and the second preset condition.
[0484] If the seventh semantic recognition result does not meet the first preset condition, the seventh semantic recognition result obtained by the first device may be relatively inaccurate, while the eighth semantic recognition result obtained by the second device may be relatively accurate. The first device can determine whether the eighth semantic recognition result meets the second preset condition, and then determine whether to end the current multi-round voice interaction.
[0485] In one possible scenario, if the first semantic recognition result satisfies the first preset condition, it means that the current multi-turn voice interaction can be a terminal-side multi-turn voice interaction. If the seventh semantic recognition result does not satisfy the first preset condition, but the eighth semantic recognition result satisfies the second preset condition, it means that the current terminal-side multi-turn voice interaction can be terminated, and the new voice interaction can be a cloud-side voice interaction.
[0486] If the seventh semantic recognition result does not meet the second preset condition, it means that the first device can determine the operation to be performed without relying on the seventh semantic recognition result. If the first semantic recognition result does not meet the first preset condition, it may mean that the current multi-turn voice interaction belongs to cloud-side voice interaction. In this case, the first device can obtain the eighth semantic recognition result from the second device to continue the current cloud-side multi-turn voice interaction. The first device can determine whether the eighth semantic recognition result meets the second preset condition, and thus determine whether to end the current cloud-side multi-turn voice interaction.
[0487] For example, the eighth semantic recognition result may not meet the second preset condition, which means that the first device may not end the multi-round voice interaction, that is, determine to execute the operation indicated by the first voice information, that is, execute the above 707a.
[0488] For example, if the eighth semantic recognition result can meet the second preset condition, it means that the first device can end the multi-round voice interaction, that is, determine to execute the operation indicated by the second device, that is, execute the above 707c.
[0489] Assume that the priority of the semantic recognition result related to the operation indicated by the first voice information is the first priority, and the priority of the semantic recognition result unrelated to the operation indicated by the first voice information is the second priority.
[0490] In one possible example, if the priority of the seventh semantic recognition result is first priority, then the second voice information can be considered related to the operation indicated by the first voice information. The first device can determine to execute the operation indicated by the second voice information. If the priority of the seventh semantic recognition result is second priority, then the second voice information can be considered unrelated to the operation indicated by the first voice information. If the second priority is higher than the first priority, then the second device can end the current multi-turn voice interaction and determine to execute the operation indicated by the second voice information. If the second priority is lower than the first priority, then the second device can continue the current multi-turn voice interaction and determine to execute the operation indicated by the first voice information.
[0491] In another possible example, if the priority of the eighth semantic recognition result is first priority, then the second voice information can be considered related to the operation indicated by the first voice information. The first device can determine to execute the operation indicated by the second device. If the priority of the eighth semantic recognition result is second priority, then the second voice information can be considered unrelated to the operation indicated by the first voice information. If the second priority is higher than the first priority, then the second device can end the current multi-turn voice interaction and determine to execute the operation indicated by the second device. If the second priority is lower than the first priority, then the second device can continue the current multi-turn voice interaction and determine to execute the operation indicated by the first voice information.
[0492] Optionally, the seventh semantic recognition result satisfies the second preset condition, including: the priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result.
[0493] The first device can compare the priority of the seventh semantic recognition result with the priority of the first semantic recognition result to determine whether to end the current multi-turn voice interaction. Optionally, in step 704, the first device can determine to execute the first operation or the second operation. Optionally, if the seventh semantic recognition result meets the second preset condition, the first device can execute step 707b or step 707c.
[0494] Optionally, determining to execute the second operation indicated by the second device based on the first semantic recognition result and the first preset condition includes: obtaining a sixth semantic recognition result from the second device when the first semantic recognition result does not meet the first preset condition, the sixth semantic recognition result being used to indicate the second operation; determining to execute the second operation based on the sixth semantic recognition result; the seventh semantic recognition result meeting the second preset condition includes: the priority of the seventh semantic recognition result being higher than the priority of the sixth semantic recognition result.
[0495] In one possible example, at 704, the first device can obtain the sixth semantic recognition result from the second device to determine whether to execute the second operation instructed by the second device. In this case, the first device can compare the priority of the seventh semantic recognition result with the priority of the sixth semantic recognition result to determine whether to end the current multi-turn voice interaction.
[0496] Optionally, the eighth semantic recognition result satisfies the second preset condition, including: the priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result.
[0497] The first device can compare the priority of the eighth semantic recognition result with the priority of the first semantic recognition result to determine whether to end the current multi-turn voice interaction. Optionally, in step 704, the first device can determine to execute the first operation or the second operation.
[0498] Optionally, determining to execute the second operation indicated by the second device based on the first semantic recognition result and the first preset condition includes: obtaining a sixth semantic recognition result from the second device when the first semantic recognition result does not meet the first preset condition, the sixth semantic recognition result being used to indicate the second operation; determining to execute the second operation based on the sixth semantic recognition result; the eighth semantic recognition result meeting the second preset condition includes: the priority of the eighth semantic recognition result being higher than the priority of the sixth semantic recognition result.
[0499] In one possible example, at 704, the first device can obtain the sixth semantic recognition result from the second device to determine whether to execute the second operation instructed by the second device. In this case, the first device can compare the priority of the eighth semantic recognition result with the priority of the sixth semantic recognition result to determine whether to end the current multi-turn voice interaction.
[0500] In other words, users can end the current multi-turn voice interaction by using high-priority voice commands. In one possible example, the first device could record some high-priority commands. If the seventh or eighth semantic recognition result corresponds to a high-priority command, the first device can end the current multi-turn voice interaction.
[0501] Optionally, the priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: the function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; the intention indicated by the seventh semantic recognition result has a higher priority than the intention indicated by the first semantic recognition result; the parameter indicated by the seventh semantic recognition result has a higher priority than the parameter indicated by the first semantic recognition result.
[0502] The first device can compare the priority of a third function with the priority of a first function, and / or compare the priority of a third intent with the priority of a first intent, and / or compare the priority of a third parameter with the priority of a first parameter, to determine whether to end the current multi-turn voice interaction. That is, the first device can record some high-priority functions, and / or high-priority intents and / or high-priority parameters. If the first device identifies one or more of the high-priority functions, high-priority intents, and high-priority parameters, the first device can end the current multi-turn voice interaction and determine to execute the operation indicated by the second voice information.
[0503] For example, the seventh semantic recognition result includes a third function, which can be a "safety control function". The first semantic recognition result includes a first function, which can be a "navigation function". To ensure device safety, the "safety control function" can have a higher priority than other functions, such as the "navigation function". Therefore, the first device can determine whether the seventh semantic recognition result meets the second preset condition based on the priority of the third function and the priority of the first function, and then determine to execute the operation of the second voice information instruction.
[0504] For example, the seventh semantic recognition result includes a third intent, which can be an "accident mode intent." The first semantic recognition result includes a first intent, which can be an "audio playback intent." An "accident mode intent" can refer to the user's intent to activate accident mode. Due to the urgency of accident safety, the "accident mode intent" can have a higher priority than other intents, such as the "audio playback intent." Therefore, the first device can determine that the seventh semantic recognition result meets the second preset condition based on the priority of the third intent and the priority of the first intent, and then determine the operation to execute the second voice information instruction.
[0505] For example, the seventh semantic recognition result includes a third parameter, which can be "door lock". The first semantic recognition result includes a first parameter, which can be "song A". The user can instruct the first device to perform an operation related to "door lock" through the second voice information. Since the opening or closing of the "door lock" can relatively easily affect the driving safety and parking safety of the vehicle, the priority of "door lock" can be higher than other parameters, such as higher than "song A". Therefore, the first device can determine that the seventh semantic recognition result meets the second preset condition based on the priority of the third parameter and the priority of the first parameter, and then determine to execute the operation instructed by the second voice information.
[0506] Optionally, the priority of the seventh semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: the priority of the function indicated by the seventh semantic recognition result is higher than the priority of the function indicated by the sixth semantic recognition result; the priority of the intent indicated by the seventh semantic recognition result is higher than the priority of the intent indicated by the sixth semantic recognition result; the priority of the parameter indicated by the seventh semantic recognition result is higher than the priority of the parameter indicated by the sixth semantic recognition result.
[0507] For examples where the priority of the seventh semantic recognition result is higher than that of the sixth semantic recognition result, please refer to the examples where the priority of the seventh semantic recognition result is higher than that of the first semantic recognition result. These will not be elaborated on here.
[0508] Optionally, the priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: the function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; the intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; the parameter indicated by the eighth semantic recognition result has a higher priority than the parameter indicated by the first semantic recognition result.
[0509] For examples where the priority of the eighth semantic recognition result is higher than that of the first semantic recognition result, please refer to the examples where the priority of the seventh semantic recognition result is higher than that of the first semantic recognition result. These will not be elaborated further here.
[0510] Optionally, the priority of the eighth semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: the function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; the intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; the parameters indicated by the eighth semantic recognition result have a higher priority than the parameters indicated by the sixth semantic recognition result.
[0511] For examples where the priority of the eighth semantic recognition result is higher than that of the sixth semantic recognition result, please refer to examples where the priority of the seventh semantic recognition result is higher than that of the first semantic recognition result. These will not be elaborated further here.
[0512] The following examples, combined with those described above, illustrate some possible scenarios. It should be understood that the solutions provided in this application are not limited to the examples provided below.
[0513] Figure 8 This is a schematic flowchart of a voice interaction method 800 provided in an embodiment of this application. Figure 8 The method 800 shown can be applied, for example, to... Figure 5 The voice interaction system 500 shown is shown.
[0514] 801, The first device responds to the user's input operation and performs a voice wake-up operation.
[0515] For a specific implementation of 801, please refer to, for example. Figure 6 The 601 or shown Figure 7 The 701 shown here will not be described in detail here.
[0516] 802, The first device acquires voice information a from the voice sensor.
[0517] For a specific implementation of 802, please refer to, for example. Figure 6 The 602 shown or Figure 7 The 702 shown will not be described in detail here. Optionally, the voice information 'a' can correspond to, for example, the first voice information mentioned above.
[0518] In one possible example, the voice message 'a' could be, for instance, 'Turn on device A'.
[0519] 803, the first device sends voice information a to the second device.
[0520] For a specific implementation of 803, please refer to, for example. Figure 6 The 603b shown will not be described in detail here.
[0521] 804, The first device determines the semantic recognition result a based on the voice information a.
[0522] For a specific implementation of 804, please refer to, for example. Figure 6 The 603a or shown Figure 7 The 703 shown will not be described in detail here. Optionally, the semantic recognition result 'a' can correspond to, for example, the first semantic recognition result mentioned above.
[0523] In one possible example, the semantic recognition result 'a' may include function 'a', intent 'a', and parameter 'a'; wherein function 'a' may be "vehicle control function"; intent 'a' may be "turn on"; and parameter 'a' may be "device A". Optionally, function 'a' may correspond to the first function mentioned above. Intent 'a' may correspond to the first intent mentioned above. Parameter 'a' may correspond to the first parameter mentioned above.
[0524] 805, the first device determines to execute operation a based on the semantic recognition result a and the first preset condition.
[0525] For a specific implementation of 805, please refer to, for example. Figure 6 The 603a or shown Figure 7 The 704 shown will not be described in detail here. Optionally, operation 'a' can correspond to, for example, the first operation mentioned above.
[0526] In one possible example, the first device stores a semantic recognition list a (optionally, semantic recognition list a may correspond to the first semantic recognition list mentioned above, or to the master table mentioned above, which may contain the first semantic recognition list, the second semantic recognition list, and the third semantic recognition list mentioned above). Multiple functions of semantic recognition list a may include "vehicle control function," multiple intentions of semantic recognition list a may include "open," and multiple parameters of semantic recognition list a may include "device A." Furthermore, in semantic recognition list a, the function "vehicle control function" may correspond to the intention "open," and the intention "open" may correspond to the parameter "device A." The first device can determine, based on the semantic recognition result a and the semantic recognition list a, that the semantic recognition result a satisfies a first preset condition. Since the semantic recognition result a satisfies the first preset condition, the first device can determine to execute the operation a determined by the first device based on the semantic recognition result a. Operation a may, for example, be opening device A.
[0527] 806, the second device sends the semantic recognition result b to the first device.
[0528] For example, a specific implementation of 806 can be found in the following: Figure 6 The 603b shown or Figure 7 The 704 shown will not be described in detail here. Optionally, semantic recognition result b can correspond to the sixth semantic recognition result mentioned above.
[0529] In one possible example, semantic recognition result b may be the same as or similar to semantic recognition result a. Semantic recognition result b may include function b, intent b, and parameter b; wherein function b may be "vehicle control function"; intent b may be "turn on"; and parameter b may be "device A". Optionally, parameter b may correspond to the second parameter mentioned above. Optionally, function b may correspond to the fourth function mentioned above. Intent b may correspond to the fourth intent mentioned above. Parameter b may correspond to the fourth parameter mentioned above.
[0530] 807, the first device discards the semantic recognition result b from the second device.
[0531] Since the first device can determine to perform the operation a determined by the first device based on the semantic recognition result a, the first device can ignore the instructions of the second device for the voice information a.
[0532] Step 807 can be an optional step. For a specific implementation of step 807, please refer to [reference needed]. Figure 6 The 603b shown will not be described in detail here.
[0533] exist Figure 7In the example shown, the first device can determine whether the current semantic recognition result has a relatively high accuracy rate based on the existing semantic recognition list. If the current semantic recognition result may have a relatively high accuracy rate, the first device can choose to determine the operation to be performed, which is beneficial for responding to the user's voice commands quickly and accurately in scenarios where the first device excels.
[0534] Figure 9 This is a schematic flowchart of a voice interaction method 900 provided in an embodiment of this application. Figure 6 The method 900 shown can be applied, for example, to... Figure 6 The voice interaction system 500 shown is shown.
[0535] 901, The first device responds to the user's input operation and performs a voice wake-up operation.
[0536] For a specific implementation of 901, please refer to, for example. Figure 7 The 601 or shown Figure 6 The 701 shown here will not be described in detail here.
[0537] 902, The first device acquires voice information b from the voice sensor.
[0538] For a specific implementation of 902, please refer to, for example. Figure 6 The 602 shown or Figure 6 The 702 shown will not be described in detail here. Optionally, the voice information b may correspond to the first voice information mentioned above.
[0539] In one possible example, the voice message b could be, for instance, "Navigate to location A".
[0540] 903, the first device sends voice information b to the second device.
[0541] For a specific implementation of 903, please refer to, for example. Figure 6 The 603b shown will not be described in detail here.
[0542] 904. The first device determines the semantic recognition result c based on the voice information b.
[0543] For a specific implementation of 904, please refer to, for example. Figure 7 The 603a or shown Figure 6 The 703 shown will not be described in detail here. Optionally, the semantic recognition result c can correspond to the first semantic recognition result mentioned above.
[0544] In one possible example, the semantic recognition result c may include, for example, a function c, an intent c, and a parameter c; wherein, function c may be "cloud navigation function"; intent c may be "route planning"; and parameter c may be "location B". "Location A" and "location B" may be the same or different. Optionally, function c may correspond to the first function mentioned above. Intent c may correspond to the first intent mentioned above. Parameter c may correspond to the first parameter mentioned above. In another possible example, the semantic recognition result c may include, for example, indication information a, which indicates that semantic recognition is not supported.
[0545] 905, the first device determines to execute the operation b instructed by the second device based on the semantic recognition result c and the first preset condition.
[0546] For a specific implementation of 905, please refer to, for example. Figure 9 The 603a or shown Figure 5 The 704 shown will not be described in detail here. Optionally, operation b can correspond to the second operation mentioned above.
[0547] In one possible example, the first device stores a semantic recognition list b (optionally, semantic recognition list b may correspond to the first semantic recognition list mentioned above, or to the master table mentioned above, which may contain the first semantic recognition list, the second semantic recognition list, and the third semantic recognition list mentioned above). Multiple functions of semantic recognition list b may not include "cloud navigation function," and / or multiple parameters of semantic recognition list b may not include "location B," and / or, in semantic recognition list b, the function "cloud navigation function" may not correspond to the parameter "location B." The first device can determine, based on the semantic recognition result c and the semantic recognition list b, that semantic recognition result a does not meet the first preset condition. Since semantic recognition result a does not meet the first preset condition, the first device can determine to execute the operation b instructed by the second device.
[0548] 906, the second device sends the semantic recognition result d for the voice information b to the first device.
[0549] For a specific implementation of 906, please refer to, for example. Figure 6 The 603b shown or Figure 7 The 704 shown will not be described in detail here. Optionally, the semantic recognition result d can correspond to the sixth semantic recognition result mentioned above.
[0550] In a possible example, the semantic recognition result d may include function d, intent d, and parameter d; where function d may be "cloud navigation function"; intent d may be "route planning"; and parameter d may be "location A". Optionally, parameter d may correspond to the second parameter mentioned above. Optionally, function d may correspond to the fourth function mentioned above. Intent d may correspond to the fourth intent mentioned above. Parameter d may correspond to the fourth parameter mentioned above.
[0551] 907, the first device performs operation b based on the semantic recognition result d.
[0552] For a specific implementation of 907, please refer to, for example. Figure 6 The 603b shown or Figure 7 The 704 error shown will not be described in detail here. Operation b could be, for example, path planning to location A.
[0553] Optional, Figure 6 The example shown may also include the following steps.
[0554] 908. The first device determines the parameter d and the parameter type a based on the semantic recognition result d, and records the association between the parameter d and the parameter type a in the semantic recognition list b.
[0555] For a specific implementation of 908, please refer to, for example. Figure 6 The 603b example shown will not be described in detail here. Optionally, parameter d can correspond to the second parameter mentioned above. Parameter type a can correspond to the second parameter type mentioned above. Semantic recognition list b can correspond to the third semantic recognition list mentioned above.
[0556] In one example, parameter type 'a' can be "location". The updated semantic recognition list 'b' can include the following association: "Location A" - "Location".
[0557] 909, The first device acquires voice information c from the voice sensor.
[0558] For a specific implementation of 909, please refer to, for example. Figure 7 The 602 shown or Figure 6 The 702 shown will not be described in detail here. Optionally, the voice information c can correspond to the first voice information mentioned above.
[0559] In one possible example, the voice message c could be, for instance, "Navigate to location A".
[0560] 910, the first device sends voice information c to the second device.
[0561] For a specific implementation of 910, please refer to, for example.Figure 7 The 603b shown will not be described in detail here.
[0562] 911, the first device determines the semantic recognition result e based on the voice information c, and the semantic recognition result e includes function c and parameter d.
[0563] For a specific implementation of 911, please refer to, for example. Figure 6 The 603a shown will not be described in detail here. Optionally, the semantic recognition result e can correspond to the second semantic recognition result mentioned above. Function c can correspond to the second function mentioned above. Parameter d can correspond to the first parameter mentioned above.
[0564] In one possible example, the semantic recognition result e may include function c, intent c, and parameter d; where function c may be "cloud navigation function"; intent c may be "route planning"; and parameter d may be "location A".
[0565] 912, the first device corrects function c in semantic recognition result e to function d based on semantic recognition result e and semantic recognition list b, and obtains semantic recognition result f, where function d and function c are two different functions of the same type.
[0566] For a specific implementation of 912, please refer to, for example. Figure 7 The 603a shown will not be described in detail here. Optionally, function d may correspond to the first function mentioned above. The semantic recognition result f may correspond to the first semantic recognition result mentioned above.
[0567] In one possible example, since multiple parameters of the semantic recognition list b can include "location A", the first device can modify "cloud navigation function" in the semantic recognition result e to "local navigation function" to obtain semantic recognition result f. Semantic recognition result f can, for example, include function d, intent c, and parameter d. Function d can be "local navigation function"; intent c can be "route planning"; parameter d can be "location A". Both "cloud navigation function" and "local navigation function" belong to navigation functions, and "cloud navigation function" and "local navigation function" are two different functions. Optionally, function d can, for example, correspond to the first function mentioned above. Intent c can, for example, correspond to the first intent mentioned above. Parameter d can, for example, correspond to the first parameter mentioned above.
[0568] 913, the first device determines to execute operation c, which is determined by the first device based on the semantic recognition result f and the first preset condition.
[0569] For a specific implementation of 913, please refer to, for example. Figure 6 The 603a or shown Figure 7The 704 error shown will not be described in detail here. Optionally, operation c can correspond to the first operation mentioned above.
[0570] In one possible example, the first device stores a semantic recognition list b (optionally, semantic recognition list b may correspond to the first semantic recognition list mentioned above, or to the master table mentioned above, which may contain the first semantic recognition list, the second semantic recognition list, and the third semantic recognition list mentioned above). Multiple functions of semantic recognition list b may include a "local navigation function"; in semantic recognition list b, the function "local navigation function" may correspond to the parameter type "location"; and in semantic recognition list c, the parameter "location A" may correspond to the parameter type "location". The first device can determine, based on the semantic recognition result f, semantic recognition list b, and semantic recognition list c, that semantic recognition result a satisfies a first preset condition. Since semantic recognition result f satisfies the first preset condition, the first device can determine to execute the operation c determined by the first device based on semantic recognition result f. Operation c may, for example, be path planning to location A.
[0571] 914, the first device discards the semantic recognition result g from the second device for the voice information c.
[0572] Since the first device can determine to perform the operation c determined by the first device based on the semantic recognition result f, the first device can ignore the instructions of the second device for the voice information c.
[0573] Step 914 can be an optional step. For a specific implementation of step 914, please refer to [reference needed]. Figure 7 The 603b shown will not be described in detail here.
[0574] exist Figure 7 In the example shown, the first device can determine whether the current semantic recognition result has a relatively high accuracy rate based on an existing semantic recognition list. If the current semantic recognition result may not have a relatively high accuracy rate, the first device can choose to perform an operation according to the instructions of other devices, thereby enabling the first device to accurately respond to the user's voice commands. Furthermore, the first device can learn the user's voice commands under the instructions of other devices, thus expanding the scenarios in which the first device can respond to user voice commands autonomously.
[0575] Figure 10 is a schematic flowchart of a voice interaction method provided in an embodiment of this application. The method shown in Figure 10 can be applied, for example, to... Figure 7 The voice interaction system 500 shown in Figure 10(b) can be performed after the step shown in Figure 10(a).
[0576] 1001, The first device responds to the user's input operation and performs a voice wake-up operation.
[0577] For a specific implementation of 1001, please refer to, for example. Figure 6 The 601 or shown Figure 7 The 701 shown here will not be described in detail here.
[0578] 1002, The first device acquires voice information d from the voice sensor.
[0579] For a specific implementation of 1002, please refer to, for example. Figure 7 The 602 shown or Figure 7 The 702 shown will not be described in detail here. Optionally, the voice information d can correspond to the first voice information mentioned above.
[0580] In one possible example, the voice message d could be, for example, "navigate to location C".
[0581] 1003, the first device sends voice information d to the second device.
[0582] For a specific implementation of 1003, please refer to, for example. Figure 6 The 603b shown will not be described in detail here.
[0583] 1004, The first device determines the semantic recognition result h based on the voice information d.
[0584] For a specific implementation of 1004, please refer to, for example. Figure 7 The 603a or shown Figure 6 The 703 shown will not be described in detail here. Optionally, the semantic recognition result h can correspond to the first semantic recognition result mentioned above.
[0585] 1005, the first device determines to execute the operation d instructed by the second device based on the semantic recognition result h and the first preset condition.
[0586] For a specific implementation of 1005, please refer to, for example. Figure 6 The 603a or shown Figure 7 The 704 shown will not be described in detail here. Optionally, operation d can correspond to the second operation mentioned above, for example.
[0587] 1006, the second device sends the semantic recognition result i for the voice information d to the first device.
[0588] For a specific implementation of 1006, please refer to, for example. Figure 7 The 603b shown or Figure 6The 704 shown will not be described in detail here. Optionally, the semantic recognition result i can correspond to, for example, the sixth semantic recognition result mentioned above.
[0589] 1007, the first device performs operation d based on the semantic recognition result i.
[0590] For a specific implementation of 1007, please refer to, for example. Figure 7 The 603b shown or Figure 6 The 704 shown here will not be described in detail here.
[0591] In one possible example, operation d could be, for instance, announcing: "Several relevant destinations have been found. Please select location C-1, location C-2, or location C-3." Operation d can also be used to request the exact navigation destination.
[0592] Optionally, in one possible example, 1001 to 1007 could be the first round of cloud-side voice interaction.
[0593] 1008, the first device acquires voice information e from the voice sensor.
[0594] For a specific implementation of 1008, please refer to, for example. Figure 11 The 705 shown will not be described in detail here. Optionally, the voice information e may correspond to the second voice information mentioned above. Optionally, the voice information e may be independent of operation d.
[0595] In one possible example, the voice message d could be, for example, "Turn on device A".
[0596] 1009, the first device sends voice information e to the second device.
[0597] For a specific implementation of 1009, please refer to, for example. Figure 6 The 603b shown will not be described in detail here.
[0598] 1010, the first device determines the semantic recognition result j based on the voice information e.
[0599] For a specific implementation of 1010, please refer to, for example. Figure 6 The 706 shown will not be described in detail here. Optionally, the semantic recognition result j can correspond to the seventh semantic recognition result mentioned above.
[0600] In one example, the semantic recognition result j could indicate "turn on device A".
[0601] 1011, If the semantic recognition result j does not meet the second preset condition, the first device obtains the semantic recognition result k for the voice information e from the second device.
[0602] For a specific implementation of 1011, please refer to, for example. Figure 6 The 707a shown will not be described in detail here. Optionally, the semantic recognition result k can correspond to, for example, the eighth semantic recognition result mentioned above.
[0603] If the semantic recognition result j does not meet the second preset condition, the first device can determine not to execute the operation indicated by the semantic recognition result j. The first device can obtain the semantic recognition result of the second device for the voice information e to determine whether to execute the operation indicated by the voice information e.
[0604] In one possible example, the priority of semantic recognition result j may be lower than the priority of semantic recognition result g or semantic recognition result i.
[0605] 1012, if the semantic recognition result k does not meet the second preset condition, the first device determines to repeat the operation d.
[0606] For a specific implementation of 1012, please refer to, for example. Figure 7 The 707a shown here will not be described in detail here.
[0607] If the semantic recognition result k does not meet the second preset condition, the first device can determine not to execute the operation indicated by the semantic recognition result k. The first device can repeat the operation d so that the first round of cloud-side voice interaction can continue.
[0608] In one possible example, semantic recognition result k may have a lower priority than semantic recognition result g or semantic recognition result i.
[0609] 1013, the first device acquires voice information f from the voice sensor.
[0610] For a specific implementation of 1013, please refer to, for example. Figure 7 The 705 shown will not be described in detail here. Optionally, the voice information f may correspond to the second voice information mentioned above. Optionally, the voice information f may be related to operation d.
[0611] In one possible example, the voice information f could be, for example, “location C-1, passing location D”.
[0612] 1014, the first device sends voice information f to the second device.
[0613] For a specific implementation of 1014, please refer to, for example. Figure 5 The 603b shown will not be described in detail here.
[0614] 1015, the first device determines the semantic recognition result m based on the voice information f.
[0615] For a specific implementation of 1015, please refer to, for example. Figure 5 The 706 shown will not be described in detail here. Optionally, semantic recognition result 5 can correspond to the seventh semantic recognition result mentioned above.
[0616] In one example, the semantic recognition result m could indicate that a first preset condition is not met.
[0617] 1016, the first device obtains the semantic recognition result n for the speech information f from the second device.
[0618] For a specific implementation of 1016, please refer to, for example. Figure 12 The 707c shown will not be described in detail here. Optionally, the semantic recognition result n can correspond to, for example, the eighth semantic recognition result mentioned above.
[0619] If the semantic recognition result m does not meet the first preset condition, the first device can determine whether to execute the operation indicated by the second device or repeat the previous operation. The first device can obtain the semantic recognition result n of the second device for the voice information f to determine whether to execute the operation indicated by the voice information f.
[0620] 1017. If the semantic recognition result n is related to the operation d, determine to execute the operation e indicated by the second device.
[0621] For a specific implementation of 1017, please refer to, for example. Figure 1 The 707c shown here will not be described in detail here.
[0622] The semantic recognition result n is related to operation d, which may mean that the voice information f is a response to the previous round of voice interaction. The first device can determine to execute the operation e instructed by the second device. 1013 to 1017 can be the second round of cloud-side voice interaction. The first round of cloud-side voice interaction and the second round of cloud-side voice interaction can be two rounds of voice interaction for employees in a multi-round voice interaction system.
[0623] In one possible example, operation e could announce something like: "Several relevant waypoints have been found. Please select location D-1, location D-2, or location D-3." Operation e can also be used to request the exact navigation waypoint.
[0624] 1018, the first device acquires voice information g from the voice sensor.
[0625] For a specific implementation of 1018, please refer to, for example. Figure 1 The 602 shown or Figure 11 The 702 shown will not be described in detail here. Optionally, the voice information g can correspond to the second voice information mentioned above.
[0626] In one possible example, the voice message g could be, for instance, "A vehicle accident has occurred."
[0627] 1019, the first device sends voice information f to the second device.
[0628] For a specific implementation of 1019, please refer to, for example. Figure 12 The 603b shown will not be described in detail here.
[0629] 1020, the first device determines the semantic recognition result p based on the voice information g.
[0630] For a specific implementation of 1020, please refer to, for example. The 603a or shown The 703 shown will not be described in detail here. Optionally, the semantic recognition result p can correspond to the seventh semantic recognition result mentioned above.
[0631] In one example, the semantic recognition result p may include function e and intent e; where function e may be "safety control function" and intent c may be "activate incident mode".
[0632] 1021, if the semantic recognition result p satisfies the second preset condition, determine to execute the operation f determined by the first device based on the semantic recognition result p.
[0633] For a specific implementation of 1021, please refer to, for example. The 707b shown will not be described in detail here. Optionally, operation f can correspond to, for example, the third operation mentioned above.
[0634] If the semantic recognition result p satisfies the second preset condition, it means that the first device can end the current cloud-side multi-round voice interaction. The first device can determine whether to execute the operation indicated by the semantic recognition result p. The first device can obtain the semantic recognition result of the second device for the voice information e to determine whether to execute the operation indicated by the voice information e.
[0635] In one possible example, the priority of semantic recognition result p can be higher than the priority of semantic recognition result n or semantic recognition result m.
[0636] 1022, the second device sends the semantic recognition result q for the voice information g to the first device.
[0637] For a detailed implementation of 1022, please refer to, for example. The 603b shown or The 704 shown here will not be described in detail here.
[0638] In one possible example, the semantic recognition result q could indicate, for instance, that the speech information g does not match the operation d.
[0639] 1023, the first device discards the semantic recognition result q from the second device.
[0640] Since the first device can determine to perform the operation f determined by the first device based on the semantic recognition result p, the first device can ignore the instructions of the second device for the voice information g.
[0641] For a detailed implementation of 1023, please refer to, for example. The 603b shown will not be described in detail here.
[0642] In the example shown in Figure 10, if the current voice information does not match the previous operation, the first device can determine whether to end the current multi-round voice interaction based on preset conditions related to multi-round voice interaction. This helps to adaptively retain the advantages of multi-round voice interaction and facilitates timely response to relatively urgent and important voice commands from users in special circumstances.
[0643] This is a schematic structural diagram of a voice interaction device 1100 provided in an embodiment of this application. The device 1100 includes an acquisition unit 1101 and a processing unit 1102. The device 1100 can be used to execute the steps of the voice interaction method provided in the embodiment of this application.
[0644] For example, the acquisition unit 1101 can be used to perform... In method 600 shown, at step 602, the processing unit 1102 can be used to execute... Method 603a of the illustrated method 600. Optionally, the apparatus 1100 further includes a transmitting unit, which can be used to perform... Method 600, specifically 603b.
[0645] For example, the acquisition unit 1101 can be used to execute... In the method 700 shown, 702 and 705, the processing unit 1102 can be used to execute... Methods 703, 704, 706, and 707 shown in method 700.
[0646] The acquisition unit 1101 is used to acquire first voice information from the voice sensor.
[0647] The processing unit 1102 is configured to determine, based on the first voice information, the target operation indicated by the first voice information.
[0648] Processing unit 1102 may include, for example, The example shown includes a semantic recognition module and an operation decision module.
[0649] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine a first semantic recognition result based on the first voice information; and determine to execute a first operation determined by the first device based on the first semantic recognition result, or to execute a second operation indicated by the second device, based on the first semantic recognition result and a first preset condition.
[0650] Optionally, as an embodiment, the processing unit 1102 is specifically used to: determine to perform the first operation when the first semantic recognition result satisfies the first preset condition.
[0651] Optionally, as an embodiment, the first device has multiple preset functions, and the first semantic recognition result satisfies a first preset condition, including: the first semantic recognition result indicates a first function, and the first function belongs to the multiple functions.
[0652] Optionally, as an embodiment, the first device has multiple preset intentions, and the first semantic recognition result satisfies a first preset condition, including: the first semantic recognition result indicates a first intention, and the first intention belongs to the multiple intentions.
[0653] Optionally, as an embodiment, the first device has multiple preset parameters, and the first semantic recognition result satisfies a first preset condition, including: the first semantic recognition result indicates a first parameter, and the first parameter belongs to the multiple parameters.
[0654] Optionally, as an embodiment, the first semantic recognition result indicates a first function and indicates a first parameter, and the first semantic recognition result satisfies a first preset condition, further including: the first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
[0655] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine a second semantic recognition result based on the first voice information, wherein the second semantic recognition result indicates a second function and indicates the first parameter; and, in the case where the second function is not included in the multiple functions preset by the first device and the first parameter is included in the multiple parameters preset by the first device, correct the second function in the second semantic recognition result to the first function to obtain the first semantic recognition result, wherein the first function and the second function are two different functions of the same type.
[0656] Optionally, as an embodiment, the first semantic recognition result indicates a first intent and indicates a first parameter, and the first semantic recognition result satisfies a first preset condition, further including: the first intent indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
[0657] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine a third semantic recognition result based on the first voice information, the third semantic recognition result including a second intention and an indication of the first parameter; when the first device presets multiple intentions but does not include the second intention, and the first device presets multiple parameters but includes the first parameter, correct the second intention in the third semantic recognition result to the first intention, thereby obtaining the first semantic recognition result, wherein the first intention and the second intention are two different intentions of the same type.
[0658] Optionally, as an embodiment, the first semantic recognition result satisfying the first preset condition includes: the first semantic recognition result includes a first indicator bit, the first indicator bit indicating that the first semantic recognition result satisfies the first preset condition.
[0659] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine a fourth semantic recognition result based on the first voice information, the fourth semantic recognition result including a first function and a first parameter; determine the first semantic recognition result when the first function belongs to a plurality of functions preset by the first device, and the first parameter belongs to a plurality of parameters preset by the first device, and the first function and the first parameter correspond to the same parameter type, the semantic recognition result including the first indicator bit.
[0660] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine a fifth semantic recognition result based on the first voice information, the fifth semantic recognition result including a first intention and a first parameter; determine the first semantic recognition result when the first intention belongs to a plurality of intentions preset by the first device, and the first parameter belongs to a plurality of parameters preset by the first device, and the first intention and the first parameter correspond to the same parameter type, the semantic recognition result including the first indicator bit.
[0661] Optionally, as an embodiment, the apparatus further includes: a sending unit, configured to send the first voice information to the second device; and a processing unit 1102, configured to discard the sixth semantic recognition result from the second device.
[0662] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: when the first semantic recognition result does not meet the first preset condition, obtain a sixth semantic recognition result from the second device, the sixth semantic recognition result being used to indicate the second operation; and determine to execute the second operation based on the sixth semantic recognition result.
[0663] Optionally, as an embodiment, the processing unit 1102 is further configured to: determine the second parameter and the second parameter type based on the sixth semantic recognition result; the device further includes: a storage unit for storing the association between the second parameter and the second parameter type.
[0664] Optionally, as an embodiment, the acquisition unit 1101 is further configured to acquire second voice information from the voice sensor; the processing unit 1102 is further configured to determine a seventh semantic recognition result based on the second voice information; the processing unit 1102 is further configured to determine, based on the seventh semantic recognition result and a second preset condition, to execute an operation indicated by the first voice information, or to execute a third operation determined by the first device based on the seventh semantic recognition result, or to execute a fourth operation indicated by the second device.
[0665] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine to perform the third operation when the seventh semantic recognition result satisfies both the first and second preset conditions; determine to perform the fourth operation when the seventh semantic recognition result does not satisfy the first preset conditions but satisfies the second preset conditions; and determine to perform the operation corresponding to the first semantic recognition result when the seventh semantic recognition result does not satisfy the second preset conditions.
[0666] Optionally, as an embodiment, the seventh semantic recognition result satisfies the second preset condition, including: the priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result.
[0667] Optionally, as an embodiment, the priority of the seventh semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: the function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; the intent indicated by the seventh semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; the parameter indicated by the seventh semantic recognition result has a higher priority than the parameter indicated by the first semantic recognition result.
[0668] Optionally, as an embodiment, the processing unit is specifically configured to: when the first semantic recognition result does not meet the first preset condition, obtain a sixth semantic recognition result from the second device, the sixth semantic recognition result being used to indicate the second operation; determine to execute the second operation based on the sixth semantic recognition result; the seventh semantic recognition result meeting the second preset condition includes: the priority of the seventh semantic recognition result being higher than the priority of the sixth semantic recognition result.
[0669] Optionally, as an embodiment, the priority of the seventh semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: the function indicated by the seventh semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; the intent indicated by the seventh semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; the parameters indicated by the seventh semantic recognition result have a higher priority than the parameters indicated by the sixth semantic recognition result.
[0670] Optionally, as an embodiment, the apparatus further includes: a sending unit, configured to send second voice information to the second device; the processing unit 1102 is specifically configured to: obtain an eighth semantic recognition result from the second device when the seventh semantic recognition result does not meet the first preset condition, or when the seventh semantic recognition result does not meet the second preset condition and the first semantic recognition result does not meet the first preset condition; and determine to execute the operation indicated by the first voice information, or determine to execute the fourth operation, based on the eighth semantic recognition result and the second preset condition.
[0671] The sending unit can be, for example, The transceiver module shown in the example.
[0672] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: determine to perform the fourth operation when the eighth semantic recognition result satisfies the second preset condition; and determine to perform the operation indicated by the first voice information when the eighth semantic recognition result does not satisfy the second preset condition.
[0673] Optionally, as an embodiment, the eighth semantic recognition result satisfies the second preset condition, including: the priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result.
[0674] Optionally, as an embodiment, the priority of the eighth semantic recognition result is higher than the priority of the first semantic recognition result, including one or more of the following: the function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the first semantic recognition result; the intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the first semantic recognition result; the parameter indicated by the eighth semantic recognition result has a higher priority than the parameter indicated by the first semantic recognition result.
[0675] Optionally, as an embodiment, the processing unit 1102 is specifically configured to: when the first semantic recognition result does not meet the first preset condition, obtain a sixth semantic recognition result from the second device, the sixth semantic recognition result being used to indicate the second operation; determine to execute the second operation based on the sixth semantic recognition result; the eighth semantic recognition result meeting the second preset condition includes: the priority of the eighth semantic recognition result being higher than the priority of the sixth semantic recognition result.
[0676] Optionally, as an embodiment, the priority of the eighth semantic recognition result is higher than the priority of the sixth semantic recognition result, including one or more of the following: the function indicated by the eighth semantic recognition result has a higher priority than the function indicated by the sixth semantic recognition result; the intent indicated by the eighth semantic recognition result has a higher priority than the intent indicated by the sixth semantic recognition result; the parameters indicated by the eighth semantic recognition result have a higher priority than the parameters indicated by the sixth semantic recognition result.
[0677] Alternatively, as an embodiment, the second voice information is unrelated to the operation indicated by the first voice information.
[0678] Optionally, as an embodiment, the device further includes a wake-up module for responding to user input and performing a voice wake-up operation.
[0679] This is a schematic structural diagram of a voice interaction device 1200 provided in an embodiment of this application. The device 1200 may include at least one processor 1202 and a communication interface 1203.
[0680] Optionally, the device 1200 may also include one or more of the memory 1201 and the bus 1204. Any two or all three of the memory 1201, processor 1202, and communication interface 1203 can be interconnected via the bus 1204.
[0681] Optionally, the memory 1201 may be a read-only memory (ROM), a static storage device, a dynamic storage device, or a random access memory (RAM). The memory 1201 may store programs. When the program stored in the memory 1201 is executed by the processor 1202, the processor 1202 and the communication interface 1203 are used to execute the various steps of the voice interaction method provided in this embodiment. That is, the processor 1202 can obtain stored instructions from the memory 1201 through the communication interface 1203 to execute the various steps of the voice interaction method provided in this embodiment.
[0682] Optionally, the memory 1201 may have The memory 152 shown is configured to perform the functions of the stored program described above. Optionally, the processor 1202 may be a general-purpose central processing unit (CPU), microprocessor, application-specific integrated circuit (ASIC), graphics processing unit (GPU), or one or more integrated circuits, used to execute relevant programs to perform the functions required by the units in the voice interaction device provided in this application embodiment, or to execute the various steps of the voice interaction method provided in this application embodiment.
[0683] Optionally, the processor 1202 may have The processor 151 shown is configured to perform the aforementioned function of executing related programs.
[0684] Optionally, the processor 1202 can also be an integrated circuit chip with signal processing capabilities. In implementation, each step of the voice interaction method provided in this application embodiment can be completed by integrated logic circuits in the processor's hardware or by software instructions.
[0685] Optionally, the processor 1202 described above can also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied in the execution of a hardware decoding processor, or executed by a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. The storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the functions required by the units included in the voice interaction device of the embodiments of this application, or executes the various steps of the voice interaction method provided in the embodiments of this application.
[0686] Optionally, the communication interface 1203 can use a transceiver device, such as, but not limited to, a transceiver, to enable communication between the device and other devices or communication networks. The communication interface 1203 can also be, for example, an interface circuit.
[0687] Bus 1204 may include a pathway for transmitting information between various components of the device (e.g., memory, processor, communication interface).
[0688] This application also provides a computer program product containing instructions that, when executed by a computer, cause the computer to implement the voice interaction method provided in the above-described method embodiments.
[0689] This application embodiment also provides a terminal device, which includes any of the above-mentioned voice interaction devices, such as... or The device shown is a voice interaction device, etc.
[0690] For example, the terminal can be a vehicle. Alternatively, the terminal can also be a terminal for remotely controlling a vehicle.
[0691] The aforementioned voice interaction device can be installed on the vehicle or can be independent of the vehicle. For example, it can be controlled by using drones, other vehicles, robots, etc.
[0692] Unless otherwise defined, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this application and in its specification is for the purpose of describing particular embodiments only and is not intended to be limiting of the application.
[0693] Various aspects or features of this application may be implemented as methods, apparatus, or articles of manufacture using standard programming and / or engineering techniques. The term "article of manufacture" as used in this application may encompass a computer program accessible from any computer-readable device, carrier, or medium. For example, computer-readable media may include, but are not limited to: magnetic storage devices (e.g., hard disks, floppy disks, or magnetic tapes), optical discs (e.g., compact discs (CDs), digital versatile discs (DVDs), etc.), smart cards, and flash memory devices (e.g., erasable programmable read-only memory (EPROMs), cards, sticks, or key drives, etc.).
[0694] The various storage media described in this application may represent one or more devices and / or other machine-readable media for storing information. The term "machine-readable media" may include, but is not limited to, wireless channels and various other media capable of storing, containing and / or carrying instructions and / or data.
[0695] It should be noted that when the processor is a general-purpose processor, DSP, ASIC, FPGA, or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component, the memory (storage module) can be integrated into the processor.
[0696] It should also be noted that the memory described in this application is intended to include, but is not limited to, these and any other suitable types of memory.
[0697] As used in this specification, the terms "component," "module," "system," etc., are used to refer to computer-related entities, hardware, firmware, combinations of hardware and software, software, or software in execution. For example, a component can be, but is not limited to, a process running on a processor, a processor, an object, an executable file, an execution thread, a program, and / or a computer. As illustrated, applications running on computing devices and computing devices can both be components. One or more components may reside in a process and / or an execution thread, and components may be located on a single computer and / or distributed among two or more computers. Furthermore, these components can be executed from various computer-readable media on which various data structures are stored. Components can communicate, for example, via local and / or remote processes based on signals having one or more data packets (e.g., data from two components interacting with another component between a local system, a distributed system, and / or a network, such as the Internet interacting with other systems via signals).
[0698] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0699] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0700] In the several embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0701] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0702] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0703] If the aforementioned functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0704] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.
Claims
1. A method for voice interaction, characterized in that, Applied to a first device, the method includes: Acquire the first voice information from the voice sensor; The first voice information is sent to the second device, and the second device is used to provide feedback to the first device regarding the first voice information; Based on the first voice information, determine the first semantic recognition result; Based on the first semantic recognition result and the first preset condition, it is determined to execute a first operation determined by the first device based on the first semantic recognition result, or to execute a second operation indicated by the second device.
2. The method as described in claim 1, characterized in that, When the first preset condition is met Discard or ignore feedback from the second device; or, Skip receiving feedback from the second device regarding the first voice information.
3. The method according to claim 1 or 2, characterized in that, The determination to perform the first operation, as determined by the first device, occurs before receiving feedback from the second device regarding the first voice information.
4. The method according to any one of claims 1 to 3, characterized in that, The step of determining to execute the first operation determined by the first device based on the first semantic recognition result and the first preset condition includes: If the first semantic recognition result meets the first preset condition, the first operation is executed.
5. The method as described in claim 4, characterized in that, The first device has multiple preset functions. The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result indicates a first function, which belongs to the plurality of functions.
6. The method as described in claim 4 or 5, characterized in that, The first device has multiple preset intentions. The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result indicates a first intent, which belongs to the plurality of intents.
7. The method according to any one of claims 4 to 6, characterized in that, The first device has multiple preset parameters. The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result indicates the first parameter, which belongs to the plurality of parameters.
8. The method according to any one of claims 4 to 7, characterized in that, The first semantic recognition result indicates a first function and indicates a first parameter. The first semantic recognition result satisfies a first preset condition and further includes: The first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
9. The method as described in claim 8, characterized in that, The step of determining the first semantic recognition result based on the first voice information includes: Based on the first voice information, a second semantic recognition result is determined, wherein the second semantic recognition result indicates a second function and indicates the first parameter; If the second function is not included in the preset functions of the first device, and the first parameter is included in the preset parameters of the first device, the second function in the second semantic recognition result is corrected to the first function to obtain the first semantic recognition result. The first function and the second function are two different functions of the same type.
10. The method according to any one of claims 4 to 7, characterized in that, The first semantic recognition result indicates a first intent and indicates a first parameter. The first semantic recognition result satisfies a first preset condition and further includes: The first intention indicated by the first semantic recognition corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
11. The method as described in claim 10, characterized in that, The step of determining the first semantic recognition result based on the first voice information includes: Based on the first voice information, a third semantic recognition result is determined, wherein the third semantic recognition result indicates a second intention and indicates the first parameter; If the second intention is not included in the multiple intentions preset by the first device, and the first parameter is included in the multiple parameters preset by the first device, the second intention in the third semantic recognition result is corrected to the first intention to obtain the first semantic recognition result. The first intention and the second intention are two different intentions of the same type.
12. The method as described in claim 4, characterized in that, The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result includes a first indicator bit, which indicates that the first semantic recognition result meets the first preset condition.
13. The method as described in claim 12, characterized in that, The step of determining the first semantic recognition result based on the first voice information includes: Based on the first voice information, a fourth semantic recognition result is determined, wherein the fourth semantic recognition result indicates the first function and the first parameter; When the first function belongs to multiple functions preset by the first device, and the first parameter belongs to multiple parameters preset by the first device, and the first function and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
14. The method as described in claim 12, characterized in that, The step of determining the first semantic recognition result based on the first voice information includes: Based on the first voice information, a fifth semantic recognition result is determined, wherein the fifth semantic recognition result indicates the first intention and the first parameter; When the first intent belongs to a plurality of intents preset by the first device, and the first parameter belongs to a plurality of parameters preset by the first device, and the first intent and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
15. The method according to any one of claims 4 to 14, characterized in that, The method further includes: Send the first voice information to the second device; The sixth semantic recognition result from the second device is discarded.
16. The method according to any one of claims 1 to 15, characterized in that, The step of determining to execute the second operation instructed by the second device based on the first semantic recognition result and the first preset condition includes: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed.
17. The method as described in claim 16, characterized in that, The method further includes: Based on the sixth semantic recognition result, determine the second parameter and the type of the second parameter; Save the association between the second parameter and the second parameter type.
18. The method according to any one of claims 1 to 17, characterized in that, The method further includes: Acquire second voice information from the voice sensor; Based on the second speech information, determine the seventh semantic recognition result; Based on the seventh semantic recognition result and the second preset condition, it is determined to execute the operation indicated by the first voice information, or to execute the third operation determined by the first device based on the seventh semantic recognition result, or to execute the fourth operation indicated by the second device.
19. The method as described in claim 18, characterized in that, The step of determining to execute the operation indicated by the first voice information based on the seventh semantic recognition result and the second preset condition, or determining to execute the third operation determined by the first device based on the seventh semantic recognition result, or determining to execute the fourth operation indicated by the second device, includes: If the seventh semantic recognition result satisfies both the first preset condition and the second preset condition, the third operation is to be performed. If the seventh semantic recognition result does not meet the first preset condition but meets the second preset condition, the fourth operation is to be performed. If the seventh semantic recognition result does not meet the second preset condition, it is determined to perform the operation corresponding to the first semantic recognition result.
20. A voice interaction device, characterized in that, include: An acquisition unit is used to acquire first speech information from a speech sensor; The processing unit is configured to send the first voice information to the second device, and the second device is configured to provide feedback to the first device regarding the first voice information. The processing unit is further configured to determine a first semantic recognition result based on the first voice information; The processing unit is further configured to determine, based on the first semantic recognition result and the first preset condition, to execute a first operation determined by the first device based on the first semantic recognition result, or to determine to execute a second operation indicated by the second device.
21. The apparatus as claimed in claim 20, characterized in that, When the first preset condition is met The processing unit is also used for: Discard or ignore feedback from the second device; or, Skip receiving feedback from the second device regarding the first voice information.
22. The apparatus as claimed in claim 20 or 21, characterized in that, The determination to perform the first operation, as determined by the first device, occurs before receiving feedback from the second device regarding the first voice information.
23. The apparatus as claimed in any one of claims 20 to 22, characterized in that, The processing unit is specifically used for: If the first semantic recognition result meets the first preset condition, the first operation is executed.
24. The apparatus as claimed in claim 23, characterized in that, The first device has multiple preset functions, and the first semantic recognition result satisfies a first preset condition, including: The first semantic recognition result indicates a first function, which belongs to the plurality of functions.
25. The apparatus as claimed in claim 23 or 24, characterized in that, The first device has multiple preset intentions. The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result indicates a first intent, which belongs to the plurality of intents.
26. The apparatus as claimed in any one of claims 23 to 25, characterized in that, The first device has multiple preset parameters. The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result indicates the first parameter, which belongs to the plurality of parameters.
27. The apparatus as claimed in any one of claims 23 to 26, characterized in that, The first semantic recognition result indicates the first function and indicates the first parameter. The first semantic recognition result satisfies the first preset condition, and further includes: The first function indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
28. The apparatus as claimed in claim 27, characterized in that, The processing unit is specifically used for: Based on the first voice information, a second semantic recognition result is determined, wherein the second semantic recognition result indicates a second function and indicates the first parameter; If the second function is not included in the preset functions of the first device, and the first parameter is included in the preset parameters of the first device, the second function in the second semantic recognition result is corrected to the first function to obtain the first semantic recognition result. The first function and the second function are two different functions of the same type.
29. The apparatus as claimed in any one of claims 23 to 26, characterized in that, The first semantic recognition result indicates the first intent and the first parameter. The first semantic recognition result satisfies the first preset condition, and further includes: The first intent indicated by the first semantic recognition result corresponds to the same parameter type as the first parameter indicated by the first semantic recognition result.
30. The apparatus as claimed in claim 29, characterized in that, The processing unit is specifically used for: Based on the first voice information, a third semantic recognition result is determined, wherein the third semantic recognition result indicates a second intention and indicates the first parameter; If the second intention is not included in the multiple intentions preset by the first device, and the first parameter is included in the multiple parameters preset by the first device, the second intention in the third semantic recognition result is corrected to the first intention to obtain the first semantic recognition result. The first intention and the second intention are two different intentions of the same type.
31. The apparatus as claimed in claim 23, characterized in that, The first semantic recognition result satisfies the first preset condition, including: The first semantic recognition result includes a first indicator bit, which indicates that the first semantic recognition result meets the first preset condition.
32. The apparatus as claimed in claim 31, characterized in that, The processing unit is specifically used for: Based on the first voice information, a fourth semantic recognition result is determined, wherein the fourth semantic recognition result indicates the first function and the first parameter; When the first function belongs to multiple functions preset by the first device, and the first parameter belongs to multiple parameters preset by the first device, and the first function and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
33. The apparatus as claimed in claim 31, characterized in that, The processing unit is specifically used for: Based on the first voice information, a fifth semantic recognition result is determined, wherein the fifth semantic recognition result includes a first intent and a first parameter; When the first intent belongs to a plurality of intents preset by the first device, and the first parameter belongs to a plurality of parameters preset by the first device, and the first intent and the first parameter correspond to the same parameter type, the first semantic recognition result is determined, and the semantic recognition result includes the first indicator bit.
34. The apparatus according to any one of claims 23 to 33, characterized in that, The device further includes: A sending unit is configured to send the first voice information to the second device; The processing unit is also used to discard the sixth semantic recognition result from the second device.
35. The apparatus as claimed in any one of claims 20 to 22, characterized in that, The processing unit is specifically used for: If the first semantic recognition result does not meet the first preset condition, a sixth semantic recognition result from the second device is obtained, and the sixth semantic recognition result is used to indicate the second operation; Based on the sixth semantic recognition result, it is determined that the second operation will be performed.
36. The apparatus as claimed in claim 35, characterized in that, The processing unit is also used for: Based on the sixth semantic recognition result, determine the second parameter and the type of the second parameter; The device further includes a storage unit for storing the association between the second parameter and the second parameter type.
37. The apparatus as claimed in any one of claims 20 to 36, characterized in that, The acquisition unit is further configured to acquire second voice information from the voice sensor; The processing unit is further configured to determine the seventh semantic recognition result based on the second speech information; The processing unit is further configured to, based on the seventh semantic recognition result and the second preset conditions, determine to execute the operation indicated by the first voice information, or determine to execute the third operation determined by the first device based on the seventh semantic recognition result, or determine to execute the fourth operation indicated by the second device.
38. The apparatus as claimed in claim 37, characterized in that, The processing unit is specifically used for, If the seventh semantic recognition result satisfies both the first preset condition and the second preset condition, the third operation is to be performed. If the seventh semantic recognition result does not meet the first preset condition but meets the second preset condition, the fourth operation is to be performed. If the seventh semantic recognition result does not meet the second preset condition, it is determined to perform the operation corresponding to the first semantic recognition result.
39. A computer-readable medium, characterized in that, The computer-readable medium stores program code for execution by the device, the program code including methods for performing any one of claims 1 to 19.
40. A computer program product comprising instructions that, when run on a computer, cause the computer to perform the method as claimed in any one of claims 1 to 19.