Voice instruction processing method and device, electronic equipment, vehicle and storage medium

By parsing voice commands and selecting the target application to execute, the problem of abnormal response from multiple applications in the vehicle was solved, enabling normal execution of voice commands and improving user experience.

CN115424610BActive Publication Date: 2026-04-21BEIJING CO WHEELS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2022-06-30
Publication Date
2026-04-21

AI Technical Summary

Technical Problem

In a vehicle, a user's voice commands may simultaneously trigger multiple applications, causing these applications to malfunction and impacting the user experience.

Method used

By parsing the content of voice commands, the target application's instruction information and the user's intent are determined. Using preset correspondences and priority information, a target application is selected from multiple executable applications to execute the voice command.

Benefits of technology

This avoids anomalies caused by multiple applications responding simultaneously, ensures that voice commands are executed normally, and improves the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115424610B_ABST
    Figure CN115424610B_ABST
Patent Text Reader

Abstract

This disclosure relates to a method, apparatus, electronic device, vehicle, and storage medium for processing voice commands. The method includes: parsing the content of a target voice command to obtain target application indication information and at least two user intents; determining an application matching each of the at least two user intents according to a preset correspondence, the preset correspondence including a preset user intent and its matching application; and determining a target application from the determined applications based on the target application indication information. This method can determine a target application from at least two executable applications to execute the target voice command, enabling the target voice command to be executed correctly, thereby improving the user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of voice control technology, and in particular to a method, apparatus, electronic device, vehicle, and storage medium for processing voice commands. Background Technology

[0002] With the development of vehicles, users have increasingly higher demands for driving and entertainment experiences. This has led to the emergence of vehicles equipped with terminal devices, which can include at least two applications, such as audio applications, video applications, karaoke applications, etc. Based on different applications, users can experience different entertainment programs.

[0003] However, when a voice command issued by a user hits at least two user intents, these at least two user intents correspond to at least two applications, resulting in at least two applications responding to the voice command. This causes abnormal voice command responses, thereby affecting the user experience. Summary of the Invention

[0004] This disclosure provides a method, apparatus, electronic device, vehicle, and storage medium for processing voice commands, which can determine a target application from at least two applications that can execute the target voice command to execute the voice command, so that the target voice command can be executed normally, thereby improving the user experience.

[0005] Firstly, this disclosure provides a method for processing voice commands, including:

[0006] The content of the target voice command is parsed to obtain target application instruction information and at least two user intents; based on the at least two user intents and a preset correspondence, the application matching each user intent is determined, wherein the preset correspondence includes a preset user intent and the application matching it;

[0007] Based on the target application indication information, a target application is determined from the identified applications. Optionally, the target application indication information includes keywords from the target voice command;

[0008] The step of determining a target application from the identified applications based on the target application indication information includes:

[0009] Based on the keywords in the target voice command and the preset keyword correspondence, candidate applications are determined from the identified applications. The preset keyword correspondence includes preset keywords and their corresponding preset applications.

[0010] If there are at least two candidate applications, the target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and the preset priority information. The preset priority information includes at least two preset status information and their respective priorities.

[0011] If there is only one candidate application, then the candidate application is determined to be the target application.

[0012] Optionally, the target application indication information includes the confidence level of the user's intent;

[0013] The step of determining a target application from the identified applications based on the target application indication information includes:

[0014] Candidate applications are determined from the identified applications based on the confidence levels of at least two user intents and a preset confidence level, wherein the confidence level of the user intent corresponding to the candidate application is greater than the preset confidence level.

[0015] If there are at least two candidate applications, the target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and the preset priority information. The preset priority information includes at least two preset status information and their respective priorities.

[0016] If there is only one candidate application, then the candidate application is determined to be the target application.

[0017] Optionally, the target application information includes the confidence level of the user's intent;

[0018] The step of determining a target application from the identified applications based on the target application indication information includes:

[0019] Execute a preset instruction and determine the maximum confidence level from at least two confidence levels of the user intent;

[0020] The application corresponding to the user intent with the maximum confidence level is identified as a candidate application.

[0021] If there are at least two candidate applications, the target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and the preset priority information. The preset priority information includes at least two preset status information and their respective priorities.

[0022] If there is only one candidate application, then the candidate application is determined to be the target application.

[0023] Optionally, the target application indication information includes the default application;

[0024] The step of determining a target application from the identified applications based on the target application indication information includes:

[0025] If the identified applications include the default application, then the default application is identified as the target application.

[0026] Optionally, determining the target application from at least two candidate applications based on the status information of at least two candidate applications and preset priority information includes:

[0027] Based on the status information of at least two candidate applications and the preset priority information, determine the status information with the highest priority among the status information of at least two candidate applications;

[0028] The candidate application corresponding to the highest priority status information is determined as the target application.

[0029] Optionally, determining the highest priority status information among the status information of at least two candidate applications based on the status information of at least two candidate applications and preset priority information includes:

[0030] Select the target state information with the highest priority from the at least two preset state information;

[0031] Based on the target state information, traverse the state information of at least two candidate applications;

[0032] If it is determined that the target state information is not included in the state information of at least two candidate applications, the target state information is updated in descending order of priority in the preset priority information. Then, the process of traversing the state information of at least two candidate applications based on the target state information is repeated until it is determined that the target state information is included in the state information of at least two candidate applications. The target state information is then determined to be the highest priority state information.

[0033] Optionally, determining the target application from at least two candidate applications based on the status information of at least two candidate applications and preset priority information includes:

[0034] Based on the status information of at least two candidate applications and the preset priority information, it is determined that the status information of at least two candidate applications does not include any of the preset status information in the preset priority information;

[0035] The target application is determined by selecting a preset default application from at least two of the candidate applications.

[0036] Optionally, determining that the state information of at least two candidate applications does not include any of the preset state information in the preset priority information, based on the state information of at least two candidate applications and preset priority information, includes:

[0037] Select the target state information with the highest priority from at least two preset state information;

[0038] Based on the target state information, traverse the state information of at least two candidate applications;

[0039] If it is determined that the target state information is not included in the state information of at least two candidate applications, the target state information is updated in descending order of priority in the preset priority information, and the process of traversing the state information of at least two candidate applications based on the target state information is resumed until it is determined that the target state information with the lowest priority is not included in the state information of at least two candidate applications.

[0040] Optionally, before parsing the content of the target voice command to obtain the target application instruction information and at least two user intents, the method further includes:

[0041] If the target voice command includes terminal device identification information, the terminal device corresponding to the terminal device identification information is determined to be the target terminal device executing the target voice command.

[0042] If the target voice command does not include terminal device identification information, the terminal device where the virtual voice avatar is located is determined to be the target terminal device executing the target voice command.

[0043] Secondly, this disclosure provides a voice command processing apparatus, comprising:

[0044] The parsing module is used to parse the content of the target voice command to obtain the target application instruction information and at least two user intents.

[0045] The determining module is configured to determine an application matching each of the at least two user intentions and a preset correspondence, wherein the preset correspondence includes a preset user intention and an application matching it; and to determine a target application from the determined applications based on the target application indication information.

[0046] Thirdly, this disclosure provides an electronic device, including: a processor, the processor being configured to execute a computer program stored in a memory, the computer program being executed by the processor to implement the steps of any of the methods provided in the first aspect.

[0047] Fourthly, this disclosure provides a vehicle, including:

[0048] A controller is configured to parse a target voice command to obtain at least two user intents; determine at least two applications corresponding to the at least two user intents based on the at least two user intents and a preset correspondence, wherein the preset correspondence includes at least two preset user intents and their respective preset applications; and determine a target application that matches the preset information from the at least two applications based on at least one target information corresponding to the target voice command and preset information.

[0049] Fifthly, this disclosure provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of any of the methods provided in the first aspect.

[0050] In the technical solution provided in this disclosure, the target application instruction information and at least two user intents are obtained by parsing the content of the target voice command. Based on the at least two user intents and a preset correspondence, an application matching each user intent is determined. The preset correspondence includes a preset user intent and its matching application. Based on the target application instruction information, a target application is determined from the determined applications. A target application can be determined from at least two applications that can execute the target voice command, so that the target application executes the target voice command. This avoids response anomalies caused by at least two applications executing the target voice command, allowing the target voice command to be executed normally, thereby improving the user experience. Attached Figure Description

[0051] The accompanying drawings, which are incorporated in and form a part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure.

[0052] To more clearly illustrate the technical solutions in the embodiments of this disclosure or the prior art, the accompanying drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0053] Figure 1 A schematic diagram illustrating one application scenario provided in this disclosure;

[0054] Figure 2 A flowchart illustrating a voice command processing method provided in this disclosure;

[0055] Figure 3 A flowchart illustrating another method for processing voice commands provided in this disclosure;

[0056] Figure 4A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0057] Figure 5 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0058] Figure 6 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0059] Figure 7 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0060] Figure 8 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0061] Figure 9 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0062] Figure 10 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0063] Figure 11 A flowchart illustrating yet another voice command processing method provided in this disclosure;

[0064] Figure 12 This is a schematic diagram of the structure of a voice command processing device provided in this disclosure;

[0065] Figure 13 This is a schematic diagram of the structure of an electronic device provided in this disclosure. Detailed Implementation

[0066] To better understand the above-mentioned objectives, features, and advantages of this disclosure, the solutions disclosed herein will be further described below. It should be noted that, unless otherwise specified, the embodiments and features described herein can be combined with each other.

[0067] Numerous specific details are set forth in the following description in order to provide a full understanding of this disclosure, but this disclosure may also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only some, and not all, of the embodiments of this disclosure.

[0068] Figure 1 This is a schematic diagram illustrating one application scenario provided in this disclosure, such as... Figure 1 As shown, the application scenarios include terminal devices 110, such as smartphones, PDAs, tablets, wearable devices with displays, desktop computers, laptops, all-in-one computers, smart home devices, etc. If the application scenario is a vehicle, such as... Figure 1 As shown, terminal device 110 can be a vehicle-mounted device. The application scenario also includes a sound acquisition device 120, such as... Figure 1 As shown, the sound acquisition device can be, for example, a microphone, which can collect the user's voice signal. The terminal device can receive the sound signal sent by the sound acquisition device, and based on the sound signal, can generate voice commands triggered by the user. For example, voice commands can be "pause playback", "continue playback", "play XXX", "previous", "next", "fast forward", "open XXX", and "close XXX", etc.

[0069] Terminal devices include various types of applications. Each type may include one application or at least two applications. For example, audio applications on a terminal device may include application A and application B, while video applications may include application C. User-triggered voice commands can generally be divided into five types: resume playback, control, on-demand playback, interface opening, and interface closing. Resume playback voice commands may include "play," "continue playback," "previous," and "next," etc. Control voice commands may include "pause playback," "favorite," "remove favorite," "fast forward," and "rewind," etc. On-demand voice commands may include "search for XXX" and "play XXX," interface opening voice commands may include "open XXX," and interface closing voice commands may include "close XXX."

[0070] When the target voice command triggered by the user is a video-on-demand command, by parsing the content of the target voice command, at least two user intents can be obtained. For example, if the target voice command triggered by the user is "Play XXX", parsing the target voice command "Play XXX" yields two user intents: "Play audio" and "Play video". For each user intent, the application that can execute the target voice command is an application of the same type. For example, based on the above embodiment, for the user intent "Play audio", the application that can execute the target voice command is an audio application, i.e., application A and application B; for the user intent "Play video", the application that can execute the target voice command is a video application, i.e., application C. Thus, the applications that can execute the target voice command are A, B, and C. Therefore, for at least two user intents triggered by the target voice command, the applications that can execute the target voice command are at least two different types of applications, which can easily lead to abnormal data processing within the terminal device, i.e., abnormal response of the terminal device to voice commands.

[0071] To address the aforementioned issues, this disclosure involves parsing the content of the target voice command to obtain target application instruction information and at least two user intents. Based on the at least two user intents and a preset correspondence, an application matching each user intent is determined. The preset correspondence includes a preset user intent and its matching application. Based on the target application instruction information, a target application is selected from the determined applications. This can be done by selecting one target application from at least two applications capable of executing the target voice command, thus enabling the target application to execute the target voice command. This avoids response anomalies caused by at least two applications executing the target voice command, ensuring the target voice command can be executed normally and improving the user experience.

[0072] The technical solutions in this disclosure are illustrated below with several specific implementation methods:

[0073] Figure 2 The following is a flowchart illustrating a voice command processing method provided in this disclosure, such as... Figure 2 As shown, it includes:

[0074] S101, parse the content of the target voice command to obtain the target application instruction information and at least two user intents.

[0075] For example, the sound acquisition device can collect the sound signal emitted by the user. The terminal device receives the sound signal sent by the sound acquisition device and generates a target voice command based on the received sound signal. The target voice command is a play-on-demand voice command, such as "Play XXX" or "Search XXX". After obtaining the target voice command triggered by the user, the content of the target voice command is parsed to determine at least two user intentions that the target voice command matches. For example, if the target voice command is "Play 123", the user intentions matched by the target voice command include audio intentions and video intentions, that is, the target voice command matches at least two user intentions.

[0076] Furthermore, parsing the content of the target voice command can yield target application instruction information, which includes keywords in the target voice command, the confidence level of the user's intent, and any one of the default applications.

[0077] S102, based on the at least two user intents and the preset correspondence, determine the application that matches each of the user intents.

[0078] The preset correspondence includes preset user intentions and their matching applications.

[0079] For example, based on at least two regular voice commands that a user can trigger, at least two preset user intentions can be obtained. For each preset user intention, an application type is determined from all types of applications on the terminal device. The terminal device can include multiple types of applications, such as audio applications, video applications, karaoke applications, news applications, etc. Each type of application may include one or more applications; for example, audio applications include application A and application B, video applications include application C, karaoke applications include application D, and news applications include a preset application E. In this way, one or more corresponding applications can be obtained for each preset user intention, thereby establishing a preset correspondence.

[0080] Based on the user-triggered target voice command, at least two user intents can be obtained. For each user intent, a preset mapping relationship is traversed to check if the user intent is included in the preset mapping relationship. If the user intent is included in the preset mapping relationship, one or at least two applications corresponding to the user intent are determined. In this way, at least two applications corresponding to at least two user intents can be obtained. For example, the preset mapping relationship includes the preset user intent "play audio" and the preset user intent "play video". The preset user intent "play audio" corresponds to applications A and B, and the preset user intent "play video" corresponds to application C. If the target voice command hits at least two user intents, namely "play video" and "play audio", then the user intent "play audio" corresponds to applications A and B, and the user intent "play video" corresponds to application C. Thus, the applications corresponding to the at least two user intents hit by the target voice command are A, B, and C.

[0081] S103, Based on the target application indication information, determine a target application from the identified applications.

[0082] For example, the target application indication information is a keyword in the target voice command. Keywords can be understood as verbs that express strong intent, such as "listen," "see," or "sing." Based on the keywords in the target voice command and a preset keyword correspondence, an application matching the keyword can be determined from the preset keyword correspondence. If the application matching the keyword is one of at least two applications, then this application is determined to be the target application. In other embodiments, the target application indication information is the confidence level of the user intent. Based on the confidence levels of at least two user intents and a preset confidence level, a target application whose confidence level matches the preset confidence level can be determined from at least two applications. In still other embodiments, the target application indication information is the confidence level of the user intent. Based on the confidence levels of at least two user intents and a preset instruction indicating the maximum confidence level, a target application matching the maximum confidence level can be determined.

[0083] Based on the above embodiments, a target application in the terminal device that executes the target voice command can be identified. Then, the application's interface can be invoked to transmit the target voice command to the target application. If the target application supports the target voice command, it can execute the target voice command; if it does not support the target voice command, it returns a first prompt message indicating that the target voice command is not supported. For example, the first prompt message could be "The current playback type does not support this operation."

[0084] In this embodiment, by parsing the content of the target voice command, target application indication information and at least two user intents are obtained. Based on the at least two user intents and a preset correspondence, at least two applications matching each user intent are determined. The preset correspondence includes a preset user intent and its matching application. Based on the target application indication information, a target application is determined from the determined applications. A target application can be determined from at least two applications that can execute the target voice command, so that the target application executes the target voice command. This avoids response anomalies caused by at least two applications executing the target voice command, allowing the target voice command to be executed normally, thereby improving the user experience.

[0085] Figure 3 This is a flowchart illustrating another method for processing voice commands provided in this disclosure. Figure 3 for Figure 2 Based on the illustrated embodiment, a specific description of a possible implementation of S103 is as follows:

[0086] S201, Based on the correspondence between the keywords in the target voice command and the preset keywords, candidate applications are determined from the identified applications.

[0087] The preset keyword correspondence includes preset keywords and their corresponding preset applications.

[0088] For example, a target voice command corresponds to a target application instruction, and the target application instruction includes keywords from the target voice command. The preset keyword mapping relationship includes at least two preset keywords, which are verbs used to express strong intent, such as "listen," "see," and "sing," and each preset keyword corresponds to one or at least two preset applications. For example, based on the above embodiment, the preset keyword mapping relationship includes three preset keywords: "see," "listen," and "sing," where the preset keyword "listen" corresponds to preset applications A and B, the preset keyword "see" corresponds to preset application C, and the preset keyword "sing" corresponds to preset application D.

[0089] Based on semantic understanding of the target voice command, keywords within the command can be obtained. A preset keyword mapping relationship is then traversed to check if the keyword exists within that relationship. If the keyword exists, the corresponding preset application is identified as a candidate application. There may be one or at least two candidate applications. For example, based on the above embodiment, if the keyword in the target voice command is "listen," then preset applications A and B are both candidate applications; if the keyword is "see," then preset application C is also a candidate application.

[0090] S202, determine whether the candidate application is a single application.

[0091] If yes, execute S203; otherwise, execute S203'.

[0092] For example, after obtaining candidate applications from at least two applications, it is determined whether there is only one candidate application. For instance, based on the above embodiments, if the keyword in the target voice command is "listen", at least two candidate applications can be determined, and if the keyword in the target voice command is "see", one candidate application can be determined.

[0093] S203, determine the candidate application as the target application.

[0094] If a candidate application is determined from at least two applications, then the candidate application can be determined to be the target application. For example, based on the above embodiment, if the keyword in the target voice command is "see", and candidate application C is obtained, then candidate application C can be determined to be the target application.

[0095] S203', Based on the status information of at least two candidate applications and preset priority information, determine the target application from at least two candidate applications.

[0096] The preset priority information includes at least two preset state information and their respective priorities.

[0097] If at least two candidate applications are determined from at least two applications, the status information of at least two candidate applications is obtained. The status information may include at least one of running status, playback status, and recent playback record status. The running status may be running in the foreground, running, running in the background, or not running. The playback status may be playing with headphones, playing with amplifier, playing, or not playing. The recent playback record status information may be that there is a recent playback record, there is a recent playback record with headphones, there is a recent playback record with amplifier, or there is no recent playback record.

[0098] The preset priority information includes at least two preset state information. For example, the priority information includes three preset state information: running in the foreground, playing, and having recently played records. Each preset state information in the priority information corresponds to a priority. For example, based on the above embodiment, the priority of running in the foreground is higher than the priority of playing, and the priority of playing is higher than the priority of having recently played records. In other embodiments, the preset priority information includes two preset state information: playing through headphones and playing through amplifiers, wherein the priority of playing through headphones is higher than the priority of playing through amplifiers; or, the two preset state information are having recently played records through headphones and having recently played records through amplifiers, wherein the priority of having recently played records through headphones is higher than the priority of having recently played records through amplifiers.

[0099] It should be noted that this embodiment only uses the example of three or two preset state information in the preset priority information to illustrate the number of preset state information in the preset priority information, and does not constitute a limitation on the number of preset state information in the preset priority information. This embodiment only uses "playing," "running in the foreground," and "with recently played records" as examples to illustrate the preset state information, and does not constitute a limitation on the preset state information. This embodiment only illustrates that "running in the foreground" has a higher priority than "playing," and "playing" has a higher priority than "with recently played records." In practical applications, the priority order of "running in the foreground," "playing," and "with recently played records" can be flexibly set.

[0100] In this embodiment, candidate applications are determined from the identified applications based on keywords in the target voice command and a preset keyword correspondence. The preset keyword correspondence includes preset keywords and their corresponding preset applications. If at least two candidate applications are determined from at least two applications, the target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and preset priority information. The preset priority information includes at least two preset status information and their respective priorities. If one candidate application is determined from at least two applications, the candidate application is determined as the target application. Keywords can more accurately reflect the user's true intention. Thus, the target application determined based on keywords is more in line with the user's expectations, thereby improving user satisfaction.

[0101] Figure 4 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 4 for Figure 2 Based on the illustrated embodiment, another possible implementation of S103 is described in detail below:

[0102] S201', Based on the confidence levels of at least two user intents and a preset confidence level, determine candidate applications from the identified applications.

[0103] The confidence level of the user intent corresponding to the candidate application is greater than the preset confidence level.

[0104] For example, a target voice command corresponds to at least two target application indications. Each indication includes a confidence level for a user intent. The confidence level of the user intent can be determined based on factors such as the popularity, conversion rate, and user habits of the target words in the target voice command. The confidence level of the user intent reflects the likelihood of the corresponding user intent. For instance, if the target voice command "Play XXX" matches the user intents "Play audio" and "Play video," and the confidence level for "Play audio" is 0.58 while the confidence level for "Play video" is 0.42, then the user's actual intent is more likely to be "Play audio." Based on this, a preset confidence level is compared with the confidence levels of at least two user intents, resulting in one or more user intents with a confidence level greater than the preset confidence level. These one or more user intents with a confidence level greater than the preset confidence level are identified as candidate user intents. For example, with a preset reliability of 0.5, based on the above embodiments, the confidence level for the user intent "play audio" is greater than the preset reliability, while the confidence level for the user intent "play video" is less than the preset reliability, thus the candidate user intent can be determined to be "play audio". As another example, with a preset reliability of 0.4, based on the above embodiments, the confidence level for the user intent "play audio" is greater than the preset reliability, and the confidence level for the user intent "play video" is also greater than the preset reliability, thus the candidate user intents can be determined to be both "play audio" and "play video".

[0105] For each candidate user intent, there can be one or at least two candidate applications. For example, the candidate user intent "play audio" corresponds to candidate applications A and B; the candidate user intent "play video" corresponds to candidate application C. Thus, for all candidate user intents, there may be one candidate application or at least two candidate applications.

[0106] S202, determine whether the candidate application is a single application.

[0107] If yes, execute S203; otherwise, execute S203'.

[0108] S203, determine the candidate application as the target application.

[0109] S203', Based on the status information of at least two candidate applications and preset priority information, determine the target application from at least two candidate applications.

[0110] In this embodiment, candidate applications are determined from the identified applications based on the confidence levels of at least two user intentions and a preset confidence level. The confidence level of the user intentions corresponding to the candidate applications is greater than the preset confidence level. If at least two candidate applications are determined from at least two applications, a target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and preset priority information. The preset priority information includes at least two preset status information and their respective priorities. If only one candidate application is determined from at least two applications, the candidate application is determined as the target application. In this way, the confidence level of the user intentions corresponding to the target application is relatively high, and the target application is more in line with the user's expectations, thereby improving user satisfaction.

[0111] Figure 5 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 5 for Figure 2 Based on the illustrated embodiment, another possible implementation of S103 is described in detail below:

[0112] S2011, execute a preset instruction to determine the maximum confidence level from the confidence levels of at least two user intentions.

[0113] For example, a target voice command corresponds to at least two target application indications. Each target application indication includes a confidence level of a user intent. The confidence level of the user intent can be determined based on factors such as the popularity, conversion rate, and user habits of the target words in the target voice command. The confidence level of the user intent reflects the likelihood of the corresponding user intent. For example, if the target voice command "Play XXX" matches the user intents "Play audio" and "Play video," and the confidence level of the user intent "Play audio" is 0.58 while the confidence level of the user intent "Play video" is 0.42, then it is more likely that the user's actual intent is "Play audio." Based on this, a preset instruction for determining the maximum confidence level is executed, and the maximum confidence level is determined among the confidence levels of the at least two user intents. For example, based on the above embodiment, the maximum confidence level among the at least two user intents is determined to be 0.58.

[0114] S2012, determine the application in the application that corresponds to the user intent with the maximum confidence level as a candidate application.

[0115] Among the identified applications, the application corresponding to the user intent with the highest confidence level is designated as a candidate application. There may be one or at least two candidate applications. For example, based on the above embodiment, the user intent with a maximum confidence level of 0.58 is the user intent "play audio". If there is only one application corresponding to the user intent "play audio", then one candidate application can be identified. If there are at least two applications corresponding to the user intent "play audio", then at least two candidate applications can be identified.

[0116] S202, determine whether the candidate application is a single application.

[0117] If yes, execute S203; otherwise, execute S203'.

[0118] S203, determine the candidate application as the target application.

[0119] S203', Based on the status information of at least two candidate applications and preset priority information, determine the target application from at least two candidate applications.

[0120] In this embodiment, by executing a preset instruction, the maximum confidence level is determined from the confidence levels of at least two user intentions; at least two applications corresponding to the user intention with the maximum confidence level are identified as candidate applications; if at least two candidate applications are identified from at least two applications, a target application is identified from at least two candidate applications based on the status information of at least two candidate applications and preset priority information, wherein the preset priority information includes at least two preset status information and their respective priorities; if only one candidate application is identified from at least two applications, the candidate application is identified as the target application. In this way, the user intention with the maximum confidence level is closest to the user's true intention, and the resulting target application is more in line with the user's expectations, thereby improving user satisfaction.

[0121] Figure 6 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 6 for Figure 2 Based on the illustrated embodiment, another possible implementation of S103 is described in detail below:

[0122] S103' If the determined application includes the default application, the default application is determined as the target application.

[0123] For example, a target voice command corresponds to a target application instruction, and this target application instruction is the default application. If the default application is included in at least two identified applications, the default application can be identified as the target application.

[0124] In this embodiment, by determining that the default application is the target application among at least two identified applications, the target application can be obtained quickly, thereby improving the processing efficiency of voice commands.

[0125] Figure 7 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 7 for Figure 3 , Figure 4 and Figure 5 A specific description of a possible implementation of S203' based on any of the embodiments shown is as follows:

[0126] S301, based on the status information of at least two candidate applications and the preset priority information, determine the status information with the highest priority among the status information of at least two candidate applications.

[0127] For example, at least two preset state information from the preset priority information are sorted in descending order of priority to obtain a state information sequence. For instance, the preset priority information includes "Running in the foreground," "Playing," and "Recently Played Recording," where "Playing" has a higher priority than "Running in the foreground," and "Running in the foreground" has a higher priority than "Recently Played Recording." The resulting state information sequence in descending order of priority is: "Playing" - "Running in the foreground" - "Recently Played Recording." The first preset state information in the state information sequence is selected. Based on this first preset state information, the state information of at least two applications is traversed to determine if the first preset state information can be found in the state information of at least two applications. If the first preset state information can be found in the state information of at least two applications, then the first preset state information found is the highest priority state information among the at least two applications. If the first preset state information cannot be found in the state information of at least two applications, the second preset state information in the state information sequence is selected. Based on this second preset state information, the state information of at least two applications is traversed to determine if the second preset state information can be found in the state information of at least two applications. Therefore, if no preset status information is found, the preset status information in the status information sequence is continuously reselected until the last preset status information is selected, and the last preset status information is found in the status information of at least two applications, the highest priority status information in the status information of at least two applications can be determined.

[0128] For example, based on the above embodiments, based on the status information of at least two applications during playback, it is queried whether the status information of at least two applications includes "playing". If the status information of at least two applications includes "playing", then the highest priority status information among the at least two applications is "playing". If the status information of at least two applications does not include "playing", based on the status information of at least two applications running in the foreground, it is queried whether the status information of at least two applications includes "running in the foreground". If the status information of at least two applications includes "running in the foreground", then the highest priority status information among the at least two applications is "running in the foreground". If the status information of at least two applications does not include "running in the foreground", based on the status information of at least two applications having a recently played record, it is queried whether the status information of at least two applications includes a recently played record. If the status information of at least two applications includes a recently played record, then the highest priority status information among the at least two applications is "has a recently played record".

[0129] S302, determine the candidate application corresponding to the highest priority status information as the target application.

[0130] Among at least two candidate applications, the candidate application corresponding to the highest priority status information is determined as the target application. For example, based on the above embodiment, if the highest priority status information is "playing", then the candidate application corresponding to "playing" among at least two candidate applications is the target application; if the highest priority status information is "running in the foreground", then the candidate application corresponding to "running in the foreground" among at least two candidate applications is the target application; if the highest priority status information is "has a recent playback record", then the candidate application corresponding to "has a recent playback record" among at least two candidate applications is the target application.

[0131] Figure 8 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 8 for Figure 7 Based on the illustrated embodiment, a specific description of a possible implementation of S301 is as follows:

[0132] S401, Select the target state information with the highest priority from the at least two preset state information.

[0133] For example, among at least two preset state information in the preset priority information, the preset state information with the highest priority is selected as the target state information. For instance, based on the above embodiment, "playing" is selected as the target state information.

[0134] S402, based on the target state information, traverse the state information of at least two candidate applications to determine whether the target state information is included in the state information of at least two candidate applications.

[0135] If not, execute S403; if yes, execute S404.

[0136] Based on the target state information, the state information of at least two candidate applications is traversed, and it is queried whether the target state information is included in the state information of at least two candidate applications. For example, based on the above embodiment, based on the state information of at least two candidate applications during playback, it is queried whether the state information of at least two candidate applications includes "playing".

[0137] S403, Update the target status information according to the preset priority information in descending order of priority. Return to execute S402.

[0138] If the target state information is not found in the state information of at least two candidate applications, new target state information is selected from the preset priority information in descending order of priority, and the new target state information is used to replace the old one, thus completing the target state information update. Then, execution returns to step S402 until the result of S402 is "yes".

[0139] S404, determine the target state information as the highest priority state information.

[0140] If the target state information is found to be included in the state information of at least two candidate applications, then the target state information can be determined as the highest priority state information.

[0141] Figure 9 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 9 for Figure 3 , Figure 4 and Figure 5 A detailed description of another possible implementation of S203' based on any of the embodiments shown is as follows:

[0142] S301', Based on the status information of at least two candidate applications and the preset priority information, determine that the status information of at least two candidate applications does not include any of the preset status information in the preset priority information.

[0143] For example, at least two preset state information from the preset priority information are sorted in descending order of priority to obtain a state information sequence. For instance, the preset priority information includes "Running in the foreground," "Playing," and "Recently Played Recording," where "Playing" has a higher priority than "Running in the foreground," and "Running in the foreground" has a higher priority than "Recently Played Recording." The resulting state information sequence in descending order of priority is: "Playing" - "Running in the foreground" - "Recently Played Recording." The first preset state information in the state information sequence is selected. Based on this first preset state information, the state information of at least two applications is traversed to determine if the first preset state information can be found in the state information of at least two applications. If the first preset state information is not found in the state information of at least two applications, the second preset state information in the state information sequence is selected. Based on this second preset state information, the state information of at least two applications is traversed to determine if the second preset state information can be found in the state information of at least two applications. Therefore, if no preset status information is found, the preset status information in the status information sequence is continuously reselected until the last preset status information is selected, and the last preset status information is not found in the status information of at least two applications. At this time, it can be determined that the status information of at least two candidate applications does not include any preset status information in the preset priority information.

[0144] For example, based on the above embodiments, based on the status information of at least two applications during playback, it is queried whether the status information of at least two applications includes "playing". If the status information of at least two applications does not include "playing", the status information of at least two applications running in the foreground is queried again to see if the status information of at least two applications includes "running in the foreground". If the status information of at least two applications does not include "running in the foreground", the status information of at least two applications is queried again to see if the status information of at least two applications includes "recently played records". If the status information of at least two applications does not include "recently played records", it is determined that the status information of at least two candidate applications does not include any of the preset status information in the preset priority information.

[0145] S302', determine the preset default application of at least two of the candidate applications as the target application.

[0146] In all candidate applications for the target voice command, a default application is preset. If the status information of at least two candidate applications does not include any of the preset status information in the preset priority information, then the default application among the candidate applications can be determined as the target application.

[0147] Figure 10 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 10 for Figure 9 Based on the illustrated embodiment, a specific description of a possible implementation of S301' is as follows:

[0148] S401, Select the target state information with the highest priority from the at least two preset state information.

[0149] S402, based on the target state information, traverse the state information of at least two candidate applications to determine whether the target state information is included in the state information of at least two candidate applications.

[0150] If not, execute S501.

[0151] S501, determine whether the target state information is the target state information with the lowest priority among the state information of at least two candidate applications.

[0152] If not, execute S403; if yes, execute S404.

[0153] For example, if the target state information is not included in the state information of at least two candidate applications, determine whether the current target state information is the target state information with the lowest priority among the state information of at least two candidate applications. If yes, it means that the current target state information is the last target state information that can be selected among the state information of at least two candidate applications, and then stop selecting new target state information; if no, it means that there is still target state information that can be selected among the state information of at least two candidate applications, and then S403 can be executed to select new target state information.

[0154] S403, Update the target status information according to the preset priority information in descending order of priority. Then return to execute S402.

[0155] S404', determine that the target state information with the lowest priority is not included in the state information of at least two of the candidate applications.

[0156] If the current target state information is the lowest priority target state information among the state information of at least two candidate applications, that is, the current target state information is the last selectable target state information among the state information of at least two candidate applications, then it can be determined that the lowest priority target state information is not included among the state information of at least two candidate applications.

[0157] Figure 11 This is a flowchart illustrating yet another voice command processing method provided in this disclosure. Figure 11 for Figure 2 Based on the illustrated embodiment, before executing S101, the following steps are also included:

[0158] S601, determine whether the target voice command includes terminal device identification information.

[0159] If yes, execute S602; otherwise, execute S602'.

[0160] For example, the application scenario includes at least two terminal devices. Some voice commands include information about the user-specified terminal device, i.e., the identification information of the specified terminal device. For example, the voice command "Play XXX in the passenger seat" uses "passenger seat" as the identification information of the user-specified terminal device. However, some voice commands do not include the identification information of the specified terminal device, such as the voice command "Play XXX". Therefore, it is necessary to determine whether the target voice command triggered by the user includes the terminal device identification information of the user-specified terminal device.

[0161] S602, determine that the terminal device corresponding to the terminal device identification information is the target terminal device that executes the target voice command.

[0162] For example, if the target voice command includes terminal device identification information, the terminal device corresponding to the terminal device identification information can be determined as the target terminal device executing the target voice command. For instance, if the target voice command "Play XXX in the co-pilot's seat" includes the terminal device identification information "co-pilot's seat", then the terminal device corresponding to "co-pilot's seat" is determined as the target terminal device executing this target voice command.

[0163] S602', Determine that the terminal device where the voice virtual avatar is located is the target terminal device for executing the target voice command.

[0164] For example, if the target voice command does not include terminal device identification information, the terminal device where the virtual voice avatar is located can be obtained, and the terminal device where the virtual voice avatar is located can be determined as the target terminal device for executing the target voice command. For instance, if the target voice command "Play XXX" does not include terminal device identification information, and the terminal device where the virtual voice avatar is located is a terminal device near the passenger seat, then the terminal device near the passenger seat can be determined as the target terminal device for executing the target voice command.

[0165] In some embodiments, the target voice command includes application identification information of a specified application. For example, in the target voice command "Application A plays XXX", "Application A" is the application identification information of the specified application. If the specified application is included in the system applications of the target terminal device executing the target voice command, then the specified application is determined to be the target application for executing the target voice command. If the specified application is not included in the system applications of the target terminal device executing the target voice command, it is determined whether the specified application is included in the application store of the target terminal device executing the target voice command. If the specified application is not included in the application store of the target terminal device executing the target voice command, a prompt message is returned to the target terminal device executing the target voice command. This prompt message is used to inform the user that the specified application has not been found. For example, based on the above embodiments, the prompt message is "Application A not found". If the specified application is included in the application store of the target terminal device executing the target voice command, it is determined whether the specified application has been downloaded to the target terminal device executing the target voice command. If the specified application has been downloaded to the target terminal device executing the target voice command, then the specified application is determined to be the target application for executing the target voice command. If the specified application has not been downloaded to the target terminal device executing the target voice command, it is determined whether the specified application is being installed on the target terminal device executing the target voice command. If the specified application is being installed on the target terminal device executing the target voice command, a prompt message is returned to the target terminal device executing the target voice command. This prompt message is used to inform the user that the specified application is being downloaded. For example, based on the above embodiment, the prompt message is "Application A is being downloaded, please wait." If the specified application is not being installed on the target terminal device executing the target voice command, the application information card of the specified application is returned to the target terminal device executing the target voice command. Based on the target terminal device executing the target voice command, multiple rounds of questions are initiated to the user whether to download the specified application. For example, based on the above embodiment, the prompt message of the application information card could be "You have not installed application A yet, do you want me to install it for you?"

[0166] In some embodiments, the target voice command does not include the application identifier information of the specified application. For example, the target voice command is "Play XXX". If the target terminal device executing the target voice command does not have an application matching the target voice command, it is determined whether the application store of the target terminal device executing the target voice command contains an application matching the target voice command that is being downloaded. If the application store of the target terminal device executing the target voice command contains an application matching the target voice command that is being downloaded, a prompt message is returned to the target terminal device executing the target voice command. This prompt message is used to inform the user that the specified application is being downloaded. For example, if the application matching the target voice command that is being downloaded is B, the prompt message is "Application B is downloading, please wait". If the application store of the target terminal device executing the target voice command does not have an application matching the target voice command that is being downloaded, an application information card of the specified application is returned to the target terminal device executing the target voice command. Based on the target terminal device executing the target voice command, multiple rounds of questions are initiated to ask the user whether to download the specified application. For example, based on the above embodiments, the prompt message of the application information card could be "You have not installed application A yet, do you want me to install it for you?"

[0167] In some embodiments, when the screen of the target terminal device specified by the user is off, voice control is not supported, and a prompt message indicating that the screen is off is returned to the target terminal device. For example, if the target terminal device is terminal device C, the prompt message would be "The screen of terminal device C is not on, so you cannot control it like this." In some embodiments, when the screen of the target terminal device specified by the user is in sleep mode, video applications and karaoke applications do not support voice control, and a prompt message indicating that voice control is not supported is returned to the target terminal device. For example, if the target terminal device is terminal device C, the prompt message would be "In the current state, this control is not supported," while applications other than video applications and karaoke applications support voice control. In some embodiments, if the screen of the target terminal device specified by the user is in screen-off mode, in response to the target voice command, the target terminal device specified by the user exits the screen-off mode and then executes the target voice command.

[0168] It should be noted that the above embodiments are all illustrated using on-demand voice commands as examples. In other embodiments, the target voice command can be an interface-closing voice command, such as "Close XXX". The preset priority information corresponding to the target voice command includes: playing, running in the foreground, and default application. Among them, the priority of playing is higher than the priority of running in the foreground, and the priority of running in the foreground is higher than the priority of the default application. Thus, if there is an application playing on the target terminal device, the target application for executing the target voice command is determined to be the application playing. If there is no application playing on the target terminal device but there is an application running in the foreground, the target application for executing the target voice command is determined to be the application running in the foreground. If there are no applications playing or running in the foreground on the target terminal device, the target application for executing the target voice command is determined to be the default application. In other embodiments, the target voice command can be an interface-opening voice command, such as "Open XXX". The preset priority information corresponding to the target voice command includes: playing, running in the foreground, having recently played content, and default application. The priority of "running in the foreground" is higher than that of "playing", which is higher than "having recently played content", which is higher than "default application". Thus, if an application is running in the foreground on the target terminal device, the target application for executing the target voice command is determined to be the application running in the foreground. If no application is running in the foreground but an application is playing, the target application for executing the target voice command is determined to be the application playing. If neither "playing" nor "running" applications exist on the target terminal device, but an application has recently played content, the target application for executing the target voice command is determined to be the application with recently played content. If neither "playing", "running" nor "having recently played content" applications exist on the target terminal device, the target application for executing the target voice command is determined to be the default application.

[0169] This disclosure also provides a voice command processing apparatus. Figure 12 This is a schematic diagram of the structure of a voice command processing device provided in this disclosure, as shown below. Figure 12 As shown, the voice command processing device includes:

[0170] The parsing module 210 is used to parse the content of the target voice command to obtain the target application instruction information and at least two user intents.

[0171] The determining module 220 is configured to determine an application matching each of the at least two user intentions and a preset correspondence, wherein the preset correspondence includes a preset user intention and a corresponding application; and to determine a target application from the determined applications based on the target application indication information.

[0172] Optionally, the target application indication information includes keywords from the target voice command.

[0173] The determining module 220 is further configured to determine candidate applications from the determined applications based on the keywords in the target voice command and a preset keyword correspondence, wherein the preset keyword correspondence includes preset keywords and their corresponding preset applications; if there are at least two candidate applications, the target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and preset priority information, wherein the preset priority information includes at least two preset status information and their respective priorities; if there is only one candidate application, the candidate application is determined to be the target application.

[0174] Optionally, the target application indication information includes the confidence level of the user's intent.

[0175] The determining module 220 is further configured to determine candidate applications from the determined applications based on the confidence levels of at least two user intents and a preset confidence level, wherein the confidence level of the user intent corresponding to the candidate application is greater than the preset confidence level; if there are at least two candidate applications, the target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and preset priority information, wherein the preset priority information includes at least two preset status information and their respective priorities; if there is only one candidate application, the candidate application is determined to be the target application.

[0176] Optionally, the target application indication information includes the user.

[0177] The determining module 220 is further configured to execute preset instructions to determine the maximum confidence level from the confidence levels of at least two user intentions; determine the application corresponding to the user intention with the maximum confidence level as a candidate application; if there are at least two candidate applications, determine the target application from the at least two candidate applications based on the status information of the at least two candidate applications and preset priority information, wherein the preset priority information includes at least two preset status information and their respective priorities; if there is only one candidate application, determine the candidate application as the target application.

[0178] Optionally, the target application indication information includes the default application.

[0179] The determining module 220 is further configured to determine the default application as the target application if the determined application includes the default application.

[0180] Optionally, the determining module 220 is further configured to determine the highest priority status information among the status information of at least two candidate applications based on the status information of at least two candidate applications and the preset priority information;

[0181] The candidate application corresponding to the highest priority status information is determined as the target application.

[0182] Optionally, the determining module 220 is further configured to select the target state information with the highest priority from at least two preset state information; based on the target state information, traverse the state information of at least two candidate applications; if it is determined that the target state information is not included in the state information of at least two candidate applications, update the target state information in descending order of priority in the preset priority information, and return to execute the step of traversing the state information of at least two candidate applications based on the target state information until it is determined that the target state information is included in the state information of at least two candidate applications, and determine the target state information as the highest priority state information.

[0183] Optionally, the determining module 220 is further configured to determine, based on the status information of at least two candidate applications and the preset priority information, that the status information of at least two candidate applications does not include any of the preset status information in the preset priority information; and to determine the preset default application among the at least two candidate applications as the target application.

[0184] Optionally, the determining module 220 is further configured to select the target state information with the highest priority from the at least two preset state information; based on the target state information, traverse the state information of at least two candidate applications; if it is determined that the target state information is not included in the state information of at least two candidate applications, update the target state information in descending order of priority in the preset priority information, and return to execute the step of traversing the state information of at least two candidate applications based on the target state information until it is determined that the target state information with the lowest priority is not included in the state information of at least two candidate applications.

[0185] Optionally, the determining module 220 is further configured to, if the target voice command includes terminal device identification information, determine that the terminal device corresponding to the terminal device identification information is the target terminal device executing the target voice command; if the target voice command does not include terminal device identification information, determine that the terminal device where the voice virtual image is located is the target terminal device executing the target voice command.

[0186] The apparatus provided in this embodiment can execute the methods provided in any embodiment of this disclosure, and has the corresponding functional modules and beneficial effects for executing the methods, which will not be elaborated here.

[0187] This disclosure also provides an electronic device, including: a processor, the processor being configured to execute a computer program stored in a memory, the computer program being executed by the processor to implement the steps of the above method embodiments.

[0188] Figure 13 This is a schematic diagram of the structure of an electronic device provided in this disclosure. Figure 13 A block diagram is shown that is suitable for implementing embodiments of the present invention. Figure 13 The electronic device shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.

[0189] like Figure 13 As shown, the electronic device 12 is represented in the form of a general-purpose computing device. The components of the electronic device 12 may include, but are not limited to: one or at least two processors 16, system memory 28, and a bus 18 connecting different system components (including system memory 28 and processor 16).

[0190] Bus 18 represents one or more of several bus architectures, including a memory bus or memory controller, a peripheral bus, a graphics acceleration port, a processor, or a local bus using any of the various bus architectures. For example, these architectures include, but are not limited to, the Industry Standard Architecture (ISA) bus, the Micro Channel Architecture (MAC) bus, the Enhanced ISA bus, the Video Electronics Standards Association (VESA) local bus, and the Peripheral Component Interconnect (PCI) bus.

[0191] Electronic device 12 typically includes a variety of computer system readable media. These media can be any media that can be accessed by electronic device 12, including volatile and non-volatile media, removable and non-removable media.

[0192] System memory 28 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 30 and / or cache memory 32. Electronic device 12 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, storage system 34 may be used to read and write non-removable, non-volatile magnetic media (commonly referred to as "hard disk drives"). Disk drives for reading and writing to removable non-volatile disks (e.g., "floppy disks") and optical disk drives for reading and writing to removable non-volatile optical disks (e.g., CD-ROMs, DVD-ROMs, or other optical media) may be provided. In these cases, each drive may be connected to bus 18 via one or at least two data media interfaces. System memory 28 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of various embodiments of the present invention.

[0193] A program / utility 40 having a set (at least one) of program modules 42 may be stored, for example, in system memory 28. Such program modules 42 include, but are not limited to, an operating system, one or at least two application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Program modules 42 typically perform the functions and / or methods described in the embodiments of this invention.

[0194] The processor 16 performs various functional applications and information processing by running at least one of at least two programs stored in the system memory 28, such as implementing the method embodiments provided in the embodiments of the present invention.

[0195] This disclosure also provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the above-described method embodiments.

[0196] Any combination of one or at least two computer-readable media may be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or at least two wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this document, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in connection with an instruction execution system, apparatus, or device.

[0197] Computer-readable signal media may include data signals propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media may also be any computer-readable medium other than computer-readable storage media, capable of sending, propagating, or transmitting programs for use by or in connection with an instruction execution system, apparatus, or device.

[0198] Program code contained on a computer-readable medium may be transmitted using any suitable medium, including but not limited to wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0199] Computer program code for performing the operations of this invention can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, as well as conventional procedural programming languages ​​such as "C" or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or wide area network (WAN) domain—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0200] This disclosure also provides a vehicle, including: a voice command processing device, electronic device, or computer-readable storage medium provided in any of the above embodiments.

[0201] This disclosure also provides a computer program product that, when run on a computer, causes the computer to perform the steps of the above-described method embodiments.

[0202] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0203] The above description is merely a specific embodiment of this disclosure, enabling those skilled in the art to understand or implement it. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this disclosure. Therefore, this disclosure is not to be limited to the embodiments described herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A method for processing voice commands, characterized in that, include: Parse the content of the target voice command to obtain the target application instruction information and at least two user intents; Based on the at least two user intents and the preset correspondence, the application matching each user intent is determined, wherein the preset correspondence includes the preset user intent and the application matching it; Based on the target application indication information, candidate applications are determined from the identified applications; If there are at least two candidate applications, a target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and the preset priority information. The preset priority information includes at least two preset status information and their respective priorities. The step of determining a target application from at least two candidate applications based on the status information of at least two candidate applications and preset priority information includes: Select the target state information with the highest priority from the at least two preset state information; Based on the target state information, traverse the state information of at least two candidate applications; If it is determined that the target state information is not included in the state information of at least two candidate applications, the target state information is updated in descending order of priority in the preset priority information, and the process of traversing the state information of at least two candidate applications based on the target state information is repeated until it is determined that the target state information is included in the state information of at least two candidate applications, and the target state information is determined to be the highest priority state information. The candidate application corresponding to the highest priority status information is determined as the target application.

2. The method according to claim 1, characterized in that, The target application indication information includes keywords in the target voice command; The step of determining candidate applications from the identified applications based on the target application indication information includes: Based on the keywords in the target voice command and the preset keyword correspondence, candidate applications are determined from the identified applications. The preset keyword correspondence includes preset keywords and their corresponding preset applications. The method further includes: If there is only one candidate application, then the candidate application is determined to be the target application.

3. The method according to claim 1, characterized in that, The target application indication information includes the confidence level of the user's intent; The step of determining candidate applications from the identified applications based on the target application indication information includes: Candidate applications are determined from the identified applications based on the confidence levels of at least two user intents and a preset confidence level, wherein the confidence level of the user intent corresponding to the candidate application is greater than the preset confidence level. The method further includes: If there is only one candidate application, then the candidate application is determined to be the target application.

4. The method according to claim 1, characterized in that, The target application indication information includes the confidence level of the user's intent; The step of determining candidate applications from the identified applications based on the target application indication information includes: Execute a preset instruction and determine the maximum confidence level from at least two confidence levels of the user intent; The application corresponding to the user intent with the maximum confidence level is identified as a candidate application. The method further includes: If there is only one candidate application, then the candidate application is determined to be the target application.

5. The method according to claim 1, characterized in that, The target application indication information is used to indicate the default application; The method further includes: If the identified application includes the default application indicated by the target application indication information, then the default application is determined to be the target application.

6. The method according to any one of claims 1-4, characterized in that, The step of determining a target application from at least two candidate applications based on the status information of at least two candidate applications and preset priority information includes: Based on the status information of at least two candidate applications and the preset priority information, it is determined that the status information of at least two candidate applications does not include any of the preset status information in the preset priority information; The target application is determined by selecting a preset default application from at least two of the candidate applications.

7. The method according to claim 6, characterized in that, The step of determining, based on the status information of at least two candidate applications and the preset priority information, that the status information of at least two candidate applications does not include any of the preset status information in the preset priority information includes: Select the target state information with the highest priority from the at least two preset state information; Based on the target state information, traverse the state information of at least two candidate applications; If it is determined that the target state information is not included in the state information of at least two candidate applications, the target state information is updated in descending order of priority in the preset priority information, and the process of traversing the state information of at least two candidate applications based on the target state information is resumed until it is determined that the target state information with the lowest priority is not included in the state information of at least two candidate applications.

8. The method according to any one of claims 1-5, characterized in that, Before parsing the content of the target voice command to obtain the target application instruction information and at least two user intents, the method further includes: If the target voice command includes terminal device identification information, the terminal device corresponding to the terminal device identification information is determined to be the target terminal device executing the target voice command. If the target voice command does not include terminal device identification information, the terminal device where the virtual voice avatar is located is determined to be the target terminal device executing the target voice command.

9. A voice command processing device, characterized in that, include: The parsing module is used to parse the content of the target voice command to obtain the target application instruction information and at least two user intents. The determination module determines the application that matches each user intent based on the at least two user intents and a preset correspondence, wherein the preset correspondence includes a preset user intent and an application that matches it. Based on the target application indication information, candidate applications are determined from the identified applications; If there are at least two candidate applications, a target application is determined from the at least two candidate applications based on the status information of the at least two candidate applications and the preset priority information. The preset priority information includes at least two preset status information and their respective priorities. The step of determining a target application from at least two candidate applications based on the status information of at least two candidate applications and preset priority information includes: Select the target state information with the highest priority from the at least two preset state information; Based on the target state information, traverse the state information of at least two candidate applications; If it is determined that the target state information is not included in the state information of at least two candidate applications, the target state information is updated in descending order of priority in the preset priority information, and the process of traversing the state information of at least two candidate applications based on the target state information is repeated until it is determined that the target state information is included in the state information of at least two candidate applications, and the target state information is determined to be the highest priority state information. The candidate application corresponding to the highest priority status information is determined as the target application.

10. An electronic device, characterized in that, include: A processor for executing a computer program stored in a memory, wherein the computer program, when executed by the processor, implements the steps of the method according to any one of claims 1-8.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1-8.

12. A vehicle, characterized in that, include: The voice command processing apparatus as described in claim 9; Or, the electronic device as described in claim 10; Alternatively, the computer-readable storage medium as described in claim 11.

Citation Information

Patent Citations

  • Method and device for displaying application, computer device and storage medium

    CN107783705A

  • Data processing method and device, equipment and medium

    CN114464176A