Voice instruction execution method, system and device, vehicle end control equipment and storage medium

By identifying keyword groups in voice commands to create subtasks and execute them, the problem that existing systems cannot execute complex multi-step instructions is solved, and efficient multi-step task execution and improved user experience is achieved.

CN119993138AActive Publication Date: 2025-05-13CHONGQING SELIS PHOENIX INTELLIGENT INNOVATION TECH CO LTD

Patent Information

Application Number
CN202411219462.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-09-02
Publication Date
2025-05-13
Estimated Expiration
2044-09-02

AI Technical Summary

Technical Problem

The existing car voice assistant system cannot execute complex multi-step instructions and can only process single-step instructions, resulting in users needing to split multiple demand instructions, reducing execution efficiency and user experience.

Method used

By identifying the keyword groups in a single voice command, creating the corresponding subtask and joining the task queue to be executed, triggering the executor to execute the subtask to obtain the execution data, and saving the data to the completed task queue, and finally output the task execution result by the associated application.

Benefits of technology

The execution of multi-step instructions is realized, reducing the number of repeated instructions by users, and improving the intelligence level and user experience of the voice recognition system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993138A_ABST
    Figure CN119993138A_ABST
Patent Text Reader

Abstract

The invention relates to a voice instruction execution method, system and device, vehicle end control equipment and a storage medium, and relates to the technical field of voice recognition. The method comprises the following steps: identifying a single voice instruction to obtain at least two key phrases, creating a plurality of corresponding sub-tasks based on the key phrases, and adding the sub-tasks into the same to-be-executed task queue; when it is recognized that the to-be-executed task queue comprises the sub-tasks, triggering a preset executor to execute the sub-tasks so as to obtain execution data corresponding to the sub-tasks, and storing the execution data of the sub-tasks to a completed task queue; triggering the associated application to output a task execution result based on the execution data when it is identified that the completed task queue comprises the execution data of the at least two sub-tasks; the number of times of input and recognition of the voice instruction of the user is reduced, the intelligent level of the vehicle end control equipment carrying the voice recognition system is improved, and therefore the use experience of the user on the voice recognition function in the vehicle is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of speech recognition technology, and in particular to a method, system, device, vehicle-side control equipment and storage medium for executing a voice command. Background Art

[0002] With the development of voice recognition technology, more and more smart devices are equipped with voice assistant systems. For example, many current vehicles are already equipped with voice assistant systems. Existing in-vehicle voice assistant systems usually rely on pre-programmed instruction sets and simple semantic understanding capabilities when processing user commands. These systems are often unable to execute complex multi-step commands and can only execute single-step commands. This requires users to split multiple required commands into multiple simple commands for execution, resulting in low execution efficiency of multi-step commands and reduced user experience. Summary of the invention

[0003] Based on this, it is necessary to provide a method, system, device, vehicle-side control equipment, computer-readable storage medium and computer program product for executing voice commands that can implement multi-step commands in response to the above-mentioned technical problems.

[0004] In a first aspect, the present application provides a method for executing a voice command, comprising:

[0005] Recognize a single voice command to obtain at least two keyword groups, create at least two corresponding subtasks based on the keyword groups and add them to the same queue of tasks to be executed;

[0006] In the case of identifying that the queue of tasks to be executed includes the subtasks, triggering a preset executor to execute each of the subtasks, so as to respectively obtain execution data corresponding to each of the subtasks, and saving the execution data of each of the subtasks to the completed task queue;

[0007] In the case of identifying that the completed task queue includes the execution data of at least two of the subtasks, triggering an associated application to output a task execution result of the voice instruction based on the execution data of at least two of the subtasks.

[0008] In one embodiment, when it is identified that the queue of tasks to be executed includes the subtask, triggering a preset executor to execute each of the subtasks to obtain execution data corresponding to each of the subtasks, and saving the execution data of each of the subtasks to the completed task queue includes:

[0009] When it is identified that the queue of tasks to be executed includes the subtask, based on the order of the subtasks to be scheduled, extracting the subtask that is the first in the order to be scheduled;

[0010] Triggering a preset executor to execute the extracted subtask to obtain execution data corresponding to the subtask, and saving the execution data to a completed task queue;

[0011] The to-be-executed task queue is identified again until the execution data of each of the subtasks in the to-be-executed task queue is saved in the completed task queue.

[0012] In one embodiment, the step of saving the execution data to a completed task queue includes:

[0013] The execution data is saved in a task data model associated with the corresponding subtask, and the task data model is saved in a completed task queue; wherein the task data model is used to indicate an execution method of the execution data.

[0014] In one embodiment, when it is recognized that the completed task queue includes the execution data of at least two of the subtasks, triggering the associated application to output the task execution result of the voice instruction based on the execution data of at least two of the subtasks includes:

[0015] When it is identified that the completed task queue includes at least two of the task data models, the associated application is triggered to output the task execution result of the voice instruction based on the execution mode of the execution data respectively indicated by the at least two task data models.

[0016] In one embodiment, executing the extracted subtask to obtain execution data corresponding to the subtask includes:

[0017] Execute the extracted subtask to obtain a candidate data table output by the associated application pointed to by the subtask;

[0018] guiding the outputter of the voice instruction to select target execution data in the candidate data table through voice and / or display;

[0019] The selected target execution data is obtained as execution data corresponding to the subtask.

[0020] In one embodiment, it also includes:

[0021] When it is identified that the to-be-executed task queue does not include the subtask and when it is identified that the completed task queue does not include the execution data, prompt information is output to guide the outputter of the voice command to re-input the voice command.

[0022] In a second aspect, the present application also provides a voice command execution system, comprising:

[0023] A speech recognition system, used to recognize a single voice command to obtain at least two key words;

[0024] A task scheduler, used for creating at least two corresponding subtasks based on the keyword group and adding them to the same queue of tasks to be executed;

[0025] The task scheduler is further used to call the task executor to execute each of the subtasks when recognizing that the queue of tasks to be executed includes the subtasks, so as to respectively obtain the execution data corresponding to each of the subtasks, and the task scheduler is further used to save the execution data of each of the subtasks to the completed task queue;

[0026] The task scheduler is also used to call the task executor to trigger the associated application to output the task execution result of the voice instruction based on the execution data of at least two subtasks when recognizing that the completed task queue already includes the execution data of at least two subtasks.

[0027] In a third aspect, the present application further provides a device for executing a voice command, the device comprising:

[0028] A to-be-executed task queue creation module, used for identifying a single voice command to obtain at least two keyword groups, creating at least two corresponding subtasks based on the keyword groups and adding them to the same to-be-executed task queue;

[0029] A task execution module, for triggering a preset executor to execute each of the subtasks when identifying that the queue of tasks to be executed includes the subtasks, so as to respectively obtain execution data corresponding to each of the subtasks, and save the execution data of each of the subtasks to the completed task queue;

[0030] The result output module is used to trigger the associated application to output the task execution result of the voice instruction based on the execution data of at least two subtasks when recognizing that the completed task queue includes the execution data of at least two subtasks.

[0031] In a fourth aspect, the present application further provides a vehicle-side control device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any step in the first aspect when executing the computer program.

[0032] In a fifth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, wherein the computer program implements any step in the first aspect when executed by a processor.

[0033] In a sixth aspect, the present application further provides a computer program product, comprising a computer program, which implements any step in the first aspect when executed by a processor.

[0034] In the above-mentioned voice command execution method, system, device, vehicle-side control device, computer-readable storage medium and computer program product provided by the present application, in the voice command execution method, at least two keyword groups are obtained by identifying a single voice command, and at least two corresponding subtasks are created based on the keyword groups and added to the same queue of tasks to be executed; when it is identified that the queue of tasks to be executed includes subtasks, the preset executor is triggered to execute each subtask to obtain the execution data corresponding to each subtask respectively, and the execution data of each subtask is saved to the completed task queue; when it is identified that the completed task queue already includes the execution data of at least two subtasks, the associated application is triggered to output the task execution result of the voice command based on the execution data of at least two subtasks; it can be seen that this method can realize the recognition of a single voice command including multiple keyword groups and the display of related results, which is beneficial to reduce the number of recognition times of voice commands for users, and is beneficial to improve the intelligence level of related equipment equipped with a voice recognition system, thereby helping to improve the user experience of the voice recognition function. BRIEF DESCRIPTION OF THE DRAWINGS

[0035] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the drawings required for use in the embodiments of the present application or related technical descriptions will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other related drawings can be obtained based on these drawings without paying creative work.

[0036] Figure 1 is an application environment diagram of a method for executing a voice command in one embodiment;

[0037] Figure 2 is a flowchart of a method for executing a voice command in one embodiment;

[0038] Figure 3 is a structural block diagram of a system for executing voice commands in one embodiment;

[0039] Figure 4 is a flow chart of a method for executing a voice command in another embodiment;

[0040] Figure 5 is a flowchart of a method for executing a voice command in yet another embodiment;

[0041] Figure 6is a structural block diagram of a device for executing voice commands in one embodiment;

[0042] Figure 7 This is a diagram of the internal structure of a vehicle-side control device in one embodiment. DETAILED DESCRIPTION

[0043] In order to make the purpose, technical solution and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0044] The method for executing a voice command provided in the embodiment of the present application can be applied to Figure 1 In the application environment shown. Among them, the terminal 12 communicates with the server 14 through the network. The data storage system can store the data that the server 14 needs to process. The data storage system can be integrated on the server 14, or it can be placed on the cloud or other network servers. For example, the data storage system can store key words, task queues, etc. Among them, the terminal 12 can be, but is not limited to, various personal computers, laptops, smart phones, tablets, Internet of Things devices and portable wearable devices. The Internet of Things devices can be smart speakers, smart TVs, smart air conditioners, smart car devices (vehicle-side control devices), projection devices, etc. Portable wearable devices can be smart watches, smart bracelets, head-mounted devices, etc. The head-mounted devices can be virtual reality (VR) devices, augmented reality (AR) devices, smart glasses, etc. The server 14 can be an independent physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides cloud computing services.

[0045] In an exemplary embodiment, Figure 2 As shown, a method for executing a voice command is provided, and the method is applied to Figure 1 Taking the vehicle-side control device in the example as an example, the method includes the following steps 101 to 103. Among them:

[0046] Step 101: Recognize a single voice instruction to obtain at least two keyword groups, create at least two corresponding subtasks based on the keyword groups, and add them to the same task queue to be executed.

[0047] Among them, the voice command is the voice data spoken by the human user, or the voice data provided by the relevant computer equipment; the voice command can use the voice recognition system to understand the key words included in the voice command; that is, the voice recognition system is used to perform voice recognition and semantic understanding based on the user's voice data, and return the result of semantic understanding, that is, the above-mentioned key words. In addition, the voice recognition system also supports speech synthesis.

[0048] Among them, the voice recognition system can be implemented based on the Voice Assistant Terminal Program (Voice Assistant Terminal Program). The Voice Assistant Terminal Program refers to a software program running on terminal devices (such as smart phones, smart speakers, smart watches, car systems, etc.). It integrates voice recognition, natural language processing, speech synthesis and other technologies to allow users to interact with it using voice commands and perform various tasks and services.

[0049] Among them, the execution task queue is used to manage and maintain the tasks to be executed. A task queue is a data structure used to manage and maintain a list of tasks to be executed. In computer science, task queues are often used to implement multitasking and asynchronous operations. Task queues schedule tasks in a specific order (such as first-in-first-out or priority queues) to ensure that each task has a chance to be executed. Task queues are widely used in operating systems, concurrent programming, message queue systems, and other fields to improve the responsiveness and efficiency of the system.

[0050] When a single voice data input by the user is received, step 101 is executed to identify the various keyword groups included in the voice instruction input by the user; wherein, a single voice instruction may include one keyword group or multiple keyword groups. Then, based on the multiple keyword groups included in the identified voice instruction, subtasks corresponding to each keyword group are created respectively, and then the priorities of each subtask can be selected, and the multiple subtasks are added to the queue of tasks to be executed in order of priority. In this way, the single voice instruction of the user is converted into multiple subtasks to be executed, and added to the queue of tasks to be executed, so as to facilitate the subsequent execution of the tasks included in the queue of tasks to be executed, so as to realize the relevant functions required by the voice instruction input by the user.

[0051] Step 102, when it is identified that the queue of tasks to be executed includes subtasks, a preset executor is triggered to execute each subtask, so as to obtain execution data corresponding to each subtask respectively, and save the execution data of each subtask to the completed task queue.

[0052] Among them, the completed (executed) task queue is used to manage and maintain the tasks that have been executed.

[0053] The execution of step 102 is used to learn that there is a task that needs to be executed when it is identified that the queue of tasks to be executed includes at least one subtask; when it is identified that the queue of tasks to be executed includes multiple subtasks, it triggers the preset executor to execute each subtask in sequence based on the execution order requirements of each subtask, so as to obtain the corresponding execution data obtained by executing each subtask respectively; after each subtask is executed and the corresponding execution data is obtained, the execution data of the subtask can be saved to the completed task queue; finally, each execution data obtained after the corresponding execution of each subtask in the queue of tasks to be executed will also be saved to the completed task queue. Among them, the number of subtasks included in the queue of tasks to be executed and the number of execution data in the corresponding completed task queue can be the same.

[0054] Step 103 : when it is identified that the completed task queue includes the execution data of at least two subtasks, trigger the associated application to output the task execution result of the voice command based on the execution data of the at least two subtasks.

[0055] Among them, associated applications refer to third-party applications, and specifically can be third-party SDKs (Software Development Kits), which are used to provide certain specific capabilities. For example, map SDKs can provide location query, navigation and other capabilities.

[0056] Map SDK is a set of software development tools that provides developers with a series of interfaces, libraries and tools to integrate map display, positioning, search, route planning and other functions into applications. Map SDK is usually released by map service providers so that developers can easily use map services in their own applications.

[0057] In addition, third-party applications and speech recognition systems can choose to communicate through API (Application Programming Interface), which is a set of predefined rules and protocols for building and interacting with software applications. It allows different software applications to communicate and exchange data. The API defines how requests are made, how to send requests, the expected response format, etc. Through the API, developers can access the functionality of a software or service without having to understand its internal workings.

[0058] The execution of step 103 is used to, when recognizing that the completed task queue includes at least two corresponding execution data, learn that there are tasks in the completed task queue that need to be further processed. At this time, based on the various execution data included in the completed task queue, the associated application (third-party SDK) can be called to enable the associated application to output the task execution result corresponding to the voice command input by the user based on the various execution data; thereby realizing the implementation of relevant functions required for the voice command input by the user and the display of results.

[0059] In the above-mentioned method for executing voice commands, at least two keyword groups are obtained by identifying a single voice command, and at least two corresponding subtasks are created based on the keyword groups and added to the same queue of tasks to be executed; when it is identified that the queue of tasks to be executed includes subtasks, the preset executor is triggered to execute each subtask to obtain the execution data corresponding to each subtask respectively, and the execution data of each subtask is saved to the completed task queue; when it is identified that the completed task queue already includes the execution data of at least two subtasks, the associated application is triggered to output the task execution result of the voice command based on the execution data of at least two subtasks; it can be seen that this method can realize the recognition of a single voice command including multiple keyword groups and the display of related results, which is beneficial to reduce the number of recognition times of voice commands for users, and is beneficial to improving the intelligence level of related equipment equipped with a voice recognition system, thereby helping to improve the user experience of the voice recognition function.

[0060] Please continue to refer to Figure 2 In an exemplary embodiment, the content of step 102 can be implemented by executing steps 201 to 203, wherein:

[0061] Step 201, when it is identified that the queue of tasks to be executed includes subtasks, based on the order of multiple subtasks to be scheduled, extract the subtask with the first order to be scheduled;

[0062] Step 202, triggering a preset executor to execute the extracted subtask to obtain execution data corresponding to the subtask, and saving the execution data to a completed task queue;

[0063] Step 203 , identifying the to-be-executed task queue again until the execution data of each subtask in the to-be-executed task queue is saved in the completed task queue.

[0064] First, step 201 is executed to identify whether the queue of tasks to be executed includes subtasks to be executed. When it is identified that there are multiple subtasks included, the subtask with the first (for example, the first) subtask to be scheduled can be extracted based on the pre-matched sequence of each subtask to be scheduled, which can be called the first subtask; then step 202 is executed. After taking out the first subtask, the preset executor is triggered to execute the first subtask to obtain the execution result corresponding to the first subtask, that is, the execution data, and the execution data of the first subtask is saved in the completed task queue for waiting; then step 203 is executed again to identify whether the queue of tasks to be executed includes subtasks to be executed. Whether it includes subtasks to be executed is identified. When it is identified that at least one subtask is included, the subtask with the first (for example, the second) subtask in the to-be-executed order can be extracted based on the pre-matched order to be scheduled of the currently included subtasks, which can be called the second subtask, and then step 202 is executed again. After taking out the second subtask, the preset executor is triggered to execute the second subtask to obtain the execution result corresponding to the second subtask, that is, the execution data, and the execution data of the second subtask is saved in the completed task queue for waiting; until the execution data of each subtask in the to-be-executed task queue are saved in the completed task queue.

[0065] It can be seen that the present application provides an optional implementation method for the execution method of the subtasks included in the pending task queue, that is, the subtasks included in the pending task queue are executed in a pre-set order, and the corresponding subsequent subtask is executed only after the execution data obtained by the execution of the previous subtask is stored in the completed task queue. Based on this execution method, an effective task management and scheduling mechanism is provided for the execution of multiple subtasks, which is conducive to ensuring the execution efficiency of the pending subtasks, thereby helping to improve the user experience.

[0066] Please continue to refer to Figure 2 In an exemplary embodiment, the content of saving the execution data to the completed task queue in the above steps can be selectively executed as: saving the execution data to the task data model associated with the corresponding subtask, and saving the task data model to the completed task queue; wherein the task data model is used to indicate the execution method of the execution data.

[0067] The data model (Task Data Model) refers to an abstract model that defines the structure, attributes, relationships, and how to process the data in the application. The data model is an important part of application design, which provides a framework for the application to organize, store, and manage data.

[0068] The task data model may include the task type, the parameters required for the task, and the result information obtained after executing the task.

[0069] Specifically, the execution data obtained by executing the subtask can be first saved in the task data model associated with the subtask. For example, if a subtask is a destination search task, then the corresponding task data model is a destination search task data model. Among them, the task data model corresponding to each subtask can be configured in advance. When the type of the subtask is identified, the corresponding task data model can be found based on the mapping relationship between the subtask and the task data model. Then, the task data model with the execution data can be saved in the completed task queue, rather than directly saving the execution data in the completed task queue. This is conducive to the subsequent extraction of the content from the completed task queue being the task data model with the execution data. Each task data model contains further execution methods indicating the execution data stored therein, which is conducive to ensuring the execution accuracy of subsequent steps.

[0070] Please continue to refer to Figure 2 In an exemplary embodiment, the content of the above step 103 can be specifically executed as follows: when it is recognized that the completed task queue includes at least two task data models, trigger the associated application to output the task execution result of the voice command based on the execution method of the execution data respectively indicated by the at least two task data models.

[0071] Specifically, when it is identified that the completed task queue already includes each task data model corresponding to each subtask in the to-be-executed task queue, an associated application (third-party application) can be called to enable the associated application to further execute each execution data based on the data processing method indicated by the task data model corresponding to the execution data of each subtask, thereby outputting the task execution result corresponding to the single voice command input by the user, realizing the relevant functions required for the single voice command input by the user, and being able to receive the relevant execution result display, so that the user can get the result he wants, which is conducive to improving the user experience.

[0072] Please continue to refer to Figure 2 In an exemplary embodiment, the above step 102 is performed to execute the extracted subtask to obtain the execution data corresponding to the subtask, which can be specifically implemented by executing steps 221 to 223, wherein:

[0073] Step 221, executing the extracted subtask to obtain a candidate data table output by the associated application pointed to by the subtask;

[0074] Step 222, guiding the outputter of the voice command to select target execution data in the candidate data table through voice and / or display;

[0075] Step 223, obtaining the selected target execution data as the execution data corresponding to the subtask.

[0076] Specifically, by executing a subtask, a candidate data table output by the associated application pointed to by the subtask can be displayed in a terminal device adopting the voice recognition system, wherein the candidate data table includes multiple candidate results related to the keyword group corresponding to the subtask; and then, through voice prompts or dynamic display of display effects, the outputter of the voice command (such as the user) can be guided to select the candidate results (target execution data) he wants in the candidate data table; finally, the selected target execution data is obtained as the execution data corresponding to the subtask.

[0077] That is, the present application proposes that dynamic display of voice prompts or display effects can be used to guide users to determine the target results, which is conducive to improving the accuracy of the task execution results of the final output voice commands, and this method is also conducive to improving the user experience.

[0078] It should also be added that the output party of the voice command (such as a human user) can select the target result (target execution data) in the query result table (candidate data table) based on touch selection, voice interaction through voice selection, or some physical buttons in the device. For example, when the voice command execution method provided in the present application is applied to a smart vehicle, the above-mentioned candidate data table will be displayed on the display screen of the vehicle, and the user of the vehicle can select the target execution data in the candidate data table based on voice interaction with the vehicle, or touching the display screen, or pressing physical buttons.

[0079] In an exemplary embodiment, it also includes: when it is identified that the queue of tasks to be executed does not include a subtask, and when it is identified that the queue of completed tasks does not include execution data, a prompt message is output to guide the outputter of the voice command to re-enter the voice command.

[0080] Specifically, when executing step 102 to identify the queue of tasks to be executed, and the corresponding identification result is that there are no subtasks to be executed in the queue of tasks to be executed, there are two possibilities. The first possibility is that all subtasks included in the queue of tasks to be executed have been executed, and the relevant execution data have been saved in the completed task queue; the second possibility is that the voice instruction was not recognized successfully, that is, the keyword group was not recognized, or the process of creating subtasks based on the recognized keyword group failed, which will result in no subtasks being able to be added to the queue of tasks to be executed.

[0081] If it is identified that the queue of tasks to be executed does not include subtasks, and the completed task queue is further identified, and the identification result is that no execution data is saved in the completed task queue, it can be determined that this situation belongs to the first possibility mentioned above. Therefore, a prompt message for the user (the party outputting the voice command) can be output at this time to guide the party outputting the voice command to re-enter the voice command. This avoids the situation where the voice command input by the user is not successfully executed and the related function is not executed, but the user is unaware or unable to know it in time, which is conducive to improving the user experience.

[0082] For the execution method of the voice command provided in this application, an optional implementation method is also provided. Please refer to Figure 3 and Figure 4 , specifically including the following steps S11 to S16, wherein:

[0083] S11, the user activates the voice assistant and issues a (single) voice command;

[0084] S12, the speech recognition system performs speech recognition and semantic understanding, and returns semantic results (keyword groups);

[0085] S13, the task scheduler creates a task (subtask) according to the semantic result and adds the task to the queue of tasks to be executed;

[0086] S14, the scheduler reads the queue of tasks to be executed; wherein:

[0087] 4.1. If there are tasks to be executed, take out the tasks and hand them over to the task executor;

[0088] 4.2. If there is no task to be executed, determine whether there is a completed task;

[0089] 4.2.1. If there are completed tasks, combine them into the final action and execute them according to the completed tasks.

[0090] 4.2.2. If there is no completed task, the process ends.

[0091] S15, the task executor executes the task and returns the execution result type and result data;

[0092] 5.1. If the task is executed successfully, the scheduler saves the result data to the task data model and stores the task model in the completed task queue.

[0093] 5.2. If the task execution fails, all task queues will be cleared, the failure will be reported, and the process ends.

[0094] S16. Repeat steps S14, S15, and S16.

[0095] Based on this, the present application also provides a specific embodiment, taking the user's voice to initiate the "navigate to location A and pass through location B" command (a single voice command) as an example, and further provides a detailed execution process description. Among them, since the "location A" (keyword group) and "location B" (keyword group) may correspond to a relatively large range, there may be corresponding multiple location search results, so it is necessary to gradually guide the user to confirm which "location A" and "location B" they want to go to. Please refer to Figure 3 and Figure 5 The following steps S21 to S36 describe the relevant execution process in detail, wherein:

[0096] S21, the user wakes up the voice assistant and issues a command "navigate to location A and pass through location B";

[0097] S22, the speech recognition system performs speech recognition and semantic understanding, and returns semantic results (i.e., "location A" and "location B");

[0098] S23, the task scheduler creates a "destination confirmation task" (the first subtask) and a "waypoint confirmation task" (the second subtask) according to the semantics; wherein the task model (task data model) stores the names of the places to be queried;

[0099] S24, the task scheduler takes out the "end point confirmation task" and gives it to the task executor for execution;

[0100] S25, the task executor calls the map SDK to query the result list (candidate data table) of the destination "location A", and displays the list to guide the user to make a selection;

[0101] S26, the task executor returns the result selected by the user (target execution data);

[0102] S27, the task scheduler saves the endpoint information selected by the user into an endpoint task data model (a task data model), and saves the endpoint task data model into a completed task queue;

[0103] S28, the task scheduler checks whether there are any remaining tasks to be executed;

[0104] S29, the task scheduler takes out the "waypoint confirmation task" and gives it to the task executor for execution;

[0105] S30, the task executor calls the map SDK to query the result list of the waypoint "Location B", and displays the list to guide the user to make a selection;

[0106] S31, the task executor returns the result selected by the user;

[0107] S32, the task scheduler saves the waypoint information selected by the user into the waypoint task data model, and saves the waypoint task data model into the completed task queue;

[0108] S33, the task scheduler detects that there are no remaining tasks to be executed;

[0109] S34, the task scheduler detects that there is a completed task;

[0110] S35, the task scheduler calls the task executor according to the completed destination confirmation and waypoint confirmation tasks, and executes the instruction of navigating to location A and passing through location B;

[0111] S36, start navigation, and the process ends.

[0112] It should be noted that, for the case where the result list of the above-mentioned "A location" includes multiple A locations, for example, a restaurant chain in a city includes multiple branches. The same is true for multiple "B locations".

[0113] The voice command execution method provided in the above-mentioned application can efficiently execute multi-step tasks based on the relevant voice command execution system, reduce the number of times users repeat commands, and help improve task execution efficiency; enable users to complete complex multi-step tasks through simple voice commands, thereby improving user experience; and the relevant voice command execution system can understand complex multi-step commands, thereby improving the intelligence level of the system.

[0114] It should be understood that, although the various steps in the flowcharts involved in the above-mentioned embodiments are displayed in sequence according to the indication of the arrows, these steps are not necessarily executed in sequence according to the order indicated by the arrows. Unless there is a clear explanation in this article, the execution of these steps does not have a strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments can include multiple steps or multiple stages, and these steps or stages are not necessarily executed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a part of the steps or stages in other steps.

[0115] Based on the same inventive concept, the present invention Figure 3 and Figure 4 The figure also shows a related voice command execution system 200 provided by the present application, which includes a voice recognition system 71, a task scheduler 72 and a task executor 73. Among them:

[0116] A speech recognition system 71, used for recognizing a single speech instruction to obtain at least two keyword groups;

[0117] A task scheduler 72, used to create at least two corresponding subtasks based on the keyword group and add them to the same to-be-executed task queue;

[0118] The task scheduler 72 is further used to call the task executor 73 to execute each subtask when recognizing that the to-be-executed task queue includes subtasks, so as to obtain the execution data corresponding to each subtask respectively. The task scheduler 72 is also used to save the execution data of each subtask to the completed task queue;

[0119] The task scheduler 72 is also used to call the task executor 73 to trigger the associated application to output the task execution result of the voice command based on the execution data of at least two subtasks when recognizing that the completed task queue includes the execution data of at least two subtasks.

[0120] The task scheduler can create tasks, add tasks to the task queue, and take tasks out of the queue and hand them over to the task executor for execution. The task executor can identify the task type and execute the corresponding task according to the incoming task, and return the task result.

[0121] Furthermore, in the technical solution provided in the present application, a method for voice assistant terminal program management and scheduling multi-step tasks is provided, which enables the voice command execution system to efficiently execute multi-step tasks through the collaborative work of the voice recognition module and intent analysis module included in the voice recognition system, the task queue module and task scheduling module included in the task scheduler, and the result integration module included in the task executor, thereby helping to improve task execution efficiency and user experience.

[0122] The voice command execution method and system provided in the present application provide an effective task management mechanism and scheduling mechanism for processing voice commands of multiple keyword groups and related multi-step tasks, which is beneficial to improving the execution efficiency of voice commands of multiple keyword groups, thereby helping to improve the user experience.

[0123] Based on the same inventive concept, the embodiment of the present application also provides a voice command execution device for implementing the voice command execution method involved above. The implementation scheme for solving the problem provided by the device is similar to the implementation scheme recorded in the above method, so the specific limitations in the embodiments of the execution device for one or more voice commands provided below can refer to the limitations of the voice command execution method above, and will not be repeated here.

[0124] In an exemplary embodiment, Figure 6 As shown, a voice command execution device 300 is provided, comprising: a to-be-executed task queue creation module 81, a task execution module 82 and a result output module 83, wherein:

[0125] A to-be-executed task queue creation module 81 is used to identify a single voice instruction to obtain at least two keyword groups, create at least two corresponding subtasks based on the keyword groups and add them to the same to-be-executed task queue;

[0126] The task execution module 82 is used to trigger a preset executor to execute each subtask when identifying that the queue of tasks to be executed includes subtasks, so as to obtain execution data corresponding to each subtask respectively, and save the execution data of each subtask to the completed task queue;

[0127] The result output module 83 is used to trigger the associated application to output the task execution result of the voice command based on the execution data of at least two subtasks when recognizing that the completed task queue includes the execution data of at least two subtasks.

[0128] In an exemplary embodiment, the task execution module 82 is used to extract the subtask that is the first in the order to be scheduled based on the order to be scheduled of multiple subtasks when it is identified that the queue of tasks to be executed includes subtasks; trigger a preset executor to execute the extracted subtask to obtain the execution data corresponding to the subtask, and save the execution data to the completed task queue; identify the queue of tasks to be executed again until the execution data of each subtask in the queue of tasks to be executed is saved to the completed task queue.

[0129] In an exemplary embodiment, the task execution module 82 is used to save the execution data into a task data model associated with the corresponding subtask, and save the task data model into a completed task queue; wherein the task data model is used to indicate the execution method of the execution data.

[0130] In an exemplary embodiment, the result output module 83 is used to trigger the associated application to output the task execution result of the voice command based on the execution method of the execution data respectively indicated by the at least two task data models when it is recognized that the completed task queue includes at least two task data models.

[0131] In an exemplary embodiment, the task execution module 82 is used to execute the extracted subtask to obtain a candidate data table output by the associated application pointed to by the subtask; guide the outputter of the voice command to select the target execution data in the candidate data table through voice and / or display; and obtain the selected target execution data as the execution data corresponding to the subtask.

[0132] In an exemplary embodiment, the voice command execution device 300 also includes a message prompt module, which is used to output a prompt message to guide the outputter of the voice command to re-enter the voice command when it is identified that the task queue to be executed does not include a subtask and that the completed task queue does not include execution data.

[0133] Each module in the above-mentioned voice command execution device 300 can be implemented in whole or in part by software, hardware or a combination thereof. Each module can be embedded in or independent of a processor in a computer device in the form of hardware, or can be stored in a memory in a computer device in the form of software, so that the processor can call and execute the operations corresponding to each module above.

[0134] Figure 7 FIG. 1 is an internal structure diagram of a vehicle-side control device in an embodiment. In an exemplary embodiment, a vehicle-side control device is provided. The internal structure diagram of the vehicle-side control device can be as follows: Figure 7As shown. The vehicle-side control device includes a processor and a memory. Among them, the processor of the vehicle-side control device is used to provide computing and control capabilities. The memory of the vehicle-side control device includes a non-volatile storage medium, and the non-volatile storage medium stores a computer program. When the computer program is executed by the processor, a control method for upgrading vehicle functions is implemented.

[0135] Those skilled in the art will understand that Figure 7 The structure shown in the figure is merely a block diagram of a partial structure related to the scheme of the present application, and does not constitute a limitation on the vehicle-side control device to which the scheme of the present application is applied. The specific vehicle-side control device may include more or fewer components than those shown in the figure, or combine certain components, or have a different component arrangement.

[0136] Based on the same inventive concept, the present application also provides a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the aforementioned method for executing voice instructions is implemented. The method for executing voice instructions is any one of the methods for executing voice instructions mentioned in the embodiments of the present application. For related embodiments, please refer to the above.

[0137] Based on the same inventive concept, the present application also provides a computer program product, including a computer program, which, when executed by a processor, implements the aforementioned voice instruction execution method. The voice instruction execution method is any one of the voice instruction execution methods mentioned in the embodiments of the present application. For related embodiments, please refer to the above.

[0138] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with relevant regulations.

[0139] A person of ordinary skill in the art can understand that all or part of the processes in the above-mentioned embodiment method can be completed by instructing the relevant hardware through a computer program, and the computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to the memory, database or other medium used in the embodiments provided in the present application can include at least one of non-volatile memory and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. As an illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM). The database involved in each embodiment provided in this application may include at least one of a relational database and a non-relational database. Non-relational databases may include distributed databases based on blockchains, etc., but are not limited to this. The processor involved in each embodiment provided in this application may be a general-purpose processor, a central processing unit, a graphics processor, a digital signal processor, a programmable logic device, a data processing logic device based on quantum computing, an artificial intelligence (AI) processor, etc., but are not limited to this.

[0140] The technical features of the above embodiments may be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this application.

[0141] The above-described embodiments only express several implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for a person of ordinary skill in the art, several variations and improvements can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the attached claims.

Claims

1. A method for executing a voice command, characterized in that: include: Recognize a single voice command to obtain at least two keyword groups, create at least two corresponding subtasks based on the keyword groups and add them to the same queue of tasks to be executed; In the case of identifying that the queue of tasks to be executed includes the subtasks, triggering a preset executor to execute each of the subtasks, so as to respectively obtain execution data corresponding to each of the subtasks, and saving the execution data of each of the subtasks to the completed task queue; In the case of identifying that the completed task queue includes the execution data of at least two of the subtasks, triggering an associated application to output a task execution result of the voice instruction based on the execution data of at least two of the subtasks.

2. The method according to claim 1, characterized in that In the case of identifying that the queue of tasks to be executed includes the subtask, triggering a preset executor to execute each of the subtasks, respectively obtaining execution data corresponding to each of the subtasks, and saving the execution data of each of the subtasks to the completed task queue, including: When it is identified that the queue of tasks to be executed includes the subtask, based on the order of the subtasks to be scheduled, extracting the subtask that is the first in the order to be scheduled; Triggering a preset executor to execute the extracted subtask to obtain execution data corresponding to the subtask, and saving the execution data to a completed task queue; The to-be-executed task queue is identified again until the execution data of each of the subtasks in the to-be-executed task queue is saved in the completed task queue.

3. The method according to claim 1 or 2, characterized in that: Saving the execution data to a completed task queue includes: The execution data is saved in a task data model associated with the corresponding subtask, and the task data model is saved in a completed task queue; wherein the task data model is used to indicate an execution method of the execution data.

4. The method according to claim 3, characterized in that The step of, in the case of recognizing that the completed task queue includes the execution data of at least two of the subtasks, triggering the associated application to output the task execution result of the voice instruction based on the execution data of at least two of the subtasks, comprises: When it is identified that the completed task queue includes at least two of the task data models, the associated application is triggered to output the task execution result of the voice instruction based on the execution mode of the execution data respectively indicated by the at least two task data models.

5. The method according to claim 2, characterized in that: The executing the extracted subtask to obtain execution data corresponding to the subtask includes: Execute the extracted subtask to obtain a candidate data table output by the associated application pointed to by the subtask; guiding the outputter of the voice instruction to select target execution data in the candidate data table through voice and / or display; The selected target execution data is obtained as execution data corresponding to the subtask.

6. The method according to claim 1, characterized in that Also includes: When it is identified that the to-be-executed task queue does not include the subtask and when it is identified that the completed task queue does not include the execution data, prompt information is output to guide the outputter of the voice command to re-input the voice command.

7. A voice command execution system, characterized in that: include: A speech recognition system, used to recognize a single voice command to obtain at least two key words; A task scheduler, used for creating at least two corresponding subtasks based on the keyword group and adding them to the same queue of tasks to be executed; The task scheduler is further used to call the task executor to execute each of the subtasks when recognizing that the queue of tasks to be executed includes the subtasks, so as to respectively obtain the execution data corresponding to each of the subtasks, and the task scheduler is further used to save the execution data of each of the subtasks to the completed task queue; The task scheduler is also used to call the task executor to trigger the associated application to output the task execution result of the voice instruction based on the execution data of at least two subtasks when recognizing that the completed task queue already includes the execution data of at least two subtasks.

8. A device for executing a voice command, characterized in that: The device comprises: A to-be-executed task queue creation module, used for identifying a single voice command to obtain at least two keyword groups, creating at least two corresponding subtasks based on the keyword groups and adding them to the same to-be-executed task queue; A task execution module, for triggering a preset executor to execute each of the subtasks when identifying that the queue of tasks to be executed includes the subtasks, so as to respectively obtain execution data corresponding to each of the subtasks, and save the execution data of each of the subtasks to the completed task queue; The result output module is used to trigger the associated application to output the task execution result of the voice instruction based on the execution data of at least two subtasks when recognizing that the completed task queue includes the execution data of at least two subtasks.

9. A vehicle-side control device, comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 6 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 6 are implemented.

Citation Information

Patent Citations

  • Multi-command single utterance input method

    CN106471570A

  • Instruction execution method and device, storage medium, and electronic device

    CN108711428A

  • Task model training method, device and equipment

    CN110473521A

  • Method and device for executing interaction function, electronic equipment and storage medium

    CN110728981A

  • Vehicle-mounted multi-instruction execution method and device, electronic equipment and storage medium

    CN116168696A

Cited By

  • Information interaction method, system and device, electronic equipment and storage medium

    CN120315798A