Task execution method and device based on voice control

Through voice recognition and voiceprint recognition technology, a voice control-based task execution method is realized, solving the problems of complex user operations, low security and low task conflict handling efficiency, providing accurate and fast task control and identity verification, ensuring the consistency and efficiency of task execution.

CN120236579APending Publication Date: 2025-07-01BEITAI ZHENHUAN (CHONGQING) TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510396480.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-31
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In the prior art, user operations are complex, low security, and low task conflict handling efficiency during task execution. Especially in noisy environments or scenarios where two-handed operations are required, traditional input devices are inefficient and lack of authentication and permission control, resulting in chaotic task execution and unsmooth interaction.

Method used

Using a voice control-based task execution method, through pre-trained voice recognition and voiceprint recognition technology, voice commands are analyzed and user identity is verified, operation permissions are determined, task instructions are matched from the instruction library, and task status is monitored in real time, and operation strategies are formulated to achieve seamless conversion and efficient execution.

Benefits of technology

It realizes accurate and fast task control and user identity authentication, ensuring that only authorized users can perform operations, avoid task conflicts, improves the consistency and efficiency of task execution, and provides a rich task instruction library and intelligent operation guidance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236579A_ABST
    Figure CN120236579A_ABST
Patent Text Reader

Abstract

The invention discloses a task execution method and device based on voice control. The method comprises the following steps: receiving a voice control instruction of a target object; analyzing through a pre-trained speech recognition model to obtain a target text instruction; obtaining target identity information by using the voiceprint recognition model, and determining a target control permission corresponding to the target identity information; matching a target operation instruction conforming to the target control permission and the target text instruction from an instruction library, wherein the instruction library comprises a plurality of operation instructions and control permissions thereof; and detecting a current task execution state, determining and executing a target operation strategy according to the state and the target operation instruction, and displaying a task execution interface. The technical problems of complex user operation, low safety and low task conflict processing efficiency in the task execution control process are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical field of task management, and more particularly, to a voice control-based task execution method and apparatus. Background Art

[0002] In modern work and living environments, the rapid switching and efficient management of tasks have become key factors in improving productivity and user experience. Whether it is the switching and execution of function modules in software applications, the process transfer in device operations, or the rapid switching and execution of different tasks in daily life, an efficient and intuitive interaction method is required. Although traditional physical input methods such as keyboards and mice are reliable, they have obvious limitations in certain scenarios, such as when both hands are busy, the environment is noisy, or instant response is needed. Especially in professional application fields such as numerical calculation and simulation software, users usually need to frequently switch between different interfaces or modules to complete a series of complex tasks, such as data analysis, model construction, and result visualization. Traditional operation methods, such as using a mouse and keyboard, although intuitive, are less efficient when performing multi-step or fine operations, and are inconvenient to use in specific environments (such as a noisy laboratory or a scenario where both hands are needed). Using physical input devices for task switching and execution, especially in scenarios of multi-task parallelism or frequent switching and execution, is cumbersome and time-consuming, reducing work efficiency. In specific environments, such as industrial sites, the operation of physical devices may become difficult or even dangerous, and the prior art fails to provide sufficient control means to adapt to these environments. In addition, the current task switching and execution mechanism lacks effective authentication and permission control, and any user may cause task execution chaos due to misoperation or malicious behavior. Moreover, in the prior art, the interaction method for task switching and execution is single, lacking intelligent feedback and operation guidance, and users may need to try multiple times to successfully switch and execute tasks, reducing the smoothness and satisfaction of the interaction.

[0003] To address the above problems, no effective solution has been proposed yet. Summary of the Invention

[0004] Embodiments of this application provide a voice control-based task execution method and apparatus to at least solve the technical problems of complex user operations, low security, and low efficiency in handling task conflicts during the control process of task execution.

[0005] According to one aspect of the embodiments of the present application, a task execution method based on voice control is provided, including: receiving a voice control instruction input by a target object; analyzing the voice control instruction by using a pre-trained speech recognition model to obtain a target text instruction; analyzing the voice control instruction by using a pre-trained voiceprint recognition model to obtain the target identity information of the target object, and determining the target control authority corresponding to the target identity information; determining a target operation instruction matching the target control authority and the target text instruction from a preset instruction library, where the instruction library stores multiple operation instructions and corresponding control authorities, and at least one task to be executed is included in the operation instruction; detecting the current task execution status, and determining a target operation strategy according to the current task execution status and the target operation instruction; executing the target operation strategy, and displaying a task execution interface.

[0006] Optionally, receiving a voice control instruction input by a target object includes: receiving a voice control signal input by the target object through a voice input device; performing noise reduction processing on the voice control signal to obtain a voice control instruction, where the noise reduction processing includes at least one of the following: dynamic noise suppression, echo cancellation.

[0007] Optionally, analyzing the voice control instruction by using a pre-trained speech recognition model to obtain a target text instruction includes: using the speech recognition model to extract the Mel frequency cepstral coefficient feature vector corresponding to the voice control instruction, and analyzing the Mel frequency cepstral coefficient feature vector to obtain an initial text instruction; using transfer learning technology to correct the context of the initial text instruction to obtain the target text instruction.

[0008] Optionally, analyzing the voice control instruction by using a pre-trained voiceprint recognition model to obtain the target identity information of the target object, and determining the target control authority corresponding to the target identity information includes: using the voiceprint recognition model to extract the target voiceprint feature vector corresponding to the voice control instruction; determining the target identity information corresponding to the target voiceprint feature vector from a preset voiceprint library, where the voiceprint library stores the mapping relationship between multiple voiceprint feature vectors and multiple identity information; determining the target control authority corresponding to the target identity information from a preset authority table, where the authority table stores the control authorities corresponding to multiple identity information.

[0009] Optionally, determining a target operation instruction matching the target control authority and the target text instruction from a preset instruction library includes: respectively determining the text similarity between the target text instruction and each operation instruction in the instruction library; when the maximum text similarity is not less than a preset similarity threshold and the target control authority is not lower than the control authority of the operation instruction corresponding to the maximum text similarity, determining the operation instruction corresponding to the maximum text similarity as the target operation instruction.

[0010] Optionally, the method further includes: generating a first prompt message when the maximum text similarity is less than a preset similarity threshold, where the first prompt message is used to prompt that the voice control instruction is unclear and needs to be re-entered; generating a second prompt message when the maximum text similarity is not less than the preset similarity threshold but the target control permission is lower than the control permission of the operation instruction corresponding to the maximum text similarity, where the second prompt message is used to prompt insufficient control permission.

[0011] Optionally, detect the current task execution status and determine the target operation strategy based on the current task execution status and the target operation instruction, including: detecting the current task execution status and determining the first task to be executed indicated by the target operation instruction; when there is no task being executed currently, determining the target operation strategy as directly executing the first task; when a second task is being executed currently, determining the task relationship between the second task and the first task and determining the target operation strategy based on the task relationship, where the task relationship is at least used to reflect whether there is a conflict between the second task and the first task.

[0012] Optionally, determining the target operation strategy based on the task relationship includes: when there is a conflict between the second task and the first task, determining that the target operation strategy includes: sending an inquiry message to the target object and determining whether to continue executing the second task or stop executing the second task and start executing the first task based on the selection instruction feedback by the target object, where the inquiry message is used to prompt that there is a conflict between the second task and the first task and one of them needs to be selected for execution; when there is no conflict between the second task and the first task, determining that the target operation strategy includes: switching the second task to run in the background, starting to execute the first task and displaying the task execution interface corresponding to the first task.

[0013] Optionally, the method further includes: when there is no conflict and there is an association relationship between the second task and the first task, determining that the target operation strategy includes: executing the second task and the first task simultaneously and displaying the task execution interfaces of the second task and the first task in a preset ratio.

[0014] According to another aspect of the embodiments of the present application, there is also provided a task execution device based on voice control, including: a receiving module, configured to receive a voice control instruction input by a target object; a first analysis module, configured to analyze the voice control instruction by using a pre-trained voice recognition model to obtain a target text instruction; a second analysis module, configured to analyze the voice control instruction by using a pre-trained voiceprint recognition model to obtain the target identity information of the target object and determine the target control authority corresponding to the target identity information; an instruction determination module, configured to determine a target operation instruction that matches the target control authority and the target text instruction from a preset instruction library, where the instruction library stores a plurality of operation instructions and corresponding control authorities, and at least one of the operation instructions includes a task to be executed; a policy determination module, configured to detect the current task execution status and determine a target operation policy according to the current task execution status and the target operation instruction; an execution module, configured to execute the target operation policy and display a task execution interface.

[0015] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including: a computer program, where when the computer program is executed by a processor, it implements the above-mentioned task execution method based on voice control.

[0016] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where a computer program is stored in the memory, and the processor is configured to execute the above-mentioned task execution method based on voice control through the computer program.

[0017] In the embodiments of the present application, by integrating voice recognition technology and voiceprint recognition technology, the present invention realizes accurate and fast task control and user identity verification. Users can start, adjust or terminate tasks through natural language instructions, and at the same time ensure that only authorized individuals can execute the control. After verifying the user identity by using voiceprint recognition, it intelligently analyzes and determines the operation authority corresponding to the user, ensuring that in any task execution scenario, participants can only perform corresponding operations according to their authorities. This method presets an instruction library containing rich task instructions, covering comprehensive contents from basic commands to complex operations, and is linked to user authorities. When receiving a voice instruction, it can quickly match the corresponding task instruction, realizing a seamless conversion from voice to task operation. It can not only recognize and parse voice instructions, but also monitor the current task execution status in real time, and intelligently formulate subsequent operation strategies according to the task status and user instructions. This feature allows the system to be flexible and avoid conflicts and interruptions during task execution, ensuring the coherence and high-efficiency execution of tasks, and thus solving the technical problems of complex user operations, low security, and low efficiency in handling task conflicts during the control process of task execution. Description of the Drawings

[0018] The accompanying drawings described herein are used to provide a further understanding of the present application, and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation of the present application. In the drawings:

[0019] Figure 1 is a schematic flowchart of an optional voice control-based task execution method according to an embodiment of the present application;

[0020] Figure 2 is a schematic structural diagram of an optional voice control-based task execution device according to an embodiment of the present application;

[0021] Figure 3 is a schematic structural diagram of an optional electronic device according to an embodiment of the present application. Detailed implementation manners

[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] It should be noted that the terms "first", "second", etc. in the specification, claims and drawings of the present application are used to distinguish similar objects, and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units does not necessarily have to be limited to those clearly listed steps or units, but may include other steps or units that are not clearly listed or are inherent to these processes, methods, products or devices.

[0024] The information collected in the embodiments of the present application is information and data authorized by the user or fully authorized by all parties. And the processing of relevant data, such as collection, storage, use, processing, transmission, provision, disclosure and application, etc., all comply with the relevant laws, regulations and standards of relevant countries and regions, take necessary confidentiality measures, do not violate public order and good customs, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0025] Embodiment 1

[0026] According to an embodiment of the present application, a method for task execution based on voice control is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.

[0027] Figure 1 is a schematic flowchart of a method for task execution based on voice control provided according to an embodiment of the present application. As Figure 1 shown, the method includes the following steps:

[0028] Step S102, receiving a voice control instruction input by a target object;

[0029] Step S104, analyzing the voice control instruction by using a pre-trained voice recognition model to obtain a target text instruction;

[0030] Step S106, analyzing the voice control instruction by using a pre-trained voiceprint recognition model to obtain the target identity information of the target object, and determining the target control permission corresponding to the target identity information;

[0031] Step S108, determining a target operation instruction that matches the target control permission and the target text instruction from a preset instruction library, where the instruction library stores multiple operation instructions and corresponding control permissions, and at least one task to be executed is included in the operation instructions;

[0032] Step S110, detecting the current task execution status, and determining a target operation strategy based on the current task execution status and the target operation instruction;

[0033] Step S112, executing the target operation strategy and displaying a task execution interface.

[0034] The following explains each step of the method for task execution based on voice control in combination with a specific implementation process.

[0035] Receiving a voice control instruction input by a target object can be carried out in the following steps:

[0036] Receiving a voice control signal input by the target object through a voice input device; performing noise reduction processing on the voice control signal to obtain a voice control instruction, where the noise reduction processing includes at least one of the following: dynamic noise suppression, echo cancellation.

[0037] For example, the user's voice is captured through a voice recognition interface integrated in the software or independent voice collection hardware such as a microphone. At this time, instructions such as "Open the data analysis module" and "Switch to the 3D simulation view" issued by the user in a natural voice state will be converted into digital audio signals. The collected voice signals are often mixed with environmental noise and possible echo problems, which will greatly affect the subsequent voice recognition accuracy. It is necessary to preprocess the received voice signals. The key preprocessing technologies include: dynamic noise suppression, echo cancellation, etc.

[0038] After obtaining the voice control instruction, use the pre-trained voice recognition model to analyze the voice control instruction to obtain the target text instruction. This process can be carried out in the following steps:

[0039] Use the voice recognition model to extract the Mel Frequency Cepstral Coefficient (MFCC) feature vector corresponding to the voice control instruction, and analyze the MFCC feature vector to obtain the initial text instruction; use transfer learning technology to correct the context of the initial text instruction to obtain the target text instruction.

[0040] For example, once a clean and clear voice signal is received, the next step is to convert it into a form that can be understood by a computer. This conversion is usually achieved through a feature extraction method called Mel Frequency Cepstral Coefficients (MFCC). The MFCC extraction technology can extract the frequency components that are most sensitive to human hearing from the voice signal, thus forming a set of feature vectors. These feature vectors contain information such as the timbre, intensity, and rhythm of the voice signal. The specific steps include: segmentation, pre-value, Fourier transform, Mel filter bank, logarithmic energy conversion, cepstral analysis. After obtaining the MFCC feature vector, the next task is to input these feature vectors into a pre-trained deep learning model for voice recognition. The pre-trained model is trained on a large amount of voice data and can recognize the corresponding text instruction from the input feature vector. Through the learning of multiple layers of neurons, these models can find patterns in complex sound features and thus convert them into text instructions.

[0041] Even the most advanced speech recognition models may make errors in some cases, especially when faced with complex factors such as accents, speech rate changes, or noise. Context-sensitive error correction of the initial text instructions can be achieved using transfer learning techniques. For example, for the task control execution of numerical calculation and simulation software in a computer, if the user's original speech instruction is: "Show the 3D view", but is misrecognized as: "Show the 3C view" for some reason, context information (such as which interface the software is currently in, what instructions the user has issued before, etc.) can be used to judge that "3C view" may not make sense in the context of numerical calculation and simulation software, while "3D view" is the reasonable option. Through transfer learning techniques, the model can correct the misrecognition based on previous experience and the current context, and adjust "3C view" to "3D view".

[0042] After obtaining the voice control instruction, use the pre-trained voiceprint recognition model to analyze the voice control instruction to obtain the target identity information of the target object and determine the target control permission corresponding to the target identity information. This process can be carried out in the following steps:

[0043] Use the voiceprint recognition model to extract the target voiceprint feature vector corresponding to the voice control instruction; determine the target identity information corresponding to the target voiceprint feature vector from the preset voiceprint library, where the voiceprint library stores the mapping relationship between multiple voiceprint feature vectors and multiple identity information; determine the target control permission corresponding to the target identity information from the preset permission table, where the permission table stores the control permissions corresponding to multiple identity information.

[0044] Voiceprint recognition is a biometric recognition technology that identifies the identity of a speaker by analyzing the individual's voice characteristics. These characteristics may include timbre, pitch, speaking speed, and other voice features, which are unique to an individual to some extent. The voiceprint recognition model first preprocesses the voice control command, and then extracts the voiceprint feature vector, which contains the key information that can uniquely identify the speaker. After the voiceprint feature vector is extracted, it is matched with a preset voiceprint database to confirm the identity of the speaker. The voiceprint database is a database that pre-recorded the voiceprint feature vectors of authorized users, and each feature vector is associated with a specific identity information. An efficient matching algorithm, such as dynamic time warping or nearest neighbor algorithm, is used to compare the target voiceprint feature vector with each record in the voiceprint database to find the most similar voiceprint feature vector, so as to determine the identity information of the speaker. Once the identity information of the speaker is confirmed, the control permissions matching the identity information need to be queried from the preset permission table. The permission table stores the control permissions of each user for various functions and modules on the software, ensuring that users can only access and control the functions they are authorized to use. For example, a junior user may only be able to control basic software interface switching and data input functions, while a senior user may also be able to access and control more complex simulation algorithms and system settings. The query process of the permission table ensures that the control permissions of the software are strictly managed and executed.

[0045] In addition, a recognition threshold can be preset. Only when the matching voiceprint similarity exceeds this threshold, the user identity is confirmed, otherwise the system will reject the execution of the command to prevent misrecognition.

[0046] The establishment of the voiceprint database can be carried out as follows: When used for the first time, the user is required to read a set of specified statements through a voice input device to generate their voiceprint feature vectors. These vectors, together with the user's identity information (such as username or user ID), are stored in the voiceprint database. As the user's voice habits change, a mechanism should be provided to update the voiceprint database to ensure long-term recognition accuracy.

[0047] After obtaining the target text command and the target control permission, the target operation command matching the target control permission and the target text command is determined from the preset command library. Among them, the command library stores multiple operation commands and the corresponding control permissions, and the operation command includes at least the task to be executed.

[0048] As an alternative implementation, to determine the target operation instruction that matches the target control permission and the target text instruction from a preset instruction library, the following process can be adopted: respectively determine the text similarity between the target text instruction and each operation instruction in the instruction library; when the maximum text similarity is not less than the preset similarity threshold and the target control permission is not lower than the control permission of the operation instruction corresponding to the maximum text similarity, determine the operation instruction corresponding to the maximum text similarity as the target operation instruction.

[0049] As an alternative implementation, the method further includes: when the maximum text similarity is less than the preset similarity threshold, generating a first prompt message, where the first prompt message is used to prompt that the voice control instruction is not clear and needs to be re-entered; when the maximum text similarity is not less than the preset similarity threshold but the target control permission is lower than the control permission of the operation instruction corresponding to the maximum text similarity, generating a second prompt message, where the second prompt message is used to prompt insufficient control permission.

[0050] For example, text similarity can be calculated through various natural language processing techniques, such as cosine similarity, edit distance, or more complex semantic similarity algorithms. These techniques can quantify the similarity between two texts, helping the system find the operation that best matches the user's voice instruction. Compare the similarity between the target text instruction issued by the user and each operation instruction in the instruction library to find the operation instruction with the maximum similarity. The preset similarity threshold is used to filter out those candidate operations that are not similar enough to the user's instruction to ensure the accuracy of the instruction. After determining the operation instruction with the highest similarity, it will be checked whether the control permission corresponding to this instruction matches the user's identity permission. If the user's identity permission meets or is higher than the control permission required by the target operation instruction, then this operation instruction is regarded as the target operation instruction and can be executed.

[0051] To further improve the user experience and the clarity of system interaction, the method also provides a feedback mechanism, which specifically includes: when it is found that the maximum text similarity is lower than the preset similarity threshold, it means that the user's voice instruction is not clear enough to find the exact corresponding operation in the instruction library. At this time, a first prompt message will be generated and sent to inform the user to reissue a clear voice instruction; even if the voice instruction is clearly recognized, but if the user's identity permission is lower than the control permission required by the target operation instruction, a second prompt message will be generated and sent to indicate that the user does not have the permission to execute this instruction and may provide alternative instructions or guide the user to contact the administrator to obtain higher permissions.

[0052] For example, assume there is a numerical calculation and simulation software that includes three functional modules: "Load Model", "Run Simulation", and "Export Results". Each module requires a different permission level to access. When the user issues a voice command: "Run Simulation", the system identifies and parses the command, verifies the user's identity and control permissions, and searches the instruction library for the instruction that best matches "Run Simulation". Assume the instruction library is as follows: "Load Model", basic user permission; "Run Simulation", intermediate user permission; "Export Results", advanced user permission. If the user has at least intermediate permission, "Run Simulation" will be determined as the instruction with the highest matching degree. If the user's permission is insufficient (e.g., only basic permission), the system generates a message indicating insufficient permission and reminds the user: "You do not have permission to run the simulation. Please contact the administrator to upgrade your account permission." If the system cannot accurately identify the instruction (e.g., the user says: "Start simulation"), since the text similarity with "Run Simulation" does not reach the preset threshold, the system will prompt: "Please clarify the instruction. Do you mean 'Run Simulation'?"

[0053] After obtaining the target operation instruction, detect the current task execution status, and determine the target operation strategy based on the current task execution status and the target operation instruction.

[0054] As an alternative implementation, to detect the current task execution status and determine the target operation strategy based on the current task execution status and the target operation instruction, the following steps can be taken: Detect the current task execution status and determine the first task to be executed indicated by the target operation instruction; in the case where no task is currently being executed, determine the target operation strategy as directly executing the first task; in the case where a second task is currently being executed, determine the task relationship between the second task and the first task, and determine the target operation strategy based on the task relationship, where the task relationship is at least used to reflect whether there is a conflict between the second task and the first task.

[0055] As an alternative implementation, to determine the target operation strategy based on the task relationship, the following steps can be taken: In the case where there is a conflict between the second task and the first task, determine that the target operation strategy includes: sending an inquiry message to the target object, and determining whether to continue executing the second task or stop executing the second task and start executing the first task based on the selection instruction feedback by the target object, where the inquiry message is used to prompt that there is a conflict between the second task and the first task and one of them needs to be selected for execution; in the case where there is no conflict between the second task and the first task, determine that the target operation strategy includes: switching the second task to run in the background, starting to execute the first task, and displaying the task execution interface corresponding to the first task.

[0056] As an alternative implementation, the method further includes: when there is no conflict and there is an association relationship between the second task and the first task, determining that the target operation strategy includes: executing the second task and the first task simultaneously, and displaying the task execution interfaces of the second task and the first task according to a preset ratio.

[0057] The above steps are the core at the decision-making and execution levels. It concerns how to integrate and execute new voice instructions based on existing tasks, ensure the smooth progress of tasks and maximize the user experience, and continuously monitor its running status, including information such as the tasks being executed, the progress of tasks, and the priorities of tasks. This step is crucial for formulating subsequent operation strategies because it provides the basic information of the current system environment and helps understand the context for executing new tasks. For example, for a computer, if the current computer system environment is in an idle state, that is, no tasks are being executed, then the received operation instructions can be directly executed. For example, for the task control execution of numerical calculation and simulation software in a computer, when the user says "Open the data analysis interface", it immediately switches to the data analysis module without any additional operations; when a new task conflicts with the ongoing task, a prompt message will be sent to the user asking whether to interrupt the current task to execute the new task or continue the current task. For example, if a long-term simulation calculation is being executed and the user requests to switch to another data set for instant analysis, the user will be reminded to make a choice to avoid accidentally interrupting important tasks. If two tasks are associated with each other but do not conflict, these two tasks can be executed simultaneously, and even their execution status can be displayed on the same interface for the user to monitor in real time. For example, when the system is rendering a complex 3D model and the user requests to open the model property panel for setting, these two tasks can be processed in parallel, and the contents of the two tasks can be displayed in a certain ratio in one interface, allowing the user to adjust the model properties while viewing the 3D model rendering, achieving efficient operation.

[0058] After obtaining the target operation strategy according to the above method, execute the target operation strategy and display the task execution interface.

[0059] After determining the appropriate target operation strategy, the strategy will be executed, and the user interface will be updated accordingly. Whether directly executing a new task, handling task conflicts, or executing associated tasks in parallel, it will ensure smooth operation execution and at the same time display the latest task execution interface, allowing the user to clearly see the operation results.

[0060] In the embodiments of the present application, by integrating speech recognition technology and voiceprint recognition technology, the present invention realizes precise and fast task control and user identity authentication. Users can start, adjust, or terminate tasks through natural language instructions. At the same time, it ensures that only authorized individuals can execute the control. After verifying the user's identity using voiceprint recognition, it intelligently analyzes and determines the corresponding operation permissions of the user, ensuring that in any task execution scenario, participants can only perform corresponding operations according to their permissions. This method presets an instruction library containing rich task instructions, covering comprehensive contents from basic commands to complex operations, and is linked to user permissions. When receiving a voice instruction, it can quickly match the corresponding task instruction, realizing a seamless conversion from voice to task operation. It can not only recognize and parse voice instructions, but also monitor the execution status of the current task in real time, and intelligently formulate subsequent operation strategies according to the task status and user instructions. This feature allows the system to respond flexibly, avoiding conflicts and interruptions during task execution, ensuring the coherence and high-efficiency execution of tasks, and thus solving the technical problems of complex user operations, low security, and low efficiency in handling task conflicts during the control process of task execution.

[0061] Embodiment 2

[0062] According to the embodiments of the present application, there is also provided a voice control-based task execution device for implementing the voice control-based task execution method in Embodiment 1, as Figure 2 shown. The voice control-based task execution device at least includes: a receiving module 21, a first analysis module 22, a second analysis module 23, an instruction determination module 24, a strategy determination module 25, and an execution module 26, where:

[0063] The receiving module 21 is configured to receive a voice control instruction input by a target object;

[0064] The first analysis module 22 is configured to analyze the voice control instruction using a pre-trained speech recognition model to obtain a target text instruction;

[0065] The second analysis module 23 is configured to analyze the voice control instruction using a pre-trained voiceprint recognition model to obtain the target identity information of the target object, and determine the target control permission corresponding to the target identity information;

[0066] The instruction determination module 24 is configured to determine a target operation instruction that matches the target control permission and the target text instruction from a preset instruction library, where multiple operation instructions and corresponding control permissions are stored in the instruction library, and at least the task to be executed is included in the operation instructions;

[0067] A policy determination module 25, configured to detect a current task execution status, and determine a target operation policy according to the current task execution status and the target operation instruction;

[0068] An execution module 26, configured to execute the target operation policy and display a task execution interface.

[0069] The functions of the various modules of the task execution device based on voice control are described below in combination with a specific implementation process.

[0070] The receiving module receives a voice control instruction input by a target object, and this process can be carried out in the following steps:

[0071] Receive a voice control signal input by the target object through a voice input device; perform noise reduction processing on the voice control signal to obtain a voice control instruction, where the noise reduction processing includes at least one of the following: dynamic noise suppression, echo cancellation.

[0072] After obtaining the voice control instruction, the first analysis module analyzes the voice control instruction using a pre-trained speech recognition model to obtain a target text instruction, and this process can be carried out in the following steps:

[0073] Use the speech recognition model to extract the Mel-frequency cepstral coefficient feature vector corresponding to the voice control instruction, and analyze the Mel-frequency cepstral coefficient feature vector to obtain an initial text instruction; use transfer learning technology to correct the context of the initial text instruction to obtain the target text instruction.

[0074] After obtaining the voice control instruction, the second analysis module analyzes the voice control instruction using a pre-trained voiceprint recognition model to obtain the target identity information of the target object, and determines the target control permission corresponding to the target identity information, and this process can be carried out in the following steps:

[0075] Use the voiceprint recognition model to extract the target voiceprint feature vector corresponding to the voice control instruction; determine the target identity information corresponding to the target voiceprint feature vector from a preset voiceprint library, where the voiceprint library stores the mapping relationship between multiple voiceprint feature vectors and multiple identity information; determine the target control permission corresponding to the target identity information from a preset permission table, where the permission table stores the control permissions corresponding to multiple identity information.

[0076] After obtaining the target text instruction and the target control permission, the instruction determination module determines a target operation instruction that matches the target control permission and the target text instruction from a preset instruction library, where the instruction library stores multiple operation instructions and the corresponding control permissions, and the operation instructions at least include the task to be executed.

[0077] As an alternative implementation, to determine a target operation instruction that matches the target control permission and the target text instruction from a preset instruction library, the following process can be adopted: respectively determine the text similarity between the target text instruction and each operation instruction in the instruction library; when the maximum text similarity is not less than the preset similarity threshold and the target control permission is not lower than the control permission of the operation instruction corresponding to the maximum text similarity, determine the operation instruction corresponding to the maximum text similarity as the target operation instruction.

[0078] As an alternative implementation, the method further includes: when the maximum text similarity is less than the preset similarity threshold, generating a first prompt message, where the first prompt message is used to prompt that the voice control instruction is unclear and needs to be re-entered; when the maximum text similarity is not less than the preset similarity threshold but the target control permission is lower than the control permission of the operation instruction corresponding to the maximum text similarity, generating a second prompt message, where the second prompt message is used to prompt insufficient control permission.

[0079] After obtaining the target operation instruction, the policy determination module detects the current task execution status and determines the target operation policy based on the current task execution status and the target operation instruction.

[0080] As an alternative implementation, to detect the current task execution status and determine the target operation policy based on the current task execution status and the target operation instruction, the following steps can be adopted: detect the current task execution status and determine the first task to be executed indicated by the target operation instruction; when there is no task being executed currently, determine the target operation policy as directly executing the first task; when a second task is being executed currently, determine the task relationship between the second task and the first task and determine the target operation policy based on the task relationship, where the task relationship is at least used to reflect whether there is a conflict between the second task and the first task.

[0081] As an alternative implementation, to determine the target operation policy based on the task relationship, the following steps can be adopted: when there is a conflict between the second task and the first task, determine that the target operation policy includes: sending an inquiry message to the target object and determining whether to continue executing the second task or stop executing the second task and start executing the first task based on the selection instruction feedback by the target object, where the inquiry message is used to prompt that there is a conflict between the second task and the first task and one of them needs to be selected for execution; when there is no conflict between the second task and the first task, determine that the target operation policy includes: switching the second task to run in the background, starting to execute the first task and displaying the task execution interface corresponding to the first task.

[0082] As an alternative implementation, the method further includes: when there is no conflict and there is an association relationship between the second task and the first task, determining that the target operation strategy includes: executing the second task and the first task simultaneously, and displaying the task execution interfaces of the second task and the first task in a preset ratio.

[0083] After obtaining the target operation strategy according to the above method, the execution module executes the target operation strategy and displays the task execution interface.

[0084] It should be noted that each module in the voice control-based task execution device in the embodiments of the present application corresponds one by one to each implementation step of the voice control-based task execution method in Embodiment 1. Since Embodiment 1 has been described in detail, some details not shown in this embodiment can be referred to Embodiment 1 and will not be elaborated here.

[0085] Embodiment 3

[0086] According to an embodiment of the present application, there is also provided a computer program product, which includes a computer program. When the computer program is executed by a processor, it implements the voice control-based task execution method in Embodiment 1.

[0087] According to an embodiment of the present application, there is also provided a non-volatile storage medium, which includes a stored computer program. The device where the non-volatile storage medium is located executes the voice control-based task execution method in Embodiment 1 by running the computer program.

[0088] According to an embodiment of the present application, there is also provided a processor, which is used to run a computer program. When the computer program runs, it executes the voice control-based task execution method in Embodiment 1.

[0089] According to an embodiment of the present application, there is also provided an electronic device, which includes: a memory and a processor. A computer program is stored in the memory, and the processor is configured to execute the voice control-based task execution method in Embodiment 1 through the computer program.

[0090] Specifically, when the computer program runs, it implements the following steps: receiving a voice control instruction input by a target object; analyzing the voice control instruction by using a pre-trained voice recognition model to obtain a target text instruction; analyzing the voice control instruction by using a pre-trained voiceprint recognition model to obtain the target identity information of the target object, and determining the target control permission corresponding to the target identity information; determining a target operation instruction matching the target control permission and the target text instruction from a preset instruction library, where multiple operation instructions and corresponding control permissions are stored in the instruction library, and at least one of the operation instructions includes a task to be executed; detecting the current task execution status, and determining a target operation strategy according to the current task execution status and the target operation instruction; executing the target operation strategy, and displaying a task execution interface.

[0091] As an optional implementation manner, the above-mentioned electronic device may exist in the form of a mobile terminal, a computer terminal, or a similar computing device. Figure 3 The hardware structure block diagram of an electronic device for implementing a task execution method based on voice control is shown. As Figure 3 shown, the electronic device 30 may include one or more (shown as 302a, 302b,..., 302n in the figure) processors 302 (the processor 302 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 304 for storing data, and a transmission device 306 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 3 the structure shown is only schematic and does not limit the structure of the above-mentioned electronic device. For example, the electronic device 30 may further include more or fewer components than those Figure 3 shown, or have a different configuration from that Figure 3 shown.

[0092] It should be noted that the above one or more processors 302 and / or other data processing circuits are generally referred to as "data processing circuits" in this article. The data processing circuit may be embodied in software, hardware, firmware, or any combination thereof in whole or in part. In addition, the data processing circuit may be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the electronic device 30. As involved in the embodiments of the present application, the data processing circuit is a processor control (such as the selection of a variable resistance terminal path connected to an interface).

[0093] The memory 304 can be used to store software programs and modules of application software, such as the program instruction / data storage device corresponding to the voice control-based task execution method in the embodiments of the present application. The processor 302 executes various functional applications and data processing by running the software programs and modules stored in the memory 304, that is, implements the vulnerability detection method of the above application program. The memory 304 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 304 may further include a memory remotely disposed relative to the processor 302, and these remote memories may be connected to the electronic device 30 through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.

[0094] The transmission device 306 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the electronic device 30. In one instance, the transmission device 306 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus can communicate with the Internet. In one instance, the transmission device 306 may be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.

[0095] The display may be, for example, a touch-screen liquid crystal display (LCD), and the liquid crystal display enables a user to interact with the user interface of the electronic device 30.

[0096] The above serial numbers of the embodiments are only for description and do not represent the advantages or disadvantages of the embodiments.

[0097] In the above embodiments of the present application, the descriptions of the respective embodiments have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.

[0098] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of units can be a logical function division, and there may be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection to each other can be through some interfaces, and the indirect coupling or communication connection of units or modules can be in an electrical or other form.

[0099] The unit described as a separation component may or may not be physically separated. The component shown as a unit may or may not be a physical unit, that is, it may be located in one place or may be distributed over multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0100] In addition, each functional unit in various embodiments of the present application can be integrated into a processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0101] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to enable a computer device (which can be a personal computer, a server or a network device, etc.) to execute all or part of the steps of the methods in various embodiments of the present application. The aforementioned storage medium includes: USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks or optical discs and other various media that can store program codes.

[0102] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.

Claims

1. A task execution method based on voice control, characterized in that: include: Receiving a voice control command input by a target object; Analyzing the voice control instruction using a pre-trained speech recognition model to obtain a target text instruction; Analyzing the voice control command using a pre-trained voiceprint recognition model to obtain target identity information of the target object, and determining the target control authority corresponding to the target identity information; Determine a target operation instruction that matches the target control authority and the target text instruction from a preset instruction library, wherein the instruction library stores a plurality of operation instructions and corresponding control authorities, and the operation instruction at least includes a task to be executed; Detecting a current task execution state, and determining a target operation strategy according to the current task execution state and the target operation instruction; Execute the target operation strategy and display the task execution interface.

2. The method according to claim 1, characterized in that Receive voice control instructions input by the target object, including: Receiving a voice control signal input by a target object through a voice input device; The voice control signal is subjected to noise reduction processing to obtain the voice control instruction, wherein the noise reduction processing includes at least one of the following: dynamic noise suppression and echo cancellation.

3. The method according to claim 1, characterized in that The voice control instruction is analyzed using a pre-trained speech recognition model to obtain a target text instruction, including: Extracting a Mel-frequency cepstral coefficient feature vector corresponding to the voice control instruction using the voice recognition model, and analyzing the Mel-frequency cepstral coefficient feature vector to obtain an initial text instruction; The initial text instruction is context-corrected using transfer learning technology to obtain a target text instruction.

4. The method according to claim 1, characterized in that: The voice control command is analyzed using a pre-trained voiceprint recognition model to obtain target identity information of the target object, and the target control authority corresponding to the target identity information is determined, including: Extracting a target voiceprint feature vector corresponding to the voice control instruction using the voiceprint recognition model; Determining target identity information corresponding to the target voiceprint feature vector from a preset voiceprint library, wherein the voiceprint library stores mapping relationships between multiple voiceprint feature vectors and multiple identity information; The target control authority corresponding to the target identity information is determined from a preset authority table, wherein the authority table stores control authorities corresponding to a plurality of identity information.

5. The method according to claim 1, characterized in that Determining a target operation instruction that matches the target control authority and the target text instruction from a preset instruction library includes: Determining the text similarity between the target text instruction and each operation instruction in the instruction library respectively; When the maximum text similarity is not less than a preset similarity threshold and the target control authority is not less than the control authority of the operation instruction corresponding to the maximum text similarity, the operation instruction corresponding to the maximum text similarity is determined to be the target operation instruction.

6. The method according to claim 5, characterized in that The method further comprises: When the maximum text similarity is less than the preset similarity threshold, generating a first prompt message, wherein the first prompt message is used to prompt that the voice control instruction is unclear and needs to be re-entered; When the maximum text similarity is not less than the preset similarity threshold, but the target control authority is lower than the control authority of the operation instruction corresponding to the maximum text similarity, a second prompt information is generated, wherein the second prompt information is used to prompt that the control authority is insufficient.

7. The method according to claim 1, characterized in that Detecting the current task execution state, and determining the target operation strategy according to the current task execution state and the target operation instruction, including: Detecting the current task execution state, and determining the first task to be executed indicated by the target operation instruction; In the case where no task is currently being executed, determining the target operation strategy to be directly executing the first task; In the case where a second task is currently being executed, a task relationship between the second task and the first task is determined, and a target operation strategy is determined based on the task relationship, wherein the task relationship is at least used to reflect whether there is a conflict between the second task and the first task.

8. The method according to claim 7, characterized in that Determine the target operation strategy based on the task relationship, including: In the case that there is a conflict between the second task and the first task, determining the target operation strategy includes: sending an inquiry message to the target object, and determining whether to continue to execute the second task or stop executing the second task and start executing the first task according to a selection instruction fed back by the target object, wherein the inquiry message is used to prompt that there is a conflict between the second task and the first task, and one of them needs to be executed; In the case that there is no conflict between the second task and the first task, determining the target operation strategy includes: switching the second task to background operation, starting to execute the first task and displaying a task execution interface corresponding to the first task.

9. The method according to claim 8, characterized in that The method further comprises: When there is no conflict and an association relationship between the second task and the first task, determining the target operation strategy includes: executing the second task and the first task simultaneously, and displaying the task execution interfaces of the second task and the first task in a preset ratio.

10. A task execution device based on voice control, characterized in that: include: A receiving module, used to receive a voice control instruction input by a target object; A first analysis module, used to analyze the voice control instruction using a pre-trained speech recognition model to obtain a target text instruction; A second analysis module is used to analyze the voice control instruction using a pre-trained voiceprint recognition model to obtain the target identity information of the target object and determine the target control authority corresponding to the target identity information; An instruction determination module, used to determine a target operation instruction that matches the target control authority and the target text instruction from a preset instruction library, wherein the instruction library stores a plurality of operation instructions and corresponding control authorities, and the operation instruction at least includes a task to be executed; A strategy determination module, used to detect the current task execution state, and determine the target operation strategy according to the current task execution state and the target operation instruction; The execution module is used to execute the target operation strategy and display the task execution interface.

11. A computer program product, characterized in that include: A computer program, wherein when the computer program is executed by a processor, the task execution method based on voice control as described in any one of claims 1 to 9 is implemented.

12. An electronic device, characterized in that: include: A memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the task execution method based on voice control according to any one of claims 1 to 9 through the computer program.

Citation Information

Cited By

  • Equipment control method and device, equipment, storage medium and program product

    CN122454975A