Voice processing method and device
By processing speech, text, and state table information in parallel, and utilizing the speech processing and decision processing models of a large language model, the problem of lengthy response links in voice interaction systems is solved, enabling real-time voice response and task planning for smart terminals, and improving the coordination and reliability of voice interaction.
Patent Information
- Application Number
- CN202511136388.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-11-04
AI Technical Summary
Existing voice interaction systems have lengthy response chains, making it difficult to meet the performance requirements of real-time interaction and respond quickly to user needs.
By inputting the voice text converted from voice commands and status table information in parallel into the voice processing model and decision processing model implemented by a large language model, it is possible to generate response voice and task planning commands in real time, and to coordinate voice interaction and task execution.
It enables instant semantic response to voice commands, meets the needs of real-time voice interaction, improves the collaboration and reliability of smart terminals, and optimizes the response efficiency of voice interaction.
Smart Images

Figure CN120895042A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of artificial intelligence, in particular to a voice processing method and device. BACKGROUND
[0002] With the rapid development of artificial intelligence and robot technology, voice interaction and task execution capability have become the core demand of intelligent terminals.
[0003] At present, many voice interaction systems are simple and single in function implementation, and can only complete some basic voice responses. For specific tasks involved behind user instructions, there is a lack of efficient processing capability. In actual operation, such systems often have a long response link, which makes it difficult to quickly respond to user needs and cannot meet the requirements of real-time interaction scenarios. SUMMARY
[0004] Therefore, the present application provides a voice processing method and device, which can significantly reduce voice response delay and meet the demand of voice real-time interaction.
[0005] The present application provides the following solutions:
[0006] In a first aspect, a voice processing method is provided, which is applied to an intelligent terminal. The method comprises:
[0007] receiving a voice instruction and converting the voice instruction into voice text;
[0008] inputting the voice text and state table information into a voice processing model and a decision processing model in parallel, wherein:
[0009] generating a first reply text using the voice processing model, and converting the first reply text into a first reply voice and outputting the first reply voice;
[0010] generating a first task planning instruction using the decision processing model; and in response to the first task planning instruction indicating that there is a task to be executed, executing the task to be executed.
[0011] The voice processing model and the decision processing model are implemented by at least one large language model, and the state table information at least includes current environment information and / or current state information of the intelligent terminal.
[0012] Optionally, the inputting of the voice text and the state table information into the voice processing model and the decision processing model in parallel comprises:
[0013] generating a dialogue prompt word adapted to the voice processing model and a planning prompt word adapted to the decision processing model based on the voice text and the state table information, respectively;
[0014] inputting the dialogue prompt word into the speech processing model and inputting the planning prompt word into the decision processing model.
[0015] Optionally, the executing the to-be-executed task in response to the first task planning instruction indicating that there is a to-be-executed task comprises:
[0016] In response to the first task planning instruction indicating that there is the to-be-executed task, a task execution sequence of the to-be-executed task is constructed, the task execution sequence comprising at least one subtask and an execution order of each of the subtasks.
[0017] The task execution sequence is written into a state table to update the state table information.
[0018] The to-be-executed task is executed based on the task execution sequence in the updated state table information.
[0019] Optionally, the constructing the task execution sequence of the to-be-executed task comprises:
[0020] A directed acyclic graph of the to-be-executed task is constructed, each node in the directed acyclic graph corresponding to a subtask, and the nodes being connected by directed edges to represent the execution order of the subtasks.
[0021] Optionally, the executing the to-be-executed task based on the task execution sequence in the updated state table information comprises:
[0022] Each of the subtasks is executed in the execution order of the subtasks in the task execution sequence until all the subtasks in the task execution sequence are completed or the execution of the task execution sequence is terminated due to the execution of any one of the subtasks failing and a preset termination condition being met.
[0023] In the process of executing the to-be-executed task, the identity and execution state of a currently executed subtask are recorded in real time, and the identity and execution state of the currently executed subtask are written into the state table as task process information to update the state table information.
[0024] Optionally, the method further comprises:
[0025] In response to the to-be-executed task being executed, the state table information containing the task process information, the first reply text and the speech text are input into the speech processing model, and a second reply text is generated by using the speech processing model;
[0026] The second reply text is converted into second reply speech and output.
[0027] Optionally, the method further comprises:
[0028] input the state table information containing the task process information, the first reply text and the speech text into the decision processing model;
[0029] In response to the execution state in the task process information being failure, regenerate second task planning instructions by using the decision processing model; the second task planning instructions are different from the first task planning instructions.
[0030] In a second aspect, a computer-readable storage medium is provided, and the computer-readable storage medium stores a computer program, and the computer program is executed to implement the steps of the method in the first aspect.
[0031] In a third aspect, an electronic device is provided, and the electronic device includes one or more processors and a memory associated with the one or more processors, and the memory is configured to store program instructions, and the program instructions are executed by the one or more processors to implement the steps of the method in the first aspect.
[0032] In a fourth aspect, a computer program product is provided, and the computer program product includes a computer program, and the computer program is executed to implement the steps of the method in the first aspect.
[0033] According to the embodiments of the present application, the following technical effects are disclosed:
[0034] In the embodiments of the present application, the speech text converted from the speech instruction and the state table information are input into the speech processing model and the decision processing model implemented by the large language model in parallel, the speech processing model can generate the reply speech in real time, and the decision processing model can output the task planning instructions and execute the to-be-executed task in real time. Compared with the prior art, the scheme can not only realize the real-time semantic response to the speech instruction and meet the real-time voice interaction demand, but also can realize the cooperative processing of the real-time semantic response to the speech instruction and the automatic task planning and execution of the intelligent terminal in the voice interaction process, thereby improving the cooperation and reliability of the voice interaction of the intelligent terminal.
[0035] Of course, implementing any product of the present application does not necessarily need to achieve all the advantages described above. BRIEF DESCRIPTION OF DRAWINGS
[0036] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the drawings needed in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description only constitute some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.
[0037] Figure 1 A speech processing method flowchart provided by the embodiments of the present application;
[0038] Figure 2 An implementation diagram of a task execution sequence provided by an embodiment of the present application;
[0039] Figure 3 A framework diagram of a voice interaction system for implementing a voice processing method provided by an embodiment of the present application;
[0040] Figure 4 A schematic block diagram of a voice processing apparatus provided by an embodiment of the present application;
[0041] Figure 5 A schematic block diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0042] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0043] The terms used in the embodiments of the present application are merely for the purpose of describing specific embodiments, and are not intended to limit the present application. The singular forms “a”, “an” and “the” used in the embodiments of the present application and the appended claims are also intended to include the plural forms, unless the context clearly indicates otherwise.
[0044] It should be understood that the term “and / or” used herein is merely to describe the association relationship of the associated objects, which means that there can be three relationships, for example, A and / or B can mean that A exists alone, A and B exist together, and B exists alone. In addition, the character “ / ” herein generally represents an “or” relationship between the front and rear associated objects.
[0045] Depending on the context, the word “if” as used herein can be interpreted as “when” or “upon” or “in response to determining” or “in response to detecting”. Similarly, depending on the context, the phrase “if it is determined” or “if (a stated condition or event) is detected” can be interpreted as “when it is determined” or “in response to determining” or “when (a stated condition or event) is detected” or “in response to detecting (a stated condition or event)”.
[0046] In the field of voice interaction and task processing of intelligent terminals, the existing technical solutions generally have fragmented functions or performance defects, which are difficult to meet the actual needs of users for comprehensive service capabilities, and the specific manifestations are as follows:
[0047] Part of the technical solutions only realize a single voice interaction function, can only complete the reception, analysis and basic dialogue reply of voice instructions, so that the intelligent terminal cannot execute the task processing corresponding to the instruction, and the application scene is obviously limited.
[0048] Another technical solution only realizes a task processing function, does not configure a voice dialogue interaction module, so that user instructions or real-time feedback of task execution status cannot be received through a natural language interface, and the convenience of human-computer interaction is insufficient.
[0049] Part of the technical solutions simply integrates voice synthesis and task processing, although both dialogue interaction and task processing capabilities are provided, but due to unreasonable collaborative mechanism design of the voice interaction module and the task decision module, the voice response link is long, the response delay is high, and the performance requirements of real-time interaction cannot be met.
[0050] Therefore, the present application provides a new voice processing method. The method can be applied to an intelligent terminal, which refers to an electronic device or intelligent system with voice interaction capability, data processing capability and task execution control function, and can specifically include intelligent robots, intelligent sound boxes, intelligent home control terminals, wearable smart devices, etc. Such terminals can receive user instructions through a voice interface, complete data analysis and decision planning based on built-in processors, memories and algorithm programs or trained machine learning models, and can directly or through a communication protocol drive an execution mechanism (such as a mechanical arm, a mobile module, etc.) or a controlled device to complete a preset task. It is an intelligent carrier integrating voice interaction, intelligent decision and task execution.
[0051] As shown in Figure 1 The flowchart of the voice processing method provided by the embodiment of the present application can include the following steps:
[0052] Step 101: receiving a voice instruction and converting the voice instruction into a voice text.
[0053] Step 102: inputting the voice text and state table information into a voice processing model and a decision processing model in parallel, wherein: generating a first reply text using the voice processing model, converting the first reply text into a first reply voice and outputting; generating a first task planning instruction using the decision processing model; in response to the first task planning instruction indicating that there is a to-be-executed task, executing the to-be-executed task.
[0054] Among them, the voice processing model and the decision processing model are realized by at least one large language model, and the state table information at least includes current environment information and / or current state information of the intelligent terminal.
[0055] As can be seen from the above process, by parallel inputting the speech text converted from the speech instruction and the state table information into the speech processing model and the decision processing model implemented by the large language model, the speech processing model can generate a reply speech in real time, and the decision processing model can output a task planning instruction in real time and execute a to-be-executed task. Compared with the prior art, the scheme can not only realize real-time semantic response to the speech instruction and meet the real-time interaction demand of the speech, but also can realize the cooperative processing of the real-time semantic response to the speech instruction and the automatic task planning and execution of the intelligent terminal in the process of the speech interaction, thereby improving the cooperation and reliability of the speech interaction of the intelligent terminal.
[0056] Next, the steps in the above process and the effects that can be further produced will be described in detail with reference to the embodiments.
[0057] First, the step 101, i.e., receiving a speech instruction and converting the speech instruction into a speech text, will be described in detail with reference to the embodiments.
[0058] The intelligent terminal can receive the speech instruction issued by the user in real time through a built-in or externally connected speech collection device (such as a microphone array). After receiving the speech instruction issued by the user, the speech instruction is recognized and processed to be converted into a text form, and a speech text corresponding to the speech instruction is obtained.
[0059] Optionally, in the embodiments of the present application, the speech instruction can be converted and processed by calling a speech-to-text (STT) technology to obtain the speech text corresponding to the speech instruction. For example, a trained STT model (such as an end-to-end speech recognition model based on deep learning) is deployed on the intelligent terminal to directly convert the speech instruction into a speech text on the device side. For another example, by calling a speech recognition API provided by a third party, the collected speech instruction is uploaded to a cloud server, and the conversion of the speech instruction into a speech text is completed by the STT service of the cloud to return the result.
[0060] Next, the step 102, i.e., parallel inputting the speech text and the state table information into the speech processing model and the decision processing model, will be described in detail with reference to the embodiments.
[0061] In the embodiments of the present application, after obtaining the speech text corresponding to the speech instruction, the intelligent terminal inputs the speech text and the state table information into the speech processing model and the decision processing model through a parallel processing mechanism. The speech processing model and the decision processing model can independently start the operation process, realize the collaborative promotion of the speech interaction reply and the task planning execution, and therefore there is no time delay caused by waiting for the completion of the processing of the opposite model, so as to completely break the efficiency bottleneck of the traditional serial mode that the generation of the speech reply needs to wait for the completion of the task planning or the task planning needs to wait for the end of the speech analysis.
[0062] In the embodiments of the present application, the state table information refers to the information in the state table, which is a structured data set that is collected and dynamically updated by the intelligent terminal in real time through sensors, device interfaces and internal storage modules, and at least includes current environmental information and / or current state information of the intelligent terminal.
[0063] The current environmental information refers to real-time environmental parameters of the physical space where the intelligent terminal is located, which can include but is not limited to:
[0064] Environmental perception data: such as temperature, humidity, light intensity and other data collected through built-in sensors (temperature and humidity sensors, light sensors);
[0065] Spatial positioning data: such as real-time position coordinates of the intelligent terminal (which can be obtained based on an indoor positioning system or a GPS module), relative position of the user (which can be determined through millimeter wave radar or camera visual recognition), etc.
[0066] The current state information refers to the running state data of the intelligent terminal itself, which can include but is not limited to:
[0067] Device running parameters: such as battery level, network connection status of the intelligent terminal, etc.
[0068] State of the execution mechanism: such as joint running state, motor running state, joint health degree, etc.
[0069] When the speech text and the state table information are input into the speech processing model and the decision processing model in parallel, the state table information can provide the scene context for the speech processing model, ensure that the reply generated by the speech processing model is consistent with the actual ability of the intelligent terminal and the current environment, and thereby avoid generating a reply that cannot be implemented; and can also provide the basis for decision-making for the decision processing model, so that it can judge the task feasibility and generate corresponding task planning instructions based on the constraints such as battery level and joint health degree.
[0070] Specifically, when the speech text and the state table information are input into the speech processing model and the decision processing model in parallel, the dialog prompt words adapted to the speech processing model and the planning prompt words adapted to the decision processing model can be generated based on the speech text and the state table information, respectively, and then the two types of prompt words are transmitted to the corresponding models. This differentiated prompt word method can focus the speech processing model on natural language reply generation and the decision processing model on task planning logic, thereby reducing the understanding cost of the model for input information, improving the naturalness of the reply text and the accuracy of the task instruction, and the standardized prompt word format can also reduce the model operation error rate, further optimizing the efficiency of parallel processing.
[0071] After the speech processing model receives the dialog prompt words, the first reply text is generated, and the intelligent terminal converts the first reply text into the first reply speech and outputs it in real time through a loudspeaker or other audio output device.
[0072] Alternatively, in the embodiments of the present application, the first reply text can be converted by text-to-speech (TTS) technology to obtain the first reply speech. For example, an end-to-end deep neural network model is used to directly convert the first reply text into a speech waveform to obtain the first reply speech. For another example, a large number of speech primitives are pre-recorded, and a corresponding acoustic parameter library is established. When converting, the input first reply text is first segmented, phonetized and prosodic analyzed, then the matching primitives are selected from the speech library for splicing and smoothing processing, and finally the continuous speech is generated to obtain the first reply speech.
[0073] After the decision processing model receives the planning prompt words, the first task planning instruction is generated based on the planning prompt words, and when the first task planning instruction indicates that there is a task to be executed, the task to be executed is executed.
[0074] Specifically, after receiving the planning prompt word, the decision processing model analyzes whether there is a to-be-executed task based on the planning prompt word, and generates a first task planning instruction based on the analysis result. When the first task planning instruction indicates that there is a to-be-executed task, a task execution sequence of the to-be-executed task can be first constructed, the task execution sequence including at least one subtask and an execution order of the subtasks, and then the task execution sequence is written into a state table to update the state table information, and the to-be-executed task is executed based on the task execution sequence in the updated state table information. The subtask can be obtained by decomposing the to-be-executed task, and when the to-be-executed task is a simple task, the subtask obtained by decomposing can be one, and when the to-be-executed task is a complex task, the subtask obtained by decomposing can be multiple. For example, if the voice instruction of the user is "give me water", the to-be-executed task can be decomposed into the following five subtasks: subtask 1 (stand up), subtask 2 (walk to the table with water), subtask 3 (take water from the table), subtask 4 (walk to the person), and subtask 5 (hand the water to the person). For example, if the voice instruction of the user is "what is the temperature today?", the to-be-executed task can be decomposed into one subtask (query the temperature of the location).
[0075] When the to-be-executed task is executed based on the task execution sequence in the updated state table information, each subtask can be executed in turn according to the execution order of the subtasks in the task execution sequence, until all subtasks in the task execution sequence are completed, or the execution of the task execution sequence is terminated due to the failure of execution of any one subtask and the satisfaction of a preset termination condition. The preset termination condition can be flexibly set according to the type of the task. For example, for a task such as "give me water" that relies on continuous actions, if subtask 2 (walk to the table with water) fails (for example, due to an obstacle that cannot reach the target position), and after retrying for a preset number of times, it is still unsuccessful, the termination condition is triggered, and the execution of the subsequent subtasks is stopped. For a parallel task such as "turn off the light in the study and pull up the curtain" that can be executed independently, if the subtask of "turn off the light in the study" fails, but the subtask of "pull up the curtain" is not affected, the preset termination condition can be set to terminate only the execution of the failed subtask, and continue to execute the other subtasks.
[0076] In addition, in the process of executing the to-be-executed task, the identifier and execution state of the currently executed subtask can also be recorded in real time, and the identifier and execution state of the currently executed subtask are written as task process information into the state table to update the state table information. The execution state can include but is not limited to success, in progress, and failure. For example, when executing the task of "give me water", when the intelligent terminal starts to execute the subtask 2 (walk to the table with water), the "task process information" field in the state table is updated to "current subtask identifier: subtask 2; execution state: in progress". If the subtask 2 is executed successfully, the "task process information" field is updated to "current subtask identifier: subtask 2; execution state: success". If the execution fails, the "task process information" field is updated to "current subtask identifier: subtask 2; execution state: failure (reason: obstacle blocking)". Through this real-time updating mechanism, the state table can dynamically reflect the task execution progress, and provide data support for subsequent possible task adjustment or user query.
[0077] In the embodiments of the present application, the representation form of the task execution sequence has diversity, which can be a directed acyclic graph, or a list or a queue or other structured data format. Taking the directed acyclic graph as an example, the task execution sequence of the to-be-executed task is constructed, that is, the directed acyclic graph of the to-be-executed task is constructed, and each node in the directed acyclic graph corresponds to a subtask. The nodes are connected by directed edges to clearly indicate the execution order of the subtasks. As shown in the following figure, it is a directed acyclic graph corresponding to the voice instruction "give me water". Figure 2
[0078] When the first task planning instruction generated by the decision processing model indicates that there is no to-be-executed task, the task execution process is not triggered, and only the reply voice generated by the voice processing model is used to feed back the instruction processing result to the user. For example, when the user inputs the voice instruction "who am I", the corresponding voice text "who am I" is obtained through step 101, and the planning prompt word is generated in combination with the state table information. After the planning prompt word is input into the decision processing model, the decision processing model learns that the instruction belongs to pure question and answer interaction, and does not need to call the device to execute the task, so the first task planning instruction indicates that there is no to-be-executed task. At this time, the intelligent terminal does not start the task execution process, and only the first reply text is generated by the voice processing model based on the dialogue prompt word, and the first reply voice is converted by the TTS model and output, to complete the pure dialogue interaction process.
[0079] Further, when the to-be-executed task is executed, the state table information containing the task process information, the first reply text and the voice text can be input into the voice processing model again, the second reply text is generated by using the voice processing model, and the intelligent terminal converts the second reply text into the second reply voice and outputs. The execution completion here includes two cases: all sub-tasks in the task execution sequence are successfully completed, and the sub-task execution fails and meets the preset termination condition. For example, in the task of "give me water", when all sub-tasks are successfully completed, the task process information in the state table information records the successful execution state of each sub-task, at this time, the state table information, the first reply text "OK, I will get the water for you" and the voice text "give me water" are input into the voice processing model, and the voice processing model generates the second reply text "give you water, please hold it well"; if sub-task 2 fails and triggers the termination condition, the state table information records the failure reason, and the voice processing model generates the second reply text "I'm sorry, there is an obstacle in front of me, please wait" in combination with the information. After generating the second reply text, the intelligent terminal converts it into the second reply voice by calling the TTS model, and outputs it through the loudspeaker, so that the user can know the execution result of the task in time, which not only enhances the user's perception of the task execution state, but also improves the integrity of the voice interaction and the user experience.
[0080] Further, in the embodiments of the present application, the state table information containing the task process information, the first reply text and the voice text can be input into the decision processing model, when the execution state in the task process information is failure, the second task planning instruction is regenerated by using the decision processing model, and the second task planning instruction is different from the first task planning instruction. Specifically, when the task process information shows that there is a sub-task execution failure, the decision processing model will analyze these information, find out the failure reason and formulate a new task execution strategy. For example, in the task of "give me water", if sub-task 2 fails due to the obstacle, after receiving the state table information containing the failure information, the first reply text "OK, I will get the water for you" and the voice text "give me water", the decision processing model will re-plan the path, and the generated second task planning instruction can contain a new sub-task sequence: sub-task 1 (stand up), sub-task 2 (go around the obstacle and walk to the table with water), sub-task 3 (take water from the table), sub-task 4 (walk to the person), sub-task 5 (hand the water to the person). The second task planning instruction is different from the first task planning instruction by adjusting the execution parameters of the sub-tasks or changing the execution order of the sub-tasks, so as to improve the success rate of task execution. When the execution state in the task process information is success, the decision processing model does not generate the task planning instruction.
[0081] Furthermore, it should be noted that the speech processing model and decision processing model in this application embodiment can be implemented by at least one large language model, which can include the following two architectures: First, using different fine-tuned instances of the same basic large language model, injecting specialized training data for dialogue interaction and task planning into the same pre-trained model to form functionally differentiated model instances; second, using independent versions of large language models optimized for interaction scenarios and decision-making scenarios respectively, such as using a dialogue-type large model that emphasizes natural language fluency for the speech processing model, and a planning-type large model that emphasizes logical reasoning for the decision processing model. In both architectures, the speech processing model is fine-tuned for dialogue interaction tasks, focusing on accurately understanding the user's semantic intent and generating response text that conforms to spoken language habits; while the decision processing model is fine-tuned for task planning scenarios, focusing on breaking down user needs into tasks to be executed. The prompt word generation process of both is based on the same speech text and state table information, but the generated prompt words differ due to different task objectives. After the generated prompt words are input into the two models simultaneously, the calculation process proceeds in parallel, fundamentally shortening the overall response cycle from user command input to intelligent terminal feedback output, and meeting the performance requirements of millisecond-level real-time interaction.
[0082] The following describes the implementation process of the speech processing method provided in the embodiments of this application in practical applications.
[0083] like Figure 3 The diagram shows a framework of a voice interaction system used to implement voice processing methods. The voice interaction system consists of a STT module, a decision LLM (Large Language Model), a speech LLM, a state table module, a task construction module, a TTS module, an executor, a MCP (Mission Control and Processing) module, and an audio stream processing module working together. The speech LLM corresponds to the aforementioned voice processing model, and the decision LLM corresponds to the aforementioned decision processing model.
[0084] Specifically, when a user inputs a voice command, the STT module first converts the voice command into speech-to-text. After the speech-to-text is generated, parallel voice response paths and task planning paths are triggered:
[0085] 1) Voice Response Path: Based on the voice text and the status table information stored in the status table module, voice prompts adapted to the voice LLM are generated. After the voice prompt is input into the voice LLM, the voice LLM generates the first response text. Then, the TTS module converts the first response text into the first response voice, and then outputs it through the audio stream processing module, thereby realizing real-time interaction of voice commands.
[0086] 2) Task planning path: based on the voice text and state table information, generate a planning prompt word for the adaptive decision LLM. After the planning prompt word is input into the decision LLM, the decision LLM generates a first task planning instruction. If the first task planning instruction indicates that there is a task to be executed, the task construction module will decompose the task to be executed into a subtask sequence, construct a directed acyclic graph, and write the directed acyclic graph into the state table module to update the state table information.
[0087] The executor will monitor the state table module in real time. When it is monitored that the directed acyclic graph of the task to be executed is written, it will be executed in turn according to the subtask order in the directed acyclic graph, and the subtask identifier and execution state will be recorded in real time. The above-mentioned subtask identifier and execution state are written back to the state table module as task process information, and the state table is continuously updated. When the executor executes the task to be executed, it can interact with the underlying hardware or external devices through the MCP module.
[0088] After the task to be executed is executed, the closed-loop feedback process is triggered, and the dual model is called again in parallel:
[0089] 1) Voice LLM path: the updated state table information (including task process information), the first reply text and the original voice text are input into the voice LLM again. The voice LLM generates a second reply text. After conversion by the TTS module and output by the audio stream processing module, the user is fed back the task result, and a second voice reply is realized.
[0090] 2) Decision LLM path: the updated state table information, the first reply text and the original voice text are also input into the decision LLM. If the task process information indicates that the subtask execution fails, the decision LLM will further analyze the failure root cause, generate a second task planning instruction different from the first task planning instruction, and then trigger the task construction module to update the directed acyclic graph, so that the executor re-executes the task according to the updated directed acyclic graph, thereby realizing dynamic error correction and adaptive execution. If the task process information indicates that the execution is successful, the decision LLM determines that there is no need to generate a new instruction, maintains the state table information, and waits for the next round of user interaction.
[0091] The following will take no task output, simple task output and complex task output as examples to introduce the voice processing method proposed in the embodiments of the present application.
[0092] Embodiment one:
[0093] The voice instruction "Who are you?" of the user is received, and the voice instruction is converted into a voice text by the STT module. Then the voice LLM path and the decision LLM path are triggered in parallel.
[0094] Voice LLM path:
[0095] The voice text and the state table information are coupled into a voice prompt word, and input into the voice LLM. The voice LLM generates a first reply text "I am XXX robot", and converts the first reply text into a first reply voice through a TTS module and outputs.
[0096] Decision LLM path:
[0097] The voice text and the state table information are coupled into a planning prompt word, and input into the decision LLM.
[0098] The decision LLM outputs no planning task after analysis.
[0099] The flow ends.
[0100] Embodiment two:
[0101] A voice instruction "What is the local temperature" of a user is received, and the voice instruction is converted into a voice text through an STT module, and then a voice LLM path and a decision LLM path are triggered in parallel.
[0102] Voice LLM path:
[0103] The voice text and the state table information are coupled into a voice prompt word, and input into the voice LLM. The voice LLM generates a first reply text "The local temperature seems to be 25℃, let me help you check again", and converts the first reply text into a first reply voice through a TTS module and outputs.
[0104] Decision LLM path:
[0105] The voice text and the state table information are coupled into a planning prompt word, and input into the decision LLM.
[0106] The decision LLM outputs a first task planning instruction after analysis, which instructs a to-be-executed task "query local temperature".
[0107] A task construction module constructs the to-be-executed task "query local temperature" into a directed acyclic graph, which includes a node corresponding to "query local temperature".
[0108] An executor calls a MCP module to query the local temperature according to the directed acyclic graph, and when it is found that the local temperature is 26℃, the state table information is updated, and the updated state table information, the first reply text and the voice text are input into the voice LLM and the decision LLM again in parallel.
[0109] The voice LLM outputs a second reply text "I seem to have made a mistake, according to the query, the local temperature today is 26℃", and converts the second reply text into a second reply voice through a TTS module and outputs.
[0110] The decision LMM determines that no new task planning instruction needs to be generated.
[0111] The flow ends.
[0112] Embodiment three:
[0113] A voice instruction "give me water" of a user is received, and the voice instruction is converted into voice text through an STT module, and then a voice LLM path and a decision LLM path are triggered in parallel.
[0114] Voice LLM path:
[0115] The voice text and the state table information are coupled into a voice prompt word, and input into the voice LLM.
[0116] The voice LLM generates a first reply text "OK, I will help you get a cup of water", and converts the first reply text into a first reply voice through a TTS module and outputs.
[0117] Decision LLM path:
[0118] The voice text and the state table information are coupled into a planning prompt word, and input into the decision LLM.
[0119] The decision LLM analyzes and outputs a first task planning instruction, which indicates the to-be-executed tasks "stand up", "walk to the table with water", "get water from the table", "walk in front of the person", and "hand the water to the person".
[0120] A task construction module constructs the to-be-executed tasks into a directed acyclic graph, which includes five nodes, each of which corresponds to a subtask.
[0121] An executor calls a MCP module according to the directed acyclic graph to execute the subtasks in sequence, updates the state table information, and inputs the updated state table information, the first reply text and the voice text into the voice LLM and the decision LLM again in parallel.
[0122] The voice LLM outputs a second reply text "give you water, please hold it well", and converts the second reply text into a second reply voice through a TTS module and outputs.
[0123] The decision LLM determines that there is no need to generate a new task planning instruction.
[0124] The flow ends.
[0125] By using the voice processing method provided in the embodiments of the present application, through the core mechanism of parallel inputting of voice text and state table information into double models, millisecond-level voice interaction response can be realized.
[0126] The above describes particular embodiments of the present specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims can be performed in a different order than the order in which they are recited and still achieve desirable results. In addition, the processes depicted in the accompanying figures do not necessarily require the particular order shown, or sequential order, to achieve the desired results. In certain implementations, multitasking and parallel processing can be advantageous or necessary.
[0127] According to another aspect, embodiments provide a voice processing apparatus. Figure 4 A schematic block diagram of an apparatus provided by embodiments of the present application can be provided in a smart terminal. As shown in the figure, Figure 4 The apparatus 400 mainly includes an instruction receiving unit 401, an input unit 402, a voice processing unit 403, and a decision processing unit 404. The main functions of each component unit are as follows:
[0128] The instruction receiving unit 401 is configured to receive a voice instruction and convert the voice instruction into voice text;
[0129] The input unit 402 is configured to input the voice text and state table information into a voice processing model and a decision processing model in parallel.
[0130] The voice processing unit 403 is configured to generate a first reply text using the voice processing model and convert the first reply text into a first reply voice and output it;
[0131] The decision processing unit 404 is configured to generate a first task planning instruction using the decision processing model, and in response to the first task planning instruction indicating that there is a task to be executed, execute the task to be executed.
[0132] The voice processing model and the decision processing model are implemented by at least one large language model, and the state table information at least includes current environment information and / or current state information of the smart terminal.
[0133] Optionally, the input unit 402 is specifically configured to:
[0134] Generate a dialogue prompt word adapted to the voice processing model and a planning prompt word adapted to the decision processing model based on the voice text and the state table information, respectively;
[0135] Input the dialogue prompt word into the voice processing model and input the planning prompt word into the decision processing model.
[0136] Optionally, the decision processing unit 404 is specifically configured to:
[0137] In response to the first task planning instruction indicating that the to-be-executed task exists, a task execution sequence of the to-be-executed task is constructed, the task execution sequence comprising at least one subtask and an execution order of each of the subtasks;
[0138] The task execution sequence is written into a state table to update the state table information;
[0139] The to-be-executed task is executed based on the task execution sequence in the updated state table information.
[0140] Optionally, the decision processing unit 404, which is specifically configured to construct the task execution sequence of the to-be-executed task, comprises:
[0141] A directed acyclic graph of the to-be-executed task is constructed, each node in the directed acyclic graph corresponding to a subtask, and the nodes being connected by directed edges to represent the execution order of the subtasks.
[0142] Optionally, the decision processing unit 404, which is specifically configured to execute the to-be-executed task based on the task execution sequence in the updated state table information, comprises:
[0143] Each of the subtasks is executed in the execution order of the subtasks in the task execution sequence until all the subtasks in the task execution sequence are completed or the execution of the task execution sequence is terminated due to the execution failure of any one of the subtasks and the satisfaction of a preset termination condition;
[0144] In the process of executing the to-be-executed task, the identification and execution state of the currently executed subtask are recorded in real time, and the identification and execution state of the currently executed subtask are written into the state table as task process information to update the state table information.
[0145] Optionally, the input unit 402 is further configured to:
[0146] In response to the completion of the to-be-executed task, the state table information containing the task process information, the first reply text and the voice text are input into the speech processing model;
[0147] The speech processing unit 403 is further configured to generate a second reply text by using the speech processing model, convert the second reply text into second reply speech, and output the second reply speech.
[0148] Optionally, the input unit 402 is further configured to:
[0149] The state table information containing the task process information, the first reply text and the voice text are input into the decision processing model;
[0150] Decision processing unit 404 is also configured as follows:
[0151] In response to the execution status being deemed a failure in the task process information, a second task planning instruction is regenerated using the decision processing model; the second task planning instruction is different from the first task planning instruction.
[0152] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or device embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the description of the method embodiments. The system and device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.
[0153] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation portals are provided for users to choose to authorize or refuse.
[0154] In addition, embodiments of this application also provide a computer-readable storage medium storing a computer program thereon, which, when executed, implements the steps of the method described in any of the foregoing method embodiments.
[0155] And an electronic device, comprising:
[0156] One or more processors; and
[0157] A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method described in any of the foregoing method embodiments.
[0158] This application also provides a computer program product, including a computer program that, when executed, implements the steps of the method described in any of the foregoing method embodiments.
[0159] in,Figure 5 The schematic block diagram of the electronic device provided in the embodiments of the present application can specifically include a processor 510, a video display adapter 511, a disk drive 512, an input / output interface 513, a network interface 514, and a memory 520. The processor 510, the video display adapter 511, the disk drive 512, the input / output interface 513, the network interface 514, and the memory 520 can be communicatively connected through a communication bus 530. The input / output interface 513 can also be referred to as an I / O interface 513.
[0160] The processor 510 can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit (ASIC), or one or more integrated circuits, and is configured to execute related programs to implement the technical solutions provided in the present application.
[0161] The memory 520 can be implemented in the form of a ROM (Read Only Memory), a RAM (Random Access Memory), a static storage device, a dynamic storage device, or the like. The memory 520 can store an operating system 521 for controlling the operation of the electronic device 500, a basic input / output system (BIOS) 522 for controlling the low-level operation of the electronic device 500. In addition, a web browser 523, a data storage management system 524, and a voice processing apparatus 400, and the like can also be stored. The voice processing apparatus 400 can be an application program for implementing the foregoing steps in the embodiments of the present application. In summary, when the technical solutions provided in the present application are implemented by software or firmware, the related program codes are stored in the memory 520 and executed by the processor 510.
[0162] The input / output interface 513 is configured to connect an input / output module to implement information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. The input device can include a keyboard, a mouse, a touch screen, a microphone, various sensors, and the like, and the output device can include a display, a speaker, a vibrator, an indicator, and the like.
[0163] The network interface 514 is configured to connect a communication module (not shown in the figure) to implement communication interaction between the device and other devices. The communication module can realize communication through a wired manner (for example, a USB, a network cable, or the like) or a wireless manner (for example, a mobile network, WIFI, Bluetooth, or the like).
[0164] Bus 530 includes a path for transmitting information among the various components of the device (e.g., processor 510, video display adapter 511, disk drive 512, input / output interface 513, network interface 514, and memory 520).
[0165] It should be noted that although the above device only shows the processor 510, video display adapter 511, disk drive 512, input / output interface 513, network interface 514, memory 520, bus 530, etc., but in the process of implementation, the device can also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device can also contain only the components necessary to implement the scheme of the present application, and does not have to contain all the components shown in the figure.
[0166] From the above description of the embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software and the necessary general hardware platform. Based on such understanding, the technical solutions of the present application can be embodied in the form of a computer program product, which can be stored in a storage medium such as ROM / RAM, disk, optical disk, etc., including a number of instructions for making a computer device (which can be a personal computer, server, or network device, etc.) execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0167] The above provides a detailed introduction to the technical solutions of the present application, and the specific examples are applied to explain the principles and implementation modes of the present application. The above description of the embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those skilled in the art, according to the idea of the present application, the specific implementation mode and application range will be changed. In conclusion, the content of the present application should not be understood as a limitation.
Claims
1. A speech processing method, characterized in that, The method is applied in a smart terminal, and the method includes: Receive voice commands and convert the voice commands into voice text; The spoken text and state table information are input in parallel into the speech processing model and the decision processing model, wherein: The speech processing model is used to generate a first response text, which is then converted into a first response speech and output. The decision processing model is used to generate a first task planning instruction; in response to the first task planning instruction indicating the existence of a task to be executed, the task to be executed is executed; The speech processing model and the decision processing model are implemented by at least one large language model, and the state table information includes at least the current environment information and / or current state information of the smart terminal.
2. The method according to claim 1, characterized in that, The parallel input of speech text and state table information into the speech processing model and decision processing model includes: Based on the spoken text and state table information, dialogue prompts adapted to the spoken text model and planning prompts adapted to the decision processing model are generated respectively. The dialogue prompts are input into the speech processing model, and the planning prompts are input into the decision processing model.
3. The method according to claim 1, characterized in that, The step of responding to the first task planning instruction indicating the existence of a task to be executed and executing the task to be executed includes: In response to the first task planning instruction indicating the existence of the task to be executed, a task execution sequence of the task to be executed is constructed, the task execution sequence including at least one subtask and the execution order of each subtask; The task execution sequence is written into the status table to update the status table information; The task to be executed is executed based on the task execution sequence in the updated status table information.
4. The method according to claim 3, characterized in that, The process of constructing the task execution sequence for the task to be executed includes: Construct a directed acyclic graph of the tasks to be executed, where each node in the directed acyclic graph corresponds to a subtask, and the nodes are connected by directed edges to represent the execution order of the subtasks.
5. The method according to claim 3, characterized in that, The step of executing the task to be executed based on the task execution sequence in the updated status table information includes: The subtasks in the task execution sequence are executed sequentially according to their execution order until all subtasks in the task execution sequence are completed, or the task execution sequence is terminated because any subtask fails to execute and meets the preset termination condition. During the execution of the task to be executed, the identifier and execution status of the currently executed subtask are recorded in real time, and the identifier and execution status of the currently executed subtask are written into the status table as task process information to update the status table information.
6. The method according to any one of claims 1 to 5, characterized in that, The method further includes: In response to the completion of the task to be executed, the status table information containing the task process information, the first reply text and the voice text are input into the voice processing model, and the second reply text is generated using the voice processing model; Convert the second reply text into a second reply voice and output it.
7. The method according to claim 5, characterized in that, The method further includes: The status table information containing the task process information, the first response text, and the voice text are input into the decision processing model; In response to the execution status being deemed a failure in the task process information, a second task planning instruction is regenerated using the decision processing model; the second task planning instruction is different from the first task planning instruction.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed, it implements the steps of the method according to any one of claims 1 to 7.
9. An electronic device, characterized in that, include: One or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions that, when read and executed by the one or more processors, perform the steps of the method according to any one of claims 1 to 7.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Task planning method, device and equipment of robot and storage medium
CN117207198A
Intelligent indexing method based on large language model
CN118132669A
Sweeping robot, voice interaction method and device thereof and storage medium
CN118924197A
Vehicle voice processing method and device, electronic equipment and storage medium
CN119170011A
Method and system for controlling intelligent robot based on large cloud model, robot, readable medium and program product
CN119739172A