Method, server and system for processing voice interaction task based on arrangement

By adopting orchestration-based methods in the voice interaction system, using large language models to orchestrate and execute task processing operations, the problem of poor non-standard instruction processing in the prior art is solved, and a more natural and flexible interactive experience is achieved.

CN120071930APending Publication Date: 2025-05-30TIANJIN YAXIN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510250533.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-04
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

Existing voice interaction technologies are not effective when processing non-standard instructions, resulting in a stiff interactive experience.

Method used

The orchestration-based voice interaction task processing method is adopted, and a large language model is used to arrange and execute task processing operations based on control instructions and tool lists, generate processing result text and convert it into audio to send to the client.

Benefits of technology

It realizes flexible processing of non-standard instructions, provides a more natural and adaptable interaction method, and improves user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071930A_ABST
    Figure CN120071930A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a voice interaction task processing method based on arrangement, a server side and a system, and relates to the technical field of artificial intelligence. The method applied to a server comprises the following steps: receiving a control instruction sent by a client; and arranging and executing at least one round of task processing operation corresponding to the control instruction by utilizing a preset large language model according to the control instruction and the tool list, obtaining a processing result text of the control instruction according to an execution result text of the at least one round of task processing operation, converting the processing result text into audio, and sending the audio to the client, the method comprises the following steps: sending an audio to a client, enabling the client to output voice according to the audio, understanding the intention of a control instruction by utilizing a preset large language model, realizing arrangement of at least one round of task processing operation corresponding to the control instruction based on each tool in a tool list, and executing the task processing operation by using the corresponding tool, a user does not need to follow a preset template or format to send an instruction, and a more flexible, natural and highly adaptive interaction mode is provided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology. Specifically, the present disclosure relates to a method, a server, and a system for processing voice interaction tasks based on orchestration. Background Art

[0002] Voice interaction technology is a cutting-edge technology that has developed rapidly in recent years. It refers to using sound signals as the input and output media to achieve interaction between humans and computers, thereby realizing information transmission and task execution. It utilizes technologies such as speech recognition, speech synthesis, and natural language processing to enable computers to understand and generate natural language, and then communicate with humans, allowing users to interact with devices or systems through voice commands.

[0003] Currently, existing voice interaction technologies usually rely on certain predefined templates to process users' voice instructions, directly mapping them to fixed processing flows according to the templates. Although they can recognize users' intentions based on natural language understanding ability to improve a certain degree of flexibility, the processing effect for non-standard instructions is still not good, and the voice interaction experience is rigid. Summary of the Invention

[0004] Embodiments of the present disclosure provide a method, a server, and a system for processing voice interaction tasks based on orchestration, which are used to solve the technical problem that the voice interaction system in the prior art has a poor effect in processing non-standard instructions and a rigid voice interaction experience.

[0005] According to one aspect of the embodiments of the present disclosure, there is provided a method for processing voice interaction tasks based on orchestration, which is applied to a server. The server includes a tool list, and the tool list includes at least one tool. The method includes: Receiving a control instruction sent by a client; Using a preset large language model to orchestrate and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list, and obtaining a processing result text of the control instruction based on the execution result text of the at least one round of task processing operations; Sending the processing result text to the client so that the client outputs voice according to the processing result text; Wherein, each round of task processing operation is executed based on one tool.

[0006] According to another aspect of the embodiments of the present disclosure, there is provided a server. The server includes a tool list, and the tool list includes at least one tool. The server includes: An instruction receiving module, configured to receive a control instruction sent by a client; A task scheduling module, which is used to utilize a preset large language model to schedule and execute at least one round of task processing operations corresponding to a control instruction according to the control instruction and a tool list, and obtain a processing result text of the control instruction according to the execution result text of at least one round of task processing operations; A result output module, which is used to send the processing result text to a client so that the client outputs voice according to the processing result text; Wherein, each round of task processing operations is executed based on a tool.

[0007] According to another aspect of the embodiments of the present disclosure, a voice interaction task processing system based on scheduling is provided, including a server and a client; Wherein, when the server executes, it implements the steps of the voice interaction task processing method provided by the server in any of the above embodiments; The client is used to send a control instruction to the server, and output voice according to the audio after receiving the audio sent by the server.

[0008] According to another aspect of the embodiments of the present disclosure, an electronic device is provided. The electronic device includes a memory, a processor, and a computer program stored on the memory. The processor executes the computer program to implement the steps of the voice interaction task processing method provided by any of the above embodiments.

[0009] According to still another aspect of the embodiments of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, it implements the steps of the voice interaction task processing method provided by any of the above embodiments.

[0010] According to one aspect of the embodiments of the present disclosure, a computer program product is provided, including a computer program. When the computer program is executed by a processor, it implements the steps of the voice interaction task processing method provided by any of the above embodiments.

[0011] The beneficial effects brought by the technical solutions provided by the embodiments of the present disclosure are: After the server obtains the control instruction sent by the client, it utilizes the large language model LLM to understand the intention of the control instruction and, based on the functions that each tool in the tool list can achieve, schedules and executes at least one round of task processing operations, and obtains the processing result text of the control instruction according to the execution result text of at least one round of task processing operations, and converts the processing result text into audio and sends it to the client for the client to output voice. It can utilize the natural language understanding ability of the large language model to realize the scheduling and serial execution of the task processing operations corresponding to the control instruction under the constraint conditions of the tool functions that the server can provide, without requiring the user to issue instructions in a preset template or format, providing a more flexible, natural and adaptable interaction method. Brief Description of the Drawings

[0012] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments of the present disclosure.

[0013] Figure 1 One of the schematic structural diagrams of a voice interaction system provided by an embodiment of the present disclosure; Figure 2 Schematic diagram of the task processing operation arrangement and execution process provided by an embodiment of the present disclosure; Figure 3 Schematic diagram of the special instruction processing method flow provided by an embodiment of the present disclosure; Figure 4 Schematic diagram of a method flow for processing a text to be processed provided by an embodiment of the present disclosure; Figure 5 Schematic diagram of the structure of a server provided by an embodiment of the present disclosure; Figure 6 Schematic diagram of the structure of an electronic device provided by an embodiment of the present disclosure. Detailed Description of the Embodiments

[0014] The following describes the embodiments of the present disclosure with reference to the drawings in the present disclosure. It should be understood that the embodiments described below in conjunction with the drawings are exemplary descriptions for explaining the technical solutions of the embodiments of the present disclosure, and do not limit the technical solutions of the embodiments of the present disclosure.

[0015] Those skilled in the art of the present technology can understand that unless specifically stated, the singular forms "a", "an", "the" and "said" used herein may also include the plural forms. It should be further understood that the terms "including" and "comprising" used in the embodiments of the present disclosure mean that the corresponding features can be implemented as the presented features, information, data, steps, operations, elements and / or components, but do not exclude the implementation of other features, information, data, steps, operations, elements, components and / or their combinations supported by the art of the present technology. It should be understood that when we say an element is "connected" or "coupled" to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. In addition, the "connection" or "coupling" used here can include wireless connection or wireless coupling. The term "and / or" used here indicates at least one of the items defined by the term, for example, "A and / or B" can be implemented as "A", or implemented as "B", or implemented as "A and B".

[0016] To make the purpose, technical solutions and advantages of the present disclosure clearer, the following will further describe the embodiments of the present disclosure in detail with reference to the drawings.

[0017] The following describes several exemplary embodiments to illustrate the technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure. It should be noted that the following embodiments can refer to, draw on or combine with each other, and the same terms, similar features and similar implementation steps in different embodiments will not be described repeatedly.

[0018] The following is an introduction and explanation of the related technologies involved in this disclosure: The current intelligent voice interaction methods, based on the capabilities of currently common smart speakers, usually use the following related technologies: First, the end side (i.e., smart speaker) collects voice information to obtain the user's voice information. After obtaining the voice information on the end side, the platform uses the Application Programming Interface (API) to call the Automatic Speech Recognition (ASR) service for voice recognition, converts the user's voice information into text form to obtain text information, matches the obtained text information with the corresponding templates of each pre-set control instruction, identifies the user's intention based on template matching, determines the target control instruction indicated by the user, queries the associated service capability based on the target control instruction and calls it, and returns the result of calling the corresponding service capability to the end side response (playing voice or music).

[0019] The expansion of the corresponding intelligent voice interaction method capabilities is basically controlled by the platform. The platform provides access specifications, and developers develop according to its requirements. After approval by the platform, it is published to a place similar to a capability store, and users subscribe and activate it. Users cannot freely expand voice dialogue capabilities and can only use the capabilities provided by the platform.

[0020] In addition, in order to simplify the process of semantic parsing of user input by the backend, the platform restricts users to converse in the form of voice templates. That is, the voice template limits the structure of the conversation and requires users to communicate within an established framework. Users need to indicate their intentions with templated standard instructions. When new conversation scenarios need to be added or existing voice templates need to be modified, the corresponding voice templates need to be redeveloped. This makes it difficult to apply to complex or non-standard conversation scenarios, and the processing effect of non-standard instructions is poor, and the semantic interaction experience is stiff.

[0021] The technical solutions of the embodiments of the present disclosure and the technical effects produced by the technical solutions of the present disclosure will be described below through the description of several exemplary embodiments. It should be noted that the following embodiments can be referred to, learned from, or combined with each other. The same terms, similar features, and similar implementation steps in different embodiments will not be described repeatedly.

[0022] It can be understood that in the method for processing voice interaction tasks based on orchestration provided by the embodiments of the present disclosure, any method step can be executed by an electronic device and / or a server. All steps in the method can be independently executed by the electronic device or the server, or jointly executed by the electronic device and the server.

[0023] Among them, the server can be an independent physical server, a server cluster or a distributed system composed of multiple physical servers, or a cloud server providing cloud computing services. The electronic device can be a smart phone, a tablet computer, a notebook computer, a desktop computer, a smart voice interaction device (such as a smart speaker), a wearable electronic device (such as a smart watch), a vehicle-mounted terminal, a smart home appliance (such as a smart TV), an AR / VR device, etc., but is not limited thereto.

[0024] For example, the embodiments of the present disclosure provide a method for processing voice interaction tasks based on orchestration. This method can be implemented by a voice interaction task processing system based on orchestration. From a functional perspective, the voice interaction task processing system based on orchestration includes a server side and a client side. The specific functional modules involved in the server side and the client side can be set according to actual needs.

[0025] Among them, the client is used to obtain the user's voice information and convert the user's voice information into text information to obtain a control instruction. The server side is used to obtain the control instruction, perform dynamic process orchestration based on the control instruction, and execute at least one round of corresponding task processing operations to obtain the execution result text of the task processing operation. Further obtain the processing result text of the control instruction. The server side sends the processing result text to the client, and the client outputs voice based on the processing result text.

[0026] The embodiments of the present disclosure can be applied to the scenario where the client side and the server side are separated, that is, the client side is deployed on the terminal device and the server side is deployed on the server. It can also be applied to the scenario where the terminal device provides intelligent voice conversations for users offline, that is, both the client side and the server side are deployed on the terminal device.

[0027] In addition, it can be understood that the functions provided by the client and the server can be adaptively adjusted according to actual needs. For example, the client receives a control instruction in the form of voice from the user and sends it to the server, and the server converts the control instruction in the form of voice into text form, etc., which does not affect the core inventive concept of the embodiments of the present disclosure, and the embodiments of the present disclosure do not make any limitations in this regard.

[0028] The method for processing a voice interaction task based on orchestration provided by the embodiments of the present disclosure is applied to a server, and the server includes a tool list, and the tool list includes at least one tool.

[0029] It can be understood that in the embodiments of the present disclosure, the server internally maintains a tool list, and the tool list includes at least one tool. Among them, a tool refers to an application program or device for implementing corresponding functions, and each tool has a clear interface and function description.

[0030] For example, a tool can refer to application programs such as music playback, weather query, map navigation, etc., or devices such as smart speakers, smart homes (such as air purifiers, humidifiers, televisions).

[0031] The server can set the tool list according to the specific application scenario of the voice interaction task, and set the tools in the tool list and the related parameters of the tools, and realize the connection and call of the corresponding tools through the configured API interface. The specific method can be determined according to actual needs, and the embodiments of the present disclosure do not make any limitations in this regard.

[0032] Figure 1 One of the schematic diagrams of the structure of a voice interaction system provided by the embodiments of the present disclosure is as Figure 1 shown. When the server is used as the execution subject, the method for processing a voice interaction task based on orchestration includes: Step S101, receiving a control instruction sent by the client.

[0033] Specifically, taking the process of voice interaction with any user as an example, the process of the method for processing a voice interaction task based on orchestration provided by the embodiments of the present disclosure will be described in detail.

[0034] After the client is started, it will continuously detect the voice information of the user. When it determines that there is a control instruction related to the voice interaction task, it obtains the control instruction and sends the control instruction to the server.

[0035] For example, in the actual application process, the client usually continuously monitors the sound in the environment, continuously obtains sound fragments, and detects the sound fragments to determine whether a wake-up word preset by the user appears in the sound fragments. When it determines that the preset wake-up word appears, it collects the complete voice segment after the wake-up word as the voice control instruction of the user.

[0036] After obtaining the voice control instruction, it can be processed according to actual needs. After converting it into text to obtain the text control instruction, it is sent to the server, or the voice control instruction is directly sent to the server, and the server converts the voice control instruction into the corresponding text control instruction.

[0037] In step S101, the control instruction sent by the client is received, where the control instruction is the user's voice control instruction or text control instruction.

[0038] Step S102, use the preset large language model to arrange and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list, and obtain the processing result text of the control instruction according to the execution result text of at least one round of task processing operations.

[0039] Among them, each round of task processing operations is executed based on one tool.

[0040] Specifically, after receiving the control instruction, based on the corresponding text of the control instruction, in step S102, use the preset large language model (Large Language Model, abbreviated as LLM) to parse the intention and requirements of the user in the control instruction according to the control instruction and the tool list, and based on the capabilities that each tool included in the tool list can provide, select at least one appropriate tool, and arrange the task processing operations and the corresponding execution order executed by each tool.

[0041] Each round of task processing operations is executed based on one tool. According to the arranged at least one round of task processing operations, the interfaces corresponding to the corresponding tools are called in sequence, so that each tool executes the corresponding task processing operations in sequence according to the arranged execution order, and the execution result text of the task processing operation is obtained after each round of task processing operations is executed.

[0042] Among them, the information included in the execution result text can be determined according to actual needs, and can include information such as whether the task processing operation is executed successfully and the text content output after execution. For example, when the task processing operation is to query the weather tomorrow, the execution result text can directly be "The weather tomorrow is sunny" or "The weather for the next five days has been queried as sunny - sunny - cloudy - cloudy - shower, and the weather tomorrow is sunny".

[0043] It can be understood that when the server orchestrates and executes at least one round of task processing operations, it can be based on the LLM to orchestrate the task processing operations and the execution order of multiple tools at one time, and determine whether to continue the execution based on the execution result of each round of task processing operations until the execution ends. It can also be that only the task processing operations of this round are orchestrated in each round. After obtaining the execution result of the task processing operations of this round, the next round of task processing operations are orchestrated until it is determined that the execution of each round of task processing operations corresponding to the control instruction ends, and so on.

[0044] In the actual application of the embodiments of the present disclosure, the specific manner of orchestrating and executing at least one round of task processing operations corresponding to the control instruction can be determined according to actual needs, and the embodiments of the present disclosure do not limit this.

[0045] The processing result text of the control instruction is obtained according to the execution result text of at least one round of task processing operations. It can be understood that the processing result text of the control instruction is the text prepared for voice output to the user. It can only feedback the final result of the control instruction to the user, or feedback the corresponding results of each intermediate processing process (each task processing operation) to the user, which can be determined according to actual needs, and the embodiments of the present disclosure do not limit this.

[0046] For example, the user control instruction indicates that "if there is no haze today, then open the window for ventilation". The server orchestrates and invokes the tool weather query application program to first query the weather today, and when the execution result feedback determines that there is no haze today, it orchestrates and invokes the intelligent window system to control the window to open, and the obtained processing result text is "There is no haze today, the window has been opened" or "The window has been opened".

[0047] Step S103: Convert the processing result text into audio and send it to the client so that the client outputs voice according to the audio.

[0048] Specifically, after the server obtains the processing result text, in step S103, the server converts the processing result text into audio and sends it to the client so that the client outputs voice according to the audio after receiving the audio.

[0049] In the technical solution provided by the embodiments of the present disclosure, after the server obtains the control instruction sent by the client, it uses the large language model (LLM) to understand the intention of the control instruction and, based on the functions that each tool in the tool list can achieve, orchestrates and executes at least one round of task processing operations, obtains the processing result text of the control instruction according to the execution result text of at least one round of task processing operations, and converts the processing result text into audio and sends it to the client for the client to output the voice. It can utilize the ability of the large language model to understand natural language to achieve the orchestration and serial execution of the task processing operations corresponding to the control instruction under the constraint conditions of the tool functions provided by the server, without requiring the user to issue instructions in accordance with a preset template or format, providing a more flexible, natural, and adaptable interaction method.

[0050] In a possible implementation manner, the server includes a decision node; Using the preset large language model, according to the control instruction and the tool list, orchestrate and execute at least one round of task processing operations corresponding to the control instruction, including: Taking the control instruction as the input instruction for the first round of task processing operations. For each round of task processing operations, input the corresponding input instruction into the decision node. The decision node uses the preset large language model to execute the current round of task processing operations according to the input instruction of the current round of task processing operations and the tool list, and obtains the execution result text of the current round of task processing operations; For each round of task processing operations, take the input instruction and the execution result text of the previous round of task processing operations as the input instruction of the current round of task processing operations until it is determined that the iteration stop condition is met according to the input instruction.

[0051] Specifically, the server includes a decision node. In the embodiments of the present disclosure, each round of task processing operations is actually an iterative operation based on the execution result text of the previous round of task processing operations. Each time an iteration occurs, the decision node orchestrates each round of task processing operations corresponding to the control instruction.

[0052] Correspondingly, when the server uses the preset large language model to orchestrate and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list, for the first round of task processing operations, take the control instruction as the input instruction of the decision node. For each non-first round of task processing operations, take the input instruction of the previous round of task processing operations and the execution result text of the previous round of task processing operations together as the input instruction of the decision node for the current round of task processing operations until the decision node determines that the iteration stop condition is met according to the input instruction.

[0053] It can be understood that the specific types and quantities of iteration stop conditions can be determined according to actual requirements, including but not limited to determining that each task processing operation corresponding to the control instruction has been completed, the number of task processing operations (iteration operation rounds) corresponding to the execution of the control instruction is greater than a preset threshold, or it is impossible to execute the next round of task processing operations, etc.

[0054] For the task processing operation in the first round, the control instruction is input as the input instruction into the decision node. The decision node uses a preset large language model and a tool list to analyze the requirements indicated by the control instruction, and based on the functions of the tools in the tool list, arranges the task processing operation in the first round, and enables the corresponding tool to execute the task processing operation in the first round to obtain the execution result text of the task processing operation in the first round. The control instruction and the execution result text of the task processing operation in the first round are jointly used as the input instruction of the decision node in the second round of task processing operation.

[0055] For each round of task processing operation other than the first round, the input instruction of the corresponding current round of task processing operation is input into the decision node. The decision node uses a preset large language model to determine whether the iteration stop condition is met according to the input instruction of the current round of task processing operation and the execution result text of the previous round of task processing operation by using the ability of the large language model to understand natural language.

[0056] If it is determined that the iteration stop condition is met, the current round of task processing operation will no longer continue, and the task processing text of the output control instruction can be determined based on the actual situation at the time of iteration stop.

[0057] If it is determined that the iteration stop condition is not met, the decision node uses the preset large language model to understand the input instruction, and based on the intention indicated by the original control instruction and the execution result text of the previous round of task processing operation, determines the tool for executing the current round of task processing operation from the tool list, and calls the tool to execute the current round of task processing operation to obtain the execution result text of the current round of task processing operation.

[0058] In the technical solution provided by the embodiments of the present disclosure, the server arranges at least one round of task processing operations corresponding to the control instruction by setting a decision node. In each round of task processing operation other than the first round, the execution result text of the previous round of task processing operation is used as feedback. When arranging each round of task processing operations corresponding to the control instruction, the natural language understanding ability of the large language model can be used to complete the process arrangement of complex control instructions involving multiple processes step by step. Through the iterative promotion method, each round of task processing operations corresponding to the control instruction is gradually completed, realizing the serial arrangement and processing of complex business processes, being able to continuously confirm and refine the intention of the control instruction, improving the accuracy of speech recognition, understanding and process arrangement for voice interaction tasks based on arrangement, providing a more intelligent interaction effect, and effectively improving the user experience.

[0059] In a possible implementation, the server includes at least one execution node, and there is a one-to-one correspondence between the execution node and the tool; Each round of task processing operation includes: The decision node uses a preset large language model to determine whether the iteration stop condition is met based on the input instruction and the tool list of the current round of task processing operation; If it is determined that the iteration stop condition is not met, the decision node determines the target task of the current round of task processing operation and checks whether there is a target tool in the tool list for executing the target task; If it is determined that there is a target tool, the execution node corresponding to the target tool is called to generate a tool call instruction related to the target task using a preset large language model, and the target tool is called to execute the target task according to the tool call instruction to obtain the execution result text of the current round of task processing operation.

[0060] Specifically, the server includes a decision node and at least one execution node, and there is a one-to-one correspondence between the execution node and the tool. In each round of task processing operation, the decision node arranges the task processing operation, and the execution node executes the task processing operation and outputs an execution result text.

[0061] Figure 2 Schematic diagram of the task processing operation arrangement and execution process provided by the embodiments of the present disclosure, as Figure 2 shown, in each round of task processing operation, the decision node uses a preset large language model to determine whether the iteration stop condition is met based on the input instruction and the tool list of the current round of task processing operation.

[0062] If it is determined that the iteration stop condition is not met, the decision node uses a preset large language model to determine the target task of the current round of task processing operation based on the context information included in the input instruction and the functions of the tools in the tool list, and checks whether there is a target tool in the tool list for executing the target task.

[0063] For example, in actual applications, a prompt can be used to guide the large language model to generate content, and a prompt template is preset. When the execution node is applied in each round of task processing operation, the execution node queries the corresponding tool list and places the function description information of each tool in the tool list into the prompt template for rendering to obtain the corresponding prompt for each tool, so that the large language model determines the target task of the current round of task processing operation based on the corresponding input instruction and the prompts corresponding to each tool and determines which execution node should handle the target task.

[0064] If it is determined that there is a corresponding target tool in the tool list, then according to the correspondence between the execution node and the tool, the execution node corresponding to the target tool is called, and the execution node uses a preset large language model to generate tool call instructions related to the target task.

[0065] It can be understood that when the execution node uses a preset large language model to generate tool call instructions related to the target task, the tool call instructions are generated by the large language model and the prompt according to the parameter information of the tool. The tool call instructions indicate the specific parameters applied when calling the tool, and the type and specific information included in the parameters are determined according to the target task.

[0066] For example, users and developers can maintain the tool list by reverse registering the tools, realizing the addition, modification, and deletion of each tool in the tool list. For each tool, the parameter information of the tool and the default parameters corresponding to the parameter information can be pre-configured. Combining with the standard process execution logic, users can expand the process capabilities by simply configuring the parameters.

[0067] When executing the target task, based on the prompt of the target tool and the target task determined by the control node, the parameter information of the target tool is filled to generate access parameters. According to the access address of the tool and the generated access parameters, the tool call instructions are determined, and the tool is called based on the tool call instructions to execute the target task.

[0068] It can be understood that when there is no specific content in the target task that can indicate filling the parameter information of the tool (that is, the generated access parameters are empty), the default parameters of the tool can be used as the access parameters for subsequent processing.

[0069] For example, when the tool is a weather query application, the parameter information of the tool that needs to be configured can be set as the time (date or moment) and geographical location to be queried, and the default parameter is set to query the weather of the current location on the current day. When the target instruction can clearly indicate the corresponding date (such as querying the weather on x year x month x day, or the weather in the next few days) and the corresponding geographical location (such as City M), the parameter information is filled with the corresponding date to generate access parameters to call the tool to query the weather on the corresponding date. When the target task can only indicate querying the weather, the default parameter is used as the access parameter to call the tool to query the weather of the current location on the current day.

[0070] After obtaining the tool call instructions, the execution node calls the target tool to execute the target task according to the tool call instructions, obtains the result of the target tool executing the target task, and correspondingly obtains the execution result text of this round of task processing operation.

[0071] In addition, see again Figure 2In the embodiments of the present disclosure, the iteration stop conditions include three types: determining that all task processing operations corresponding to the control instruction have been executed, there is no corresponding target tool in the tool list, and the number of task processing operations (iteration operation rounds) corresponding to the execution of the control instruction is greater than a preset threshold.

[0072] If it is determined that the iteration stop conditions are met, the current round of task processing operations will not continue, and based on the actual situation at the time of iteration stop, different processing methods can be executed according to the types of different iteration stop conditions, and the task processing text of the output control instruction can be determined.

[0073] For example, firstly, when it is determined that all task processing operations corresponding to the control instruction have been executed (i.e., there is no need to continue to call specific tools), the process is directly confirmed to end, and based on the execution result text of each round of task processing operations during the entire iteration execution process, the task processing text of the control instruction is determined and output.

[0074] Secondly, to improve the user interaction experience, the embodiments of the present disclosure also set up a bottom node. The bottom node is used to transfer the target task to the bottom node in any round of task processing operations when it is determined that there is no corresponding target tool in the tool list (i.e., all tools in the tool list do not meet the user's needs). The bottom node uses a preset large language model to generate and output the corresponding response text according to the prompt and input instruction.

[0075] It can be understood that the types of preset large language models adopted by different nodes in the embodiments of the present disclosure can be set to be the same or different according to actual needs, or the large language model can be fine-tuned based on the specific requirements applied by the node to improve the processing effect. The embodiments of the present disclosure do not limit this.

[0076] Thirdly, when it is determined that the number of task processing operations corresponding to the execution of the control instruction is greater than the preset threshold, it can be determined that the iteration operation round corresponding to the control instruction is exceeded, and the response time corresponding to the control instruction is too long. At this time, it is considered that the control instruction processing fails, and a corresponding text can be returned to indicate to the user that the processing fails, and the user is required to re-enter the instruction, which can stop losses in time when a failure, speech recognition deviation, or other unexpected situations occur.

[0077] In the technical solution provided by the embodiments of the present disclosure, the server sets at least one decision node and at least one execution node. In each round of task processing operations, the decision node arranges the task processing operations to determine the target tasks to be executed in this round of task processing operations and the corresponding target tools for executing the target tasks. The execution node generates a tool call instruction according to the target tasks and calls the corresponding tools to execute the target tasks, and obtains the execution result text of this round of task processing operations. By separating the decision-making logic and the execution logic in each round of task processing operations, optimizing the arrangement and execution methods of at least one round of task processing operations corresponding to the control instructions, a flexible, efficient, and reliable mechanism is designed, enabling the server to flexibly expand the processing capacity by configuring the tool list and adding execution nodes, effectively applicable to the serial arrangement processing and execution of different requirements and complex business processes, without the need for users to configure complex process flow logic, and effectively improving the user experience.

[0078] In a possible implementation manner, the processing result text of the control instruction includes: the execution result text of each round of task processing operations in at least one round of task processing operations corresponding to the control instruction, or the execution result text of the last round of task processing operations in at least one round of task processing operations corresponding to the control instruction.

[0079] Specifically, since the control instruction corresponds to at least one round of task processing operations, the processing result text of the corresponding control instruction, that is, the text finally output to the user by voice, can output different types according to different actual requirements.

[0080] More specifically, the processing result text of the control instruction can be the execution result text of each round of task processing operations in at least one round of task processing operations corresponding to the control instruction.

[0081] The processing result text of the control instruction can also be the execution result text of the last round of task processing operations in at least one round of task processing operations corresponding to the control instruction.

[0082] For example, when at least one round of task processing operations corresponding to the control instruction is A - B - C - D in the processing order.

[0083] First, the processed result text can be the execution result text corresponding to each of the task processing operations A, B, C, and D, that is, the execution result text of the intermediate processing process when the output control instruction is executed. The output method can be to output the execution result text by voice every time an execution result text is obtained after each round of task processing operations, or to output each execution result text together when all task processing operations are completed. Outputting the execution result text of each round of task processing operations can improve transparency, enable users to perceive that the processing process of the control instruction is traceable and controllable, obtain feedback in a timely manner, be suitable for users who prefer to know the progress of task processing in a timely manner and need to monitor the task processing process in detail, and enhance the user experience.

[0084] Second, the processed result text can be the execution result text of the last round of task processing operation D, without outputting the execution result text related to the intermediate processing process. Only outputting the execution result of the last round of task processing operation can save time and resources, avoid unnecessary communication and data processing, be suitable for users who prefer concise output, and enable users to quickly obtain the required information.

[0085] It can be understood that the specific information contained in the execution result text of each round of task processing operations is determined according to the target tool of the actual application and the target task executed by the target tool.

[0086] The technical solution provided by the embodiments of the present disclosure provides different output methods for the processed result text of the control instruction, which can be determined based on different user requirements, application scenarios, or the iteration rounds of the task processing operations involved in the control instruction. Adjusting the type of the processed result text of the control instruction according to actual needs can better meet the needs of users and improve the voice interaction experience.

[0087] In a possible implementation manner, the server includes multiple tool lists; at least one list switching instruction is included in the preset switching instruction set, and each list switching instruction has a unique tool list that has a mapping relationship with the list switching instruction; the list switching instruction is used to instruct the server to apply the tool list that has a mapping relationship with the list switching instruction; Using the preset large language model according to the control instruction and the tool list, previously included: The server determines whether the control instruction includes a target list switching instruction in the preset switching instruction set; If it is determined that there is a target list switching instruction, the target tool list that has a mapping relationship with the target list switching instruction is determined, and the server uses the target tool list as the tool list applied by the preset large language model.

[0088] Specifically, in the embodiments of the present disclosure, the server provides a function for switching different tool lists. The server includes multiple tool lists, and the tools included in each tool list can be different. Users or developers can configure different tools for the tool lists based on different application scenarios applicable to different tool lists.

[0089] And a switching instruction set is preset in advance. The switching instruction set includes at least one list switching instruction. Each list switching instruction has a unique tool list that has a mapping relationship with the list switching instruction. The list switching instruction is used to instruct the server to apply the tool list that has a mapping relationship with the list switching instruction. The list switching instruction is used to instruct the server to apply the tool list that has a mapping relationship with the list switching instruction.

[0090] Users can use the list switching instruction to cause the server to switch the tool list of the preset large language model application, so as to switch the process execution ability provided by the server when arranging and executing the task processing operation of the control instruction.

[0091] In the embodiments of the present disclosure, after the server obtains the control instruction, it identifies the control instruction and determines whether any list switching instruction in the preset switching instruction set is included in the control instruction.

[0092] If any list switching instruction in the preset switching instruction set is included in the control instruction, the list switching instruction included in the control instruction is used as the target list switching instruction, and the target tool list that has a mapping relationship with the target list switching instruction is determined. The server uses the target tool list as the tool list of the preset large language model application.

[0093] It is possible to realize the dynamic switching of the tool list based on the actual needs of users during the process of the server providing services, so as to provide different process execution capabilities and take effect in real time without affecting the use of other users.

[0094] For example, the server can configure Tool List 1 and Tool List 2, which are respectively applicable to two different application scenarios of whole-house intelligent voice control and mobile terminal voice control. The tools in Tool List 1 can include devices such as intelligent lamps, monitoring devices, intelligent door and window systems, intelligent speakers, and air conditioners. The tools in Tool List 2 can include application programs such as communication, music playback, weather query, and map navigation in the mobile terminal.

[0095] The list switching instruction is set to perform keyword matching through the name of the tool list to determine the tool list to be switched. For example, the template of the list switching instruction is "Switch the process, switch to XXX". Users can replace XXX in the list switching instruction with the specific tool list name to achieve the switching of the tool list.

[0096] After the user returns home from going out, the server can be instructed by the list switching command "Switch the process to Tool List 2" to switch the tool list 1 providing services to Tool List 2 to meet the service requirements of the user in the whole-house intelligent voice control application scenario.

[0097] In addition, it can be understood that the order of each tool list can be preset, and the list switching command is used to indicate switching from the current tool list to the next tool list in a sequential rotation manner until returning to the first tool list after switching to the last tool list. Accordingly, when determining that the user instructs to switch to the next tool list by the list switching command, the server can feedback the name of the switched tool list to the user to facilitate the user to determine the required work list.

[0098] The following combines application examples of tool configuration and tool list configuration to elaborate in detail on the relevant steps of the tool list switching method provided by the embodiments of the present disclosure: Figure 3 It is a schematic flowchart of the special instruction processing method provided by the embodiments of the present disclosure. As Figure 3 shown, the embodiments of the present disclosure provide two types of special instructions: a list switching instruction and a list query instruction. Among them, the list switching instruction is used to switch the tool list provided by the server to the tool list indicated by the list switching instruction. The list query instruction is used to query the name of the tool list currently provided by the server for services.

[0099] The special instruction processing method is implemented through the Hypertext Transfer Protocol (HTTP) service. After receiving the control instruction, the server determines whether the control instruction includes a special instruction. If it is determined that the control instruction does not include a special instruction, the server directly arranges and executes at least one round of task processing operations of the control instruction based on the obtained control instruction and the tool list currently provided by the server for services.

[0100] If it is determined that the control instruction includes a special instruction, it enters the corresponding special processing flow. First, it is determined whether the special instruction is a list switching instruction. If it is determined that the special instruction is a list switching instruction, the tool list provided by the server for services is switched to the tool list indicated by the list switching instruction, and the server saves the switched tool list to arrange and execute at least one round of task processing operations of the control instruction (the control instruction after removing the special instruction).

[0101] If it is determined that the special instruction is not a list switching instruction, it means that the special instruction is a list query instruction. The name of the tool list currently provided by the server for services is queried, and the name of the tool list is fed back to the user.

[0102] For example, when developing software for a server, the model of the tool list can be set through Object-Oriented Programming. The model of the tool list includes the name of the tool list and the description of the tool list. The corresponding interfaces of different subclasses can be implemented through the standardized base class interface, and the unified business logic of different tool lists can be achieved through the interface calls of the base class. Among them, the interfaces of the tool list include: initialization interface, prompt service template rendering interface, process orchestration inference interface, configuration acquisition interface, and result parsing interface.

[0103] Moreover, a tool management function and a process management function are provided. Among them, the tool management function provides functions such as tool registration, tool modification, tool viewing, and tool deletion, and users can maintain the tools in each tool list based on the tool management function. The process management function provides functions such as process registration, process modification, process viewing, process deletion, and process execution. The model of the process includes the process name and the tool list description. In the embodiments of the present disclosure, the process refers to the process of implementing the orchestration and execution of at least one round of task processing operations of the control instruction, and users can maintain (add, change, and delete) each tool list and the name of the tool list based on the process management function.

[0104] The process orchestration can provide a fixed process code. According to the base class interface of the tool list, the application mode of the local execution service and the binding of the intelligent voice logic are realized. The tool capabilities can be adjusted and the tools included in the process can be modified through the configuration page of the process orchestration, and it will take effect dynamically when processing the next voice interaction task based on the orchestration, so as to achieve the effect of dynamically modifying the response logic of the intelligent voice dialogue.

[0105] Users can maintain tools based on a pre-designed tool model. The tool model includes the tool name, tool description, the uniform resource locator (URL) for accessing the tool, the parameter description of the tool, the default parameter description of the tool, and the tool execution timeout description.

[0106] Among them, the access URL of the tool is used to describe how to access an external tool (external interface access address), such as http: / / 27.128.115.153:8890 / weather. The parameter description of the tool is used to describe how to call an external tool (the input parameter specification of the external interface), such as {"city":"str"}. The default parameter description of the tool is the default input parameter of the external interface if the user does not specify a parameter, such as {"city":"City M"}. The tool execution timeout means that if the tool does not return within a certain period of time when it is called, it is considered an execution failure.

[0107] Accordingly, when the tool list is used on the server to implement the scheduling and execution of at least one round of task processing operations of the control instruction, two parameters, the process number (returned by the create process interface) and the control instruction, are required. The process includes the following steps: Initialization includes: calling the configuration acquisition interface to initialize LLM during initialization; calling the prompt service template rendering interface to fill the control instructions input by the user into the variable position of the prompt template, and saving the integrated input to realize the message format initialization.

[0108] After successful initialization, repeat the following two steps: calling the process orchestration reasoning interface on the server side to return the orchestration result of the task processing operation in the task processing operation process indicated by the control instruction, and calling the result parsing interface to convert the result returned by LLM into a standard interface until the execution is determined to be completed.

[0109] The technical solution server provided by the embodiment of the present disclosure includes multiple tool lists, and can implement switching of the tool list of a large language model application in the server based on the list switching instruction, allowing each user to maintain multiple tool instructions and each tool in each tool instruction, expand the process capabilities through a simple configuration parameter method, and customize the functions of the server according to their own needs without reconfiguring or restarting the service, providing the server with powerful scenario adaptability and personalized service capabilities, so that the server can more flexibly adapt to different user needs and environments, provide more customized services and ease of use of services, and enhance the user experience.

[0110] In a possible implementation, the processing result text is converted into audio and sent to the client, including Split the processed result text into multiple text segments, process the text segments one by one, and execute two parallel execution threads; Thread 1 is used to splice each text segment into the text to be processed; Thread 2 is used to recognize each character in the text to be processed in sequence according to the character order of the text to be processed until each character in the text to be processed is recognized; Among them, each character in the text to be processed is recognized, including: Determine whether the currently recognized character is a preset punctuation mark; If the character is a preset punctuation mark, the character is taken as the target character; The text before the target character in the text to be processed is intercepted as the output text, and the output text is deleted from the text to be processed; The output text is converted into an audio clip, and the audio clip is sent to the client, so that the client adds the audio clip to an audio playlist, and the audio clips in the audio playlist are played in sequence according to the order in which they are added.

[0111] Specifically, in order to optimize the process of real-time voice conversations, in the embodiments of the present disclosure, after obtaining the processed result text, the server uses a streaming response method to split the processed result text into multiple text segments, and processes each text segment one by one, executing two threads that execute in parallel.

[0112] Among them, Thread 1 is used to splice each text segment to the text to be processed for each text segment. That is, when the server processes the text segment, it splices each text segment to obtain the text to be processed, and adds each processed text segment after the text to be processed obtained from the previous processing. Thread 1 ensures the coherence of the acquired text information.

[0113] Thread 2 is used to sequentially recognize each character in the text to be processed in the order of the characters in the text to be processed until each character in the text to be processed is recognized.

[0114] Among them, recognizing each character in the text to be processed includes: Determining whether the currently recognized character is a preset punctuation mark. The specific type of the preset punctuation mark can be set according to actual needs, and is used to indicate whether there is a pause in the text to be processed and indicate the natural end of a sentence or short sentence.

[0115] For example, the specific types of the preset punctuation marks can be set to four types:,,?,!.

[0116] If the currently recognized character is a preset punctuation mark, then the character is used as the target character.

[0117] The text to be processed is segmented by the target character, and the text before the target character in the text to be processed (from the first character of the text to be processed to the target character) is intercepted as the output text, and the output text is deleted from the text to be processed.

[0118] The output text is converted into an audio segment, and the audio segment is sent to the client so that the client adds the audio segment to the end of the audio playlist.

[0119] Thread 2 is responsible for parsing the text to be processed, recognizing the preset punctuation marks in the text to be processed, and using the recognized preset punctuation marks as the segmentation points of the text to be processed, ensuring that each audio segment is complete or relatively complete in grammar and semantics. Thus, on the basis of being able to divide a long text into smaller texts for voice playback to improve the interaction response speed, the comprehensibility of the played audio segments is guaranteed, making the voice output closer to natural language.

[0120] The audio segments in the audio playlist are played in the order of addition, ensuring the order and coherence of the voice output. Moreover, since the text segment processing, the acquisition of the output text, the audio conversion, and the audio playback are performed in a streaming parallel manner simultaneously, the delay from obtaining the processed result text to playing the audio can be reduced, making the voice output more rapid.

[0121] For example, the client detects whether there are audio segments in the audio playlist. If it is determined that there are audio segments, the audio segments are sequentially extracted from the audio playlist for playback. If it is determined that there are no audio segments, the client waits for a period of time and then detects again whether there are audio segments in the audio playlist, repeating the above steps.

[0122] It can be understood that the server splits the processed result text into multiple text segments, and the number of characters included in each text segment (the size of the text segment) can be determined according to actual requirements.

[0123] In the technical solution provided by the embodiments of the present disclosure, after obtaining the processed result text, the server splits the processed result text into multiple text segments, processes the text segments one by one, and immediately starts the text segment processing, the acquisition of the output text, and the audio conversion by executing parallel threads 1 and 2, without waiting for the entire processed result text to be completely generated. Thus, the response time for the first voice playback is significantly reduced, and the resource pressure that may be brought about by processing a large amount of data at one time can be avoided, making the response faster and smoother, enabling the user to obtain faster feedback during the voice interaction process, making the interaction more natural, and enhancing the overall interaction experience.

[0124] In a possible implementation manner, after each character in the text to be processed is recognized, it includes: If it is determined that the text to be processed does not include a preset punctuation mark, the text to be processed is converted into an audio segment, and the audio segment is added to the audio playlist.

[0125] Specifically, Figure 4 is a schematic flowchart of a method for processing a text to be processed provided by the embodiments of the present disclosure. As Figure 4 shown, when processing the text to be processed, starting from the first character of the text to be processed, each character in the text to be processed is recognized in sequence. For the currently recognized character, whenever the currently recognized character is a preset punctuation mark, the current character is used as the target character, and the text before the target character in the text to be processed is split and intercepted as the output text, and the output text is converted into an audio segment, and the audio segment is added to the audio playlist.

[0126] That is, whenever a target character is recognized, the output text is intercepted and deleted from the text to be processed. After that, the first character of the updated text to be processed is the first character after the target character in the original text to be processed. For the updated text to be processed, the recognition of each character in the text to be processed is restarted from the first character of the text to be processed in sequence.

[0127] If it is determined that the recognition of the last character in the text to be processed has not been completed, the recognition steps for each character are continued in sequence.

[0128] If it is determined that the recognition of the last character in the text to be processed has been completed in sequence, it means that the recognition of each character in the text to be processed has been completed, and the recognition of the text to be processed has been completed. If it is determined that the text to be processed does not include a preset punctuation mark, the text to be processed is converted into an audio segment, the audio segment is added to the audio playlist, and the text to be processed is deleted.

[0129] In the technical solution provided by the embodiments of the present disclosure, after the recognition of each character in the text to be processed is completed, for the case where the text to be processed does not include a preset punctuation mark, the text to be processed still needs to be converted into an audio segment and output, so that the server can convert the text into an audio output completely for different situations, ensuring that the user receives complete information and maintaining the fluency and continuity of the voice interaction.

[0130] According to another aspect of the embodiments of the present disclosure, a voice interaction task processing system based on choreography is provided, including a server and a client; Wherein, when the server executes, it implements the steps of the method for processing a voice interaction task based on choreography provided by the server in any of the above embodiments; The client is used to send a control instruction to the server, and output voice according to the audio after receiving the audio sent by the server.

[0131] Specifically, the embodiments of the present disclosure provide a voice interaction task processing system based on choreography, including a server and a client. Wherein, when the server executes, it implements the steps of the method for processing a voice interaction task based on choreography provided by the server in any of the above embodiments. The client is used to send a control instruction to the server, and output voice according to the audio after receiving the audio sent by the server, and jointly provide a semantic interaction service for the user with the server.

[0132] The following combines a specific application example to illustrate the data flow transmission steps in the technical solution provided by the embodiments of the present disclosure in combination with the specific implementation manners of each functional module in the client and the server: The orchestration-based voice interaction task processing system includes a client and a server. The client includes: a client management module, a recording module, a wake word module, a Voice Activity Detection (VAD) module, and a playback module. The server includes: a server management module, an Automatic Speech Recognition (ASR) module, a Text-to-Speech (TTS) module, and an orchestration processing module.

[0133] Among them, the client management module is responsible for managing the recording module, the wake word module, the VAD module, and the playback module. The recording module is used to continuously obtain sound fragments of recordable hardware. For example, on the linux system, the arecord command can be used to continuously obtain recording data. The wake word module determines whether the recorded sound is the user's preset wake word. If it is, subsequent actions are taken; if not, it is discarded to avoid interference from invalid sounds during recording and to ensure the normal process. The VAD module is used to detect voice activity, determine which parts of the audio stream contain human speech, which parts are silent or noisy, and detect whether there is an interruption in the sound. During a voice conversation, it automatically determines when the user's conversation ends.

[0134] The server management module is responsible for managing the ASR module, the TTS module, and the orchestration processing module. The ASR module is used to perform speech recognition on the user's input speech and convert the speech into text. The TTS module is used to generate a voice file based on the provided text for playback. The main function of the orchestration processing module is to generate output text based on the input text. The orchestration processing module supports streaming output. Streaming output means that when the local processing module generates a segment of output text, it is returned to the server management module, but the service is not ended. When there is subsequent output, it continues to be returned to the server management module, achieving the effect of quick response to the user by calling the client to output voice in small steps.

[0135] The specific steps for implementing the orchestration-based voice interaction task processing based on the message pipeline are as follows: Step 1, initialization. After the client starts, the client management module is responsible for initializing the recording module, the wake word module, and the VAD module; the server management module is responsible for initializing the ASR module, the TTS module, and the orchestration processing module.

[0136] Step 2, the client management module establishes a connection (establishes a websocket connection) with the server management module and periodically sends heartbeat messages. After the connection is successfully established, the heartbeat keeps the connection alive to avoid frequent interruptions and repeated connection establishment.

[0137] Step 3: The user starts speaking, and the recording module starts sending recording data messages to the client management module. The client management module forwards the recording data messages to the wake word module. The wake word module determines whether the recording data contains a wake word according to the preset wake word. If there is a wake word, the wake word module returns a Detect message.

[0138] Step 4: After receiving the Detect message, the client management module modifies its local status. When subsequent recording data is received, it will be forwarded to the server management module.

[0139] Step 5: The client management module sends a start pipeline message to the server management module, notifying the server that subsequent recording data will arrive and to get ready.

[0140] Step 6: The server management module receives the start pipeline message and starts initializing the context and state machine. Here, the context refers to the address of the client management module, the local pipeline execution id, the current pipeline execution step and status.

[0141] Step 7: The client management module starts continuously forwarding newly received recording data messages to the server management module in a pipeline manner, and at the same time copies a copy of the recording data message and sends it to the VAD module.

[0142] Step 8: When the VAD module detects a pause / blank in the user's speech, it sends a VAD Detect message to the client management module.

[0143] Step 9: After receiving the VAD Detect message, the client management module modifies its local status. When subsequent recording data is received, it will no longer be forwarded to the server management module, and it sends a chunk stop message to the server management module to notify that the recording data has stopped being sent.

[0144] Step 10: The server management module sends an ASR recognition message to the ASR module. The ASR recognition message includes the locally cached recording data.

[0145] Step 11: The ASR module recognizes the recording data and, after the recognition is completed, returns the ASR recognition result to the server management module.

[0146] Step 12: The server management module forwards the ASR recognition result to the client management module.

[0147] Step 13: The server management module sends a control instruction to the orchestration processing module. The input parameter of the control instruction is the ASR recognition result.

[0148] Step 14: The orchestration processing module uses a large language model and a tool list to orchestrate and execute at least one round of task processing operations corresponding to the control instruction, obtains the processing result text of the control instruction based on the execution result text of at least one round of task processing operations, and streams back the processing result text. For the specific task processing operation orchestration and execution method, refer to the above description and will not be elaborated here.

[0149] Step 15: The server management module processes and parses the text fragments of the processing result text streamed back, and sends the parsing result to the TTS module. For the parsing and processing steps of the text fragments of the specific processing result text, refer to the above description and will not be elaborated here.

[0150] Step 16: The TTS module converts the parsing result from text to audio and returns a TTS chunk to the server management module.

[0151] Step 17: The server management module forwards the TTS chunk to the client management module.

[0152] Step 18: The client management module caches the TTS chunk locally and sends a TTS play message to the play module.

[0153] Step 19: The play module starts voice playback. After playback is completed, it returns a TTS play completion message to the client management module, completing the entire process of processing the voice interaction task based on orchestration.

[0154] It can be understood that the embodiments of the present disclosure are only used as a specific example to illustrate the present application in detail. The described embodiments are only intended to facilitate the understanding of the present application and do not impose any limitations on it. When the present application is actually applied, the modules included in the client and the server and the functions of each module can be adjusted according to actual needs.

[0155] Figure 5 FIG. is a schematic structural diagram of a server provided by an embodiment of the present disclosure. The server includes a tool list, and the tool list includes at least one tool, such as Figure 5 As shown, the server 50 includes: An instruction receiving module 501, configured to receive a control instruction sent by a client; A task orchestration module 502, configured to use a preset large language model to orchestrate and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list, and obtain the processing result text of the control instruction based on the execution result text of at least one round of task processing operations; A result output module 503, configured to send the processing result text to the client so that the client outputs voice according to the processing result text; Among them, each round of task processing operation is executed based on a tool.

[0156] In the technical solution provided in this embodiment, after the server obtains the control instruction sent by the client, it uses the large language model LLM to understand the intention of the control instruction and, based on the functions that each tool in the tool list can achieve, arranges and executes at least one round of task processing operations, obtains the processing result text of the control instruction according to the execution result text of at least one round of task processing operations, and converts the processing result text into audio and sends it to the client for the client to output the voice. It can utilize the ability of the large language model to understand natural language and realize the arrangement and serial execution of the task processing operations corresponding to the control instruction under the constraint conditions of the tool functions provided by the server, without requiring the user to issue instructions in accordance with a preset template or format, providing a more flexible, natural, and adaptable interaction method. The device of the embodiments of the present disclosure can execute the methods provided by the embodiments of the present disclosure, and their implementation principles are similar. The actions performed by each module in the devices of the embodiments of the present disclosure correspond to the steps in the methods of the embodiments of the present disclosure. For the detailed function descriptions of each module of the device, reference can specifically be made to the descriptions in the corresponding methods shown above, and details are not described herein again.

[0157] In a possible implementation manner, the server includes a decision-making node; Using a preset large language model to arrange and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list, including: Taking the control instruction as the input instruction for the first round of task processing operation. For each round of task processing operation, input the corresponding input instruction into the decision-making node. The decision-making node uses the preset large language model to execute the current round of task processing operation according to the input instruction of the current round of task processing operation and the tool list, and obtains the execution result text of the current round of task processing operation; For each round of task processing operation, taking the input instruction and execution result text of the previous round of task processing operation as the input instruction for the current round of task processing operation until it is determined that the iteration stop condition is met according to the input instruction.

[0158] In a possible implementation manner, the server includes at least one execution node, and there is a one-to-one correspondence between the execution node and the tool; Each round of task processing operation includes: The decision-making node uses the preset large language model to determine whether the iteration stop condition is met according to the input instruction of the current round of task processing operation and the tool list; If it is determined that the iteration stop condition is not met, the decision-making node determines the target task of the current round of task processing operation and determines whether there is a target tool in the tool list for executing the target task; If it is determined that there is a target tool, the execution node corresponding to the target tool is called to generate a tool call instruction related to the target task using a preset large language model, and the target tool is called according to the tool call instruction to execute the target task, and the execution result text of the current round of task processing operation is obtained.

[0159] In a possible implementation, the processing result text of the control instruction includes: The execution result text of each round of task processing operation in at least one round of task processing operation corresponding to the control instruction, or The execution result text of the last round of task processing operation in at least one round of task processing operation corresponding to the control instruction.

[0160] In a possible implementation, the server includes multiple tool lists; at least one list switching instruction is included in the preset switching instruction set, and each list switching instruction has a unique tool list that has a mapping relationship with the list switching instruction; the list switching instruction is used to instruct the server to apply the tool list that has a mapping relationship with the list switching instruction; Using the preset large language model according to the control instruction and the tool list, includes: The server determines whether the control instruction includes a target list switching instruction in the preset switching instruction set; If it is determined that there is a target list switching instruction, the target tool list that has a mapping relationship with the target list switching instruction is determined, and the server uses the target tool list as the tool list applied by the preset large language model.

[0161] In a possible implementation, converting the processing result text into audio and sending it to the client includes Splitting the processing result text into multiple text fragments, processing the text fragments one by one, and executing two threads that execute in parallel; Among them, thread 1 is used to splice each text fragment to the text to be processed for each text fragment; Thread 2 is used to sequentially identify each character in the text to be processed in the order of the characters of the text to be processed until each character in the text to be processed is identified; Among them, identifying each character in the text to be processed includes: Judging whether the current identified character is a preset punctuation mark; If the character is a preset punctuation mark, the character is used as the target character; Intercepting the text before the target character in the text to be processed as the output text, and deleting the output text from the text to be processed; Converting the output text into an audio segment, sending the audio segment to the client, so that the client adds the audio segment to the audio playlist, and the audio segments in the audio playlist are played in the order of addition.

[0162] In a possible implementation, after each character in the text to be processed has been recognized, it includes: If it is determined that the text to be processed does not include a preset punctuation mark, the text to be processed is converted into an audio segment, and the audio segment is added to the audio playlist.

[0163] An embodiment of the present disclosure provides an electronic device (computer device / equipment / system), including a memory, a processor, and a computer program stored on the memory. The processor executes the above computer program to implement the steps of the method provided in any optional embodiment of the present disclosure and achieve the corresponding technical effects.

[0164] In an optional embodiment, an electronic device is provided. Figure 6 It is a schematic structural diagram of an electronic device provided by an embodiment of the present disclosure. As Figure 6 shown, the electronic device 60 includes: a processor 601 and a memory 603. Among them, the processor 601 and the memory 603 are connected, such as through a bus 602. Optionally, the electronic device 600 may further include a transceiver 604, and the transceiver 604 may be used for data interaction between the electronic device and other electronic devices, such as data sending and / or data receiving, etc. It should be noted that in practical applications, the transceiver 604 is not limited to one, and the structure of the electronic device 600 does not constitute a limitation to the embodiments of the present disclosure.

[0165] The processor 601 may be a CPU (Central Processing Unit, central processor), a general-purpose processor, a DSP (Digital Signal Processor, digital signal processor), an ASIC (Application Specific Integrated Circuit, application-specific integrated circuit), an FPGA (Field Programmable Gate Array, field programmable gate array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute various exemplary logical blocks, modules, and circuits described in combination with the content disclosed in the present disclosure. The processor 601 may also be a combination that implements a computing function, such as a combination including one or more microprocessors, a combination of a DSP and a microprocessor, etc.

[0166] The bus 602 may include a path for transmitting information between the above components. The bus 602 may be a PCI (Peripheral Component Interconnect) bus, an EISA (Extended Industry Standard Architecture) bus, or the like. The bus 602 may be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, Figure 6 only a thick line is used to represent it in Figure 6 , but it does not mean that there is only one bus or one type of bus.

[0167] The memory 603 may be a ROM (Read Only Memory), or other types of static storage devices that can store static information and instructions, a RAM (Random Access Memory), or other types of dynamic storage devices that can store information and instructions. It may also be an EEPROM (Electrically Erasable Programmable Read Only Memory), a CD-ROM (Compact Disc Read Only Memory), or other optical disc storage, optical disc storage (including compact discs, laser discs, optical discs, digital versatile discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium that can be used to carry or store computer programs and can be read by a computer, which is not limited herein.

[0168] The memory 603 is used to store the computer program for implementing the embodiments of the present disclosure and is controlled by the processor 601 to execute. The processor 601 is used to execute the computer program stored in the memory 603 to implement the steps shown in the foregoing method embodiments.

[0169] The electronic device in the embodiments of the present disclosure may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Tablet Computers), PMPs (Portable Multimedia Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), wearable devices, etc., and fixed terminals such as digital TVs, desktop computers, etc.

[0170] The embodiments of the present disclosure provide a computer-readable storage medium on which a computer program is stored. When the computer program is executed by a processor, the steps and corresponding contents of the foregoing method embodiments can be implemented.

[0171] Embodiments of the present disclosure also provide a computer program product, including a computer program, which when executed by a processor can implement the steps and corresponding content of the foregoing method embodiments.

[0172] It should be noted that the computer-readable storage medium in the present disclosure may be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0173] In the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, which can be used by or in conjunction with an instruction execution system, apparatus, or device. In the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer-readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium may be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0174] Computer program code for performing the operations of the present disclosure may be written in one or more programming languages or combinations thereof. The programming languages include, but are not limited to, object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., connected through the Internet using an Internet service provider).

[0175] Terms such as "first", "second", "third", "fourth", "1", "2", and "target" (if any) in the specification, claims, and the above-mentioned drawings of the present disclosure are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances so that the embodiments of the present disclosure described herein can be implemented in an order other than that shown or described in words.

[0176] It should be understood that although the flowchart of the embodiments of the present disclosure indicates each operation step by an arrow, the execution order of these steps is not limited to the order indicated by the arrow. Unless explicitly stated herein, in some implementation scenarios of the embodiments of the present disclosure, the implementation steps in each flowchart may be executed in other orders according to requirements. In addition, some or all of the steps in each flowchart may include multiple sub-steps or multiple stages based on the actual implementation scenario. Some or all of these sub-steps or stages may be executed at the same time, and each sub-step or stage of these sub-steps or stages may also be executed at different times. In the scenario where the execution times are different, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and the embodiments of the present disclosure do not limit this.

[0177] The above are only optional implementation manners of some implementation scenarios of the present disclosure. It should be noted that for those of ordinary skill in the art, without departing from the technical concept of the solution of the present disclosure, using other similar implementation means based on the technical idea of the present disclosure also belongs to the protection scope of the embodiments of the present disclosure.

Claims

1. A method for processing voice interaction tasks based on choreography, characterized in that: Applied to a server, the server includes a tool list, the tool list includes at least one tool, and the method includes: Receive control instructions sent by the client; Using a preset large language model, according to the control instruction and the tool list, arranging and executing at least one round of task processing operations corresponding to the control instruction, and obtaining a processing result text of the control instruction according to the execution result text of the at least one round of task processing operations; Convert the processing result text into audio and send it to the client, so that the client outputs speech according to the audio; Each round of task processing operations is executed based on one tool.

2. The method for processing voice interaction tasks based on choreography according to claim 1, characterized in that: The server includes a decision node; The using of the preset large language model to arrange and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list includes: The control instruction is used as the input instruction of the first round of task processing operation. For each round of task processing operation, the corresponding input instruction is input into the decision node. The decision node uses a preset large language model to execute the task processing operation of this round according to the input instruction of the task processing operation of this round and the tool list, and obtains the execution result text of the task processing operation of this round; For each round of task processing operations, the input instructions and execution result text of the previous round of task processing operations are used as the input instructions of the current round of task processing operations until it is determined that the iteration stop condition is met according to the input instructions.

3. The method for processing voice interaction tasks based on choreography according to claim 2, characterized in that: The server includes at least one execution node, and the execution node has a one-to-one correspondence with the tool; Each round of task processing operations includes: The decision node uses a preset large language model to determine whether an iteration stop condition is met according to the input instructions of the current round of task processing operations and the tool list; If it is determined that the iteration stop condition is not met, the decision node determines the target task of the current round of task processing operation, and determines whether there is a target tool for executing the target task in the tool list; If it is determined that the target tool exists, the execution node corresponding to the target tool is called to use the preset large language model to generate a tool calling instruction related to the target task, and the target tool is called according to the tool calling instruction to execute the target task, and the execution result text of this round of task processing operation is obtained.

4. The method for processing voice interaction tasks based on choreography according to claim 3, characterized in that: The processing result text of the control instruction includes: The control instruction corresponds to the execution result text of each round of task processing operation in at least one round of task processing operation, or The control instruction corresponds to the execution result text of the last round of task processing operation in at least one round of task processing operation.

5. The method for processing voice interaction tasks based on choreography according to claim 1, characterized in that: The server includes a plurality of tool lists; the preset switching instruction set includes at least one list switching instruction, each list switching instruction has a unique tool list that has a mapping relationship with the list switching instruction; the list switching instruction is used to instruct the server to apply a tool list that has a mapping relationship with the list switching instruction; The using of a preset large language model according to the control instruction and the tool list previously includes: The server determines whether the control instruction includes a target list switching instruction in a preset switching instruction set; If it is determined that there is a target list switching instruction, a target tool list that has a mapping relationship with the target list switching instruction is determined, and the server uses the target tool list as the tool list of the preset large language model application.

6. The method for processing voice interaction tasks based on choreography according to any one of claims 1 to 5, characterized in that: The processing result text is converted into audio and sent to the client, including Splitting the processing result text into multiple text segments, processing the text segments one by one, and executing two parallel execution threads; Thread 1 is used to splice each text segment into the text to be processed; Thread 2 is used to recognize each character in the text to be processed in sequence according to the character order of the text to be processed until each character in the text to be processed is recognized; The step of identifying each character in the text to be processed includes: Determine whether the currently recognized character is a preset punctuation mark; If the character is a preset punctuation mark, the character is used as the target character; intercepting the text before the target character of the text to be processed as output text, and deleting the output text from the text to be processed; The output text is converted into an audio segment, and the audio segment is sent to a client, so that the client adds the audio segment to an audio playlist, and the audio segments in the audio playlist are played in sequence according to the order in which they are added.

7. The method for processing voice interaction tasks based on choreography according to claim 6, characterized in that: The process until every character in the text to be processed is recognized includes: If it is determined that the text to be processed does not include the preset punctuation mark, the text to be processed is converted into an audio segment, and the audio segment is added to an audio playlist.

8. A server, characterized in that: The server includes a tool list, the tool list includes at least one tool, and the server includes: An instruction receiving module, used for receiving control instructions sent by a client; A task scheduling module, used to schedule and execute at least one round of task processing operations corresponding to the control instruction according to the control instruction and the tool list by using a preset large language model, and obtain a processing result text of the control instruction according to the execution result text of the at least one round of task processing operations; A result output module, used for converting the processing result text into audio and sending it to the client, so that the client outputs speech according to the audio; Each round of task processing operations is executed based on one tool.

9. A voice interaction task processing system based on choreography, characterized in that: Including server and client; The server implements the method described in any one of claims 1 to 7 when executed; The client is used to send a control instruction to the server, and after receiving the audio sent by the server, output voice according to the audio.

10. An electronic device comprising a memory, a processor and a computer program stored in the memory, characterized in that: The processor executes the computer program to implement the steps of the method according to any one of claims 1 to 7.

11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

12. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Cited By

  • Voice control method and voice control system

    CN121171225A