A voice service interaction method and device, a storage medium and an electronic device
By transforming unstructured speech data into standardized instructions through a large language model architecture, scheduling and executing sub-models and generating synchronous feedback, the problem of existing voice interaction technologies being unable to handle complex tasks and synchronous interface rendering is solved, realizing synchronous collaboration between voice output and interface rendering and multi-task execution.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- CHONGQING ANT CONSUMER FINANCE CO LTD
- Filing Date
- 2025-12-19
- Publication Date
- 2026-05-29
AI Technical Summary
Existing voice interaction technologies struggle to handle the multi-parameter, multi-step, and cross-system data call requirements of complex tasks. Furthermore, voice feedback is disconnected from interface rendering, making it impossible to achieve coordinated interaction between voice input, intent recognition, task planning, execution sub-model scheduling, and fusion output.
A large language model architecture, comprising a master control sub-model, an execution sub-model, and a fusion sub-model, is adopted to transform unstructured speech data into standardized natural language instructions. Multiple execution sub-models are scheduled through collaborative planning to generate intermediate processing results. Finally, a composite feedback content containing natural language text information and interactive component control information is generated through the fusion sub-model to ensure that the speech output is synchronized with the client interface rendering.
It enables controllable execution of multiple business logics, ensures synchronous coordination between voice output and client interface rendering, and can stably execute multiple sub-tasks, satisfying the complete link control of semantic understanding and business actions.
Smart Images

Figure CN121354564B_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of computer technology, and in particular to a voice service interaction method, apparatus, storage medium and electronic device. Background Technology
[0002] Voice interaction technology is widely used in smart terminals and online services, but existing solutions mostly only complete speech recognition and simple command parsing, lacking the ability to break down complex tasks and struggling to handle data call requirements involving multiple parameters, multiple steps, and cross-systems. Furthermore, traditional large-scale model outputs are mainly natural language text, which cannot achieve accurate time synchronization with front-end components, leading to a disconnect between voice feedback and interface rendering. In addition, speech recognition errors, uncertain intents, and missing business rule validation often cause execution deviations or service interruptions. Existing technologies cannot achieve coordinated linkage between voice input, intent recognition, task planning, execution sub-model scheduling, deterministic business logic validation, and fused output, nor can they accurately drive front-end interactive components during voice broadcasting. Therefore, this invention proposes a voice service interaction method to address the above deficiencies. Summary of the Invention
[0003] This specification provides a voice service interaction method, apparatus, storage medium, and electronic device, the technical solution of which is as follows:
[0004] Firstly, this specification provides a voice service interaction method, the method comprising: responding to a service interaction request, generating a standardized instruction, wherein the service interaction request includes unstructured voice data, and the standardized instruction is a standardized natural language text instruction; inputting the standardized instruction into a large language model to obtain composite feedback information; and feeding the composite feedback information back to a client so that the client synchronously outputs voice feedback and renders the corresponding interactive component page; wherein the large language model includes a master control sub-model, a fusion sub-model, and a pre-set functional component library, and the pre-set functional component library contains multiple execution sub-models; the step of inputting the standardized instruction into the large language model to obtain composite feedback information specifically includes: inputting the standardized instruction into the master control sub-model for semantic analysis to generate a collaborative planning instruction, wherein the collaborative planning instruction includes multiple sub-tasks; scheduling the corresponding execution sub-model according to the collaborative planning instruction, so that the execution sub-model calls the underlying data interface to execute the corresponding sub-tasks and generate intermediate processing results; and inputting the intermediate processing results into the fusion sub-model to generate composite feedback information.
[0005] Secondly, this specification provides a voice service interaction device, comprising: an instruction generation module for generating standardized instructions in response to a service interaction request, wherein the service interaction request includes unstructured voice data and the standardized instructions are standardized natural language text instructions; an instruction input module for inputting the standardized instructions into a large language model to obtain composite feedback information; and an information feedback module for feeding back the composite feedback information to a client so that the client can synchronously output voice feedback and render the corresponding interactive component page; wherein the large language model includes a main control sub-model, a fusion sub-model, and a pre-set functional component library, the pre-set functional component library containing multiple execution sub-models; the instruction input module specifically includes: a semantic analysis sub-module for inputting the standardized instructions into the main control sub-model for semantic analysis to generate collaborative planning instructions, wherein the collaborative planning instructions contain multiple sub-tasks; an instruction scheduling sub-module for scheduling the corresponding execution sub-models according to the collaborative planning instructions, so that the execution sub-models call the underlying data interface to execute the corresponding sub-tasks and generate intermediate processing results; and a result fusion sub-module for inputting the intermediate processing results into the fusion sub-model to generate composite feedback information.
[0006] Thirdly, this specification provides a computer storage medium storing a plurality of instructions adapted for loading by a processor and executing the above-described method steps.
[0007] Fourthly, this specification provides an electronic device that may include: a processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and to execute the above-described method steps.
[0008] Fifthly, this specification provides a computer program product that stores at least one instruction, which is loaded by a processor and executes the above-described method steps.
[0009] The beneficial effects of the technical solutions provided in some embodiments of this specification include at least the following:
[0010] In one or more embodiments of this specification, unstructured speech data is converted into standardized natural language instructions, enabling subsequent models to perform semantic understanding and reasoning based on a unified format. A large language model architecture comprising a master control sub-model, an execution sub-model, and a fusion sub-model is adopted to process user speech intent in multiple stages, enabling the model to stably execute multiple sub-tasks. Multiple execution sub-models are scheduled based on collaborative planning instructions, and intermediate processing results are generated in conjunction with underlying data interfaces to achieve controllable execution of multiple business logics. The fusion sub-model generates composite feedback content containing natural language text information and interactive component control information to support the synchronous collaboration between client-side voice feedback playback and business component rendering. This achieves complete link control from voice input to execution feedback, ensuring that model output satisfies both semantic understanding and drives business actions, while also ensuring synchronization between voice output and client-side interface rendering. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in this specification or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a schematic diagram of a voice service interaction system provided in this manual.
[0013] Figure 2 This is a flowchart illustrating a voice service interaction method provided in this manual.
[0014] Figure 3 It is based on Figure 2 A flowchart illustrating a specific implementation of step S200 in the voice service interaction method shown in the corresponding embodiment.
[0015] Figure 4 It is based on Figure 2 The system architecture flowchart shown in the corresponding embodiment.
[0016] Figure 5 This is a schematic diagram of the structure of a voice service interaction device provided in this specification.
[0017] Figure 6 This is a schematic diagram of the structure of an electronic device provided in this specification.
[0018] Figure 7 This is a schematic diagram of the operating system and user space provided in this manual.
[0019] Figure 8 yes Figure 7Architecture diagram of the Android operating system in China.
[0020] Figure 9 yes Figure 7 Architecture diagram of the iOS operating system. Detailed Implementation
[0021] The technical solutions in this specification will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this specification, and not all embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this specification.
[0022] In the description of this specification, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. In the description of this specification, it should be noted that, unless otherwise expressly specified and limited, "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or devices. Those skilled in the art can understand the specific meaning of the above terms in this specification based on the specific circumstances. Furthermore, in the description of this specification, unless otherwise stated, "multiple" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship.
[0023] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the object characteristics, interactive behavior characteristics, and user information involved in this specification were all obtained under full authorization.
[0024] The present specification will now be described in detail with reference to specific embodiments.
[0025] Please see Figure 1 This is a schematic diagram of a voice service interaction system provided in this specification. Figure 1As shown, the voice service interaction system may include at least a client cluster and a service platform 100.
[0026] The client cluster may include at least one client, such as Figure 1 As shown, it specifically includes client 1 corresponding to user 1, client 2 corresponding to user 2, ..., client n corresponding to user n, where n is an integer greater than 0.
[0027] Each client in a client cluster can be an electronic device with communication capabilities, including but not limited to: wearable devices, handheld devices, personal computers, tablets, in-vehicle devices, smartphones, computing devices, or other processing devices connected to a wireless modem. Electronic devices may have different names in different networks, such as: user equipment, access terminal, user unit, user station, mobile station, mobile station, remote station, remote terminal, mobile device, user terminal, terminal, wireless communication equipment, user agent or user device, cellular phone, cordless phone, personal digital assistant (PDA), and electronic devices in 5G networks or future evolved networks.
[0028] The service platform 100 can be a standalone server device, such as a rack-mount, blade, tower, or cabinet-type server device, or a workstation, mainframe, or other hardware device with strong computing power; or it can be a server cluster composed of multiple servers. The servers in the service cluster can be composed in a symmetrical manner, wherein each server is functionally and hierarchically equivalent in the transaction chain, and each server can provide services independently. The independent provision of services can be understood as not requiring the assistance of other servers.
[0029] In one or more embodiments of this specification, the service platform 100 can establish a communication connection with at least one client in the client cluster, and complete the data exchange during the voice service interaction based on the communication connection, such as transaction data exchange of service interaction requests, including but not limited to various types of transaction request data exchange, and the specific transaction service type is determined based on the actual application situation.
[0030] It should be noted that the service platform 100 establishes a communication connection with at least one client in the client cluster via a network for interactive communication. This network can be a wireless network or a wired network. Wireless networks include, but are not limited to, cellular networks, wireless LANs, infrared networks, or Bluetooth networks. Wired networks include, but are not limited to, Ethernet, universal serial bus (USB), or controller area networks. In one or more embodiments of the specification, technologies and / or formats including Hyper Text Markup Language (HTML), Extensible Markup Language (XML), etc., are used to represent data exchanged over the network (such as target compressed packets). Furthermore, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), and Internet Protocol Security (IPsec) can be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can be used to replace or supplement the aforementioned data communication technologies.
[0031] The voice service interaction system embodiments provided in this specification and the voice service interaction methods described in one or more embodiments belong to the same concept. The execution entity corresponding to the voice service interaction method involved in one or more embodiments of this specification can be the aforementioned service platform 100; the execution entity corresponding to the voice service interaction method involved in one or more embodiments of this specification can also be the electronic device corresponding to the client, specifically determined based on the actual application environment. The specific implementation process of the voice service interaction system embodiments can be found in the following method embodiments, and will not be repeated here.
[0032] based on Figure 1 The following is a detailed description of the voice service interaction method provided by one or more embodiments of this specification, illustrated in the scenario diagram.
[0033] Please see Figure 2 This document provides a flowchart illustrating a voice service interaction method according to one or more embodiments. This method can be implemented using a computer program and can run on a voice service interaction device based on the von Neumann architecture. The computer program can be integrated into an application or run as a standalone utility application. The voice service interaction device can be a service platform.
[0034] Specifically, the voice service interaction method includes:
[0035] S100, in response to a service interaction request, a standardized instruction is generated, wherein the service interaction request contains unstructured voice data and the standardized instruction is a standardized natural language text instruction.
[0036] S200, the standardized instructions are input into the large language model to obtain composite feedback information.
[0037] S300, the composite feedback information is fed back to the client so that the client can synchronously output voice feedback and render the corresponding interactive component page.
[0038] This specification describes the transformation of unstructured speech data into standardized natural language instructions, enabling subsequent models to perform semantic understanding and reasoning based on a unified format. A large language model architecture, comprising a master control sub-model, an execution sub-model, and a fusion sub-model, is employed to process user speech intent in multiple stages, ensuring stable execution of multiple sub-tasks. Multiple execution sub-models are scheduled based on collaborative planning instructions, and intermediate processing results are generated using underlying data interfaces to achieve controllable execution of multiple business logics. The fusion sub-model generates composite feedback content containing natural language text information and interactive component control information to support the synchronous collaboration between client-side speech feedback playback and business component rendering. This achieves complete link control from speech input to execution feedback, ensuring that model output satisfies both semantic understanding and drives business actions, while maintaining synchronization between speech output and client-side interface rendering.
[0039] In one embodiment of this specification, a user issues a voice command to the remote operating system of a smart device to "open the router's advanced settings and check the network status." The remote operating system receives the voice data and converts it into a standardized text command. The main control sub-model identifies the sub-tasks to be executed as entering the advanced settings page and performing network status detection. The remote operating system schedules two execution sub-models, which respectively call the device configuration interface and the network diagnostic interface to generate intermediate processing results. The fusion sub-model combines the intermediate processing results to generate composite feedback information, including voice text indicating that advanced settings have been entered and the current network status is normal, as well as page component control anchors, enabling the client to simultaneously open the corresponding page and display the diagnostic results.
[0040] In another embodiment, the user inputs the command "I want to borrow 3,000 yuan and view the installment plan" via voice. The credit system converts the voice data into a standardized text command: "Borrow 3,000 yuan and view the installment plan." The main control sub-model identifies two sub-tasks: loan amount verification and installment plan calculation. The execution sub-model sequentially calls the loan amount verification interface and the plan calculation interface to generate intermediate processing results of loan availability information and a list of available installments. The fusion sub-model generates composite feedback information including voice text: "You can borrow 3,000 yuan, and you can choose 3, 6, or 12 installment plans," along with the corresponding component rendering identifier. The client renders the loan plan page while playing the voice message.
[0041] In S100, voice data uploaded by the client is converted into standardized text instructions, giving it a parsable semantic structure.
[0042] Specifically, upon receiving a service interaction request, the system parses the voice data in the request and identifies unstructured voice data segments. The server then invokes the speech recognition module to extract acoustic features, segment the speech, and decode the language model to generate text-based content. This text content undergoes standardization to conform to the semantic format of the model input, such as unifying grammatical structure, removing noise symbols, and completing logical subjects or verbs, thus transforming it into a natural language text instruction with complete sentence meaning.
[0043] The S100 converts unstructured speech information into standardized input that the model can use to perform semantic parsing, thus improving the accuracy and stability of semantic understanding.
[0044] Specifically, in some embodiments, the specific implementation of step S100 can be found in the following embodiments. This embodiment is based on... Figure 2 According to the detailed description of step S100 in the information acquisition method shown in the corresponding embodiment, step S100 in the information acquisition method may include the following steps:
[0045] In response to a service interaction request, the service interaction request is parsed to obtain unstructured voice data.
[0046] The speech recognition module is invoked to decode and convert unstructured speech data, generating standardized instructions.
[0047] In this embodiment, when a service interaction request arrives, the unstructured speech data is extracted by parsing the request to identify the original speech content and enter a processable state. The unstructured speech data is decoded and semantically reconstructed by a dedicated speech recognition module, converting the original speech into standardized natural language text instructions. This provides a recognizable and inferable input format for the large language model, realizing the standardized conversion of speech data from acoustic signals to semantic text, and providing basic data support for the subsequent model inference, task planning, and component rendering in claim 1.
[0048] Specifically, upon receiving a service interaction request from the client, the request body is parsed in a structured manner. This request may contain various data payloads, including device identification information, terminal status, network transmission parameters, and audio data. Audio data segments are extracted from the request body; these audio data segments represent the unstructured speech data input by the user through the microphone. The parsing process may include steps such as packet splitting, audio stream format recognition, container structure parsing, and speech segment reconstruction, ensuring that the extracted audio content is complete and undamaged. The speech data is then input to the speech recognition module. The speech recognition module may include components such as an acoustic model, a pronunciation dictionary, a language model, and a post-processing unit. The speech recognition module first extracts acoustic features from the speech signal, obtaining acoustic feature vectors including Mel-frequency cepstral coefficients and endpoint detection results. Subsequently, it infers the phonetic correspondences using the acoustic model and combines this with the language model to infer semantically coherent word sequences. The decoded initial text may contain modal particles, pause markers, or omitted expressions. The text is formatted by a standardization processing unit, including removing redundant modal particles, completing missing semantic components, unifying numerical expressions, and organizing unstructured natural language into complete sentences. The processed text serves as standardized instructions that can be directly input into a large language model for semantic analysis and task planning.
[0049] In one embodiment of this specification, a user issues a voice command to the communication system to "open server status," and their client uploads recorded audio to the server. The server parses the service interaction request packet and extracts unstructured speech data (such as PCM audio data). The speech recognition module decodes the audio and converts it into text content indicating that the server status is open. The communication system formats the text to standardize the command to query the server status. This command is then passed to the main control sub-model for further analysis.
[0050] In another embodiment of this specification, the user inputs "Check my available credit limit today" via voice. After the client uploads the audio, the server parses the audio data in the interaction request and identifies unstructured speech segments. The speech recognition module decodes the speech signal, converting the speech content into natural language text to query today's available credit limit. The system performs standardization processing, unifying the expression to "query available credit limit." This standardized instruction is then fed into the subsequent execution flow of the large language model, enabling the model to further perform credit limit verification, data query, and generate feedback content based on the instruction.
[0051] In S200, standardized natural language commands are taken as input and submitted to the large language model. To ensure stable processing of complex business logic, the large language model adopts a hierarchical structure consisting of a main control sub-model, multiple execution sub-models, and a fusion sub-model.
[0052] Specifically, in some embodiments, the specific implementation of step S200 can be found in [reference needed]. Figure 3 . Figure 3 It is based on Figure 2 The detailed description of step S200 in the information acquisition method shown in the corresponding embodiment is as follows: In the information acquisition method, the large language model includes a main control sub-model, a fusion sub-model, and a preset functional component library. The preset functional component library contains multiple execution sub-models. Step S200 may include the following steps:
[0053] S210, the standardized instructions are input into the master control sub-model for semantic analysis to generate collaborative planning instructions, which contain multiple sub-tasks.
[0054] S220, according to the collaborative planning instruction, the corresponding execution sub-model is scheduled, so that the execution sub-model calls the underlying data interface to execute the corresponding sub-task and generate intermediate processing results.
[0055] S230, the intermediate processing results are input into the fusion sub-model to generate composite feedback information.
[0056] In S210, the master control sub-model first performs semantic analysis on the standardized instructions, generating collaborative planning instructions based on the semantic features of the instructions, the key entities involved, and the logical conditions. These collaborative planning instructions characterize the multiple subtasks required by the input instructions and their dependency structures.
[0057] Specifically, in some embodiments, the specific implementation of step S210 can be found in the following embodiments. This embodiment is based on... Figure 3 According to the detailed description of step S210 in the information acquisition method shown in the corresponding embodiment, step S210 in the information acquisition method may include the following steps:
[0058] The standardized instructions are semantically analyzed by the main control sub-model to parse the key entities in the standardization and determine the user's operation type.
[0059] Based on the operation type and key entities, collaborative planning instructions are generated, which contain the dependencies between their respective tasks.
[0060] In this embodiment, the master control sub-model can not only parse key entities from standardized instructions and determine the user's operation type accordingly, but also construct collaborative planning instructions based on key entities and operation types, and clarify the dependencies between sub-tasks, enabling subsequent execution sub-models to execute tasks in a structured order. This embodiment defines a semantic parsing architecture for multi-task execution chains, realizing the transformation from single-sentence natural language to a clear multi-task structure.
[0061] Specifically, after receiving standardized instructions, the master control sub-model performs semantic segmentation, syntactic analysis, and semantic role labeling on the instructions. Key entities identified by the master control sub-model can include numerical entities, operation object entities, modifier entities, and action entities. Numerical entities include, for example, amount, quantity, date, and duration; operation object entities include, for example, credit limit, reminder, and network status; modifier entities include, for example, today, three periods, and advanced settings; and action entities include, for example, query, borrow, adjust, and open.
[0062] After parsing the key entities, the master control sub-model determines the user's operation type based on entity distribution and semantic structure. For example, if the instruction contains numerical entities and the action "borrow," the type is a loan application; if the instruction contains query terms, the type is an information query; if the instruction contains modification terms, the type is parameter adjustment. This determination provides the semantic basis for the next step of task planning.
[0063] The master control sub-model determines the list of business tasks to be executed based on the operation type and sets the dependency order for each task according to the relationships between key entities. The collaborative planning instruction generally includes task identifiers, task parameters, dependencies between tasks, and the control method for task execution. For example, querying network status might only include one sub-task; borrowing 3,000 yuan and selecting six installments would include a dependency chain of credit limit verification, scheme generation, and installment adjustment; opening a page and performing diagnostics would include a dependency chain of page loading and diagnostic execution. This collaborative planning instruction serves as the basis for scheduling the execution sub-model, thereby ensuring the controllability of the execution chain and the correctness of the task completion order.
[0064] In some embodiments, a user speaks the command "Check the switch port and update the configuration" to the communication equipment management system. In the preceding steps, the communication equipment management system converts the speech into text to check the switch port and update the configuration. The main control sub-model performs semantic parsing on this text, identifying the key entities as switch port and update configuration, and the action entities as check and update. Based on internal model rules, the operation type is determined to be device detection and configuration update. The main control sub-model generates collaborative planning instructions based on entity relationships. Specifically, subtask 1 could be port status diagnosis, subtask 2 could be configuration distribution, and the dependency relationship could be that configuration update must be executed after port status diagnosis is completed. Through this planning, the execution sub-model can complete the device detection and configuration tasks in the correct order.
[0065] In other embodiments, the user inputs via voice, "Borrow me 2,000 yuan and change it to three installments." After processing by the preceding steps, a standardized text "Borrow 2,000 yuan and adjust to three installments" is obtained. The main control sub-model identifies the key entities as 2,000 yuan and three installments. The action entities are "borrow" and "change." The main control sub-model determines the operation type as loan application and scheme adjustment. Subsequently, the main control sub-model generates a task chain, including sub-task 1, loan amount verification; sub-task 2, installment scheme generation; sub-task 3, installment number adjustment; and dependencies: loan amount verification, scheme generation, and installment number adjustment. The resulting collaborative planning instruction clarifies the execution order and dependencies of each sub-task, enabling the execution sub-model to complete the credit process according to the instruction.
[0066] Specifically, in some other embodiments, the specific implementation of step S210 can be found in the following embodiments. This embodiment is based on... Figure 3 According to the detailed description of step S210 in the information acquisition method shown in the corresponding embodiment, step S210 in the information acquisition method may include the following steps:
[0067] The standardized instructions are semantically analyzed using the master control sub-model to calculate the intent uncertainty index.
[0068] If the intent uncertainty index is higher than the preset index threshold, then a collaborative planning instruction is generated in the form of a multi-step inference chain.
[0069] If the intent uncertainty index is lower than the preset index threshold, then collaborative planning instructions are generated in the form of a direct execution chain.
[0070] In this embodiment, the intent uncertainty index is used to quantify the semantic ambiguity of natural language input. Based on this index and a threshold judgment, dynamic selection between two different planning methods—a multi-step inference chain or a direct execution chain—can be achieved. The multi-step inference chain is used to handle complex tasks with high semantic ambiguity or those requiring supplementary inference, while the direct execution chain is used for semantically clear scenarios, improving processing efficiency and solving the problem of execution deviation caused by the inability of existing technologies to distinguish between ambiguous and clear semantics.
[0071] Specifically, after receiving standardized instructions, the master control sub-model performs multi-dimensional semantic analysis on them. This multi-dimensional semantic analysis includes sentence completeness analysis, key entity coverage analysis, ambiguity analysis, and context dependency analysis. Sentence completeness analysis determines whether the sentence structure lacks key semantic components; key entity coverage analysis determines whether parameters are missing, such as time, quantity, or object; ambiguity analysis determines whether the instruction contains unclear referents, polysemous words, or structures that can produce more than one interpretation; and context dependency analysis determines whether the instruction depends on historical context for complete understanding. Based on the above analysis results, the master control sub-model calculates an intent uncertainty index using its internal weighting system. A higher index indicates more ambiguous semantics, while a lower index indicates clearer semantics.
[0072] When the intent uncertainty index exceeds a preset threshold, the master control sub-model determines that the instruction contains incomplete semantics, ambiguous entities, or missing parameters. To avoid misjudgment, the master control sub-model uses a multi-step inference chain to generate collaborative planning instructions. The specific form of the multi-step inference chain can be breaking down the task into multiple inference stages, with each stage progressively completing the semantic information; supplementing missing parameters when necessary, such as completing dates, quantities, or objects; allowing the introduction of context-related completion logic; and generating collaborative planning instructions containing multiple sub-tasks, each used to clarify a specific semantic element. For example, in the user's voice command "Help me find that swimming pool," "that swimming pool" cannot be directly determined and requires an inference chain to complete the entity.
[0073] When the intent uncertainty index is below the threshold, it indicates that the semantics are complete, the entities are clear, and no further reasoning is needed. The master control sub-model generates collaborative planning instructions for the direct execution chain.
[0074] The specific form of a direct execution chain can be a small number of subtasks, often a single-chain structure, which does not require semantic completion or context verification, does not require inference iteration, and is suitable for instructions with complete parameters and clearly defined actions. For example, "query today's network status" and "borrow 3000 yuan for 6 installments" are both examples of direct execution chains.
[0075] In one embodiment of this specification, the user gives the voice command "Check that status". The system generates standardized text to view the status in the preceding steps. The master control sub-model analysis reveals a lack of specific object descriptions, meaning the status could refer to multiple network states, leading to high ambiguity. The intent uncertainty index exceeds a threshold, therefore the master control sub-model generates a multi-step inference chain. Subtask 1: Inferring possible objects in the context (e.g., ports, links, routers, etc.); Subtask 2: Prompting the execution sub-model to infer the most likely object; Subtask 3: Constructing the final task chain based on the inference results. If the user inputs "Check port 1 status", the master control sub-model will recognize this as an explicit command, the intent uncertainty index will be below the threshold, and a direct execution chain will be generated: Subtask: Port 1 status query.
[0076] In another embodiment, a user inputs "borrow some money" via voice. The system converts this into a standardized instruction to apply for a loan. Semantic analysis by the main control sub-model reveals that the input does not include the loan amount or loan period, indicating that "borrow some money" could refer to a large range of amounts. The semantics depend on the user's historical behavior or context, therefore the intent uncertainty index is higher than the threshold, generating a multi-step inference chain. Subtask 1: Infer a typical loan amount from the user's historical behavior; Subtask 2: Prompt for missing parameters to complete the amount; Subtask 3: Call the execution sub-model to generate candidate loan periods; Subtask 4: Finally, construct the loan task chain. If the user's instruction is "borrow 2000 yuan for three periods," the main control sub-model parses the semantics clearly, the intent uncertainty index is lower than the threshold, and a direct execution chain is generated: Subtask 1: Credit limit verification; Subtask 2: Scheme generation; Subtask 3: Loan period confirmation.
[0077] In S220, the corresponding execution sub-model is scheduled according to the collaborative planning instructions. The execution sub-model calls the underlying database interface, business engine interface, or configuration interface, etc., to complete the processing of its respective sub-tasks and form a structured intermediate processing result.
[0078] Specifically, in some embodiments, the specific implementation of step S220 can be found in the following embodiments. This embodiment is based on... Figure 3 According to the detailed description of step S220 in the information acquisition method shown in the corresponding embodiment, step S220 in the information acquisition method may include the following steps:
[0079] Set an upper limit on the number of iterations based on the collaborative planning instructions.
[0080] Within the upper limit of the number of iterations, the corresponding execution sub-models are scheduled in sequence according to the task dependencies, so that each sub-model executes its corresponding sub-tasks in sequence and generates intermediate processing results.
[0081] In this embodiment, a control mechanism is introduced to limit the number of iterations for the collaborative planning instructions, ensuring that task execution is completed within the controllable range of the model and avoiding processing delays or logical deviations caused by excessive model inference or abnormal loops. Multiple execution sub-models are scheduled according to task dependencies, enabling each sub-model to execute its subtasks in the order specified by the collaborative planning instructions, thereby ensuring the integrity of the task chain and the consistency of execution logic. Intermediate processing results are generated by the execution sub-models calling the underlying data interface, allowing the entire method to output intermediate data in a controllable manner within a structured process, providing raw processing results for subsequent fusion sub-models to synthesize composite feedback information.
[0082] Specifically, the collaborative planning instructions generated by the master control sub-model are first read. An upper limit for the number of iterations is set based on factors such as the type of multi-step inference chain or direct execution chain, the number and complexity of subtasks, the uncertainty index of the input semantics, and the current resource utilization of the system. This upper limit is generally used to limit the maximum number of execution rounds of the entire execution chain. For example, a value of 1 indicates only one loop, disallowing repeated inference; a value greater than 1 indicates that the system allows inference compensation or error retries within a reasonable range. The upper limit for the number of iterations serves as a control parameter throughout the execution process, ensuring that the system maintains execution stability and avoids unlimited loops under complex task scenarios.
[0083] In each iteration, based on the task dependency structure in the collaborative planning instructions, the corresponding execution sub-models are scheduled sequentially. Specifically, the dependencies are first read to determine the first executable subtask, and the corresponding execution sub-model is called to begin executing the subtask. After the subtask is completed, the execution result is passed to the next execution sub-model. If a subtask depends on a preceding task, it can only be executed after its dependent task is completed. If there are situations in the execution chain such as insufficient inference data, parameters to be completed, or business validation failure, the inference is re-inferred or the execution is backtracked within the iteration count. Through this sequential scheduling mechanism, each sub-model can strictly execute tasks according to the business logic chain, thereby avoiding logic overwriting, incorrect data references, or task conflicts caused by parallel execution.
[0084] Each execution sub-model, upon receiving a task instruction, completes the corresponding task by calling underlying data interfaces. For example, it might call a database interface to retrieve data records, invoke the business logic engine to perform rule validation, call a device interface or external API to obtain real-time status information, and calculate the structured output required for the task. The execution sub-model then standardizes the execution results into intermediate processing results for use by subsequent fusion sub-models. These intermediate processing results are typically structured data objects containing task identifiers, execution status, processing results, timestamps, etc. The system ultimately generates a complete chain of intermediate processing results within the number of iterations.
[0085] In one embodiment of this specification, the user issues a voice command to "check network issues and refresh configuration." In the preliminary steps, the main control sub-model generates a collaborative planning instruction, which includes two subtasks, Subtask 1 and Subtask 2. Subtask 1 performs network status detection, serving as the starting point of its dependency chain; Subtask 2 performs configuration refresh, depending on the detection results of Subtask 1. An upper limit of two iterations is set based on task complexity for fault tolerance compensation. After the dependencies are determined, the following steps are performed: In the first iteration, sub-model A is scheduled to perform network detection and obtain status data; sub-model B is scheduled to perform configuration refresh and obtain execution confirmation information; intermediate processing results are generated, including structured network detection data and configuration refresh status information. If data anomalies are detected in the first iteration, the detection process is allowed to be executed again within the upper limit of the number of iterations to ensure accurate results.
[0086] In another embodiment of this specification, the user verbally commands "Borrow 2,000 and view the three-phase plan." The main control sub-model generates a collaborative planning instruction, which includes the following task chain: sub-task 1, credit limit verification; sub-task 2, plan generation; sub-task 3, phase number matching; the dependency relationship is credit limit verification, plan generation, and phase number matching. The maximum number of iterations is set to 3 according to the planning instruction, so that predictions can be retried or parameters can be supplemented if credit limit verification fails or plan generation is abnormal. The execution process is as follows: the scheduled execution sub-model A calls the underlying credit limit interface to generate the credit limit verification result; the scheduled execution sub-model B calls the phase calculation interface to generate candidate phase plans; the scheduled execution sub-model C executes the phase number matching task according to the plan and the "three phases" parameters; the output of each task execution is integrated to form an intermediate processing result. The final intermediate processing result includes credit limit information, candidate phase data, and phase number matching status, which are then processed by the subsequent fusion sub-model.
[0087] Specifically, in some other embodiments, the specific implementation of step S220 can be found in the following embodiments. This embodiment is a detailed description of step S220 in the information acquisition method shown in the above embodiments. In the information acquisition method, the large language model includes a deterministic business logic engine, and step S220 may further include the following steps:
[0088] Structured parameters extracted from the standardized instructions through a deterministic business logic engine.
[0089] The structured parameters are processed and verified to obtain a reference processing result.
[0090] The intermediate processing result is updated based on the reference processing result.
[0091] In this embodiment, the deterministic business logic engine can extract structured parameters from standardized instructions, making the parameters clear in their boundaries and directly usable. Furthermore, it can perform deterministic validation on these structured parameters, including logic validation, rule validation, and range validation, ensuring business rigor. After intermediate processing results are generated, they are updated by referencing the previous results, preventing numerical errors, rule conflicts, or business inconsistencies in the final output caused by large model inference biases. This enhances the determinism of the large model in business task execution, making model inference and business rules complementary.
[0092] Specifically, after parsing standardized instructions, the large language model submits the key entities within the instructions to the deterministic business logic engine. The engine performs structured parameter operations based on parameter types (such as numbers, times, periods, configuration item names, etc.). Specifically, it converts natural language numbers into standard numerical formats; standardizes time expressions (such as "3 PM") into 24-hour timestamps; normalizes object identifiers into the encoding format used internally by the system; and converts the business operation fields involved in the instructions into predefined enumeration types. This ultimately generates a specific set of parameters.
[0093] Then, the deterministic business logic engine performs deterministic validation on the structured parameters according to predefined rules. Validation may include range validation, logical validation, conflict validation, and computational calculations. Range validation determines whether the value is within a valid range; logical validation determines whether the parameter combination conforms to regulations, such as the time limit not exceeding the maximum available time limit; conflict validation determines whether there is a conflict with existing system states, such as the inability to configure VLANs if a network port is not enabled; computational calculations perform calculations based on the parameters, such as interest rate calculations and capacity verification. The validation results will form a reference processing result, including information such as parameter validity, correction rules, and calculation outputs.
[0094] If the intermediate processing result generated by the execution sub-model is inconsistent with the reference processing result, the intermediate processing result will be corrected. This mainly includes correcting parameter values, such as adjusting illegal values to the closest legal value range; covering calculation results that do not conform to business rules; deleting invalid task nodes generated during execution; and adding necessary but missing structured parameters. The corrected intermediate processing result will be submitted to the fusion sub-model to generate composite feedback information.
[0095] In some embodiments of this specification, the user voice inputs "adjust the port to 30 Mbps". The system generates a standardized instruction to adjust the port rate to 30 Mbps through preceding steps. The structured parameters include the port identifier: Port1 and the target rate: 30 Mbps. Further, the deterministic business logic engine looks up the legal rate range supported by this port in the device configuration rule base, which is 10 Mbps, 100 Mbps, and 1000 Mbps. Therefore, the reference processing result is that the rate is invalid and it is recommended to adjust it to 100 Mbps. If the intermediate processing result generated by the execution sub-model is still 30 Mbps, the system corrects the intermediate processing result to 100 Mbps according to the reference result and writes the correction mark into the intermediate processing data for use in the subsequent fusion stage.
[0096] In some other embodiments of this specification, the user voice inputs "borrow 15,000 yuan". The system, after previous steps, obtains a standardized instruction to borrow 15,000 yuan. The structured parameters include the loan amount: 15,000 yuan and the user's credit limit: 12,000 yuan. Further, the deterministic business logic engine, based on business rules, verifies that the loan amount exceeds the credit limit, requiring correction to a maximum loan amount of 12,000 yuan. Therefore, the suggested loan amount is 12,000 yuan, and the user is prompted about the credit limit. If the intermediate processing result generated by the execution sub-model contains the amount 15,000 yuan, the system updates this value to 12,000 yuan and adds a description field to the intermediate processing result, indicating that the value has been verified and corrected. The corrected intermediate processing result is submitted to the fusion sub-model to generate composite feedback information with voice prompts and visual components.
[0097] In S230, the intermediate processing results are input into the fusion sub-model, which identifies the key semantic content to be presented and generates corresponding natural language text feedback based on the content. At the same time, the front-end component mapping information related to the semantics is added to the output content to form composite feedback information.
[0098] Specifically, in some embodiments, the specific implementation of step S230 can be found in the following embodiments. This embodiment is based on... Figure 3 According to the detailed description of step S230 in the information acquisition method shown in the corresponding embodiment, step S230 in the information acquisition method may include the following steps:
[0099] Identify key semantic entities in the intermediate processing results and retrieve the front-end interaction component identifiers associated with the key semantic entities.
[0100] Based on the duration prediction model of speech synthesis, the time offset of the key semantic entity in the speech stream is calculated, and the corresponding control anchor is inserted into the data stream of the natural language text instruction response. The control anchor includes a trigger timestamp and the corresponding interactive component control instruction.
[0101] Generate composite feedback information containing text data and the control anchor point.
[0102] In this embodiment, the fusion sub-model can identify key semantic entities from intermediate processing results and map them to corresponding front-end interactive component identifiers. Furthermore, it can calculate the time offset of key semantic entities in the speech stream based on a speech synthesis duration prediction model. Control anchors containing trigger timestamps and interactive component control instructions are inserted into the natural language response data stream based on these time offsets. Finally, composite feedback information containing text data and control anchors is generated, achieving time alignment between the speech response and the interface display. This allows the language model output to directly drive dynamic changes in the front-end interface, a key capability that traditional text-output-based LLM models cannot achieve.
[0103] Specifically, after receiving the intermediate processing results, the fusion sub-model performs semantic annotation and entity recognition on these results. Key semantic entities may include: amount entities, status entities, time entities, result entities, configuration entities, etc. The fusion sub-model accesses a pre-built component mapping table based on the entity type to retrieve the corresponding component identifier. For example, amount corresponds to a quota display component, status corresponds to a status card component, time corresponds to a schedule display component, and configuration corresponds to a device status component. Through component identifier retrieval, the fusion sub-model can bind semantic outputs to specific interactive components.
[0104] Based on the characteristics of TTS synthesis, a duration prediction model is used to estimate the pronunciation duration of each word or entity in natural language text. The prediction model input includes the text in the key semantic entity, the speech feature parameters used by the TTS model, and the influence factor of the context text on the pronunciation duration. The final output is the time offset of the key entity in the speech stream and the duration of the entity's continuous pronunciation. The time offset is in milliseconds and is used to determine the triggering time of the component. The duration of the entity's continuous pronunciation is used for accurate synchronization of the interface display.
[0105] Based on the time offset, control anchors are inserted into the sequence of natural language text. Each control anchor contains a trigger timestamp and interactive component control instructions. The trigger timestamp identifies a point in time when voice playback reaches that point, triggering a component update. Interactive component control instructions include behaviors such as component rendering, display, refreshing, or highlighting. It's important to note that the basic structure of a control anchor includes control code, a target component identifier, a trigger condition, and an action type, such as rendering, expanding, or updating. Using control anchors, the client can accurately synchronize component rendering behavior during voice playback.
[0106] In some embodiments of this specification, if the intermediate processing result includes the current rate of port 1 being 100Mbps, the fusion sub-model identifies the key semantic entities "port 1" and "100Mbps". Based on the entity type, it retrieves the component identifier corresponding to port 1 (port status component) and the component identifier corresponding to 100Mbps (rate display component). Using a duration prediction model, it is estimated that port 1 will appear at 500ms in the audio stream, and 100Mbps will appear at 1200ms. Therefore, the port status component is rendered at 500ms, and the rate display component is refreshed at 1200ms. The fusion sub-model inserts control anchors into the text stream to generate composite feedback information. The client automatically and synchronously displays the port status and rate information during audio playback.
[0107] In some other embodiments of this specification, the intermediate processing result includes a loan amount of 5,000 yuan, with the option of 3, 6, or 12 installments.
[0108] The fusion sub-model identifies semantic entities, including the monetary entity "5000 yuan" and installment option entities "3 months," "6 months," and "12 months." Based on entity type, it retrieves the corresponding component identifiers: the credit limit display component for the monetary entity and the installment plan component for the installment entity. The duration prediction model predicts that the text "5000 yuan" appears approximately 850ms after voice playback, while "3 months," "6 months," and "12 months" appear sequentially between approximately 1300ms and 1600ms. Therefore, control anchors are inserted based on the time offset: the credit limit display component is triggered to render at 850ms; the installment plan component is triggered to expand at 1300ms; and all installment options are updated at 1600ms. The final composite feedback information includes natural language text and precisely aligned interactive control anchors, ensuring that when the voice announces "You can borrow 5000 yuan," the interface simultaneously displays the credit limit card; and when the 3, 6, and 12 installment plans are announced, the interface simultaneously displays the installment plan.
[0109] In the S300, the composite feedback information generated by the fusion sub-model is encapsulated and transmitted to the client. This composite feedback information includes natural language text content, speech synthesis markers synchronized with the text, and identifiers for front-end interactive components. Upon receiving this information, the client reads the feedback text aloud through the speech synthesis module and renders the corresponding page components based on the component identifiers and control commands in the composite feedback. This ensures that the speech feedback and interface changes are presented synchronously, enabling users to interact with the system using natural speech and instantly obtain an immersive and visual interactive experience.
[0110] In some embodiments of this specification, the above method is applied to a network operation and maintenance and equipment management system, which includes a voice acquisition and recognition module, a main control sub-model, an execution sub-model, a deterministic business logic engine, a fusion sub-model, a front-end interaction module, etc.
[0111] The system comprises several sub-models: a voice acquisition and recognition module that collects user voice data, performs noise reduction and endpoint detection, and decodes the voice data into standardized text; a main control sub-model responsible for semantic understanding, task decomposition, dependency planning, and generating collaborative planning instructions for user text commands; an execution sub-model including port query, device configuration, alarm analysis, and bandwidth calculation sub-models; a deterministic business logic engine responsible for logical validation, range validation, and rule calculation of structured parameters; and a fusion sub-model responsible for integrating intermediate processing results and inserting them into front-end component control anchors. The front-end interaction module drives interface component updates and synchronizes voice output based on these control anchors.
[0112] After the user issues a voice command to "check port 1 status and adjust the rate to 30Mbps", the voice recognition module converts the user's voice into a text command to check the port 1 status and adjust the rate to 30Mbps. If the confidence level is insufficient, a fallback mechanism for text input is triggered. The main control sub-model parses the text, identifies two sub-tasks: querying the port 1 status and adjusting the port 1 rate to 30Mbps, and identifies the dependency relationship, i.e., task 2 depends on task 1 to complete. The main control sub-model generates a collaborative planning instruction and sets an upper limit for the number of iterations (e.g., 2 times). The system schedules the corresponding execution sub-models according to the task dependency order. First, the port query sub-model is scheduled to call the underlying device interface to obtain the port status (such as UP / DOWN, current rate, number of error packets, etc.) and generate the query result; then, the port configuration sub-model is scheduled to execute the rate adjustment logic according to the planning instruction, forming an intermediate processing result: {port status information: UP, current rate: 100Mbps, target rate: 30Mbps}. Then, the deterministic logic engine verifies the validity of the speed. If the device supports speeds of 10 / 100 / 1000 Mbps, but the user's requested speed of 30 Mbps is not in the valid list, the verification outputs a reference processing result, suggesting adjustment to the most recent valid speed of 100 Mbps. The main control system updates the intermediate processing result based on the reference result. Finally, the fusion sub-model identifies key semantic entities, including port 1 being in a normal state and supporting a speed of 100 Mbps. The component mapping table is queried to obtain the state component corresponding to the state and the configuration component corresponding to the speed. The time offset of the entity in the speech is calculated using the TTS duration prediction model; for example, port 1 being in a normal state is at 500ms, and 100 Mbps is at 1200ms. The fusion sub-model inserts control anchors, rendering the port state component at 500ms and the speed configuration component at 1200ms, generating composite feedback information and sending it to the client. The client uses the control anchors to drive interface updates, ensuring complete synchronization between the interface and the speech.
[0113] In other embodiments of this specification, such as Figure 4 As shown, the above method is applied to a credit system, which also includes a voice acquisition and recognition module, a main control sub-model, an execution sub-model, a deterministic business logic engine, a fusion sub-model, a front-end interaction module, etc. However, its execution sub-model types include credit limit calculation sub-model, installment recommendation sub-model, preferential rights sub-model, risk verification sub-model, and repayment date calculation sub-model, etc.
[0114] The user issues a voice command, "Borrow 2,000 and recommend the cheapest option." The voice recognition module decodes the voice as "Request a loan of 2,000 yuan and recommend the most favorable option." If the recognition confidence is low, a text fallback mechanism is triggered. The main control sub-model recognizes the loan amount as 2,000 yuan and the intention to recommend the most favorable option, which depends on information such as the user's credit limit, available loan period, interest rate, and discount records. The task is broken down into credit limit verification, option generation, discount filtering, and recommendation of the cheapest option. A dependency chain is established and an iteration limit is set (e.g., 3 times). The execution sub-models are executed in the dependency order of the credit limit verification sub-model, installment plan sub-model, discount rights sub-model, and option recommendation sub-model, outputting intermediate processing results, for example, {credit limit: 1,500, requested amount: 2,000, insufficient available amount}.
[0115] The logic engine identifies and adjusts insufficient credit limits. Since the recommended maximum borrowable limit is 1500, it marks the need to inform the customer of the insufficient limit and updates the intermediate processing results. The fusion sub-model identifies key entities, including the borrowable limit of 1500 yuan, the recommended installment plan of 1500 yuan × 6 months, and the 20 yuan discount coupon, etc. These entities are mapped to the credit limit display component, installment plan component, and discount component, respectively. The entity appearance time is predicted using duration (e.g., 700ms, 1400ms, 1800ms), and anchor points are inserted. Composite feedback information is output and synchronously displayed on the client side.
[0116] The following will combine Figure 5 This manual provides a detailed description of the voice service interaction device provided. It should be noted that... Figure 5 The voice service interaction device shown is used to execute this specification. Figures 1-6 The methods of the embodiments shown are illustrated only in connection with this specification for ease of explanation. For specific technical details not disclosed, please refer to this specification. Figures 1-4 The example shown.
[0117] Please see Figure 5 This diagram illustrates the structure of the voice service interaction device 500 described in this specification. The voice service interaction device 500 can be implemented as all or part of a user terminal through software, hardware, or a combination of both. According to some embodiments, the voice service interaction device 500 includes an instruction generation module 510, an instruction input module 520, and an information feedback module 530.
[0118] The instruction generation module 510 is used to generate standardized instructions in response to service interaction requests, wherein the service interaction requests include unstructured voice data and the standardized instructions are standardized natural language text instructions; the instruction input module 520 is used to input the standardized instructions into a large language model to obtain composite feedback information; and the information feedback module 530 is used to feed the composite feedback information back to the client so that the client can synchronously output voice feedback and render the corresponding interactive component page.
[0119] The aforementioned large language model includes a master control sub-model, a fusion sub-model, and a pre-built functional component library. The pre-built functional component library contains multiple execution sub-models, and the instruction input module 520 specifically includes: a semantic analysis sub-module 521, an instruction scheduling sub-module 522, and a result fusion sub-module 523.
[0120] The semantic analysis submodule 521 is used to input the standardized instructions into the main control submodel for semantic analysis and generate collaborative planning instructions, which include multiple sub-tasks; the instruction scheduling submodule 522 is used to schedule the corresponding execution submodel according to the collaborative planning instructions, so that the execution submodel calls the underlying data interface to execute the corresponding sub-tasks and generate intermediate processing results; the result fusion submodule 523 is used to input the intermediate processing results into the fusion submodel to generate composite feedback information.
[0121] Optionally, the instruction generation module 510 specifically includes: a request parsing submodule, used to respond to a service interaction request, parse the service interaction request, and obtain unstructured speech data; and a decoding and conversion submodule, used to call the speech recognition module to decode and convert the unstructured speech data to generate standardized instructions.
[0122] Optionally, the semantic analysis submodule 521 specifically includes: a semantic analysis unit, used to perform semantic analysis on the standardized instructions through the main control sub-model, parse the key entities in the standardization, and determine the user's operation type; and a collaborative planning unit, used to generate collaborative planning instructions based on the operation type and key entities, wherein the collaborative planning instructions contain the dependencies between their respective tasks.
[0123] Optionally, the semantic analysis submodule 521 specifically includes: an index calculation unit, used to perform semantic analysis on the standardized instructions through the master control sub-model and calculate the intent uncertainty index; a multi-step reasoning unit, used to generate collaborative planning instructions in the form of a multi-step reasoning chain if the intent uncertainty index is higher than a preset index threshold; and a direct execution unit, used to generate collaborative planning instructions in the form of a direct execution chain if the intent uncertainty index is lower than a preset index threshold.
[0124] Optionally, the device further includes a fallback trigger module, used to trigger a fallback expression template and output guiding text through the fusion sub-model if the intermediate processing result is missing a key field, so as to guide the user to supplement the missing information and conduct subsequent interactions.
[0125] Optionally, the instruction scheduling submodule 522 specifically includes: an iteration count unit, used to set an upper limit for the number of iterations according to the collaborative planning instruction; and a sequence scheduling unit, used to schedule the corresponding execution sub-models in sequence according to the task dependencies within the upper limit of the number of iterations, so that each sub-model executes its corresponding sub-tasks in sequence and generates intermediate processing results.
[0126] Optionally, the large language model includes a deterministic business logic engine; the instruction scheduling submodule 522 further includes: a parameter extraction unit, used to extract structured parameters from the standardized instructions through the deterministic business logic engine; an operation verification unit, used to perform operation verification on the structured parameters to obtain a reference processing result; and a result update unit, used to update the intermediate processing result according to the reference processing result.
[0127] Optionally, the result fusion submodule 523 specifically includes: an entity recognition unit, used to identify key semantic entities in the intermediate processing results and retrieve the front-end interactive component identifiers associated with the key semantic entities; a duration prediction unit, used to calculate the time offset of the key semantic entities in the speech stream based on a duration prediction model of speech synthesis, and insert corresponding control anchors into the data stream of the natural language text instruction response, wherein the control anchors include a trigger timestamp and the corresponding interactive component control instructions; and a composite feedback unit, used to generate composite feedback information containing text data and the control anchors.
[0128] It should be noted that the voice service interaction device provided in the above embodiments is only illustrated by the division of the above functional modules when executing the voice service interaction method. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the voice service interaction device and the voice service interaction method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0129] The serial numbers in this specification are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0130] This specification describes the transformation of unstructured speech data into standardized natural language instructions, enabling subsequent models to perform semantic understanding and reasoning based on a unified format. A large language model architecture, comprising a master control sub-model, an execution sub-model, and a fusion sub-model, is employed to process user speech intent in multiple stages, ensuring stable execution of multiple sub-tasks. Multiple execution sub-models are scheduled based on collaborative planning instructions, and intermediate processing results are generated using underlying data interfaces to achieve controllable execution of multiple business logics. The fusion sub-model generates composite feedback content containing natural language text information and interactive component control information to support the synchronous collaboration between client-side speech feedback playback and business component rendering. This achieves complete link control from speech input to execution feedback, ensuring that model output satisfies both semantic understanding and drives business actions, while maintaining synchronization between speech output and client-side interface rendering.
[0131] This specification also provides a computer storage medium capable of storing multiple instructions adapted to be loaded and executed by a processor as described above. Figures 1-4 The voice service interaction method described in the illustrated embodiment can be found in the following document for a detailed execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.
[0132] This specification also provides a computer program product storing at least one instruction, which is loaded and executed by the processor as described above. Figures 1-4 The voice service interaction method described in the illustrated embodiment can be found in the following documentation for its specific execution process. Figures 1-4 The specific details of the illustrated embodiments will not be elaborated here.
[0133] Please refer to Figure 6 This diagram illustrates a structural block diagram of an electronic device provided in an exemplary embodiment of this specification. The electronic device in this specification may include one or more components such as a processor 110, a memory 120, an input device 130, an output device 140, and a bus 150. The processor 110, memory 120, input device 130, and output device 140 may be connected via the bus 150.
[0134] Processor 110 may include one or more processing cores. Processor 110 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in memory 120, and by calling data stored in memory 120. Optionally, processor 110 may be implemented using at least one hardware form of digital signal processing (DSP), field-programmable gate array (FPGA), or programmable logic array (PLA). Processor 110 may integrate one or more of a central processing unit (CPU), graphics processing unit (GPU), and modem. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the displayed content; and the modem handles wireless communication. It is understood that the modem may also not be integrated into processor 110 and may be implemented separately using a communication chip.
[0135] The memory 120 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 120 may include a non-transitory computer-readable storage medium. The memory 120 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 120 may include a program storage area and a data storage area. The program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as touch functionality, sound playback functionality, image playback functionality, etc.), instructions for implementing the various method embodiments described below, etc. The operating system may be the Android system, including systems deeply developed based on the Android system, the iOS system developed by Apple Inc., including systems deeply developed based on the iOS system, or other systems. The data storage area may also store data created by the electronic device during use, such as phonebook data, audio and video data, chat log data, etc.
[0136] See Figure 7As shown, the memory 120 can be divided into operating system space and user space. The operating system runs in the operating system space, while native and third-party applications run in the user space. To ensure that different third-party applications can achieve good running performance, the operating system allocates corresponding system resources for each application. However, different application scenarios within the same third-party application have different requirements for system resources. For example, in local resource loading scenarios, third-party applications have high requirements for disk read speed; in animation rendering scenarios, third-party applications have high requirements for GPU performance. Since the operating system and third-party applications are independent of each other, the operating system often cannot promptly perceive the current application scenario of a third-party application, resulting in the operating system's inability to adapt system resources accordingly to the specific application scenario of the third-party application.
[0137] In order for the operating system to distinguish the specific application scenarios of third-party applications, it is necessary to establish data communication between the third-party applications and the operating system. This would allow the operating system to obtain the current scenario information of the third-party applications at any time, and then perform targeted system resource adaptation based on the current scenario.
[0138] Taking the Android operating system as an example, the programs and data stored in memory 120 are as follows: Figure 8As shown, the memory 120 can store the Linux kernel layer 320, the system runtime library layer 340, the application framework layer 360, and the application layer 380. The Linux kernel layer 320, system runtime library layer 340, and application framework layer 360 belong to the operating system space, while the application layer 380 belongs to the user space. The Linux kernel layer 320 provides low-level drivers for various hardware components of the electronic device, such as display drivers, audio drivers, camera drivers, Bluetooth drivers, Wi-Fi drivers, and power management. The system runtime library layer 340 provides support for key features of the Android system through several C / C++ libraries. For example, the SQLite library provides database support, the OpenGL / ES library provides 3D graphics support, and the Webkit library provides browser kernel support. The system runtime library layer 340 also provides the Android runtime library, which mainly provides core libraries that allow developers to write Android applications using the Java language. The Application Framework Layer 360 provides various APIs that may be used when building applications. Developers can also use these APIs to build their own applications, such as activity management, window management, view management, notification management, content provider, package management, call management, resource management, and location management. At least one application runs in the Application Layer 380. These applications can be native applications that come with the operating system, such as contacts, SMS, clock, and camera apps; or third-party applications developed by third-party developers, such as games, instant messaging, and photo editing apps.
[0139] Taking the operating system as an example (iOS), the programs and data stored in memory 120 are as follows: Figure 9As shown, the iOS system includes: Core OS layer 420, Core Services layer 440, Media layer 460, and Cocoa Touch layer 480. Core OS layer 420 includes the operating system kernel, drivers, and low-level program frameworks. These low-level program frameworks provide hardware-level functionality for use by the program frameworks located in Core Services layer 440. Core Services layer 440 provides system services and / or program frameworks required by applications, such as Foundation framework, account framework, advertising framework, data storage framework, network connectivity framework, geolocation framework, motion framework, etc. Media layer 460 provides applications with audiovisual interfaces, such as interfaces related to graphics and images, audio technology, video technology, and wireless playback (AirPlay) interfaces. Cocoa Touch layer 480 provides various commonly used interface-related frameworks for application development and is responsible for user touch interaction on electronic devices. Examples include local notification services, remote push services, advertising frameworks, game tool frameworks, message user interface (UI) frameworks, UIKit user interface frameworks, map frameworks, and so on.
[0140] exist Figure 9 The framework shown includes, but is not limited to, the base framework in the core service layer 440 and the UIKit framework in the touchable layer 480. The base framework provides many basic object classes and data types, offering the most basic system services to all applications, and is independent of the UI. The UIKit framework, on the other hand, provides a basic UI class library for creating touch-based user interfaces. iOS applications can use the UIKit framework to provide their UI, thus providing the application's infrastructure for building user interfaces, drawing, handling user interaction events, responding to gestures, and so on.
[0141] The methods and principles for implementing data communication between third-party applications and the operating system in the iOS system can be found in the Android system, and will not be repeated here.
[0142] The input device 130 is used to receive input instructions or data, and includes, but is not limited to, a keyboard, mouse, camera, microphone, or touch device. The output device 140 is used to output instructions or data, and includes, but is not limited to, a display device and a speaker. In one example, the input device 130 and the output device 140 can be combined into a touch screen, which is used to receive touch operations from the user using a finger, stylus, or any suitable object on or near it, and to display the user interface of various applications. The touch screen is usually located on the front panel of the electronic device. The touch screen can be designed as a full-screen, curved screen, or irregularly shaped screen. The touch screen can also be designed as a combination of a full-screen and a curved screen, or a combination of an irregularly shaped screen and a curved screen; this specification does not limit this.
[0143] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more or fewer components than shown, or combine certain components, or have different component arrangements. For example, the electronic device may also include radio frequency circuits, input units, sensors, audio circuits, wireless fidelity (WiFi) modules, power supplies, Bluetooth modules, etc., which will not be described in detail here.
[0144] In this specification, the entity executing each step can be the electronic device described above. Optionally, the entity executing each step can be the operating system of the electronic device. The operating system can be Android, iOS, or other operating systems; this specification does not limit this.
[0145] The electronic device described in this manual may also be equipped with a display device. This display device can be any device capable of displaying information, such as a cathode ray tube display (CR), a light-emitting diode display (LED), an e-ink screen, a liquid crystal display (LCD), or a plasma display panel (PDP). Users can use the display device on the electronic device to view displayed text, images, videos, and other information. The electronic device may include smartphones, tablets, gaming devices, AR (Augmented Reality) devices, automobiles, data storage devices, audio playback devices, video playback devices, laptops, desktop computing devices, and wearable devices such as electronic watches, electronic glasses, electronic helmets, electronic bracelets, electronic necklaces, and electronic clothing.
[0146] exist Figure 6 In the illustrated electronic device, which can be a terminal, the processor 110 can be used to call the network optimization application stored in the memory 120 and specifically perform the following operations: In response to a service interaction request, it generates standardized instructions, wherein the service interaction request contains unstructured speech data and the standardized instructions are standardized natural language text instructions; it inputs the standardized instructions into a large language model to obtain composite feedback information; it feeds the composite feedback information back to the client so that the client can synchronously output speech feedback and render the corresponding interactive component page; wherein the large language model includes a main control sub-model, a fusion sub-model, and a pre-set functional component library, and the pre-set functional component library contains multiple execution sub-models; when the processor 110 executes the operation of inputting the standardized instructions into the large language model to obtain composite feedback information, it specifically performs the following operations: it inputs the standardized instructions into the main control sub-model for semantic analysis to generate collaborative planning instructions, wherein the collaborative planning instructions contain multiple sub-tasks; it schedules the corresponding execution sub-models according to the collaborative planning instructions, so that the execution sub-models call the underlying data interface to execute the corresponding sub-tasks and generate intermediate processing results; it inputs the intermediate processing results into the fusion sub-model to generate composite feedback information.
[0147] In one embodiment, when the processor 110 generates standardized instructions in response to a service interaction request, it specifically performs the following operations: in response to the service interaction request, it parses the service interaction request to obtain unstructured speech data; and it calls the speech recognition module to decode and convert the unstructured speech data to generate standardized instructions.
[0148] In one embodiment, when the processor 110 performs semantic analysis on the standardized instructions input into the main control sub-model to generate collaborative planning instructions, it specifically performs the following operations: performs semantic analysis on the standardized instructions through the main control sub-model, parses the key entities in the standardization, and determines the user's operation type; generates collaborative planning instructions based on the operation type and key entities, wherein the collaborative planning instructions contain the dependencies between their respective tasks.
[0149] In one embodiment, when the processor 110 performs semantic analysis on the standardized instructions input into the master control sub-model to generate collaborative planning instructions, it specifically performs the following operations: performs semantic analysis on the standardized instructions through the master control sub-model to calculate the intent uncertainty index; if the intent uncertainty index is higher than a preset index threshold, then generates collaborative planning instructions in the form of a multi-step inference chain; if the intent uncertainty index is lower than the preset index threshold, then generates collaborative planning instructions in the form of a direct execution chain.
[0150] In one embodiment, after the processor 110 executes the execution sub-model scheduled according to the collaborative planning instruction, causing the execution sub-model to call the underlying data interface to execute the corresponding sub-task and generate intermediate processing results, it also performs the following operations: if the intermediate processing results are missing key fields, the processor 110 triggers the fallback expression template through the fusion sub-model and outputs guiding text to guide the user to supplement the missing information and conduct subsequent interactions.
[0151] In one embodiment, when the processor 110 executes the corresponding execution sub-models scheduled according to the collaborative planning instructions, and causes the execution sub-models to call the underlying data interface to execute the corresponding sub-tasks and generate intermediate processing results, the processor 110 specifically performs the following operations: according to the collaborative planning instructions, it sets an upper limit for the number of iterations; within the upper limit of the number of iterations, according to the task dependencies, it schedules the corresponding execution sub-models in sequence, so that each sub-model executes the corresponding sub-tasks in sequence and generates intermediate processing results.
[0152] In one embodiment, the large language model includes a deterministic business logic engine; when the processor 110 executes the corresponding execution sub-model scheduled according to the collaborative planning instruction, causing the execution sub-model to call the underlying data interface to execute the corresponding sub-task and generate intermediate processing results, it also performs the following operations: extracting structured parameters from the standardized instructions through the deterministic business logic engine; performing calculation verification on the structured parameters to obtain a reference processing result; and updating the intermediate processing result according to the reference processing result.
[0153] In one embodiment, when the processor 110 executes the process of inputting the intermediate processing result into the fusion sub-model to generate composite feedback information, it specifically performs the following operations: identifying key semantic entities in the intermediate processing result and retrieving the front-end interactive component identifier associated with the key semantic entity; calculating the time offset of the key semantic entity in the speech stream based on the duration prediction model of speech synthesis, and inserting the corresponding control anchor point into the data stream of the natural language text instruction response, wherein the control anchor point includes a trigger timestamp and the corresponding interactive component control instruction; and generating composite feedback information containing text data and the control anchor point.
[0154] This specification describes the transformation of unstructured speech data into standardized natural language instructions, enabling subsequent models to perform semantic understanding and reasoning based on a unified format. A large language model architecture, comprising a master control sub-model, an execution sub-model, and a fusion sub-model, is employed to process user speech intent in multiple stages, ensuring stable execution of multiple sub-tasks. Multiple execution sub-models are scheduled based on collaborative planning instructions, and intermediate processing results are generated using underlying data interfaces to achieve controllable execution of multiple business logics. The fusion sub-model generates composite feedback content containing natural language text information and interactive component control information to support the synchronous collaboration between client-side speech feedback playback and business component rendering. This achieves complete link control from speech input to execution feedback, ensuring that model output satisfies both semantic understanding and drives business actions, while maintaining synchronization between speech output and client-side interface rendering.
[0155] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory, or random access memory, etc.
[0156] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of related data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the object characteristics, interactive behavior characteristics, and user information involved in this specification were all obtained under full authorization.
[0157] The above-disclosed embodiments are merely preferred embodiments of this specification and should not be construed as limiting the scope of this specification. Therefore, any equivalent variations made in accordance with the claims of this specification shall still fall within the scope of this specification.
Claims
1. A voice service interaction method, characterized in that, The method includes: In response to a service interaction request, a standardized instruction is generated, wherein the service interaction request contains unstructured voice data and the standardized instruction is a standardized natural language text instruction. The standardized instructions are input into the large language model to obtain composite feedback information; The composite feedback information is fed back to the client so that the client can synchronously output voice feedback and render the corresponding interactive component page; The large language model includes a main control sub-model, a fusion sub-model, and a pre-built functional component library, wherein the pre-built functional component library contains multiple execution sub-models. The step of inputting the standardized instructions into the large language model to obtain composite feedback information specifically includes: The standardized instructions are input into the main control sub-model for semantic analysis to generate collaborative planning instructions, which contain multiple sub-tasks. According to the collaborative planning instructions, the corresponding execution sub-model is scheduled, and the execution sub-model calls the underlying data interface to execute the corresponding sub-task and generate intermediate processing results. The intermediate processing results are input into the fusion sub-model to generate composite feedback information; The step of inputting the intermediate processing result into the fusion sub-model to generate composite feedback information specifically includes: Identify key semantic entities in the intermediate processing results, and retrieve the front-end interactive component identifiers associated with the key semantic entities by accessing the preset component mapping table according to the entity type. The key semantic entities include: amount entity, status entity, time entity, result entity, and configuration entity. The duration prediction model based on speech synthesis calculates the time offset of the key semantic entity in the speech stream and the duration of the entity's continuous pronunciation. It then inserts corresponding control anchors into the data stream of the natural language text instruction response. Each control anchor includes a trigger timestamp and corresponding interactive component control instructions. The trigger timestamp identifies a point in time when speech playback reaches that point, triggering a component update. Interactive component control instructions include component rendering, display, refresh, or highlighting. The basic structure of the control anchor includes control code, target component identifier, trigger condition, and execution action type. The input to the duration prediction model includes the text in the key semantic entity, the speech feature parameters used by the TTS model, and the influence factor of the context text on the pronunciation duration. Generate composite feedback information containing text data and the control anchor point.
2. The method according to claim 1, characterized in that, The process of generating standardized instructions in response to service interaction requests specifically includes: In response to a service interaction request, the service interaction request is parsed to obtain unstructured voice data; The speech recognition module is invoked to decode and convert unstructured speech data, generating standardized instructions.
3. The method according to claim 1, characterized in that, The step of inputting the standardized instructions into the master control sub-model for semantic analysis to generate collaborative planning instructions specifically includes: The standardized instructions are semantically analyzed by the main control sub-model to parse the key entities in the standardization and determine the user's operation type. Based on the operation type and key entities, collaborative planning instructions are generated, which contain the dependencies between their respective tasks.
4. The method according to claim 1, characterized in that, The step of inputting the standardized instructions into the master control sub-model for semantic analysis to generate collaborative planning instructions specifically includes: The standardized instructions are semantically analyzed using the master control sub-model to calculate the intent uncertainty index; If the intent uncertainty index is higher than the preset index threshold, then a collaborative planning instruction is generated in the form of a multi-step inference chain. If the intent uncertainty index is lower than the preset index threshold, then collaborative planning instructions are generated in the form of a direct execution chain.
5. The method according to claim 1, characterized in that, The method further includes: If the intermediate processing result is missing a key field, the fallback expression template is triggered through the fusion sub-model and guiding text is output to guide the user to fill in the missing information and conduct subsequent interactions.
6. The method according to claim 1, characterized in that, The step of scheduling the corresponding execution sub-model according to the collaborative planning instruction, causing the execution sub-model to call the underlying data interface to execute the corresponding sub-task and generate intermediate processing results, specifically includes: Based on the collaborative planning instructions, set an upper limit on the number of iterations; Within the upper limit of the number of iterations, the corresponding execution sub-models are scheduled in sequence according to the task dependencies, so that each sub-model executes its corresponding sub-tasks in sequence and generates intermediate processing results.
7. The method according to claim 6, characterized in that, The large language model includes a deterministic business logic engine; The step of scheduling the corresponding execution sub-model according to the collaborative planning instruction, causing the execution sub-model to call the underlying data interface to execute the corresponding sub-task and generate intermediate processing results, also includes: Structured parameters extracted from the standardized instructions through a deterministic business logic engine; The structured parameters are processed and validated to obtain a reference processing result; The intermediate processing result is updated based on the reference processing result.
8. A voice service interaction device, characterized in that, The voice service interaction device includes: The instruction generation module is used to generate standardized instructions in response to service interaction requests, wherein the service interaction requests contain unstructured voice data and the standardized instructions are standardized natural language text instructions. The instruction input module is used to input the standardized instructions into the large language model to obtain composite feedback information. The information feedback module is used to send the composite feedback information back to the client so that the client can synchronously output voice feedback and render the corresponding interactive component page; The large language model includes a main control sub-model, a fusion sub-model, and a pre-built functional component library, wherein the pre-built functional component library contains multiple execution sub-models. The instruction input module specifically includes: The semantic analysis submodule is used to input the standardized instructions into the main control sub-model for semantic analysis and generate collaborative planning instructions, which contain multiple sub-tasks. The instruction scheduling submodule is used to schedule the corresponding execution sub-model according to the collaborative planning instruction, so that the execution sub-model calls the underlying data interface to execute the corresponding sub-task and generate intermediate processing results. The result fusion submodule is used to input the intermediate processing results into the fusion submodel to generate composite feedback information; The result fusion submodule specifically includes: An entity recognition unit is used to identify key semantic entities in the intermediate processing results and retrieve the front-end interactive component identifier associated with the key semantic entity by accessing a preset component mapping table according to the entity type. The key semantic entities include: amount entity, status entity, time entity, result entity, and configuration entity. The duration prediction unit is used to calculate the time offset of the key semantic entity in the speech stream and the duration of the entity's continuous pronunciation based on the speech synthesis duration prediction model. It also inserts the corresponding control anchor point into the data stream of the natural language text instruction response. The control anchor point includes a trigger timestamp and the corresponding interactive component control instruction. The trigger timestamp identifies a time point when the speech playback reaches that time point, triggering the component update. The interactive component control instruction includes component rendering, display, refresh, or highlighting. The basic structure of the control anchor point includes control code, target component identifier, trigger condition, and execution action type. The input of the duration prediction model includes the text in the key semantic entity, the speech feature parameters used by the TTS model, and the influence factor of the context text on the pronunciation duration. A composite feedback unit is used to generate composite feedback information that includes text data and the control anchor point.
9. A computer storage medium, characterized in that, The computer storage medium stores a plurality of instructions, which are adapted to be loaded by a processor and executed as method steps as claimed in any one of claims 1 to 7.
10. A computer program product, characterized in that, The computer program product stores at least one instruction, which is loaded by a processor and executed as the method steps of any one of claims 1 to 7.
11. An electronic device, characterized in that, include: A processor and a memory; wherein the memory stores a computer program adapted to be loaded by the processor and executed the method steps as claimed in any one of claims 1 to 7.