Voice interaction system interaction interrupting method and device, medium and electronic equipment
By detecting the current state of the voice interaction system, an adaptive interruption strategy was adopted to solve the flexibility problem of in-vehicle voice assistants under a multi-model collaborative architecture, improve user experience and response speed, and ensure the continuity of critical tasks.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- XG TECHNOLOGIES PTE LTD
- Filing Date
- 2026-04-20
- Publication Date
- 2026-07-14
AI Technical Summary
In smart cockpit scenarios, the interruption strategy of in-vehicle voice assistants is not flexible enough, resulting in a poor user experience, especially in multi-model collaborative architectures where it is difficult to respond to user voice commands in a timely manner.
By detecting the current interaction state of the voice interaction system, the corresponding interruption strategy is determined, and different interruption strategies are adopted for different states to execute different schemes to interrupt the current action of the system, including temporary storage, delayed processing, or direct response to user input.
It improves the flexibility of interruption operations and user experience, ensures the continuity and real-time response of critical tasks, and reduces gaps in dialogue history and system stability issues.
Smart Images

Figure CN122392516A_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of human-computer interaction technology, and in particular to a method, device, medium and electronic device for interrupting interaction in a voice interaction system. Background Technology
[0002] In smart cockpit scenarios, if multiple tasks such as navigation, multimedia, and air conditioning are in progress, the in-vehicle voice assistant needs to support users interrupting the ongoing interaction at any time via voice commands so that the system can respond promptly to user voice commands. However, current interruption strategies lack flexibility and result in a poor user experience. Summary of the Invention
[0003] This disclosure provides a method, apparatus, medium, and electronic device for interrupting interaction in a voice interaction system, so as to improve the flexibility of interruption and enhance the user experience.
[0004] The first aspect of this disclosure provides a method for interrupting the interaction of a voice interaction system, comprising: in response to detecting a voice input event, determining the current interaction state of the voice interaction system; determining a target interruption strategy for the voice interaction system based on the current interaction state, wherein the voice interaction system includes multiple interaction states, the multiple interaction states correspond to multiple interruption strategies, the interruption strategy is used to indicate the interruption method of the current action of the voice interaction system, the current action includes the current voice output action and / or the current task execution action; and interrupting the current action of the voice interaction system based on the target interruption strategy.
[0005] A second aspect of this disclosure provides an interaction interruption device for a voice interaction system, comprising: a determining module, configured to determine the current interaction state of the voice interaction system in response to detecting a voice input event; and to determine a target interruption strategy for the voice interaction system based on the current interaction state, wherein the voice interaction system includes multiple interaction states, the multiple interaction states correspond to multiple interruption strategies, the interruption strategy is used to indicate the interruption method of the current action of the voice interaction system, the current action includes the current voice output action and / or the current task execution action; and an interruption execution module, configured to interrupt the current action of the voice interaction system based on the target interruption strategy.
[0006] A third aspect of this disclosure provides a computer-readable storage medium storing a computer program for executing the interaction interruption method of the voice interaction system proposed in the first aspect.
[0007] A fourth aspect of this disclosure provides an electronic device comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the interactive interruption method of the voice interaction system proposed in the first aspect above.
[0008] The interactive interruption method disclosed herein first responds to a voice input event, then the system determines its current state to match a corresponding interruption strategy, and then executes the interruption operation according to the strategy, thereby interrupting the system's current action. This allows for timely response to user voice commands after the interruption operation is completed, improving user interaction efficiency and responsiveness. Furthermore, different interruption strategies can be adopted for different system states, executing different schemes to interrupt the system's current action. This solution improves the flexibility of interruption operations, enhances the user experience, and improves real-time response while ensuring the continuity of critical tasks. Attached Figure Description
[0009] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0010] Figure 1 This is a schematic diagram of a scenario where the interaction of a voice interaction system is interrupted, provided by an exemplary embodiment of this disclosure. Figure 2 This is a flowchart illustrating an exemplary embodiment of a voice interaction system for interrupting interaction. Figure 3 This is a flowchart illustrating step S200 of an interaction interruption method for a voice interaction system provided in an exemplary embodiment of this disclosure. Figure 4 This is a flowchart illustrating step S210 of an interaction interruption method for a voice interaction system provided in an exemplary embodiment of this disclosure. Figure 5 This is a flowchart illustrating step S210 of an interaction interruption method for a voice interaction system provided in another exemplary embodiment of this disclosure. Figure 6 This is a flowchart illustrating a method for interrupting the interaction of a voice interaction system provided in another exemplary embodiment of this disclosure; Figure 7 This is a flowchart illustrating step S200 of an interaction interruption method for a voice interaction system provided in another exemplary embodiment of this disclosure. Figure 8 This is a flowchart illustrating a method for interrupting the interaction of a voice interaction system, provided in yet another exemplary embodiment of this disclosure. Figure 9This is a flowchart illustrating an interaction interruption method for a voice interaction system provided in the fourth exemplary embodiment of this disclosure; Figure 10 This is a flowchart illustrating step S200 of an interaction interruption method for a voice interaction system provided in yet another exemplary embodiment of this disclosure. Figure 11 This is a flowchart illustrating a method for interrupting the interaction of a voice interaction system provided in the fifth exemplary embodiment of this disclosure; Figure 12 This is a flowchart illustrating step S600 of an interaction interruption method for a voice interaction system provided in an exemplary embodiment of this disclosure. Figure 13 This is a flowchart illustrating step S600 of an interaction interruption method for a voice interaction system provided in another exemplary embodiment of this disclosure. Figure 14 This is a schematic diagram of the structure of an interactive interruption device for a voice interaction system provided in an exemplary embodiment of the present disclosure; Figure 15 This is a schematic diagram of the structure of an interaction interruption device for a voice interaction system provided in another exemplary embodiment of this disclosure; Figure 16 This is a schematic diagram of the structure of an interaction interruption device for a voice interaction system provided in yet another exemplary embodiment of this disclosure; Figure 17 This is a structural diagram of an electronic device provided as an exemplary embodiment of the present disclosure. Detailed Implementation
[0011] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0012] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0013] Application Overview In the context of smart cockpits, the in-vehicle voice assistant needs to interrupt the current interaction when the user inputs a new voice command in order to respond to the new command. In other words, the in-vehicle voice assistant needs to support the user in interrupting the current interaction at any time via voice.
[0014] In traditional single-model voice assistants, a simple interruption mechanism can be used to respond to new user input. This means that after detecting new user voice input, the current interaction is immediately stopped and the new request is started.
[0015] However, in a multi-model collaborative architecture, multiple models may simultaneously handle different tasks. For example, a speech model might be used to handle casual conversation, a task agent to handle vehicle control commands, and an intent arbitrator to make routing decisions. In this case, interruption operations must consider different interruption strategies implemented by the system in different states. Multiple concurrently running models must coordinate to respond to interruption signals. If a model is performing an important task, the interruption should be ignored to protect the integrity of critical operations. Furthermore, the relevant parts generated by the model during runtime must be saved to avoid gaps in the dialogue history.
[0016] To address the aforementioned technical problems, this disclosure provides a method, apparatus, medium, and electronic device for interrupting interactions in a voice interaction system. This method detects the current interaction state of the voice interaction system, determines the corresponding interruption strategy, and employs different interruption strategies for different system states, executing different schemes to interrupt the system's current action. This solution improves the flexibility of interruption operations, enhances the user experience, and ensures both the continuity of critical tasks and real-time response.
[0017] Exemplary scenario Figure 1 This is a schematic diagram of a scenario where the interaction of a voice interaction system is interrupted, provided by an exemplary embodiment of this disclosure.
[0018] like Figure 1 As shown, after a user issues a voice command, the voice interaction system (such as an in-vehicle voice assistant) must respond to the voice command. In other words, after receiving a voice command, the voice interaction system performs the corresponding operation to meet the user's needs.
[0019] against Figure 1 In the scenario shown, the interaction interruption system of the voice interaction system provided in this embodiment can analyze the current interaction state of the voice interaction system to determine the interruption strategy corresponding to the current interaction state, and then execute the interruption strategy to interrupt the current action of the system. After the current action is interrupted, the recognition module of the voice interaction system converts the speech into text, and then the natural language understanding module parses the intent to execute the voice command.
[0020] The interaction interruption system of the voice interaction system in this embodiment of the disclosure executes executable instructions in the memory to analyze the current interaction state of the voice interaction system, determine the interruption strategy corresponding to the current interaction state, and then interrupt the current action of the system by executing the interruption strategy.
[0021] Exemplary methods Figure 2This is a flowchart illustrating an interactive interruption method for a voice interaction system provided in an exemplary embodiment of this disclosure. This embodiment can be applied to electronic devices, such as... Figure 2 As shown, the method includes the following steps: S100, in response to detecting a voice input event, determines the current interaction state of the voice interaction system.
[0022] In one implementation, the electronic device can monitor user voice. When user voice is detected, the electronic device can identify the content of the user's voice through a voice interaction system, and determine whether the content of the user's voice contains a wake word and / or an interaction intent. If a wake word and / or interaction intent are identified from the content of the user's voice, it is determined that a user voice input event has been detected.
[0023] In one implementation, the electronic device may include a button for triggering a voice input event, wherein a user voice input event is determined to have been detected when the user clicks the button.
[0024] For example, when a user says, "Help, pause music playback," the voice interaction system recognizes the wake word "helper," identifies the user's intention as "pause music playback," and can respond by pausing music playback. The voice interaction system can also analyze the current interaction state in real time to determine the corresponding target interruption strategy, thereby interrupting the current operation of the voice interaction system and executing the new action required by the user's command.
[0025] For example, the current interaction state of a voice interaction system can be determined by querying the state machine of the voice interaction system.
[0026] By adopting this approach, after a user issues a new voice event, the current system state of the voice interaction system can be determined. This allows for a decision on whether to interrupt the interaction and the specific interruption strategy, thereby improving the user experience.
[0027] S200 determines the target interruption strategy of the voice interaction system based on the current interaction state.
[0028] The voice interaction system includes multiple interaction states, each corresponding to a different interruption strategy. The interruption strategy is used to indicate how to interrupt the current action of the voice interaction system. The current action includes the current voice output action and / or the current task execution action.
[0029] In one implementation, the voice output action includes the voice interaction system playing a voice message. For example, the voice interaction system engages in casual conversation with the user. Alternatively, the voice interaction system provides feedback on the result of task execution via voice announcement, such as the voice interaction system announcing "Air conditioner is on" when the user instructs it to turn on. In another implementation, the task execution action includes the voice interaction system performing a task. For example, the voice interaction system executes a tool call corresponding to a voice command, or the voice interaction system arbitrates the command to determine its intent.
[0030] In one implementation, the multiple interaction states include at least a task execution state, an idle dialogue state, and / or an arbitration state. For example, the task execution state indicates that the voice interaction system is performing a specific task; the idle dialogue state indicates that the voice interaction system is in standby mode, performing only basic speech recognition and response; and the arbitration state indicates that the voice interaction system is waiting for the arbitration result of the current action output.
[0031] For example, a targeted interruption policy could instruct the current action of the voice interaction system to be interrupted to process the input voice input event. Alternatively, a targeted interruption policy could instruct the current action of the voice interaction system not to be interrupted to ignore the input voice input event.
[0032] By adopting the implementation method of this step, the voice interaction system can use different interruption strategies for different interaction states, and then use different interruption methods to interrupt the current action of the voice interaction system, so that the interruption action matches the current interaction state of the voice interaction system.
[0033] S300 interrupts the current action of the voice interaction system based on a target interruption strategy.
[0034] In one implementation, if the target interruption policy indicates that the current action should be interrupted, the current action can be interrupted, and the voice input event can be processed after the interruption is completed. If the target interruption policy indicates that the current action should not be interrupted, the system can either wait for the current action to finish executing before processing the user-input voice event, or directly discard the voice event, thereby improving the operational stability of the voice interaction system.
[0035] The technical solution of this disclosure first responds to a voice input event, then determines the current system state to match a corresponding interruption strategy, and then executes an interruption operation according to the interruption strategy to interrupt the current system action. This allows for timely response to user voice commands after the interruption operation is completed, improving user interaction efficiency and response smoothness. Furthermore, different interruption strategies can be adopted for different system states, executing different schemes to interrupt the current system action. This solution improves the flexibility of interruption operations, enhances the user experience, and ensures both the continuity of critical tasks and real-time response.
[0036] Figure 3 This is a flowchart illustrating step S200 of an interaction interruption method for a voice interaction system provided in an exemplary embodiment of this disclosure. Figure 3 As shown above, in the above Figure 2 Based on the illustrated embodiment, step S200 may include the following steps: S210, if the current interaction state is the task execution state, determine the current sub-state corresponding to the task execution state.
[0037] In one implementation, the task execution state includes multiple sub-states, each corresponding to a different stage of task execution and a different interruption strategy. This allows for different interruption strategies to be applied to different sub-states, thus interrupting the current action of the voice interaction system.
[0038] For example, if the current interaction state is a task execution state, it indicates that the voice interaction system is currently performing a specific task. In this case, it is necessary to further identify its sub-states so as to match the most appropriate interruption strategy based on the sub-states.
[0039] For example, multiple sub-states may include, but are not limited to, a tool execution sub-state and a voice broadcast sub-state. In the tool execution sub-state, interrupting the current action will affect tool invocation and has a significant impact on the system. In the voice broadcast sub-state, interrupting the current action has a smaller impact on the system. Therefore, different target interruption strategies can be adopted for different sub-states to reduce the impact of interruption operations on the system and improve system stability.
[0040] By adopting this approach, we can further refine and confirm the specific sub-states in which the task is executed, and determine the corresponding target interruption strategy based on the different sub-states.
[0041] S220, Based on the current sub-state, determine the target interruption strategy.
[0042] In one implementation, a target interruption strategy is determined based on different current sub-states to enable further refined interruption of the system's current operations. The technical solution of this disclosure, when the current interaction state is a task execution state, further determines the current sub-state to select a matching target interruption strategy for different current sub-states.
[0043] Figure 4 This is a flowchart illustrating step S210 of an interaction interruption method for a voice interaction system provided in an exemplary embodiment of this disclosure. Figure 4 As shown above, in the above Figure 3 Based on the illustrated embodiment, step S210 may include the following steps: S211, if the current sub-state is the tool execution sub-state, determine the third interruption strategy corresponding to the tool execution sub-state.
[0044] In one implementation, the tool execution sub-state includes invoking the target tool to perform the target operation. In the tool execution sub-state, interrupting the current action will affect the tool invocation. Therefore, when the voice interaction system is in the tool execution sub-state, when the user inputs a new command, the current action should not be interrupted directly; instead, the current action should be allowed to complete to improve the atomicity and completeness of the tool invocation.
[0045] For example, the tool execution sub-state includes, but is not limited to, executing device control commands such as turning on the air conditioner or adjusting the volume. At this time, the voice interaction system is communicating in real time with external hardware or service interfaces to invoke the corresponding tool.
[0046] In one implementation, the third interruption strategy includes: temporarily storing the voice input event; continuing to execute the tool execution sub-state; and clearing the voice input event in response to the completion of the tool execution sub-state.
[0047] By using this method, it is possible to determine whether the current system state is in the tool execution sub-state. In this sub-state, the corresponding third interruption strategy can be activated to continue the work and ignore the user's voice input.
[0048] S212, the third interrupt strategy is determined as the target interrupt strategy.
[0049] By adopting the implementation method in this step, the third interruption strategy can be used to temporarily store and delay the processing of user-input voice events, thereby avoiding interruption of critical tool links and affecting the stability of system operation.
[0050] For example, based on the third interruption strategy, the current action being performed by the voice interaction system can be protected from being interrupted, reducing the risk of device status abnormalities or service response failures caused by interrupting the current action, and improving system robustness.
[0051] Thus, the third interruption strategy described above is executed, prioritizing the current tool operation and clearing the voice input event after it is completed, thereby ignoring the user's newly input voice event.
[0052] The technical solution of this disclosure first determines whether the current sub-state is a tool execution sub-state. If the current sub-state is a tool execution sub-state, then the third interruption strategy is enabled to ensure the integrity and reliability of the tool chain.
[0053] For example, if the current sub-state is not the tool execution sub-state, it can switch to the interruption strategy branch corresponding to other sub-states, so that the system can take into account both response speed and task stability in various interaction scenarios.
[0054] Figure 5 This is a flowchart illustrating step S210 of a voice interaction system interruption method provided in another exemplary embodiment of this disclosure. Figure 5 As shown above, in the above Figure 3 Based on the illustrated embodiment, step S210 may include the following steps: S213, when the current sub-state is the voice broadcast sub-state, determine the fourth interruption strategy corresponding to the voice broadcast sub-state.
[0055] In one implementation, the voice broadcast sub-state includes the voice interaction system broadcasting feedback on the execution result of the current action. For example, when the current sub-state is the voice broadcast sub-state, it indicates that the processing result of the task agent is being broadcast using text-to-speech (TTS) technology.
[0056] For example, the voice broadcast sub-state corresponds to a non-critical process of the current action, and the current voice broadcast can be interrupted in response to a voice input event.
[0057] In one implementation, the fourth interruption strategy includes: stopping voice playback; discarding unplayed voice text; and saving the played voice text to the dialogue history, which is used by the model to generate audio. For example, the fourth interruption strategy can interrupt the current action in response to user-input voice events. For instance, it can be used to truncate the TTS output stream in real time, clear the buffer to be played, and store the played text content in the dialogue history, avoiding gaps in the dialogue history. This allows for rapid response to new user input and ensures a coherent and complete dialogue context, providing an accurate semantic foundation for subsequent intent understanding and task integration.
[0058] For example, after saving the spoken text to the dialogue history, it may also include: synchronizing the spoken content to the dialogue context of the speech model (dialogue context synchronization interface) to make the context of multiple models consistent.
[0059] S214, the fourth interrupt strategy is determined as the target interrupt strategy.
[0060] The technical solution of this disclosure first determines whether the current sub-state is a voice broadcast sub-state. If the current sub-state is a voice broadcast sub-state, the fourth interruption strategy is enabled to interrupt the current operation corresponding to the voice broadcast sub-state and save the TTS output to reduce the problem of gaps in the dialogue history.
[0061] For example, if the current sub-state is not the voice broadcast sub-state, then the system switches to the interruption strategy branch corresponding to other sub-states, so that the system can balance response speed and task stability in various interaction scenarios.
[0062] Figure 6 This is a flowchart illustrating a method for interrupting the interaction of a voice interaction system, provided in another exemplary embodiment of this disclosure. For example... Figure 6 As shown, in Figure 5 Based on the illustrated embodiment, after step S220, the method further includes: S221, Extract memory information from the already broadcast voice text.
[0063] For example, through step S221, after interrupting the TTS output based on the fourth interruption strategy, memory retrieval processing can be performed to extract the memory information of the output text.
[0064] S222, clear the downlink transmission queue.
[0065] In one implementation, the downlink transmission queue is used to send information to the client. If there are other messages to be sent in the downlink transmission queue before the interrupt signal is sent, the interrupt signal will be placed at the end of the downlink transmission queue after it is generated and enters the queue. This causes the client's TTS to still execute the interrupt operation after playing a lot of content. To solve this delay problem, this embodiment of the disclosure simultaneously clears all messages to be sent in the downlink transmission queue when generating the interrupt signal, and then sends the interrupt signal to the downlink transmission queue. Since the downlink transmission queue is empty at this time, the client will receive the interrupt signal first and execute it immediately.
[0066] S223 sends an interruption signal to the client through the downlink transmission queue.
[0067] The technical solution of this disclosure embodiment, after executing the fourth interruption strategy, further performs semantic parsing and key information extraction on the already broadcast text to facilitate accurate understanding of user intent and contextual continuity. Furthermore, it sequentially clears the downlink transmission queue and sends an interruption signal to improve the response speed of the interruption signal and avoid delayed responses caused by pending messages after the interruption signal has been sent.
[0068] It is understandable that the above steps S221-S223 can also be performed after step S300.
[0069] Figure 7 This is a flowchart illustrating step S200 of a voice interaction system interruption method provided in another exemplary embodiment of this disclosure. For example... Figure 7 As shown above, in the above Figure 2 Based on the illustrated embodiment, step S200 may include the following steps: S230, if the current interaction state is an idle dialogue state, determine the first interruption strategy corresponding to the idle dialogue state.
[0070] For example, idle conversation states include: the voice interaction system is engaging in casual conversation with the user, or the voice interaction system is answering open-ended questions but not performing a specific task.
[0071] For example, the first interruption strategy includes: stopping audio playback; discarding unplayed audio text; and saving the played audio text to the dialogue history, which is used by the model to generate audio.
[0072] S240, the first interrupt strategy is determined as the target interrupt strategy.
[0073] This approach allows for immediate interruption of audio playback and context clearing based on the first interruption strategy. Furthermore, it saves the played audio to the dialogue context, preventing gaps in the dialogue history. This enables rapid response to new voice input from the user, improving the user experience.
[0074] The technical solution of this disclosure, when determining that the current system state is an idle dialogue state, can directly interrupt the current operation of the system based on a first interruption strategy to respond promptly to newly initiated voice events by the user. In this way, it can quickly respond to changes in user intent in the idle dialogue state, immediately terminate the current dialogue process using the first interruption strategy, and re-parse the new voice events input by the user, responding to new voice events with low latency and improving user experience.
[0075] For example, if a user asks the system "How's the weather today?" without triggering any specific task, and only maintains an open-ended question-and-answer interaction, the system is in an idle dialogue state. If the user inputs a new voice event, the system can promptly interrupt the previous response process and switch to understanding and responding to the new intent, without waiting for the current response to end.
[0076] Figure 8 This is a flowchart illustrating a method for interrupting interaction in a voice interaction system, provided in yet another exemplary embodiment of this disclosure. (Refer to...) Figure 8 As shown, when the current interaction state is an idle dialogue state, after step S240, the following steps are also included: S241, Clear the downlink transmission queue. The downlink transmission queue is used to send information to the client.
[0077] S242 sends an interruption signal to the client through the downlink transmission queue.
[0078] The technical solution of this disclosure embodiment can sequentially clear the downlink transmission queue and send an interruption signal to improve the response speed of the interruption signal, reduce the delayed response caused by the interruption signal being sent but there are still pending messages, and improve the user experience.
[0079] Understandably, the specific execution methods of steps S241 and S242 can be referred to steps S222 and S223, and will not be repeated here.
[0080] Figure 9 This is a flowchart illustrating a method for interrupting interaction in a voice interaction system according to a fourth exemplary embodiment of this disclosure. (Refer to...) Figure 9 As shown, after sending an interrupt signal to the client through the downlink transmission queue (step S223 or step S242), the process further includes: S400 resets the markers for the voice interaction system.
[0081] In one implementation, resetting the markers of the voice interaction system may include resetting runtime state markers, such as streaming response state markers, dialogue text generation state markers, and task-based text generation state markers. The streaming response state marker indicates whether the system is currently in a streaming response state; the dialogue text generation state marker indicates whether dialogue text generation has ended; and the task-based text generation state marker indicates whether task-based text generation has ended.
[0082] In this way, by resetting the above flags, the system can be fully restored to an idle state after being interrupted, which means it can directly process the state of new events and reduce the problem of subsequent response logic being abnormal due to residual flag states.
[0083] For example, step S400 can be achieved by setting the streaming response status flag to false and both the dialogue text generation status flag and the task-based text generation status flag to true.
[0084] The technical solution of this disclosure can immediately trigger a state reset process after the interruption signal is sent, so that the states of each module in the system are synchronously reset to zero, providing a clean and definite initial environment for the next round of interaction. In this way, in multi-model concurrent scenarios, the risk of response disorder or repeated triggering caused by state residue can be reduced, and the stability and response consistency of multi-model systems can be improved.
[0085] It is worth noting that when interrupting the current action, a consistent interruption boundary must be established between the server and the client. Since the data flow of a voice interaction system is bidirectional and asynchronous—that is, the server sends TTS audio and events to the client through a downlink transmission queue, and the client uploads the voice stream and tool call results to the server—if strict and orderly post-processing is not performed after an interruption, it will cause the old data to pollute the state of the new data stream.
[0086] For example, in this embodiment of the disclosure, step S400 is used to reset the voice interaction system during the current operation to solve the above-mentioned pollution problem.
[0087] Understandably, the server can clear the residual data in the downlink transmission queue through the above steps S222-S223 or S241-S242, and then instruct the client to perform the corresponding interruption operation through the downlink transmission queue.
[0088] For example, clearing the downlink transmission queue on the server side includes: discarding accumulated TTS audio data blocks in the downlink transmission queue; and / or discarding streaming text events and intermediate state notifications awaiting transmission in the downlink transmission queue; canceling the server-sent events (SSE) streaming connection of the task agent; and stopping the writing of new data to the transmission queue.
[0089] For example, upon receiving an interrupt signal, the client stops playing the current TTS audio and clears the local audio buffer. It also stops uploading the execution results of any unfinished tool calls in the current round (such as an air conditioning control callback being executed by the client). Furthermore, it discards all received but not yet played residual audio data packets.
[0090] In one implementation, before step S400, the server can switch the audio cache to discard mode to intercept all possible residual downlink audio events. Furthermore, it emptys the event queue of pending events belonging to older rounds. This further clears the relevant content from older rounds, reduces state pollution for new rounds, and improves the purity of the next round of model interaction.
[0091] Figure 10 This is a flowchart illustrating step S200 of a voice interaction system interruption method provided in another exemplary embodiment of this disclosure. For example... Figure 10 As shown above, in the above Figure 2 Based on the illustrated embodiment, step S200 may include the following steps: S250, if the current interaction state is arbitration state, determine the second interruption strategy corresponding to the arbitration state.
[0092] For example, the arbitration status includes waiting for the arbitration result of the current action.
[0093] For example, the arbitration state includes the state in which the voice interaction system is waiting for the intent arbitrator to make a judgment on the routing results in order to determine the direction of subsequent actions.
[0094] For example, the second interruption strategy includes: determining the intent type corresponding to the arbitration result of the arbitration state; determining, based on the intent type, whether to continue executing the current action corresponding to the arbitration result or to execute the voice input event; and interrupting the current action if the duration of the arbitration state is greater than or equal to a time threshold.
[0095] S260, the second interrupt strategy is determined as the target interrupt strategy; The technical solution of this disclosure first waits for the arbitration result of the current action when the current interaction state is in arbitration state. After the arbitration result is output, based on the second interruption strategy, it is dynamically determined whether the current action can be interrupted according to the arbitration result. This realizes the determination of whether to interrupt immediately or delay processing based on the intent type of the old request, reducing the risk of system chaos caused by interrupting before the previous intent type is determined, and improving the accuracy of intent understanding and system robustness.
[0096] Furthermore, if the duration of the arbitration status is greater than or equal to the time threshold, it indicates that the current arbitration process may be erroneous, which can trigger the exception handling mechanism to forcibly terminate the current arbitration process and directly handle newly input voice events.
[0097] Figure 11 This is a flowchart illustrating a method for interrupting interaction in a voice interaction system according to a fifth exemplary embodiment of this disclosure. (Refer to...) Figure 11 As shown, the methods for interrupting interaction in a voice interaction system also include: S500, in response to detecting a task insertion event, determines the task type corresponding to the currently executing task of the voice interaction system when the current interaction state is the tool execution sub-state.
[0098] In one implementation, the task insertion event is used to insert a task into the task queue.
[0099] For example, the priority of the task inserted by the task insertion event is higher than the priority of the currently executing task.
[0100] S600 determines the target interruption strategy based on the task type corresponding to the currently executing task.
[0101] By adopting this approach, an appropriate interruption strategy can be dynamically matched based on the task type corresponding to the currently executing task, reducing the interference of interruptions on the user experience and improving the response speed of high-priority tasks.
[0102] The technical solution of this disclosure supports not only voice interruption but also event interruption (such as system messages). Furthermore, considering the high timeliness of newly inserted tasks, if an event message is received during continuous task execution, the event message will interrupt the entire task flow but will not interrupt the currently executing task. In other words, after the current task is completed, the system immediately takes over the event message and starts a new task flow, improving the timeliness of high-priority event response.
[0103] Figure 12 This is a flowchart illustrating step S600 of an interaction interruption method in a voice interaction system provided in an exemplary embodiment of this disclosure. (Refer to...) Figure 12 As shown, step S600 includes: S610, if the task type corresponding to the currently executing task is a single task type, determine the fifth interruption strategy corresponding to the tool execution sub-state.
[0104] In one implementation, a single-task type refers to a task type that does not require feedback on the task execution result. In other words, a single-task type is a task with only a single-step execution period. When executing a single-task type, a newly inserted task can be executed immediately after the currently executing task has finished.
[0105] In one implementation, the fifth interruption strategy includes: executing an event after the current task execution state has been completed.
[0106] S620, the fifth interrupt strategy is determined as the target interrupt strategy.
[0107] By adopting the implementation method of this step, the fifth interruption strategy can be used to allow the system to execute a new task directly after a single task is completed, avoiding logical conflicts or state abnormalities caused by interrupting the current operation.
[0108] The technical solution of this disclosure employs a fifth interruption strategy when the currently executing task is a single-task type. This strategy allows the system to continue executing the currently executing task until it is completed, and then directly executes the newly inserted task after completion. In this way, the fifth interruption strategy allows the system to directly execute a new task after the single task is completed, protecting the atomicity of the currently executing single-step tool call, reducing the occurrence of half-complete operations, and improving system stability.
[0109] Figure 13 This is a flowchart illustrating step S600 of a voice interaction system interruption method provided in another exemplary embodiment of this disclosure. (Refer to...) Figure 13 As shown, step S600 includes: S630, when the task type corresponding to the currently executing task is a multi-task type, determine the sixth interruption strategy corresponding to the tool execution sub-state.
[0110] In one implementation, a multi-task type refers to a task type that requires feedback on the task execution results. In other words, a multi-task type is a task type that has at least two execution steps and depends on the feedback of intermediate results.
[0111] In one implementation, the sixth interruption strategy includes: interrupting the feedback action corresponding to the currently executing task after the currently executing task has been completed. The feedback action is used to provide feedback on the execution result of the currently executing task.
[0112] For example, during the execution of multiple tasks, the system needs to wait for intermediate results to be returned before proceeding to subsequent steps. Therefore, the sixth interruption strategy is adopted: after the current step is completed, the subsequent processing steps are interrupted, and a new task flow is started directly. S640, the sixth interrupt strategy is determined as the target interrupt strategy.
[0113] The technical solution of this disclosure employs a sixth interruption strategy when the currently executing task is a multi-task type. This strategy allows the current step to continue execution until it is completed, after which the newly inserted task is executed directly. This protects the integrity of single-step operations within the currently executing multi-task, preventing premature interruption that could lead to loss of intermediate states or data inconsistency, thus improving system stability.
[0114] Figure 14 This is a schematic diagram of the structure of an interactive interruption device for a voice interaction system provided in an exemplary embodiment of this disclosure.
[0115] like Figure 14 As shown, in one embodiment, the interaction interruption device 1400 of the voice interaction system includes: a determination module 1410 and an interruption execution module 1420.
[0116] The determination module 1410 is configured to: determine the current interaction state of the voice interaction system in response to the detection of a voice input event; and determine the target interruption strategy of the voice interaction system based on the current interaction state. The voice interaction system includes multiple interaction states, and the multiple interaction states correspond to multiple interruption strategies. The interruption strategy is used to indicate the interruption method of the current action of the voice interaction system. The current action includes the current voice output action and / or the current task execution action.
[0117] The interruption execution module 1420 is used to interrupt the current action of the voice interaction system based on the target interruption strategy.
[0118] In one possible implementation, the determining module 1410 is used to: determine the current sub-state corresponding to the task execution state when the current interaction state is the task execution state, wherein the task execution state includes multiple sub-states, the multiple sub-states correspond to different stages of task execution, and the multiple sub-states correspond to different interruption strategies; and determine the target interruption strategy based on the current sub-state.
[0119] In one possible implementation, the determining module 1410 is configured to: determine a first interruption strategy corresponding to the idle dialogue state when the current interaction state is an idle dialogue state; and determine the first interruption strategy as the target interruption strategy; wherein the first interruption strategy includes: stopping audio playback; discarding unplayed audio text, and saving the played audio text to the dialogue history, the dialogue history being used by the model to generate audio.
[0120] In one possible implementation, the determining module 1410 is configured to: determine a second interruption strategy corresponding to the arbitration state when the current interaction state is an arbitration state; the arbitration state includes waiting for the arbitration result of the current action; determine the second interruption strategy as the target interruption strategy; wherein the second interruption strategy includes: determining the intent type corresponding to the arbitration result of the arbitration state; determining, based on the intent type, whether to continue executing the current action corresponding to the arbitration result or to execute a voice input event; and interrupting the current action if the duration of the arbitration state is greater than or equal to a time threshold.
[0121] In one possible implementation, the determining module 1410 is configured to: determine a target interruption strategy based on the current sub-state, including: when the current sub-state is a tool execution sub-state, determining a third interruption strategy corresponding to the tool execution sub-state; the tool execution sub-state includes the current action calling the target tool to perform the target operation; determining the third interruption strategy as the target interruption strategy; wherein the third interruption strategy includes: temporarily storing the voice input event; continuing to execute the tool execution sub-state; and clearing the voice input event in response to the completion of the tool execution sub-state.
[0122] In one possible implementation, the determining module 1410 is configured to: determine a fourth interruption strategy corresponding to the voice broadcast sub-state when the current sub-state is a voice broadcast sub-state; the voice broadcast sub-state includes the voice interaction system broadcasting feedback on the execution result of the current action; determine the fourth interruption strategy as the target interruption strategy; wherein the fourth interruption strategy includes: stopping voice broadcast; discarding unplayed voice text; saving the broadcast voice text to the dialogue history, the dialogue history being used by the model to generate audio.
[0123] Figure 15 This is a schematic diagram of the structure of an interaction interruption device for a voice interaction system provided in another exemplary embodiment of this disclosure.
[0124] like Figure 15 As shown, in one embodiment, the interaction interruption device of the voice interaction system further includes: a sending module 1430, used to: clear the downlink sending queue, which is used to send information to the client; and send an interruption signal to the client through the downlink sending queue.
[0125] In one possible implementation, the sending module 1430 is used to: extract memory information from the broadcast voice text; clear the downlink sending queue, which is used to send information to the client; and send an interrupt signal to the client through the downlink sending queue.
[0126] In one possible implementation, the sending module 1430 is used to: reset the markers of the voice interaction system.
[0127] Figure 16 This is a schematic diagram of the structure of an interaction interruption device for a voice interaction system provided in another exemplary embodiment of this disclosure.
[0128] like Figure 16 As shown, in one embodiment, the interaction interruption device of the voice interaction system further includes: a task detection module 1440, configured to: in response to detecting a task insertion event, determine the task type corresponding to the currently executing task of the voice interaction system when the current interaction state is the tool execution sub-state, wherein the task insertion event is used to insert a task into the task queue; and a determination module 1410, configured to: determine a target interruption strategy based on the task type corresponding to the currently executing task.
[0129] In one possible implementation, the determining module 1410 is used to: determine the fifth interruption strategy corresponding to the tool execution sub-state when the task type corresponding to the currently executing task is a single task type, wherein the single task type refers to a task type that does not require feedback of task execution results; determine the fifth interruption strategy as the target interruption strategy; wherein the fifth interruption strategy includes: executing an event after the current execution task execution state is completed.
[0130] In one possible implementation, the determining module 1410 is configured to: determine the sixth interruption strategy corresponding to the tool execution sub-state when the task type corresponding to the currently executed task is a multi-task type, wherein the multi-task type refers to the task type that requires feedback on the task execution result; and determine the sixth interruption strategy as the target interruption strategy; wherein the sixth interruption strategy includes: interrupting the feedback action corresponding to the currently executed task after the currently executed task is completed, the feedback action being used to provide feedback on the execution result of the currently executed task.
[0131] Exemplary electronic devices Figure 17 A structural diagram of an electronic device provided for an exemplary embodiment of the present disclosure includes at least one processor 1710 and a memory 1720.
[0132] The processor 1710 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 1700 to perform desired functions.
[0133] The memory 1720 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 1710 may execute one or more computer program instructions to implement the interactive interruption method and / or other desired functions of the voice interaction system of the various embodiments of this disclosure described above.
[0134] In one example, the electronic device 1700 may also include an input device 1730 and an output device 1740, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0135] The input device 1730 may also include, for example, a keyboard, a mouse, etc.
[0136] The output device 1740 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0137] Of course, for the sake of simplicity, Figure 17Only some of the components of the electronic device 1700 relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device 1700 may include any other suitable components depending on the specific application.
[0138] Exemplary computer program products and computer-readable storage media In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the interactive interruption methods of a voice interaction system according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0139] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0140] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the interactive interruption method of a voice interaction system according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0141] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0142] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0143] The block diagrams of devices, apparatuses, devices, and systems disclosed herein are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0144] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0145] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0146] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.
Claims
1. A method for interrupting interaction in a voice interaction system, comprising: In response to the detection of a voice input event, the current interaction state of the voice interaction system is determined; Based on the current interaction state, a target interruption strategy for the voice interaction system is determined. The voice interaction system includes multiple interaction states, each corresponding to a multiple interruption strategy. The interruption strategy is used to indicate the interruption method for the current action of the voice interaction system. The current action includes the current voice output action and / or the current task execution action. Based on the target interruption strategy, the current action of the voice interaction system is interrupted.
2. The method according to claim 1, wherein, The step of determining the target interruption strategy of the voice interaction system based on the current interaction state includes: When the current interaction state is a task execution state, determine the current sub-state corresponding to the task execution state, wherein the task execution state includes multiple sub-states, the multiple sub-states correspond to different stages of task execution, and the multiple sub-states correspond to different interruption strategies; Based on the current sub-state, the target interruption strategy is determined.
3. The method according to claim 1, wherein, The step of determining the target interruption strategy of the voice interaction system based on the current interaction state includes: If the current interaction state is an idle dialogue state, determine the first interruption strategy corresponding to the idle dialogue state; The first interruption strategy is determined as the target interruption strategy; The first interruption strategy includes: stopping audio playback; discarding unplayed audio text and saving played audio text to the dialogue history, which is used by the model to generate the audio.
4. The method according to claim 1, wherein, The step of determining the target interruption strategy of the voice interaction system based on the current interaction state includes: If the current interaction state is an arbitration state, a second interruption strategy corresponding to the arbitration state is determined; the arbitration state includes waiting for the arbitration result of the current action. The second interruption strategy is determined as the target interruption strategy; The second interruption strategy includes: determining the intent type corresponding to the arbitration result of the arbitration state; determining, based on the intent type, whether to continue executing the current action corresponding to the arbitration result or to execute the voice input event; and interrupting the current action if the duration of the arbitration state is greater than or equal to a time threshold.
5. The method according to claim 2, wherein, The step of determining the target interruption strategy based on the current sub-state includes: If the current sub-state is a tool execution sub-state, a third interruption strategy corresponding to the tool execution sub-state is determined; the tool execution sub-state includes the current action calling the target tool to perform the target operation; The third interruption strategy is determined as the target interruption strategy; The third interruption strategy includes: temporarily storing the voice input event; continuing to execute the tool execution sub-state; and clearing the voice input event in response to the completion of the tool execution sub-state.
6. The method according to claim 2, wherein, The step of determining the target interruption strategy based on the current sub-state includes: If the current sub-state is a voice broadcast sub-state, determine the fourth interruption strategy corresponding to the voice broadcast sub-state; the voice broadcast sub-state includes the voice interaction system broadcasting feedback on the execution result of the current action; The fourth interruption strategy is determined as the target interruption strategy; The fourth interruption strategy includes: stopping voice broadcasting; discarding unplayed voice text; and saving the broadcast voice text to the dialogue history, which is used by the model to generate the audio.
7. The method according to claim 3, wherein, When the current interaction state is an idle dialogue state, after interrupting the current action of the voice interaction system based on the target interruption strategy, the method further includes: Clear the downlink transmission queue, which is used to send information to the client; An interrupt signal is sent to the client through the downlink transmission queue.
8. The method according to claim 6, wherein, When the current sub-state is the voice broadcast sub-state, after interrupting the current action of the voice interaction system based on the target interruption strategy, the method further includes: Extract memory information from the already broadcast voice text; Clear the downlink transmission queue, which is used to send information to the client; An interrupt signal is sent to the client through the downlink transmission queue.
9. The method according to claim 7 or 8, wherein, After sending the interrupt signal to the client through the downlink transmission queue, the method further includes: The labels of the voice interaction system are reset.
10. The method according to claim 5, wherein, Also includes: In response to the detection of a task insertion event, if the current interaction state is the tool execution sub-state, the task type corresponding to the current execution task of the voice interaction system is determined, wherein the task insertion event is used to insert a task into the task queue; The target interruption strategy is determined based on the task type corresponding to the currently executing task.
11. The method according to claim 10, wherein, Determining the target interruption strategy based on the task type corresponding to the currently executing task includes: If the task type corresponding to the currently executing task is a single task type, determine the fifth interruption strategy corresponding to the tool execution sub-state, wherein the single task type refers to a task type that does not require feedback of task execution results; The fifth interruption strategy is determined as the target interruption strategy; The fifth interruption strategy includes: executing the event after the current task execution state is completed.
12. The method according to claim 11, wherein, The step of determining the target interruption strategy based on the task type corresponding to the currently executing task includes: If the task type corresponding to the currently executing task is a multi-task type, determine the sixth interruption strategy corresponding to the tool execution sub-state, wherein the multi-task type refers to the task type that requires feedback on the task execution result; The sixth interruption strategy is determined as the target interruption strategy; The sixth interruption strategy includes: interrupting the feedback action corresponding to the current task after the current task is completed, wherein the feedback action is used to provide feedback on the execution result of the current task.
13. An interaction interruption device for a voice interaction system, comprising: A determination module is used to determine the current interaction state of the voice interaction system in response to the detection of a voice input event; Based on the current interaction state, a target interruption strategy for the voice interaction system is determined. The voice interaction system includes multiple interaction states, each corresponding to a multiple interruption strategy. The interruption strategy is used to indicate the interruption method for the current action of the voice interaction system. The current action includes the current voice output action and / or the current task execution action. The interruption module is used to interrupt the current action of the voice interaction system based on the target interruption strategy.
14. A computer-readable storage medium storing a computer program for executing the interaction interruption method of the voice interaction system according to any one of claims 1-12.
15. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the interactive interruption method of the voice interaction system according to any one of claims 1-12.