Intelligent voice interaction real-time interruption processing method and system

By leveraging the end-to-end real-time processing link and the computing power of the voice AI chip, the system accurately distinguishes between user voice and echo signals, solving the unnatural problem of real-time interruption in AI voice interaction devices and achieving natural interruption and efficient interaction without wake words.

CN121641018APending Publication Date: 2026-03-10SHANGHAI IND U TECH RES INST
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-28
Publication Date
2026-03-10

AI Technical Summary

Technical Problem

Existing AI voice interaction devices have problems in real-time interruption functions, such as the inability to accurately distinguish between device-played voice and user voice, which leads to the need to wait for playback to end or force a wake word, violating natural communication habits. In addition, traditional AEC algorithms have slow convergence speed and obvious residual echo in nonlinear distortion scenarios, making it difficult to achieve natural interruption without a wake word.

Method used

An end-to-end real-time processing link is formed by the AEC module, the residual echo nonlinear suppression module, and the AI-VAD module. By leveraging the computing power of the voice AI chip and combining the VSS-NLMS adaptive filtering algorithm and the Wiener filtering spectrum enhancement algorithm, the user's voice and echo signal are accurately distinguished, and natural interruption without wake word is achieved.

Benefits of technology

It achieves low-latency user voice recognition, reduces the rate of false interruptions, improves interaction efficiency, meets the needs of natural communication, and reduces the user's operational burden.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121641018A_ABST
    Figure CN121641018A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent voice interaction real-time interruption processing method and system, and the method comprises the steps: a preprocessing module synchronously collects microphone mixed signals and loudspeaker far-end reference signals, and stores the collected signals; inputting the preprocessed signal into an AEC module based on a VSS-NLMS adaptive filtering algorithm for real-time processing, and eliminating linear echoes; the residual echo nonlinear suppression module processes nonlinear residual echo in real time based on a spectrum enhancement algorithm of Wiener filtering; and inputting the processed signal into an AI-VAD module for detection, and if a user voice signal of which the duration is greater than or equal to 100ms exists, judging that an interruption instruction exists, triggering an interruption mechanism, and interrupting the current playing. The computing power advantage of the voice AI chip is fully utilized, an end-to-end real-time processing link is formed through the AEC module, the residual echo nonlinear suppression module, the AI-VAD module and the interrupt mechanism, low-delay cooperative operation of all the modules is ensured, user voice and echo signals and environment noise are accurately distinguished, natural interruption without wake-up words is achieved, and the user experience is improved. The problem that the interaction mode is not natural is solved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of speech processing, and particularly relates to a method and system for real-time interrupting of intelligent speech interaction. BACKGROUND

[0002] In an AI speech interaction system, the real-time interrupting function is crucial for achieving a natural and smooth conversation, and can enhance user control, optimize experience, and improve efficiency. However, there are obvious problems in current implementations, such as the user having to wait for the device to complete the current speech playback before interacting again, which makes it impossible to interrupt at any time. Alternatively, when the device is playing a long news, the user has difficulty in asking for details in real time. While some devices allow interrupting, they require a fixed wake-up word such as "Xiaoyue" to wake up before giving instructions, which goes against the natural habit of communication and increases the operational burden.

[0003] One of the core problems lies in the shortcomings of Acoustic Echo Cancellation (AEC) technology. When the loudspeaker plays speech, the microphone will pick up the output sound of the loudspeaker and environmental noise, making it difficult for the speech recognition system to accurately recognize the user's interrupting instructions. Traditional AEC algorithms often have slow convergence speed and obvious residual echo when dealing with nonlinear distortion scenarios, which makes it difficult for these devices to accurately distinguish between the device's own played speech and the user's real-time interrupting speech, and thus unable to achieve natural interrupting without a wake-up word, and also difficult to capture the user's non-wake-up word instructions in real time during device playback.

[0004] In existing technologies, most AEC solutions for intelligent speech devices rely on cloud processing. The hardware end only sends the collected audio signal packets and reference signal packets to the cloud server for echo cancellation. However, AEC algorithms are extremely sensitive to the delay jitter between the collected signals and the reference signals. Delay fluctuations during network transmission can cause AEC algorithms to frequently re-converge, which can easily cause missed sounds and false interrupts. The AEC cloud processing solution for intelligent speech devices cannot accurately distinguish between the device's own played speech and the user's interrupting speech due to insufficient processing capacity, and thus has to rely on waiting for the playback to end or using a forced wake-up word for interaction, making it difficult to achieve a natural and smooth real-time interrupting function. At the same time, traditional solutions generally lack effective Linear Predictive Coding (LPC) residual processing, and the accuracy of speech signal feature extraction is insufficient, further restricting the accuracy of interrupt recognition.

[0005] The existing cloud processing scheme has the disadvantages of high cost, poor real-time performance, and dependence on network, which directly leads to many pain points in the implementation of the real-time interrupt function of the mainstream devices on the market when realizing AI voice interaction, such as inaccurate voice recognition: due to the insufficient AEC technology, the device cannot accurately distinguish between its own played voice and the user's interrupt voice, making it difficult to realize the real-time interrupt function, and the user often needs to wait for the device to complete the current voice playing before interacting again, and when the device plays long news, the user cannot ask for details halfway through; or unnatural interaction: some devices allow interruption, but require shouting a fixed wake-up word, such as waking up before instructing to cut songs, which goes against the habit of natural communication and increases the user's operation burden.

[0006] Therefore, how to accurately distinguish between user voice and echo / noise and realize natural interruption without a wake-up word has become a problem to be solved. SUMMARY

[0007] The present application provides a kind of intelligent voice interaction real-time interrupt processing method and system, make full use of the computing power advantage of voice AI chip, through AEC module, residual echo nonlinear suppression module, AI-VAD module and interrupt mechanism form end-to-end real-time processing link, Ensure that each module runs in low delay, accurately distinguish between user voice and echo signal and environmental noise, realize natural interruption without a wake-up word, solve the problem of unnatural interaction.

[0008] Other purposes and advantages of the present application can be further understood from the technical features disclosed in the present application.

[0009] To achieve one or part or all of the above purposes or other purposes, the present application provides an intelligent voice interaction real-time interrupt processing method and system.

[0010] An intelligent voice interaction real-time interrupt processing method, comprising: Step S1: The preprocessing module synchronously collects the microphone mixed signal and the loudspeaker far-end reference signal, and stores the collected signals; Step S2: input the preprocessed signal into the AEC module based on the VSS-NLMS adaptive filtering algorithm for real-time processing to eliminate linear echo; Step S3: the residual echo nonlinear suppression module processes the nonlinear residual echo in real time based on the Wiener filter spectrum enhancement algorithm; Step S4: input the processed signal into the AI-VAD module for detection, if there is a user voice signal with a duration of ≥100ms, it is determined as an interrupt instruction, and the interrupt mechanism is triggered to interrupt the current playing.

[0011] The interrupt mechanism is set as: Pausing the current playing of the loudspeaker, recognizing the user voice signal, performing instruction analysis and execution, and selecting to continue playing or starting a new interactive process according to the user instruction.

[0012] The microphone mixed signal includes: The echo signal formed after the loudspeaker far-end reference signal is coupled through the acoustic echo path impulse response, the user voice signal, and the environmental noise.

[0013] The mathematical model of the microphone mixed signal is: Wherein, y(n) is the microphone mixed signal, x(n) is the loudspeaker far-end reference signal, h(n) is the acoustic echo path impulse response, s(n) is the user voice signal, and d(n) is the environmental noise.

[0014] The sampling rate in the step S1 is fixed at 48 kHz, and the quantization precision is 16 bits; The microphone mixed signal and the loudspeaker far-end reference signal are respectively collected by the MIC input collection chip in the step S1. The iteration formula of the AEC module is: Wherein, w(n) is the filter coefficient vector, μ is the step factor, e(n) is the error signal, is a regularization parameter.

[0015] The step formula of the VSS-NLMS adaptive filtering algorithm is: Wherein, e(n) is the current error signal, μmax is the initial maximum step, which is set to 0.1, is a minimum value, which is set to 0.001 to avoid a denominator of 0.

[0016] The filter coefficient is dynamically updated by the VSS-NLMS adaptive filtering algorithm, and the filter length is extended to 256-1024 orders.

[0017] The transfer function of the residual echo nonlinear suppression module is: Wherein, P_s(k) is the user voice power spectrum, and P_e(k) is the residual echo power spectrum calculated in real time through short-term Fourier transform.

[0018] An intelligent voice interaction real-time interruption processing system for implementing the intelligent voice interaction real-time interruption processing method, comprising: A pre-processing module that collects and stores a microphone mixed signal and a loudspeaker far-end reference signal; An AEC module that performs linear echo real-time cancellation based on a VSS-NLMS adaptive filtering algorithm; A residual echo nonlinear suppression module that performs real-time processing of nonlinear residual echo through a spectral enhancement algorithm of Wiener filtering; An AI-VAD module that detects a user voice signal and determines whether to interrupt the current playback.

[0019] Compared with the prior art, the beneficial effects of the present application mainly include: The present application makes full use of the computing power advantage of the voice AI chip, forms an end-to-end real-time processing link through the AEC module, the residual echo nonlinear suppression module, the AI-VAD module and the interrupt mechanism, ensures the low-delay cooperative operation of each module, accurately distinguishes the user voice from the echo signal and the environmental noise, realizes the natural interruption without the wake-up word, and solves the problem of unnatural interaction mode.

[0020] In order to make the above and other objects, features and advantages of the present application more apparent, the following will describe a preferred embodiment in detail, and the accompanying drawings will be described as follows. BRIEF DESCRIPTION OF DRAWINGS

[0021] In order to more clearly illustrate the technical solutions in the specific embodiments of the present application, the following will briefly introduce the drawings needed to be used in the embodiment description. Obviously, the drawings described below are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0022] Figure 1 The intelligent voice interaction real-time interruption processing method flowchart provided by the embodiment of the present application.

[0023] Figure 2 The intelligent voice interaction real-time interruption processing system block diagram provided by the embodiment of the present application. Figure 1 .

[0024] Figure 3 The intelligent voice interaction real-time interruption processing system block diagram provided by the embodiment of the present application. Figure 2 .

[0025] Figure 4 The intelligent voice interaction real-time interruption processing process flowchart provided by the embodiment of the present application. DETAILED DESCRIPTION

[0026] The foregoing and other technical contents, features, and effects of the present invention will be clearly presented in the following detailed description of a preferred embodiment with reference to the accompanying drawings. The directional terms mentioned in the following embodiments, such as up, down, left, right, front, or back, are merely for reference to the accompanying drawings. Therefore, the directional terms used are for illustrative purposes and not for limiting the present invention.

[0027] The embodiments of this application will now be described in detail with reference to the accompanying drawings. However, those skilled in the art will understand that many technical details have been provided in the embodiments of this application to facilitate a better understanding of the application. However, the technical solutions claimed in this application can be implemented even without these technical details and various variations and modifications based on the following embodiments.

[0028] Example 1 like Figure 1 As shown, a method for handling real-time interruptions in intelligent voice interaction includes: Step S1: The preprocessing module synchronously acquires the microphone mixed signal and the far-end reference signal of the speaker, and stores the acquired signals; In this embodiment, the preprocessing module is a MIC input acquisition chip, and the preprocessing process includes voice acquisition and storage. Specifically, the MIC input acquisition chip acquires the microphone mixed signal and the speaker far-end reference signal, and stores the acquired signals. The MIC input acquisition chip is a general-purpose chip that interacts with the AEC processor. Step S2: Input the preprocessed signal into the AEC module based on the VSS-NLMS adaptive filtering algorithm for real-time processing to eliminate linear echo; Step S3: The residual echo nonlinear suppression module processes nonlinear residual echoes in real time based on the Wiener filter-based spectrum enhancement algorithm; Step S4: Input the processed signal into the AI-VAD module for detection. If there is a user voice signal with a duration of ≥100ms, it is determined to be an interrupt command, triggering the interrupt mechanism to interrupt the current playback.

[0029] Specifically, the core reason for echo generation lies in the coupling of the acoustic path. The microphone mixed signal includes: the echo signal formed by coupling the far-end reference signal x(n) of the loudspeaker through the impulse response h(n) of the acoustic echo path, the user's voice signal, and environmental noise. Its mathematical model is as follows: Where y(n) is the microphone mixed signal, x(n) is the far-end reference signal of the loudspeaker, h(n) is the acoustic echo path impulse response, s(n) is the user's voice signal, and d(n) is the ambient noise.

[0030] likeFigure 2 As shown, the pre-processing module synchronously collects the microphone mixed signal and the loudspeaker far-end reference signal, the sound played by the loudspeaker (i.e. the loudspeaker far-end reference signal x(n)) is coupled into the microphone through the acoustic echo path impulse response h(n) (such as air propagation, device shell vibration, etc.) to form the microphone mixed signal y(n) mixed with the user voice signal s(n) and the environmental noise d(n); wherein the sound emitted by the loudspeaker as an echo signal is the core factor to interfere with the recognition of the interrupt command, if it cannot be effectively eliminated, the voice recognition system will misjudge the sound emitted by the loudspeaker as the user input, or cover the real user voice, resulting in the failure of the real-time interrupt function.

[0031] The present application collects the microphone mixed signal in real time through the MIC input acquisition chip, the sampling rate is fixed at 48 kHz, and the quantization precision is 16 bits, ensuring high fidelity of the signal; at the same time, the MIC input acquisition chip collects the loudspeaker far-end reference signal in real time, synchronously acquires the far-end reference signal x(n) played by the loudspeaker (i.e. the voice signal output by the device itself), as the reference for echo cancellation.

[0032] As shown in Figure 3 The pre-processing module synchronously collects the microphone mixed signal and the loudspeaker far-end reference signal, and stores the collected signals; the AEC module performs real-time processing on the pre-processed signals based on the VSS-NLMS adaptive filtering algorithm to eliminate linear echo; the residual echo nonlinear suppression module performs real-time processing on the signals processed by the AEC module based on the frequency spectrum enhancement algorithm of Wiener filtering to eliminate nonlinear residual echo; the AI-VAD module detects the user voice signal and judges whether to interrupt the current playing.

[0033] The signal processed by the pre-processing module is transmitted to the AEC module based on the VSS-NLMS adaptive filtering algorithm for real-time processing, and the AEC algorithm mostly adopts the adaptive filtering based on the normalized least mean square (NLMS), and its iteration formula is: Wherein w(n) is the filter coefficient vector, μ is the step factor, e(n) is the error signal, is the regularization parameter.

[0034] Since the algorithm is not robust enough in complex acoustic environment, the present application based on the VSS-NLMS adaptive filtering algorithm adopts the variable step size normalized least mean square algorithm to dynamically update the filter coefficients, and extends the filter length to 256-1024 orders to adapt to complex acoustic environment, and the step size adjustment formula is: Where e(n) is the current error signal; μmax is the initial maximum step size, set to 0.1; To minimize the value, we set it to 0.001 to avoid a denominator of 0.

[0035] In a preferred embodiment of the present invention, the filter length is extended to 256-1024 order (adaptively adjusted according to the complexity of the acoustic environment, such as 256 order for indoor static environment and 1024 order for noisy mobile environment), accurately modeling the acoustic echo path impulse response h(n), and initially eliminating linear echo.

[0036] This application improves the linear echo cancellation efficiency by accurately modeling the acoustic echo path impulse response h(n) through dynamic adjustment of step size and filter length, providing core technical support for "echo cancellation amount reaching 30-45dB" and "adapting to complex environments".

[0037] The Residual Echo Nonlinear Suppression (RES) module uses a Wiener filter-based spectrum enhancement algorithm to process nonlinear residual echoes in real time. Its transfer function is: Where Ps(k) is the user's speech power spectrum, and Pe(k) is the residual echo power spectrum calculated in real time through short-term Fourier transform.

[0038] like Figure 4 As shown, the sound played by the loudspeaker (i.e., the loudspeaker's far-end reference signal x(n)) is coupled through the acoustic echo path impulse response h(n) to form an echo signal. The echo signal and the ambient noise d(n) are acquired by the MIC input acquisition chip. After the echo signal is noise-reduced, it first passes through the AEC module to eliminate linear echo, and then passes through the residual echo nonlinear suppression module to eliminate nonlinear residual echo. At this point, the voice data has completely eliminated the echo signal and ambient noise.

[0039] This application utilizes a residual echo nonlinear suppression (RES) module based on the power spectrum (… Wiener filter transfer function ( Further suppressing residual echoes caused by nonlinear distortion, the total echo cancellation amount reaches 30-45dB, enhancing the echo cancellation effect, directly improving the purity of speech recognition, and ensuring the achievement of "low false interruption rate (≤2%)". The echo-cancelled speech signal is then input into the AI-VAD module for detection. At this point, the speech data has eliminated the echo signal and environmental noise, leaving only the user's speech signal, which serves to interrupt the AI ​​speech. The interruption is triggered by detecting user speech segments that last ≥100ms, and the system is linked to stop the current playback.

[0040] Specifically, the AI-VAD module outputs a real-time detection result, if a user voice signal (duration > 100 ms, excluding false triggering) is detected, it is determined as a "interrupt instruction"; the system immediately triggers the interrupt mechanism, the interrupt mechanism is set as: pause the current voice playing of the loudspeaker, the user voice signal collected by the microphone is preferentially sent into the voice recognition engine, and the instruction analysis and execution are completed; after the execution is completed, it can be selected according to the user instruction to continue playing the interrupted content or start a new interactive process.

[0041] To sum up, the application makes full use of the computing power advantage of the voice AI chip, forms an end-to-end real-time processing link through the AEC module, the residual echo nonlinear suppression module, the AI-VAD module and the interrupt mechanism, ensures the low-delay cooperative operation of each module, accurately distinguishes the user voice from the echo signal and the environmental noise, realizes the natural interruption without the wake-up word, and solves the unnatural problem of the interactive mode.

[0042] Some commonly used English names or letters used by the application for the purpose of clear description are only used for exemplary reference, and are not limited to the protection scope of the application.

[0043] It should also be noted that, in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between the entities or operations.

Claims

1. A method for intelligent real-time voice interaction interrupt processing, characterized in that, Comprising: Step S1: a pre-processing module synchronously collects a microphone mixed signal and a loudspeaker far-end reference signal, and stores the collected signals; Step S2: the pre-processed signals are input into an AEC module based on a VSS-NLMS adaptive filtering algorithm for real-time processing to eliminate linear echo; Step S3: a residual echo nonlinear suppression module performs real-time processing on nonlinear residual echo based on a Wiener filtering-based spectral enhancement algorithm; Step S4: the processed signals are input into an AI-VAD module for detection, and if there is a user voice signal with a duration of ≥100 ms, it is determined as an interrupt instruction, triggering an interrupt mechanism to interrupt the current playback. 2.The intelligent voice interaction real-time interrupt processing method of claim 1, wherein, The interrupt mechanism is set as: Pausing the current playback of the loudspeaker, recognizing the user voice signal, performing instruction analysis and execution, and selecting to continue playing or starting a new interactive process according to the user instruction. 3.The intelligent voice interaction real-time interrupt processing method of claim 1, wherein, The microphone mixed signal includes: An echo signal formed after the loudspeaker far-end reference signal is coupled through an acoustic echo path impulse response, a user voice signal, and environmental noise.

4. The method of claim 3, wherein, The mathematical model of the microphone mixed signal is: Where y(n) is the microphone mixed signal, x(n) is the loudspeaker far-end reference signal, h(n) is the acoustic echo path impulse response, s(n) is the user voice signal, and d(n) is the environmental noise.

5. The method of claim 3, wherein the method further comprises: The sampling rate in step S1 is fixed at 48 kHz, and the quantization accuracy is 16 bits; The microphone mixed signal and the loudspeaker far-end reference signal are collected by a MIC input collection chip in step S1.

6. The method of claim 1, wherein, The iteration formula of the AEC module is: where w(n) is a filter coefficient vector, μ is a step factor, e(n) is an error signal, is a regularization parameter.

7. The method of claim 1, wherein the method further comprises: The step size formula of the VSS-NLMS adaptive filtering algorithm is: Wherein, e(n) is the current error signal, μmax is the initial maximum step size, set to 0.1, is a minimum value, set to 0.001 to avoid a denominator of 0.

8. The method of claim 7, wherein, The filter coefficients are dynamically updated by the VSS-NLMS adaptive filtering algorithm, and the filter length is extended to 256-1024 orders.

9. The method of claim 1, wherein, The transfer function of the residual echo nonlinear suppression module is: Where P_s(k) is the user voice power spectrum, and P_e(k) is the residual echo power spectrum calculated in real time by short-time Fourier transform.

10. An intelligent voice interaction real-time interrupt processing system, used to implement the intelligent voice interaction real-time interrupt processing method of any one of claims 1-9, characterized in that, Comprising: A pre-processing module that collects and stores microphone mixed signals and loudspeaker far-end reference signals; An AEC module that performs real-time elimination of linear echo based on a VSS-NLMS adaptive filtering algorithm; A residual echo nonlinear suppression module that performs real-time processing of nonlinear residual echo through a Wiener filtering-based spectral enhancement algorithm; An AI-VAD module that detects user voice signals and determines whether to interrupt the current playback.