Cascade echo cancellation method and system based on double-end cooperation

By collaborating between the peer and local echo processing, combined with frequency domain dynamic spectrum subtraction algorithm and adaptive filtering technology, the problem of incomplete single-end processing and lack of coordination mechanism in the existing echo cancellation technology is solved, and the quality and user experience of voice communication are improved.

CN120452465APending Publication Date: 2025-08-08SHANGHAI LONGCHEER INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635005.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The existing echo cancellation technology has problems such as incomplete single-ended processing, lack of dual-ended collaboration mechanism, poor adaptability in complex scenarios, and excessive reliance on peer-end devices, resulting in residual echo affecting downlink voice quality, voice distortion, and poor user experience.

Method used

The cascading echo cancellation method based on dual-end collaboration is adopted, and the echo is processed through the joint and local ends, and the residual echo is initially eliminated. Combined with the frequency domain dynamic spectral subtraction algorithm and adaptive filtering technology, the enable status and parameters of the suppression function are dynamically adjusted to balance the echo cancellation effect and speech fidelity.

Benefits of technology

Effectively reduce residual echoes, improve downlink voice quality and user experience, adapt to different call scenarios and device capabilities, and achieve a balance between echo cancellation effect and voice fidelity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120452465A_ABST
    Figure CN120452465A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice signal processing, and discloses a cascade echo cancellation method and system based on double-end cooperation. According to the method, a double-end cooperative cascade echo cancellation architecture is constructed, and an opposite end and a home end cooperate together to process echoes. The opposite end firstly carries out preliminary elimination on the echo and marks a part with relatively high residual echo; after a home terminal receives related information, secondary processing is carried out on a downlink signal, a frequency domain dynamic spectrum subtraction algorithm is adopted to emphatically process a high-residual echo frame, and the echo cancellation effect is improved. And meanwhile, according to a call scene or whether the two ends support a coordination mechanism, dynamically adjusting the starting state and parameters of the suppression function of the home terminal, and dynamically adapting the suppression strength of the home terminal based on the residual echo identification information so as to balance the echo cancellation effect and the voice fidelity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech signal processing, and in particular relates to a cascade echo cancellation method and system based on double-end collaboration. Background Art

[0002] With the rapid development of modern voice communication technology and the widespread adoption of various voice interaction devices such as TWS headsets, smart speakers, and video conferencing terminals, people are increasingly demanding voice call quality. Echo cancellation technology, as a core component in ensuring clear and smooth voice communication, has a direct impact on the user experience. However, existing echo cancellation technology faces numerous challenges in practical applications, urgently requiring innovative solutions.

[0003] First, traditional echo cancellation solutions often use a single-ended approach, typically processing the echo only on the remote end. However, in real-world voice communication environments, interference factors such as nonlinear distortion, device differences, and complex and variable environmental noise levels exist. These factors result in a certain amount of residual echo being transmitted to the local end via the downlink even after echo cancellation has been performed on the remote end. During a call, users will noticeably hear a delayed echo of their own voice, significantly reducing the clarity and quality of the downlink voice, and severely affecting the fluency and intelligibility of the call.

[0004] Secondly, under the current technology system, there is a lack of effective collaborative working mechanisms between the remote end and the local end. The reference signal and related information obtained by the remote end during echo processing cannot be efficiently transmitted to the local end. This also makes it difficult for the local end to optimize downlink signal processing by leveraging the remote end's processing results. The independent processing modes of both ends result in suboptimal overall echo cancellation, limiting the potential for improvement in voice communication quality.

[0005] Furthermore, in duplex calls, where voice signals are transmitted in both directions, the generation and propagation of echoes are more complex, making them difficult to effectively address with traditional single-ended processing solutions. Furthermore, in noisy environments, such as multi-person conferences or noisy outdoor streets, background noise can severely interfere with the proper functioning of echo cancellation algorithms, leading to a sharp decline in voice quality. Furthermore, in multi-device environments, the remote device may mistake local voice for echo and incorrectly cancel it, causing voice distortion. Alternatively, echoes from multiple devices may overlap, further degrading call quality and significantly reducing the user experience.

[0006] Finally, existing echo cancellation solutions rely heavily on the performance of the peer device. In scenarios such as hands-free calls, if the peer device experiences processing errors, the local user will inevitably experience echo interference. Furthermore, traditional single-ended solutions lack effective solutions for residual echo caused by nonlinear distortion and device differences, which has become a key bottleneck hindering the development of voice communication technology. Summary of the Invention

[0007] The present invention aims to provide a cascaded echo cancellation method and system based on dual-end collaboration, aiming to solve the problems of traditional single-end echo cancellation solutions, such as incomplete single-end processing, lack of a dual-end collaboration mechanism, poor adaptability to complex scenarios, and excessive dependence on peer devices, which in turn lead to residual echo affecting downlink voice quality, voice distortion, and poor user experience.

[0008] To achieve the above object, in a first aspect of the present invention, a cascade echo cancellation method based on dual-end collaboration is provided, characterized in that it includes the following steps:

[0009] The peer end uses the voice signal sent by the local end as a reference signal, generates an echo estimation signal through adaptive filtering, performs preliminary echo cancellation, and outputs the preliminarily processed uplink signal; calculates the residual echo energy ratio; and generates residual echo identification information marking the frame as a high residual frame if the residual echo energy ratio exceeds a preset threshold;

[0010] The local end receives the reference signal and the residual echo identification information; generates a secondary echo estimation signal through secondary adaptive filtering based on the reference signal and the downlink signal; performs frequency domain optimization processing on the downlink signal based on the residual echo identification information, and outputs a speech signal after the residual echo is eliminated;

[0011] Dynamically adjust the enabled state and parameters of the local suppression function based on the call scenario or whether both ends support the collaborative mechanism; dynamically adapt the local suppression strength based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

[0012] Furthermore, in the dual-end collaboration-based cascade echo cancellation method, the adaptive filtering is implemented using a minimum mean square error algorithm or a recursive least squares algorithm.

[0013] Furthermore, in the dual-end collaboration-based cascade echo cancellation method, performing frequency domain optimization processing on the downlink signal according to the residual echo identification information includes:

[0014] Converting the downlink signal into a frequency domain signal;

[0015] Based on the residual echo identification information, performing a spectral subtraction operation only on the frequency domain signal marked as a high residual frame;

[0016] The optimized frequency domain signal is converted back into a time domain signal for output.

[0017] Furthermore, in the dual-end collaboration-based cascade echo cancellation method, dynamically adjusting the activation state of the suppression function in the local end according to the call scenario includes:

[0018] Activating the suppression function of the local end in a two-way call scenario;

[0019] In a one-way call scenario, whether to disable the suppression function of the local end is determined according to the residual echo identification information.

[0020] Furthermore, in the dual-end collaboration-based cascade echo cancellation method, dynamically adjusting the activation state of the suppression function in the local end according to whether the two ends support the collaboration mechanism includes:

[0021] If the local end does not support the cooperative mechanism, but the peer end supports the cooperative mechanism, the peer end adjusts the primary echo cancellation strength to the maximum to ensure the overall effect;

[0022] If the opposite end does not support the coordination mechanism, but the local end supports the coordination mechanism, the suppression function of the local end is activated.

[0023] Furthermore, in the dual-end collaboration-based cascade echo cancellation method, the dynamic adaptation of the local-end suppression strength based on the residual echo identification information is specifically:

[0024] When the residual echo energy ratio exceeds a preset threshold, increasing the suppression strength of the local end;

[0025] When the residual echo energy ratio is lower than a preset threshold, the local suppression strength is reduced.

[0026] In another aspect of the present invention, a cascade echo cancellation system based on dual-end collaboration is also proposed, comprising a peer-end processing module, a local-end processing module, and a collaborative control module;

[0027] The peer processing module includes: a reference signal acquisition unit, configured to acquire the voice signal sent by the local processing module as a reference signal; a preliminary echo cancellation unit, configured to generate an echo estimation signal based on the reference signal, perform preliminary echo cancellation, and output an uplink signal after preliminary processing; and a residual echo analysis unit, configured to calculate a residual echo energy ratio and generate residual echo identification information for marking a high residual frame.

[0028] The local processing module includes: a data receiving unit for receiving the reference signal and the residual echo identification information; a secondary echo cancellation unit for generating a secondary echo estimation signal through secondary adaptive filtering based on the reference signal and the downlink signal; and a frequency domain optimization processing unit for performing frequency domain optimization processing on the downlink signal based on the residual echo identification information and outputting a speech signal after the residual echo is eliminated;

[0029] The collaborative control module includes: a scene recognition unit, which is used to dynamically adjust the activation status of the suppression function in the local processing module according to the call scene or whether the two ends support the collaborative mechanism; and a parameter adaptation unit, which is used to dynamically adapt the suppression strength of the local processing module based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

[0030] Furthermore, in the cascaded echo cancellation system based on dual-end collaboration, the frequency domain optimization processing unit performs the following operations: converts the downlink signal into a frequency domain signal; based on the residual echo identification information, performs spectral subtraction operation only on the frequency domain signal marked as a high residual frame; and converts the optimized frequency domain signal back into a time domain signal output.

[0031] Furthermore, in the dual-end collaboration-based cascade echo cancellation system, when the scene recognition unit dynamically adjusts the activation state of the suppression function in the local processing module according to the call scene:

[0032] activating the suppression function of the local processing module in a two-way call scenario;

[0033] In a one-way call scenario, whether to disable the suppression function of the local processing module is determined according to the residual echo identification information.

[0034] Furthermore, in the dual-end collaboration-based cascade echo cancellation system, when the scene recognition unit dynamically adjusts the activation state of the suppression function in the local processing module according to whether the two ends support the collaboration mechanism:

[0035] If the local processing module does not support the collaborative mechanism, but the peer processing module supports the collaborative mechanism, the peer processing module adjusts the primary echo cancellation strength to the maximum to ensure the overall effect;

[0036] If the opposite-end processing module does not support the collaborative mechanism, but the local-end processing module supports the collaborative mechanism, the suppression function of the local-end processing module is activated.

[0037] Compared with the prior art, the present invention has at least the following technical effects:

[0038] The present invention constructs a dual-end collaborative cascade echo cancellation architecture, where the peer end and the local end collaborate to process the echo. The peer end first performs preliminary echo cancellation and marks the portion with higher residual echo. After receiving the relevant information, the local end performs secondary processing on the downlink signal, using a frequency domain dynamic spectrum subtraction algorithm to focus on processing high residual echo frames to improve the echo cancellation effect. At the same time, based on the call scenario or whether both ends support the collaborative mechanism, the enabled state and parameters of the local end's suppression function are dynamically adjusted, and the local end's suppression strength is dynamically adapted based on the residual echo identification information to balance the echo cancellation effect and voice fidelity. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 4 is a flowchart of a method for cascade echo cancellation based on dual-end collaboration in one embodiment of the present invention;

[0040] Figure 2 This is a flowchart of frequency domain dynamic spectrum subtraction in one embodiment of the present invention;

[0041] Figure 3 FIG. 4 is a flow chart of an echo cancellation system for a near-end device and a far-end device in another embodiment of the present invention. DETAILED DESCRIPTION

[0042] The following is a more detailed description of a dual-end collaborative cascade echo cancellation method and system of the present invention, with reference to a schematic diagram. This diagram illustrates a preferred embodiment of the present invention. It should be understood that those skilled in the art may modify the present invention described herein while still achieving the beneficial effects of the present invention. Therefore, the following description should be understood as generally known to those skilled in the art and is not intended to limit the present invention.

[0043] For the sake of clarity, not all features of actual embodiments are described. In the following description, well-known functions and structures are not described in detail because they would obscure the present invention with unnecessary detail. It should be understood that in the development of any actual embodiment, numerous implementation details must be made to achieve the developer's specific goals, such as adapting from one embodiment to another to accommodate system or business constraints. Furthermore, it should be understood that such development work may be complex and time-consuming, but is nevertheless a routine undertaking for those skilled in the art.

[0044] The following paragraphs describe the present invention in more detail by way of example with reference to the accompanying drawings. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are provided solely for the purpose of assisting in the description of the embodiments of the present invention.

[0045] Based on the teachings of this specification, those skilled in the art may form new technical solutions by cross-combining different implementation methods without generating technical contradictions. Such variations should be deemed to fall within the scope of protection of this patent.

[0046] Existing echo cancellation technologies have shortcomings such as single-ended processing affected by factors such as nonlinear distortion, resulting in residual echo; a lack of a coordinated mechanism between the two ends; poor adaptability to complex scenarios; and excessive reliance on peer devices. These shortcomings make it impossible to resolve the residual echo problem caused by nonlinear distortion and device differences, severely restricting the quality of voice communications.

[0047] Based on this, Figure 1 As shown, in one embodiment of the present invention, a cascade echo cancellation method based on dual-end collaboration is proposed, comprising the following steps:

[0048] Step S1: The peer end uses the voice signal sent by the local end as a reference signal, generates an echo estimation signal through adaptive filtering, performs preliminary echo cancellation, and outputs an uplink signal after preliminary processing; calculates the residual echo energy ratio, and generates residual echo identification information marking it as a high residual frame if the residual echo energy ratio exceeds a preset threshold.

[0049] Step S2: The local end receives the reference signal and the residual echo identification information; generates a secondary echo estimation signal through secondary adaptive filtering based on the reference signal and the downlink signal; performs frequency domain optimization processing on the downlink signal based on the residual echo identification information, and outputs a speech signal after the residual echo is eliminated.

[0050] Step S3: Dynamically adjust the enabled state and parameters of the local suppression function according to the call scenario or whether both ends support the collaborative mechanism; dynamically adapt the local suppression strength based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

[0051] It should be noted that in this embodiment, the local end refers to the voice communication device directly used by the user (such as a mobile phone, computer, or conference terminal), which is primarily responsible for receiving voice signals sent by the remote end and reducing residual echo through secondary echo cancellation and frequency domain optimization processing. The remote end refers to the device at the other end of the voice communication (such as the terminal of the other party), whose main function is to perform preliminary echo cancellation on the voice signals sent by the local end, mark voice frames with high residual echo, and transmit the relevant information to the local end. The two work together through a low-latency communication link, forming a cascaded architecture of "initial noise cancellation and marking of high residual frames by the remote end—targeted secondary processing by the local end." This dynamically adapts to different call scenarios and device capabilities, achieving a balance between echo cancellation effectiveness and voice fidelity.

[0052] Specifically, in step S1, first, the other end collects a voice signal from the remote end of the communication link through a microphone or other audio input device. The voice signal sent by the local end serves as a reference template for subsequent adaptive filtering to estimate the echo generated by the sound played by the local speaker at the microphone.

[0053] Furthermore, an echo estimation signal is generated through adaptive filtering. The adaptive filtering is implemented using at least one of a minimum mean square error algorithm or a recursive least squares algorithm. Using adaptive algorithms such as the minimum mean square error (LMS) or recursive least squares (RMS) algorithm, the filter coefficients are dynamically adjusted based on a reference signal to generate an echo estimation signal that is highly similar to the actual echo. The adaptive nature of the filter enables real-time tracking of echo path changes (such as device position movement and changes in ambient acoustics). Compared to fixed filters, adaptive filtering can more effectively process nonlinear echoes.

[0054] Afterwards, preliminary echo cancellation is performed on the mixed signal. Specifically, the echo estimation signal is subtracted from the mixed signal (containing the near-end speech signal and the echo signal) to obtain the preliminarily processed uplink signal. The mathematical expression is: y(n) = x(n) - d^(n). Here, x(n) is the mixed signal, d^(n) is the echo estimation signal, and y(n) is the preliminarily processed signal. This preliminarily processed uplink signal can significantly reduce echo energy, improve uplink speech clarity, and minimize interference to the listener at the other end.

[0055] Furthermore, the residual echo energy ratio (RER) is calculated by comparing the energy of the signal after preliminary processing with the energy of the original mixed signal: Residual echo energy ratio RER = E[y(n) 2 ] / E[x(n) 2 ], where E[·] represents the energy expectation operation. The residual echo energy ratio is compared with a preset threshold. If it exceeds the preset threshold, the frame is marked as a high-residual frame, and a residual echo identifier containing information such as the timestamp and energy ratio is generated. It should be noted that by marking frames exceeding the preset threshold as high-residual frames and only initiating enhancement processing for high-residual frames, it is possible to avoid uniform high-intensity processing of all frames, reduce the system's computational load, minimize the risk of speech distortion, and balance echo cancellation with speech fidelity.

[0056] Specifically, in step S2, the local end first receives a reference signal and residual echo identification information for high-residual frames from the peer end via a communication channel. This communication channel must ensure data synchronization and time alignment between the reference signal and the downlink signal. The communication channel can be a low-latency wireless communication protocol for achieving data synchronization between the peer end and the local end. Using the residual echo identification, the local end can focus on high-risk high-residual frames and reduce excessive processing of normal frames.

[0057] Next, the local end uses the received reference signal and the local downlink signal to apply an adaptive filtering algorithm (such as LMS or RLS) to generate a secondary echo estimate signal. This step forms a cascade structure with the initial adaptive filtering performed by the remote end, compensating for deficiencies in the remote end's processing. This further refines the echo model, particularly for nonlinear distortion or rapidly changing echo paths. Combining the remote end's reference signal with the local downlink signal, a secondary echo estimate signal is generated that more closely resembles the actual echo, improving cancellation accuracy.

[0058] Furthermore, based on the residual echo identification information, the downlink signal is subjected to frequency domain optimization processing, and the specific steps are as follows: the downlink signal is converted into a frequency domain signal; based on the residual echo identification information, a spectral subtraction operation is performed only on the frequency domain signal marked as a high residual frame; and the optimized frequency domain signal is converted back into a time domain signal for output.

[0059] Specifically, such as Figure 2 As shown, frame windowing is a preprocessing step that converts the downlink signal into the frequency domain. Each windowed frame undergoes a Fourier transform (FFT) to convert the time domain to the frequency domain, generating a frequency domain signal. The system then dynamically calculates a coefficient α based on the residual echo energy ratio (RER) flag. For example, when the RER exceeds a preset threshold, indicating strong residual echo, the value of α is adjusted to increase echo suppression. If the RER falls below the threshold, the value of α is adjusted accordingly to reduce the processing intensity of the speech signal to ensure speech fidelity. The RER flag and the calculated α are then combined to process the frequency domain signal. Specifically, spectral subtraction is performed only on frames marked as high residual. This subtraction subtracts echo-related spectral components from the amplitude spectrum of the frequency domain signal (the amount of subtraction is determined by α) to suppress the echo. The RER flag guides this process, determining the intensity of the spectral subtraction based on the degree of residual echo. This effectively eliminates the echo while minimizing damage to the original speech signal.

[0060] Finally, the frequency-domain signal after spectral subtraction is converted back to the time domain via an inverse fast Fourier transform (IFFT), yielding s_out. The signal s_out, restored to the time domain after this series of processing, is output as the optimized speech signal after dynamic spectral subtraction in the frequency domain effectively suppresses echoes. This signal has eliminated residual echoes to the greatest extent possible while preserving the original speech characteristics.

[0061] For step S3, dynamically adjusting the activation status of the suppression function in the local end according to the call scenario includes: activating the suppression function of the local end in a two-way call scenario; in a one-way call scenario, determining whether to turn off the suppression function of the local end according to the residual echo identification information.

[0062] Specifically, by analyzing the voice activity detection results, energy distribution characteristics, etc., it is possible to identify whether the current scenario is a two-way call (double talk) or a one-way call (single talk). In the two-way call scenario, the local suppression function can be forcibly activated. Even if the residual echo energy ratio does not reach the threshold, a certain intensity of processing is maintained to prevent echo accumulation. In the one-way call scenario, whether to turn off the local suppression function is determined based on the residual echo identification information and the actual echo cancellation effect of the other end. When the echo cancellation effect of the other end is good and the residual echo energy ratio is low, the local suppression function can be turned off to reduce power consumption and voice distortion. This can avoid excessive processing in scenarios where strong suppression is not required, such as retaining more original voice details in one-way calls. In addition, actively strengthening suppression in two-way calls can effectively prevent howling problems caused by echo feedback.

[0063] Furthermore, in step S3, the activation state of the suppression function at the local end may be dynamically adjusted according to whether the two ends support the coordination mechanism, specifically including:

[0064] If the local end does not support the cooperative mechanism, but the peer end supports the cooperative mechanism, the peer end adjusts the primary echo cancellation strength to the maximum to ensure the overall effect;

[0065] If the opposite end does not support the coordination mechanism, but the local end supports the coordination mechanism, the suppression function of the local end is activated.

[0066] Specifically, in the dual-end collaborative cascade echo cancellation method, the collaborative mechanism involves information exchange between the two communicating ends (the local end and the remote end) to achieve coordinated echo cancellation. This allows for flexible adjustment of the noise cancellation strategy based on the processing capabilities of both ends, avoiding duplicate processing or processing blind spots. This reduces the computational burden on a single device while ensuring effective echo cancellation, thereby improving voice fidelity.

[0067] Scenario 1: When the local device (such as an old terminal) is unable to receive or process the collaborative information sent by the other end, but the other end device has the collaborative capability.

[0068] Because the local end cannot enable the suppression function (due to lack of collaborative information support), the remote end must assume full echo cancellation responsibility. By enhancing the primary noise cancellation strength (for example, by increasing the convergence speed of adaptive filtering and increasing the filter order), the remote end can complete most of the echo cancellation as much as possible, reducing the amount of residual echo transmitted to the local end. This ensures that the overall echo cancellation effect is not significantly reduced. However, it should be noted that excessive noise cancellation on the remote end may cause voice distortion. Therefore, it is necessary to set an upper limit on the strength of the algorithm or implement voice activity detection (VAD) to prevent false cancellation.

[0069] Scenario 2: When the peer device cannot generate or send collaborative information, but the local device has collaborative processing capabilities.

[0070] Due to the lack of collaborative capability of the other end, its primary echo cancellation may be insufficient (such as the residual echo energy ratio is higher than the threshold), but it cannot feedback information to this end. At this time, this end needs to actively enable the suppression function and set the secondary echo cancellation of this end to an independent working mode. Based on the received reference signal, this end calculates the residual echo energy ratio in the downlink signal by itself, and automatically identifies high residual frames through spectrum analysis and energy threshold comparison. Afterwards, the identified high residual frames are subjected to enhanced secondary elimination processing and enhanced noise reduction as in step S2. The lack of collaborative capability of the other end is compensated by independent processing of this end, avoiding echo leakage caused by incomplete noise reduction of the other end.

[0071] Furthermore, the dynamic adaptation of the local suppression strength based on the residual echo identification information is specifically: increasing the local suppression strength when the residual echo energy ratio exceeds a preset threshold; and decreasing the local suppression strength when the residual echo energy ratio is lower than the preset threshold.

[0072] Specifically, the spectral subtraction factor α is dynamically adjusted according to the residual echo energy ratio to change the suppression strength at the local end. A larger α value is used to enhance suppression for high residual frames, while weak processing is maintained for low-risk frames, achieving "noise reduction on demand", responding to sudden changes in echo energy, and maintaining stability.

[0073] In another embodiment, a cascade echo cancellation system based on dual-end collaboration is proposed, including a peer-end processing module, a local-end processing module, and a collaborative control module.

[0074] The peer processing module includes a reference signal acquisition unit, a preliminary echo cancellation unit, and a residual echo analysis unit. The reference signal acquisition unit is configured to acquire the speech signal sent by the local processing module as a reference signal; the preliminary echo cancellation unit is configured to generate an echo estimation signal based on the reference signal, perform preliminary echo cancellation, and output an uplink signal after preliminary processing; and the residual echo analysis unit is configured to calculate the residual echo energy ratio and generate residual echo identification information for marking high residual frames.

[0075] The local processing module includes a data receiving unit, a secondary echo cancellation unit, and a frequency domain optimization processing unit. The data receiving unit is configured to receive the reference signal and the residual echo identification information; the secondary echo cancellation unit generates a secondary echo estimation signal through secondary adaptive filtering based on the reference signal and the downlink signal; and the frequency domain optimization processing unit performs frequency domain optimization processing on the downlink signal based on the residual echo identification information and outputs a speech signal after residual echo elimination.

[0076] The collaborative control module includes a scene recognition unit and a parameter adaptation unit. The scene recognition unit is used to dynamically adjust the activation state of the suppression function in the local processing module based on the call scenario or whether both ends support the collaborative mechanism. The parameter adaptation unit is used to dynamically adapt the suppression strength of the local processing module based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

[0077] Furthermore, in the dual-end collaboration-based cascade echo cancellation system, the opposite-end processing module and the local-end processing module communicate with each other via a low-latency wireless communication protocol to synchronize the reference signal and the residual echo identification information.

[0078] Furthermore, in the cascaded echo cancellation system based on dual-end collaboration, the frequency domain optimization processing unit performs the following operations: converts the downlink signal into a frequency domain signal; based on the residual echo identification information, performs spectral subtraction operation only on the frequency domain signal marked as a high residual frame; and converts the optimized frequency domain signal back into a time domain signal output.

[0079] Furthermore, in the dual-end collaboration-based cascade echo cancellation system, when the scene recognition unit dynamically adjusts the activation state of the suppression function in the local processing module according to the call scene:

[0080] activating the suppression function of the local processing module in a two-way call scenario;

[0081] In a one-way call scenario, whether to disable the suppression function of the local processing module is determined according to the residual echo identification information.

[0082] Furthermore, in the dual-end collaboration-based cascade echo cancellation system, when the scene recognition unit dynamically adjusts the activation state of the suppression function in the local processing module according to whether the two ends support the collaboration mechanism:

[0083] If the local processing module does not support the collaborative mechanism, but the peer processing module supports the collaborative mechanism, the peer processing module adjusts the primary echo cancellation strength to the maximum to ensure the overall effect;

[0084] If the opposite-end processing module does not support the collaborative mechanism, but the local-end processing module supports the collaborative mechanism, the suppression function of the local-end processing module is activated.

[0085] like Figure 3The figure illustrates an exemplary echo cancellation system flow for a near-end device and a far-end device. It is primarily divided into a near-end uplink echo module, a near-end downlink echo algorithm module, and relevant components of the far-end device. A series of signal processing steps are used to reduce the impact of echo on voice communications. In the near-end uplink echo module, a microphone collects sound signals, which are processed by the uplink echo algorithm module. The signal is then combined with the uplink echo algorithm output and the expected deviation signal from the microphone flow field. This signal is then output after passing through the uplink algorithm. In the near-end downlink echo algorithm module, an adaptive filter combined with a parameter cache processes the signal. This signal is combined with the far-end downlink signal and, after mute detection and other operations, is output as the downlink echo algorithm output signal. The final signal is processed by the sound effect algorithm and output as the actual downlink output. The far-end device processes the signal through the far-end uplink algorithm module and outputs it. This figure intuitively illustrates the signal flow and processing relationships between the various modules, which helps to understand the implementation process of the present invention.

[0086] In summary, the present invention constructs a dual-end collaborative cascade echo cancellation architecture, in which the opposite end and the local end jointly process the echo. The opposite end first performs a preliminary elimination of the echo and marks the part with higher residual echo; after receiving the relevant information, the local end performs a secondary processing on the downlink signal, and uses the frequency domain dynamic spectrum subtraction algorithm to focus on processing high residual echo frames to improve the echo cancellation effect. At the same time, according to the call scenario or whether the two ends support the collaborative mechanism, the activation state and parameters of the local end suppression function are dynamically adjusted, and the suppression strength of the local end is dynamically adapted based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

[0087] The above description is merely a preferred embodiment of the present invention and does not limit the present invention in any way. Any person skilled in the art who, without departing from the scope of the present invention, makes any equivalent substitution, modification, or other changes to the technical solution and technical content disclosed in the present invention shall be deemed to be within the scope of the present invention and still fall within the scope of protection of the present invention.

Claims

1. A cascade echo cancellation method based on dual-end collaboration, characterized in that: The following steps are involved: The peer end uses the voice signal sent by the local end as a reference signal, generates an echo estimation signal through adaptive filtering, performs preliminary echo cancellation, and outputs the preliminarily processed uplink signal; calculates the residual echo energy ratio; and generates residual echo identification information marking the frame as a high residual frame if the residual echo energy ratio exceeds a preset threshold; The local end receives the reference signal and the residual echo identification information; generates a secondary echo estimation signal through secondary adaptive filtering based on the reference signal and the downlink signal; performs frequency domain optimization processing on the downlink signal based on the residual echo identification information, and outputs a speech signal after the residual echo is eliminated; Dynamically adjust the enabled state and parameters of the local suppression function based on the call scenario or whether both ends support the collaborative mechanism; dynamically adapt the local suppression strength based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

2. The cascade echo cancellation method based on dual-end collaboration according to claim 1, characterized in that: The adaptive filtering is implemented by using at least one of a minimum mean square error algorithm and a recursive least squares algorithm.

3. The cascade echo cancellation method based on dual-end collaboration according to claim 1, characterized in that: The performing frequency domain optimization processing on the downlink signal according to the residual echo identification information includes: Converting the downlink signal into a frequency domain signal; Based on the residual echo identification information, performing a spectral subtraction operation only on the frequency domain signal marked as a high residual frame; The optimized frequency domain signal is converted back into a time domain signal for output.

4. The cascade echo cancellation method based on dual-end collaboration according to claim 1, characterized in that: The dynamically adjusting the activation state of the suppression function in the local end according to the call scenario includes: Activating the suppression function of the local end in a two-way call scenario; In a one-way call scenario, whether to disable the suppression function of the local end is determined according to the residual echo identification information.

5. The cascade echo cancellation method based on dual-end collaboration according to claim 1, characterized in that: The dynamically adjusting the activation state of the suppression function at the local end according to whether the two ends support the coordination mechanism includes: If the local end does not support the cooperative mechanism, but the peer end supports the cooperative mechanism, the peer end adjusts the primary echo cancellation strength to the maximum to ensure the overall effect; If the opposite end does not support the coordination mechanism, but the local end supports the coordination mechanism, the suppression function of the local end is activated.

6. The cascade echo cancellation method based on dual-end collaboration according to claim 1, characterized in that: The dynamic adaptation of the local suppression strength based on the residual echo identification information is specifically as follows: When the residual echo energy ratio exceeds a preset threshold, increasing the local end suppression strength; When the residual echo energy ratio is lower than a preset threshold, the local suppression strength is reduced.

7. A cascade echo cancellation system based on dual-end collaboration, characterized in that: It includes peer processing module, local processing module and collaborative control module; The peer processing module includes: a reference signal acquisition unit, configured to acquire the voice signal sent by the local processing module as a reference signal; a preliminary echo cancellation unit, configured to generate an echo estimation signal based on the reference signal, perform preliminary echo cancellation, and output an uplink signal after preliminary processing; and a residual echo analysis unit, configured to calculate a residual echo energy ratio and generate residual echo identification information for marking a high residual frame. The local processing module includes: a data receiving unit for receiving the reference signal and the residual echo identification information; a secondary echo cancellation unit for generating a secondary echo estimation signal through secondary adaptive filtering based on the reference signal and the downlink signal; and a frequency domain optimization processing unit for performing frequency domain optimization processing on the downlink signal based on the residual echo identification information and outputting a speech signal after the residual echo is eliminated; The collaborative control module includes: a scene recognition unit, which is used to dynamically adjust the activation status of the suppression function in the local processing module according to the call scene or whether the two ends support the collaborative mechanism; and a parameter adaptation unit, which is used to dynamically adapt the suppression strength of the local processing module based on the residual echo identification information to balance the echo cancellation effect and voice fidelity.

8. The cascade echo cancellation system based on dual-end collaboration according to claim 7, characterized in that: The frequency domain optimization processing unit performs the following operations: converting the downlink signal into a frequency domain signal; performing a spectrum subtraction operation only on the frequency domain signal marked as a high residual frame based on the residual echo identification information; and converting the optimized frequency domain signal back into a time domain signal for output.

9. The cascade echo cancellation system based on dual-end collaboration according to claim 7, characterized in that: When the scene recognition unit dynamically adjusts the activation state of the suppression function in the local processing module according to the call scene: activating the suppression function of the local processing module in a two-way call scenario; In a one-way call scenario, whether to disable the suppression function of the local processing module is determined according to the residual echo identification information.

10. The cascade echo cancellation system based on dual-end collaboration according to claim 7, characterized in that: When the scene recognition unit dynamically adjusts the activation state of the suppression function in the local processing module according to whether the two ends support the cooperation mechanism: If the local processing module does not support the collaborative mechanism, but the peer processing module supports the collaborative mechanism, the peer processing module adjusts the primary echo cancellation strength to the maximum to ensure the overall effect; If the opposite-end processing module does not support the collaborative mechanism, but the local-end processing module supports the collaborative mechanism, the suppression function of the local-end processing module is activated.