Self-adaptive echo cancellation method and device based on mutual inductance mechanism, equipment and medium
By employing an adaptive echo cancellation method based on mutual inductance mechanism, combined with closed-loop optimization of signal preprocessing, dual-talk state detection, and nonlinear compensation strategy, the problem of insufficient adaptability and robustness in existing technologies is solved, achieving efficient echo cancellation and speech signal optimization.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-02-12
- Publication Date
- 2026-06-26
AI Technical Summary
Existing adaptive echo cancellation technology has poor adaptability and robustness in complex and ever-changing audio communication environments, and is prone to incomplete echo cancellation and speech distortion. Furthermore, dual-talk detection and nonlinear compensation strategies cannot be optimized in a coordinated manner.
An adaptive echo cancellation method based on mutual inductance mechanism is adopted. Through a mutual inductance feedback closed loop consisting of signal synchronization preprocessing, dual-talk state detection, adaptive filter update and nonlinear compensation strategy, the judgment threshold and compensation strategy are dynamically adjusted to form a closed loop optimization.
It improves the adaptability and robustness of the adaptive echo cancellation system, accurately and efficiently eliminates echoes in audio communication, ensures the transmission quality of voice signals, and optimizes the call experience of audio communication.
Smart Images

Figure CN122290619A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of digital signal processing technology, and in particular to an adaptive echo cancellation method, apparatus, device and medium based on mutual inductance mechanism. Background Technology
[0002] In audio communication systems, the far-end reference signal played by the speaker is picked up by the microphone due to acoustic coupling effects such as air propagation and device reflection. This causes the near-end signal collected by the microphone to be mixed with far-end echo, which seriously affects the clarity and naturalness of voice calls. Therefore, acoustic echo cancellation has become a core key technology for real-time audio communication systems, and the adaptive filter method is currently the most widely used implementation scheme in this field.
[0003] Currently, mainstream adaptive echo cancellation solutions typically integrate adaptive filters, dual-talk detection, and nonlinear processing modules to achieve echo cancellation. However, existing technologies have some shortcomings in complex audio communication environments. On the one hand, the fixed judgment threshold of dual-talk detection cannot be dynamically adjusted according to environmental changes or the actual state of signal processing. In scenarios with high noise or abrupt echo path changes, it is prone to false positives or false negatives, leading to abnormal updates of adaptive filter weights, introducing residual echoes, or causing speech distortion. On the other hand, the independent operation mode of each core processing module makes it impossible for the nonlinear compensation strategy to be adaptively adjusted according to the dual-talk detection results. Dual-talk detection also cannot optimize its judgment criteria with the processing results of nonlinear compensation. The entire system cannot form a collaborative optimization closed loop, making it difficult to adapt to the ever-changing real-time audio communication scenarios. It is evident that existing technologies have poor adaptability and robustness in complex and ever-changing audio communication environments, and are prone to problems such as incomplete echo cancellation and speech distortion.
[0004] The preceding description is intended to provide general background information and does not necessarily constitute prior art. Summary of the Invention
[0005] Based on this, this application provides an adaptive echo cancellation method, apparatus, device, and medium based on mutual inductance mechanism, which can improve the adaptability and robustness of the adaptive echo cancellation system, accurately and efficiently eliminate echoes in audio communication systems, ensure the transmission quality of voice signals, and thus optimize the call experience of audio communication.
[0006] In a first aspect, embodiments of this application provide an adaptive echo cancellation method based on a mutual inductance mechanism, comprising: Acquire a far-end reference signal and a near-end microphone signal containing echo, and perform synchronous preprocessing on the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal; Dual-talk status detection is performed based on the preprocessed far-end signal and the preprocessed near-end signal, the current dual-talk status detection result is output, and the judgment threshold used for dual-talk status detection is dynamically adjusted based on feedback information. Based on the dual-talk state detection results, the adaptive filter is selectively updated, and the preprocessed far-end signal is filtered using the adaptive filter to output an estimated echo signal. The system receives the dual-talk state detection result and the estimated echo signal, selects a corresponding nonlinear compensation strategy based on the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, outputs the compensated echo signal, and generates feedback information for adjusting the judgment threshold. The compensated echo signal is subtracted from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
[0007] Optionally, in some embodiments of this application, the step of synchronously preprocessing the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal includes: Calculate the time delay between the far-end reference signal and the near-end microphone signal, and perform delay alignment on the far-end reference signal based on the time delay to obtain the delay-aligned far-end reference signal; High-pass filtering is applied to the delayed-aligned far-end reference signal and the near-end microphone signal respectively to remove low-frequency noise; The amplitude normalization process is performed on the high-pass filtered far-end reference signal and the near-end microphone signal to obtain the pre-processed far-end signal and the pre-processed near-end signal.
[0008] Optionally, in some embodiments of this application, the step of performing dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, outputting the current dual-talk status detection result, and dynamically adjusting the judgment threshold used for dual-talk status detection based on feedback information includes: Calculate the acoustic features corresponding to the preprocessed far-end signal and the preprocessed near-end signal respectively; the acoustic features include at least one of signal power ratio, spectral similarity and temporal envelope correlation; The calculated acoustic features are compared with the current judgment threshold. Based on the comparison result, it is determined whether the current frame is in dual-talk state, and the dual-talk state detection result is output. The decision threshold used for dual-talk determination is dynamically updated based on the feedback information generated by environmental noise estimation or mutual inductance nonlinear compensation.
[0009] Optionally, in some embodiments of this application, the step of selectively updating the adaptive filter based on the dual-talk state detection result, and using the adaptive filter to filter the preprocessed far-end signal to output an estimated echo signal, includes: The preprocessed far-end signal is input into the adaptive filter to calculate the estimated echo signal; The preprocessed near-end signal and the estimated echo signal are subtracted to obtain the error signal; Based on the error signal and the preprocessed remote signal, the weights of the adaptive filter are updated according to the normalized least mean square algorithm; wherein, when the dual-talk state detection result indicates a dual-talk state, the update of the adaptive filter weights is paused or slowed down.
[0010] Optionally, in some embodiments of this application, the step of receiving the dual-talk state detection result and the estimated echo signal, selecting a corresponding nonlinear compensation strategy based on the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, outputting the compensated echo signal, and generating feedback information for adjusting the judgment threshold includes: The estimated echo signal is subjected to distortion assessment, and distortion assessment results are generated; Based on the distortion assessment results and the dual-talk state detection results, a corresponding nonlinear compensation function is selected, and the selected nonlinear compensation function is applied to the estimated echo signal to generate an intermediate compensation signal. Based on the intermediate compensation signal, feedback information for adjusting the determination threshold is generated, and the intermediate compensation signal is output as the compensated echo signal.
[0011] Optionally, in some embodiments of this application, it further includes: When the dual-talk state detection result indicates a dual-talk state, the compensation strength of the selected nonlinear compensation function is lower than that of the nonlinear compensation function selected when the dual-talk state is not in a dual-talk state.
[0012] Optionally, in some embodiments of this application, the step of subtracting the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo-cancelled signal, followed by post-processing to output the final speech signal, includes: Subtracting the compensated echo signal from the preprocessed near-end signal yields the preliminary echo cancellation signal. Residual echoes are detected in the preliminary echo cancellation signal. If residual echoes are present, the preliminary echo cancellation signal is subjected to post-filtering processing. The initial echo cancellation signal, after residual echo suppression processing, is injected with comfortable noise in the silent section to generate the final speech signal.
[0013] Optionally, in some embodiments of this application, the adaptive filter is a finite impulse response filter, and the order of the finite impulse response filter is dynamically determined based on the time delay calculated in the signal synchronization preprocessing.
[0014] Secondly, embodiments of this application also provide an adaptive echo cancellation device based on a mutual inductance mechanism, comprising: The preprocessing module is used to acquire a far-end reference signal and a near-end microphone signal containing echo, and to perform synchronous preprocessing on the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal. The dual-talk detection module is used to perform dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, output the current dual-talk status detection result, and dynamically adjust the judgment threshold used for dual-talk status detection based on feedback information. The echo estimation module is used to selectively update the adaptive filter based on the dual-talk state detection result, and use the adaptive filter to filter the preprocessed far-end signal to generate an estimated echo signal. The compensation module is used to receive the dual-talk state detection result and the estimated echo signal, select the corresponding nonlinear compensation strategy according to the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, output the compensated echo signal, and generate feedback information for adjusting the judgment threshold. The signal synthesis module is used to subtract the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
[0015] Thirdly, embodiments of this application provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the adaptive echo cancellation method based on mutual inductance mechanism as described in the first aspect.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, implements the steps of the adaptive echo cancellation method based on mutual inductance mechanism as described in the first aspect.
[0017] This application provides an adaptive echo cancellation method, apparatus, device, and medium based on a mutual inductance mechanism. First, the acquired far-end reference signal and near-end microphone signal are synchronously preprocessed to provide a precise and unified signal foundation for subsequent two-way communication detection, echo path estimation, and nonlinear compensation, thus avoiding processing deviations caused by unstandardized signals at the source. Then, two-way communication status detection is performed based on the preprocessed signal, outputting accurate detection results and dynamically adjusting the judgment threshold based on feedback information, ensuring that the two-way communication judgment adapts to actual audio communication scenarios. Simultaneously, the adaptive filter is selectively updated based on the detection results to avoid erroneous filter updates in the two-way communication state, guaranteeing the accuracy of echo path estimation. This process involves several steps: First, the estimated echo signal is more closely aligned with the actual echo characteristics. Second, the nonlinear compensation stage receives the dual-talk detection result and selects the corresponding compensation strategy to adapt the nonlinear compensation to the actual call state, avoiding over- or under-compensation issues caused by a single compensation strategy. Simultaneously, feedback information is generated to adjust the dual-talk detection threshold, constructing a mutual feedback loop between dual-talk detection and nonlinear compensation, thus improving the system's adaptability to complex and changing communication environments. Finally, the preprocessed near-end signal is subtracted from the compensated echo signal to obtain the initial echo cancellation signal, which is then post-processed to output the final voice signal, achieving accurate echo cancellation and output signal optimization, effectively eliminating echo components in audio communication. In summary, this application improves the adaptability and robustness of the adaptive echo cancellation system, accurately and efficiently eliminates echoes in audio communication systems, ensures the transmission quality of voice signals, and thus optimizes the call experience in audio communication. Attached Figure Description
[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] in: Figure 1 This is a flowchart of the adaptive echo cancellation method based on mutual inductance mechanism provided in the embodiments of this application; Figure 2 This is a schematic diagram of the structure of the adaptive echo cancellation device based on mutual inductance mechanism provided in the embodiments of this application; Figure 3 This is a schematic diagram of the structure of the computer device provided in this application. Detailed Implementation
[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0021] This application provides an adaptive echo cancellation method, apparatus, device, and medium based on mutual inductance mechanism.
[0022] The adaptive echo cancellation device based on mutual inductance mechanism provided in this application can be integrated into a computer device. This computer device can be directly or indirectly connected to the device via wired or wireless communication. The computer device can be a smartphone, tablet, laptop, desktop computer, smart speaker, or smartwatch, but is not limited to these. Furthermore, the computer device can also connect to a server. The server can be a standalone physical server, a server cluster or distributed system composed of multiple physical servers, or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN, and big data and artificial intelligence platforms. This application does not impose any restrictions on these aspects.
[0023] This application discloses an adaptive echo cancellation method based on mutual inductance mechanism, comprising: acquiring a far-end reference signal and a near-end microphone signal containing echo, and synchronously preprocessing the far-end reference signal and the near-end microphone signal to obtain a preprocessed far-end signal and a preprocessed near-end signal; performing dual-talk state detection based on the preprocessed far-end signal and the preprocessed near-end signal, outputting the current dual-talk state detection result, and dynamically adjusting the judgment threshold used for dual-talk state detection based on feedback information; selectively updating an adaptive filter according to the dual-talk state detection result, and using the adaptive filter to filter the preprocessed far-end signal to output an estimated echo signal; receiving the dual-talk state detection result and the estimated echo signal, selecting a corresponding nonlinear compensation strategy according to the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, outputting a compensated echo signal, and generating feedback information for adjusting the judgment threshold; subtracting the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, performing post-processing, and outputting the final speech signal.
[0024] Please see Figure 1 , Figure 1 A flowchart illustrating the adaptive echo cancellation method based on mutual inductance mechanism provided in this application is shown. The specific process of this adaptive echo cancellation method based on mutual inductance mechanism can be as follows: S1. Acquire the far-end reference signal and the near-end microphone signal containing the echo, and perform synchronous preprocessing on the far-end reference signal and the near-end microphone signal to obtain the preprocessed far-end signal and the preprocessed near-end signal; Specifically, for step S1, the far-end reference signal and the near-end microphone signal are first acquired from the corresponding acquisition device of the audio communication system. The far-end reference signal is the audio signal transmitted from the far end to the local source and output by the local playback device during audio communication. The near-end microphone signal is a mixed signal captured by the local microphone, typically containing local speech, echoes from the far-end reference signal, and environmental noise. After signal acquisition, synchronous preprocessing is performed on both signals. The core of synchronous preprocessing is to achieve matching and calibration of the two signals in terms of timing and basic signal characteristics, eliminating timing deviations and basic signal interference caused by signal acquisition, transmission, and playback. This ensures that the processed far-end and near-end signals have a unified signal processing foundation, providing an accurate and suitable signal source for subsequent steps such as dual-talk detection and echo estimation.
[0025] For example, in a VoIP voice call scenario, the voice of the distant caller transmitted from the network to the local mobile phone and played by the mobile phone speaker is collected as the distant reference signal, and the mixed signal captured by the mobile phone microphone, which includes the echo of the distant voice, the voice of the local caller, and ambient noise, is collected as the near-end microphone signal. The two signals are time-calibrated and basic signal conditioning is performed to eliminate the signal asynchrony problem caused by network transmission delay and device acquisition timing differences, so as to obtain the preprocessed distant signal and near-end signal.
[0026] S2. Perform dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, output the current dual-talk status detection result, and dynamically adjust the judgment threshold used for dual-talk status detection based on the feedback information. Specifically, in step S2, based on the preprocessed far-end and near-end signals, dual-talk status detection is performed. The core of dual-talk status detection is to determine whether there is simultaneous voice input from the local and far-end during audio communication. By analyzing the features of the two signals, the dual-talk status detection result for the current communication scenario is output. The detection result is mainly divided into two categories: dual-talk status (local and far-end speaking simultaneously) and non-dual-talk status (only one end speaking or no voice at either end). While completing the dual-talk status detection and outputting the result, the judgment threshold used for dual-talk status detection is dynamically adjusted based on the feedback information from subsequent steps, rather than using a fixed threshold. This allows the threshold to adapt to the actual changes in the audio communication scenario, improving the accuracy of dual-talk status detection.
[0027] For example, in a teleconference scenario, based on the preprocessed voice signal of the remote participant and the signal collected by the local microphone, it is analyzed to determine whether there is a situation where the local and remote participants are speaking at the same time. If the voice features of the local signal are significantly enhanced, it is determined to be a dual-talk state and the result is output. If the echo compensation effect is not as expected in subsequent steps, it is suggested that there may be a false positive for dual-talk. Based on this, the threshold for dual-talk detection is adjusted appropriately to reduce the probability of false positives.
[0028] S3. Based on the dual-talk status detection results, selectively update the adaptive filter, and use the adaptive filter to filter the preprocessed far-end signal to output the estimated echo signal; Specifically, in step S3, the preprocessed far-end signal is first input into an adaptive filter. The filter simulates the echo path in audio communication, and then outputs an estimated echo signal that matches the actual echo characteristics. The core function of the adaptive filter is to dynamically simulate the echo path, achieving accurate estimation of the echo signal. Simultaneously, based on the output dual-talk status detection results, a selective update operation is performed on the adaptive filter. That is, the update rhythm and method of the filter are determined according to the actual call status, thereby avoiding coefficient deviations due to improper updates and ensuring the effectiveness of the echo estimation.
[0029] For example, in the voice call scenario of a smart speaker, when the dual-talk detection result is a non-dual-talk state, the coefficients of the adaptive filter are updated normally, allowing the filter to adapt to the changes in the echo path in real time, accurately simulate the echo, and output the estimated echo signal; when the detection result is a dual-talk state, the coefficient update of the adaptive filter is paused to avoid the local voice signal interfering with the echo path simulation of the filter and to prevent the estimated echo signal from deviating.
[0030] S4. Receive the dual-talk status detection result and the estimated echo signal, select the corresponding nonlinear compensation strategy according to the dual-talk status detection result to perform nonlinear compensation processing on the estimated echo signal, output the compensated echo signal, and generate feedback information for adjusting the judgment threshold. Specifically, in step S4, the dual-talk status detection result and the estimated echo signal are received simultaneously. Based on the dual-talk status detection result, a corresponding nonlinear compensation strategy is selected for the estimated echo signal. Then, nonlinear compensation processing is performed on the estimated echo signal according to the selected strategy to eliminate the nonlinear distortion problem of the echo signal caused by equipment distortion, transmission loss, etc. After processing, the compensated echo signal is output. While completing the nonlinear compensation processing and outputting the compensated echo signal, corresponding feedback information is also generated. This feedback information is transmitted to the dual-talk status detection stage as the basis for dynamically adjusting the dual-talk judgment threshold.
[0031] For example, in a video conferencing scenario, if the received dual-talk detection result indicates a dual-talk state, a low-intensity nonlinear compensation strategy is selected to process the estimated echo signal, avoiding overcompensation that could interfere with the local speech. If the detection result indicates a non-dual-talk state, a conventional compensation strategy is selected to fully eliminate the nonlinear distortion of the echo. After compensation, the power and characteristics of the compensated echo signal are used as feedback information and transmitted to the dual-talk state detection stage for dynamic adjustment of its judgment threshold.
[0032] S5. Subtract the compensated echo signal from the preprocessed near-end signal to obtain the preliminary echo cancellation signal, then perform post-processing to output the final speech signal; Specifically, for step S5, the core subtraction operation for echo cancellation is first performed. The preprocessed near-end signal is subtracted from the compensated echo signal. This operation removes the echo component from the near-end signal, resulting in a preliminary echo-cancelled signal. The preliminary echo-cancelled signal is then post-processed. The core of post-processing is to optimize the output quality of the signal, ensuring it meets the auditory and transmission requirements of audio communication. Finally, the post-processed signal is used as the final audio signal and output to the remote playback device of the audio communication system.
[0033] For example, in a vehicle voice call scenario, the pre-processed mixed signal from the vehicle microphone is subtracted from the compensated echo signal to remove the echo component of the remote caller, resulting in a preliminary echo-cancelled signal containing the local driver's voice and a small amount of in-vehicle ambient noise. This signal is then subjected to minor noise suppression, signal gain optimization, and other post-processing before a clear local voice signal is finally output to the remote call device.
[0034] This embodiment constructs a mutual inductance feedback mechanism for dual-talk detection and nonlinear compensation, and coordinates with the orderly connection of signal synchronous preprocessing, selective filter updating and echo cancellation postprocessing to form a closed-loop optimization of the entire echo cancellation system. This improves the adaptability and robustness to complex audio communication environments, effectively eliminates echoes in audio communication, reduces speech distortion, and ensures the clarity and naturalness of the final output speech signal, thereby optimizing the overall audio communication call experience.
[0035] Optionally, in some embodiments, step S1, "simultaneously preprocessing the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal," may specifically include: S11. Calculate the time delay between the far-end reference signal and the near-end microphone signal, and perform delay alignment on the far-end reference signal based on the time delay to obtain the delay-aligned far-end reference signal; Specifically, for step S11, the acquired far-end reference signal and near-end microphone signal are first subjected to temporal feature analysis to calculate the time delay between the two signals. This delay is usually caused by factors such as signal transmission, device playback, and timing differences in acquisition. Based on the calculated time delay value, a delay alignment operation is performed on the far-end reference signal. The timing of the far-end reference signal is adjusted through signal buffering and other methods to match the far-end reference signal with the near-end microphone signal in the time dimension after delay alignment. This eliminates the deviation caused by timing asynchrony in subsequent signal processing, giving the two signals a unified processing basis in the time dimension.
[0036] For example, in a teleconference scenario, the voice signal of a remote participant is transmitted over the network with a 30ms time delay. After the local speaker plays the signal, it is quickly picked up by the microphone. The 30ms time delay is calculated through signal analysis, and then the remote reference signal is aligned with the 30ms delay to synchronize the aligned remote reference signal with the near-end signal picked up by the microphone in time.
[0037] S12. High-pass filtering is applied to the delayed-aligned far-end reference signal and near-end microphone signal respectively to remove low-frequency noise; Specifically, in step S12, the delayed-aligned far-end reference signal and the original near-end microphone signal are used as processing objects, and high-pass filtering is performed on both signals respectively. The core of high-pass filtering is to retain the effective audio components above a certain frequency in the signal and filter out low-frequency interference noise. This type of low-frequency noise is mostly low-frequency background noise in the environment, low-frequency current noise generated by equipment operation, etc., which will be superimposed on the effective speech signal and affect the accuracy of subsequent processing. The same high-pass filtering rule is used to process the two signals to ensure the consistency of signal filtering and avoid introducing new signal deviations due to differences in filtering methods.
[0038] For example, a high-pass filter with a specific cutoff frequency is used to filter the delayed and aligned far-end reference signal and near-end microphone signal respectively, filtering out interference components such as low-frequency humming of air conditioners and equipment current noise below 50Hz in the two signals, and retaining only the effective high-frequency audio components related to speech.
[0039] S13. Perform amplitude normalization processing on the high-pass filtered far-end reference signal and near-end microphone signal respectively to obtain the pre-processed far-end signal and the pre-processed near-end signal. Specifically, in step S13, amplitude normalization is performed on both the high-pass filtered far-end reference signal and the near-end microphone signal. The core of amplitude normalization is to adjust the amplitude values of the two signals to a uniform and reasonable range, eliminating amplitude differences caused by factors such as device acquisition, signal transmission, and playback distance. For example, the far-end signal may experience amplitude attenuation during network transmission, while the near-end signal may fluctuate due to varying microphone acquisition distances. Irregular amplitude variations can lead to calculation errors in subsequent dual-talk detection and filter weight updates. By normalizing the two signals separately, their amplitudes are brought to the same baseline, ultimately yielding the pre-processed far-end signal and the pre-processed near-end signal.
[0040] For example, the amplitudes of the high-pass filtered far-end reference signal and the near-end microphone signal are normalized to the standard numerical range of [-1, 1]. The amplitude values of the signals that exceed this range are reasonably compressed to prevent audio distortion caused by signal amplitude overflow, while keeping the amplitudes of the two signals at a unified reference.
[0041] This embodiment solves the timing asynchrony problem between far-end and near-end signals by calculating the time delay and aligning the delay with the far-end reference signal. Then, high-pass filtering is used to remove low-frequency noise and amplitude normalization is used to standardize the signal amplitude. The original signal is calibrated and optimized from three dimensions: time, noise, and amplitude. This provides a precise and unified signal foundation for all subsequent steps such as dual-talk detection and echo estimation, avoiding subsequent processing deviations caused by inconsistent signal characteristics from the source and improving the accuracy of subsequent signal processing.
[0042] Optionally, in some embodiments, step S2, "performing dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, outputting the current dual-talk status detection result, and dynamically adjusting the judgment threshold used for dual-talk status detection based on feedback information," may specifically include: S21. Calculate the acoustic features of the preprocessed far-end signal and the preprocessed near-end signal respectively; the acoustic features include at least one of signal power ratio, spectral similarity and temporal envelope correlation; Specifically, in step S21, the far-end signal and near-end signal after synchronization preprocessing are taken as the processing objects, and their corresponding acoustic features are calculated respectively. The selected acoustic features are at least one of signal power ratio, spectral similarity, and temporal envelope correlation. Each type of acoustic feature is analyzed around different dimensions of the audio signal, providing multi-dimensional or targeted judgment criteria for determining the two-way communication status. Among them, the signal power ratio is used to reflect the power value comparison between the far-end signal and the near-end signal, the spectral similarity is used to analyze the degree of feature similarity between the two in the frequency domain, and the temporal envelope correlation is used to measure the degree of correlation between the two in the temporal signal envelope. Depending on the actual audio communication scenario requirements, a single feature or a combination of multiple features can be selected as the judgment basis for two-way communication detection.
[0043] For example, in the audio processing scenario of video conferencing, based on the pre-processed voice signal of the remote participant and the signal collected by the local microphone, two acoustic features, namely the signal power ratio and spectral similarity, are calculated simultaneously to provide a basis for judgment of dual-talk status detection from the two dimensions of power comparison and frequency domain similarity.
[0044] S22. Compare the calculated acoustic features with the current judgment threshold, determine whether the current frame is in dual-talk state based on the comparison result, and output the dual-talk state detection result. Specifically, in step S22, the calculated acoustic features are compared one by one or comprehensively with the currently set dual-talk determination threshold. Based on the comparison results, the dual-talk state of the current audio signal frame is determined. During audio signal processing, continuous signals are processed in frames with a fixed frame length. The determination in this step is based on the signal frame. Through the quantitative comparison of acoustic features and thresholds, it is determined whether the current frame is in a dual-talk state where local and remote speakers are speaking simultaneously. After the determination is completed, the corresponding dual-talk state detection result is output immediately, providing a real-time state reference for subsequent filter updates and nonlinear compensation strategy selection.
[0045] For example, the calculated signal power ratio is compared with the current power ratio judgment threshold of 0.5, and the spectral similarity is compared with the current spectral similarity judgment threshold of 0.8. If the signal power ratio is greater than 0.5 and the spectral similarity is less than 0.8, the current signal frame is determined to be in a dual-talk state, and the dual-talk state detection result is output to the subsequent processing stage; if the condition is not met, it is determined to be in a non-dual-talk state and the corresponding result is output.
[0046] S23. Based on feedback information generated by environmental noise estimation or mutual inductance nonlinear compensation, dynamically update the judgment threshold used for dual-talk determination; Specifically, in step S23, the threshold used for determining the dual-talk state is dynamically updated and adjusted based on two types of feedback information. The first type is information obtained from environmental noise estimation, and the second type is feedback information generated by the mutual inductance nonlinear compensation stage. Through this dynamic adjustment method, the dual-talk determination threshold is no longer fixed but can adapt to the actual environmental changes in audio communication and the actual state of subsequent signal processing. The dynamic update of the threshold will be adjusted in a targeted manner according to the characteristics of the feedback information. For example, the threshold is increased when the environmental noise increases, and the threshold is decreased when the nonlinear compensation feedback shows insufficient detection sensitivity, so that the determination threshold always matches the actual processing requirements.
[0047] For example, if environmental noise estimation reveals a significant increase in background noise power in the current communication environment, the signal power ratio threshold is raised from 0.5 to 0.6 to reduce misjudgment of dual-talk in high-noise environments. If feedback information from the nonlinear compensation stage is received, indicating a significant reduction in echo signal power after compensation, the dual-talk detection sensitivity needs to be improved, so the signal power ratio threshold is lowered from 0.5 to 0.4.
[0048] This embodiment extracts at least one acoustic feature from signal power ratio, spectral similarity, and temporal envelope correlation as the basis for dual-talk determination. After quantitative comparison, it outputs accurate detection results. Then, based on the feedback information of environmental noise estimation and nonlinear compensation, it dynamically adjusts the determination threshold to improve the accuracy, adaptability, and robustness of dual-talk status detection and reduce false positives and false negatives.
[0049] Optionally, in some embodiments, step S3, "selectively updating the adaptive filter based on the dual-talk state detection result, and using the adaptive filter to filter the preprocessed far-end signal to output an estimated echo signal," may specifically include: S31. Input the preprocessed far-end signal into the adaptive filter to calculate the estimated echo signal; Specifically, in step S31, the far-end signal after synchronization preprocessing is used as input and imported into an adaptive filter for filtering. The core function of the adaptive filter is to simulate the actual acoustic echo path in an audio communication scenario. By performing path fitting and filtering calculations on the input far-end signal, it reconstructs the echo characteristics formed after the far-end signal is played by a speaker, propagates through the air, or is reflected by the device. Finally, it calculates and outputs an estimated echo signal that highly matches the characteristics of the actual echo signal. This signal is the core reference for subsequent echo cancellation, and its accuracy directly affects the effect of subsequent echo cancellation.
[0050] For example, in a vehicle-mounted voice call scenario, the pre-processed voice signal of the remote caller is input into an adaptive filter. The filter simulates the acoustic echo path from the vehicle speaker to the vehicle microphone, and the output is an estimated echo signal that is consistent with the actual echo characteristics inside the vehicle through filtering calculation.
[0051] S32. Subtract the preprocessed near-end signal from the estimated echo signal to obtain the error signal; Specifically, in step S32, the preprocessed near-end signal and the estimated echo signal are processed, and a subtraction operation is performed between them. The preprocessed near-end signal is a mixed signal containing local speech, actual echo, and environmental noise. By subtracting the estimated echo signal, the echo component in the near-end signal can be initially removed. The error signal obtained after the operation reflects the remaining signal components of the near-end signal after removing the estimated echo. This signal is not only an important basis for subsequently judging the accuracy of echo estimation, but also the core data support for iteratively updating the weights of the adaptive filter.
[0052] For example, in a video conferencing scenario, the preprocessed local microphone mixed signal is subtracted from the estimated echo signal output by the adaptive filter to obtain an error signal containing the local participant's voice, a small amount of environmental noise, and a small amount of residual echo, which is then used for subsequent filter weight updates and adjustments.
[0053] S33. Based on the error signal and the preprocessed remote signal, update the weights of the adaptive filter according to the normalized least mean square algorithm; wherein, when the dual-talk state detection result indicates dual-talk state, pause or slow down the update of the adaptive filter weights. Specifically, in step S33, using the error signal and the preprocessed far-end signal as core data, the normalized least mean square algorithm is used to iteratively update the weights of the adaptive filter. This algorithm continuously adjusts and optimizes the filter weights, ensuring that the estimated echo signal output by the filter increasingly matches the actual echo path characteristics, thus improving the accuracy and convergence of echo estimation. Simultaneously, this step performs selective weight update operations based on the dual-talk state detection results. When the detection result indicates a dual-talk state, the update of the filter weights is paused or slowed down to avoid interference from the local speech signal causing the filter weights to deviate from the actual echo path, thus preventing distortion of the estimated echo signal.
[0054] For example, in a voice call using a smart speaker, in a non-dual-talk state, the filter weights are updated iteratively using the normalized least mean square algorithm based on the error signal and the preprocessed remote signal, so that the estimated echo signal continuously adapts to the changes in the echo path from the speaker to the microphone. When dual-talk state is detected (the user is speaking into the speaker while there is voice input from a remote location), the filter weight update is paused to avoid interference from the user's local voice with the weight adjustment, thus ensuring the accuracy of the estimated echo signal.
[0055] This embodiment generates an estimated echo signal through an adaptive filter. An error signal is obtained by subtracting the far-end and near-end signals. The filter weights are then iteratively updated based on the error signal and the far-end signal using a normalized least mean square algorithm. The weight update is paused or slowed down in dual-talk mode to effectively avoid interference from dual-talk on the filter. This allows the filter to dynamically adapt to changes in the actual echo path, improving the accuracy of echo estimation and algorithm convergence, and ensuring a high degree of matching between the output estimated echo signal and the actual echo characteristics.
[0056] Optionally, in some embodiments, step S4, "receiving the dual-talk state detection result and the estimated echo signal, selecting the corresponding nonlinear compensation strategy according to the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, outputting the compensated echo signal, and generating feedback information for adjusting the judgment threshold," may specifically include: S41. Evaluate the distortion of the estimated echo signal and generate the distortion evaluation result; Specifically, in step S41, the estimated echo signal is the object of this processing. This signal is the echo reference signal obtained after simulating the acoustic echo path. During its formation and transmission, it is prone to nonlinear distortion due to factors such as playback device distortion and acoustic propagation loss. This type of distortion affects the accuracy of subsequent echo cancellation. The core of this step is to comprehensively analyze and quantitatively evaluate the degree of nonlinear distortion of the estimated echo signal. By using professional signal analysis methods, key information such as the type and value of distortion are determined, and finally, a corresponding distortion evaluation result is generated. This result is the core quantitative basis for the subsequent selection of the nonlinear compensation function, ensuring that the selection of the compensation strategy closely matches the actual distortion state of the estimated echo signal and avoiding signal processing deviations caused by blind compensation.
[0057] For example, in the audio processing scenario of video conferencing, nonlinear distortion analysis is performed on the estimated echo signal output by the adaptive filter. The total harmonic distortion value is calculated by extracting the harmonic characteristics of the signal. The estimated echo signal is determined to be of different distortion levels, namely low, medium, and high, based on the magnitude of the value. The result of the level determination and the specific distortion value are used as the distortion evaluation result.
[0058] S42. Based on the distortion assessment results and the dual-talk state detection results, select the corresponding nonlinear compensation function, and apply the selected nonlinear compensation function to the estimated echo signal to generate an intermediate compensation signal; Specifically, in step S42, two sets of data are processed simultaneously: the distortion assessment result and the dual-talk status detection result output from the dual-talk status detection stage. Different distortion assessment results correspond to different compensation requirements, and different dual-talk statuses also have differentiated adaptation requirements for compensation intensity and compensation method. This step selects a matching nonlinear compensation function from a preset nonlinear compensation function library based on the combined characteristics of these two results. After selecting the function, it is applied to the estimated echo signal, and the nonlinear distortion in the signal is specifically eliminated and compensated through function operations. After processing, an intermediate compensated signal is obtained, which is the core transition signal after nonlinear compensation is completed.
[0059] For example, if the distortion assessment result shows that the estimated echo signal has high nonlinear distortion and the dual-talk state detection result is not dual-talk state, a strong compensation nonlinear compensation function of the soft-limiting type is selected; if the distortion assessment result is low distortion and the detection result is dual-talk state, a lightweight nonlinear compensation function is selected, and the selected function is applied to the corresponding estimated echo signal to complete the distortion compensation and generate an intermediate compensation signal.
[0060] S43. Based on the intermediate compensation signal, generate feedback information for adjusting the judgment threshold, and output the intermediate compensation signal as the compensated echo signal; Specifically, for step S43, two parallel processing tasks are carried out with the intermediate compensation signal as the core. First, the signal characteristics and compensation effect of the intermediate compensation signal are extracted and analyzed. Based on this information, corresponding feedback information is generated. This feedback information is transmitted to the dual-talk status detection stage as an important basis for its dynamic adjustment of the judgment threshold, realizing information interaction between this stage and the dual-talk status detection stage. Second, the intermediate compensation signal after nonlinear compensation processing is directly output as the compensated echo signal. This signal is the final echo reference signal after completing nonlinear distortion compensation, which will provide a precise processing target for the subsequent echo cancellation core stage.
[0061] For example, the power and distortion compensation rate of the intermediate compensation signal are extracted, and these quantified data are integrated into feedback information and transmitted to the dual-talk state detection stage. At the same time, the intermediate compensation signal is directly output as the compensated echo signal, which is then used to perform subtraction with the near-end signal to eliminate the echo.
[0062] This embodiment quantitatively evaluates the distortion of the estimated echo signal, and selects a suitable nonlinear compensation function by combining the distortion evaluation results and the dual-talk state detection results. This achieves targeted elimination of nonlinear distortion in the estimated echo signal. At the same time, feedback information is generated based on the compensated intermediate signal to feed back into the threshold adjustment of dual-talk detection, establishing a mutual feedback channel with dual-talk detection, thereby improving the adaptability and effectiveness of nonlinear compensation processing.
[0063] Optionally, in some embodiments, the method further includes: When the dual-talk status detection result indicates dual-talk status, the compensation strength of the selected nonlinear compensation function is lower than that of the nonlinear compensation function selected when the dual-talk status is not in dual-talk status.
[0064] Specifically, the system receives the detection results from the dual-talk status detection stage, identifies these results in real time, and clarifies whether the current audio communication call status is dual-talk or non-dual-talk. This identification operation is a crucial prerequisite for subsequently selecting nonlinear compensation functions of different intensities. Only by accurately determining the call status can the compensation intensity be adjusted accordingly, ensuring that the nonlinear compensation processing matches the actual call scenario.
[0065] Based on the determined call state, a suitable nonlinear compensation function is selected from a preset nonlinear compensation function library. When the call is determined to be in a two-way talk state, the compensation intensity of the selected nonlinear compensation function must be lower than that of the nonlinear compensation function selected when the call is determined to be in a non-two-way talk state. The compensation intensity reflects the correction magnitude of the nonlinear distortion of the estimated echo signal. A lower compensation intensity means a smaller correction magnitude, while a higher compensation intensity can more fully correct the nonlinear distortion. This differentiated setting is to adapt to the processing needs of different call states.
[0066] The nonlinear compensation function, selected based on the call state and corresponding compensation intensity, is directly applied to the estimated echo signal. Nonlinear compensation processing is performed on the signal according to the function's operational rules. The processing strictly adheres to the principle of low intensity for dual-talk and conventional intensity for non-dual-talk calls, without altering the core compensation logic of the function. Only the compensation intensity is adjusted differentially to ensure the nonlinear compensation process fits the current call state. After compensation, a corresponding intermediate compensation signal is generated, providing the foundation for subsequent output of the compensated echo signal and the generation of feedback information.
[0067] This embodiment achieves differentiated adaptation of compensation intensity by selecting a nonlinear compensation function with lower compensation intensity in dual-talk mode. This avoids interference and distortion of local speech signals caused by overcompensation in dual-talk mode, ensuring the integrity of local speech. In non-dual-talk mode, it can fully eliminate the nonlinear distortion of the estimated echo signal through normal compensation intensity, reducing problems such as speech distortion and incomplete echo cancellation caused by improper compensation intensity.
[0068] Optionally, in some embodiments, step S5, "subtracting the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, followed by post-processing to output the final speech signal," may specifically include: S51. Subtract the compensated echo signal from the preprocessed near-end signal to obtain the preliminary echo cancellation signal; Specifically, in step S51, a subtraction operation is performed on the near-end signal after synchronous preprocessing and the compensated echo signal after nonlinear compensation processing, subtracting the compensated echo signal from the preprocessed near-end signal. The preprocessed near-end signal contains components such as local speech, echo, and environmental noise. The compensated echo signal is a signal that highly matches the actual echo characteristics and has completed nonlinear distortion compensation. Through this subtraction operation, the main echo components in the preprocessed near-end signal can be directly eliminated, initially achieving the core goal of echo cancellation. The preliminary echo-cancelled signal obtained after the operation becomes the basic processing object for subsequent post-processing stages.
[0069] For example, in a VoIP voice call scenario, the local microphone mixed signal that has completed synchronous preprocessing is subtracted from the echo signal after nonlinear compensation, and most of the far-end echo components in the mixed signal are removed to obtain a preliminary echo cancellation signal containing local speech, a small amount of environmental noise and a trace of residual echo.
[0070] S52. Detect residual echo in the preliminary echo cancellation signal. If residual echo exists, perform post-filtering processing on the preliminary echo cancellation signal. Specifically, in step S52, taking the preliminary echo cancellation signal as the processing object, residual echo detection is first performed on it. Signal feature analysis is used to determine whether any residual echo components that have not been removed still exist in the signal. If the detection result indicates the presence of residual echo, post-filtering is immediately performed on the preliminary echo cancellation signal. Filtering methods are used to specifically suppress and eliminate residual echo, further removing echo interference from the signal and improving signal purity. If the detection result indicates no residual echo, the process proceeds directly to the next processing step without post-filtering.
[0071] For example, in audio processing for video conferencing, by calculating the correlation and residual power ratio between the preliminary echo cancellation signal and the far-end reference signal, and detecting the presence of residual echo in the preliminary echo cancellation signal, an adapted post-filtering algorithm is used to process the signal, suppressing and eliminating the residual echo components, resulting in a signal without obvious echo.
[0072] S53. For the preliminary echo cancellation signal after residual echo suppression processing, inject comfortable noise into the silent segment to generate the final speech signal; Specifically, in step S53, the preliminary echo-cancelled signal after residual echo suppression is first subjected to speech activity detection to accurately identify silent segments in the signal, i.e., signal segments without effective speech input. For the identified silent segments, low-level, comfortable noise that fits the environmental characteristics is injected. The intensity and characteristics of this noise are based on the principle of not causing auditory interference and filling the gaps in the silent segments, avoiding the auditory abruptness and unnaturalness caused by directly outputting silent segments. After the noise injection is completed, the final output speech signal is generated.
[0073] For example, in the context of in-vehicle voice calls, the signal with residual echo suppression is subjected to frame-by-frame detection to identify silent frames without driver voice. Low-level white noise is injected into these silent frames as comfort noise, ultimately generating a clear in-vehicle voice call signal with a natural auditory experience.
[0074] This embodiment achieves basic echo cancellation through subtraction, then performs residual echo detection on the initial echo-cancelled signal and performs post-filtering as needed to remove residual echo interference to the greatest extent. Comfortable noise is then injected into the silent segments of the signal to fill the auditory gaps. The audio signal is optimized layer by layer, ensuring the clarity of the final output speech signal while solving the problem of abruptness in the silent segments, thus balancing the transmission quality and auditory naturalness of the speech signal.
[0075] Optionally, in some embodiments, the adaptive filter in this embodiment is a finite impulse response filter, wherein the order of the finite impulse response filter is dynamically determined based on the time delay calculated in the signal synchronization preprocessing.
[0076] Specifically, the adaptive filter used in the echo path simulation in this embodiment is a finite impulse response filter. This type of filter has the characteristics of linear phase and strong stability, and can accurately match the signal propagation characteristics of the acoustic echo path in the audio communication system. It meets the needs of real-time audio signal processing and is the core carrier for realizing echo path simulation and calculating and estimating echo signals. After selecting this filter, it is used as the core device for subsequent filtering of the preprocessed far-end signal.
[0077] From the signal synchronization preprocessing stage, the calculated time delay data between the far-end reference signal and the near-end microphone signal is extracted. This time delay data is the core calculation result of the synchronization preprocessing stage, directly reflecting the timing deviation between the far-end and near-end signals. It is also a key timing characteristic of the actual acoustic echo path, directly determining the propagation span of the echo signal in the time dimension. Therefore, this data is used as the sole core basis for dynamically determining the filter order, ensuring that the filter order configuration matches the timing characteristics of the actual echo path. Based on the extracted time delay data and combined with the signal propagation laws of the acoustic echo path, the order of the finite impulse response filter is dynamically determined, rather than using a fixed order configuration. The order is positively correlated with the time delay; the larger the time span of the echo path reflected by the time delay, the higher the corresponding filter order. This ensures that the number of filter taps can completely cover the actual echo path, accurately simulating the echo signal formation process. After completing the order calculation, the finite impulse response filter is immediately configured. The configured filter can then be used for subsequent filtering of the preprocessed far-end signal.
[0078] This embodiment selects a finite impulse response filter as an adaptive filter and dynamically determines its order based on the time delay obtained from the signal synchronization preprocessing stage. This allows the filter order to accurately match the timing characteristics of the actual acoustic echo path, avoiding both increased computational overhead and processing delay caused by an excessively high order, and simulation distortion caused by an excessively low order that cannot fully cover the echo path. This approach controls computational costs while ensuring the accuracy of echo estimation, thereby improving the filter's adaptability, robustness, and processing efficiency.
[0079] To facilitate understanding of the adaptive echo cancellation method based on mutual inductance mechanism provided in this embodiment, a specific implementation method is also provided. The core of this method lies in establishing a mutual inductance mechanism between dual-talk detection (DTD) and nonlinear compensation (NLP). Specifically, this mutual inductance mechanism feeds back the DTD results to the adjustment strategy of nonlinear compensation (e.g., reducing the compensation intensity in dual-talk mode to avoid overcompensation), and feeds back the nonlinear compensation output to the optimization of DTD (e.g., adjusting the detection threshold using the compensated signal power), thereby achieving dynamic coupling between the two and improving the robustness and adaptability of the overall system. The specific steps are as follows: Step 1: Signal Acquisition and Dual-Domain Preprocessing Acquire the far-end signal x(n) (the signal played by the speaker) and the near-end signal d(n) (the signal captured by the microphone, including local speech s(n), echo y(n), and noise v(n)) from the audio input device.
[0080] To ensure time synchronization between the far-end signal x(n) and the near-end signal d(n), a delay estimation method is employed: The cross-correlation function R_{xd}(τ) = sum_{n} x(n) * d(n+τ) of x(n) and d(n) is calculated. The delay τ_max corresponding to the correlation peak is found, and x(n) is aligned with the corresponding delay using a buffer. For example, a FIFO buffer queue is used to store x(n), and samples are output with a delay of τ_max. If the delay exceeds a preset upper limit (e.g., 100ms), path reestimation is triggered.
[0081] Signal preprocessing: A high-pass filter is applied to remove low-frequency noise. Specifically, a 4th-order Butterworth high-pass filter is used with a cutoff frequency of 50Hz and a transfer function H(z) = (b0 + b1 z^{-1} + ... + b4 z^{-4}) / (1+ a1 z^{-1} + ... + a4 z^{-4}), where the coefficients are calculated using standard design tools (such as the MATLAB butter function) to ensure phase linearity.
[0082] Signal amplitude is normalized to prevent overflow. Specifically, the frame-by-frame normalization process involves dividing the signal into frames of length N (e.g., 512 samples), calculating the maximum amplitude for each frame as max_frame = max(|x_frame|), and then x_norm_frame(n) = x_frame(n) / max_frame; similarly, d(n) is processed. Global normalization is an alternative: x_norm(n) = x(n) / max(|x|).
[0083] Step 2: Dynamic Dual-Talk Detection Based on Nonlinear Mapping Feedback Two-way communication is detected to avoid incorrect filter updates when local speech is active. A power ratio method is used: calculate the far-end signal power P_x = sum(x^2(n)) / N and the near-end signal power P_d = sum(d^2(n)) / N (where N is the frame length, e.g., 512 samples).
[0084] If P_d / P_x > threshold (e.g., 0.5), it is determined to be a two-way signal, and filter updates are paused; otherwise, updates continue. This step prevents filter coefficients from deviating during two-way signaling. Besides the power ratio, other features can be detected, such as spectral features (calculating the FFT power spectrum similarity of x(n) and d(n), using cosine similarity cos_sim = (sum S_x * S_d) / (||S_x|| * ||S_d||), if the similarity > 0.8, it may not be a two-way signal); or time-domain features (such as envelope correlation, using Hilbert transform to extract the signal envelope and calculating the Pearson correlation coefficient, if > 0.7, it is determined to be echo-dominated). These features can be fused with the power ratio using weighted voting (weights such as 0.6 power ratio + 0.2 spectrum + 0.2 time domain).
[0085] The triggering condition for the mutual inductance mechanism is as follows: when dual-talk detection determines dual-talk for three consecutive frames, feedback to nonlinear compensation is triggered. The threshold is initially set to 0.5, and is dynamically adjusted if the ambient noise changes. The dynamic threshold adjustment algorithm estimates the ambient noise power P_v = min(P_d over silent frames), then sets the threshold to 0.5 + α * (P_v / P_x), where α = 0.2 is an adjustment factor to ensure that the threshold is increased in high-noise environments to reduce false positives. Noise estimation uses a VAD (Voice Activity Detection) algorithm, such as a simple VAD based on an energy threshold.
[0086] Step 3: Echo path estimation using NLMS filters Initialize the M-order FIR filter weights w(0) = [0, 0, ..., 0] (M is typically 128-512). Specifically, M is determined based on the echo delay: initially M = 256. The peak position τ_peak is found using the cross-correlation function R_{xd}(τ) by real-time detection of the echo path length, and M is set = τ_peak + margin (margin = 64 to cover tail attenuation). Real-time detection: R_{xd} is recalculated and M is adjusted every 10 seconds or when a path change is detected (e.g., when the error e(n) power suddenly increases > 20%).
[0087] For each sample n, calculate the filter output (estimated echo): y(n) = sum_{k=0}^{M-1} w(k) *x(nk).
[0088] Calculate the error signal: e(n) = d(n) - y(n).
[0089] The filter weights are updated using the NLMS algorithm: w(n+1) = w(n) + (μ / (sum(x^2(nk)) + ε))* e(n) * x(nk), where μ is the step size (0 < μ < 2, usually 0.1), and ε is a small positive number (to prevent division by zero, such as 1e-6). NLMS improves convergence speed and stability by normalizing the step size.
[0090] Step 4: Nonlinear Compensation Handling nonlinear echoes: A nonlinear function is applied to compensate for speaker distortion. Soft clipping is used by default: y_comp(n) = tanh(y(n)) * max_y, where max_y is the maximum amplitude of the speaker. Distortion is calculated using Total Harmonic Distortion (THD) = sqrt(sum_{k=2}^∞ |A_k|^2) / |A_1|, where A_k is the harmonic amplitude extracted by FFT. Compensation is triggered if THD > 5%.
[0091] Triggering conditions for different compensation functions: If the distortion is mainly saturated (high THD and significant high-frequency harmonics), use tanh; if it is asymmetric distortion (odd harmonics are detected as dominant), switch to polynomial compensation y_comp(n) = y(n) + βy^2(n) + γy^3(n) (β, γ are estimated by least squares fitting); if the distortion is low (THD < 2%), skip compensation to save computation.
[0092] Feedback strategy for nonlinear compensation based on dual-talk detection results: In dual-talk mode, reduce the compensation intensity (e.g., reduce the tanh scaling factor to 0.5) to avoid compensating for local speech; Optimization scheme for dual-talk detection based on nonlinear compensation output: Recalculate P_d = sum((d(n) - y_comp(n))^2 / N) using the compensated signal y_comp(n), and adjust the threshold accordingly (if P_d decreases by >10% after compensation, reduce the threshold to improve detection sensitivity).
[0093] Step 5: Echo Cancellation and Post-processing Output the echo-canceled signal: s_est(n) = e(n) (the error signal is the estimated local speech).
[0094] Residual echo detection algorithm: Calculate the residual power ratio P_res = sum(e^2(n)) / sum(d^2(n)). If P_res>0.1, then residual echo is determined to exist; or use correlation detection corr_res = max(R_{xe}(τ)). If >0.2, then trigger additional post-filtering (such as Wiener filter H_w(ω) = (P_s(ω)) / (P_s(ω) + P_v(ω)) applied to e(n)).
[0095] Apply Comfort Noise Generation (CNG) to fill residual noise gaps: If a silent frame is detected, inject low-level white noise to avoid unnatural silence.
[0096] Finally, output s_est(n) to the audio output.
[0097] The core of this embodiment is the introduction of the NLMS algorithm, which adapts the step size to the input signal power, improving robustness to noise and path changes. Simultaneously, the mutual inductance mechanism between DTD and NLP achieves closed-loop optimization through the aforementioned feedback strategy. The entire process is implemented in the time domain, making it suitable for real-time applications. The computational complexity is O(M) per sample, and it can run on a DSP chip.
[0098] Alternatively, a Frequency-Domain Adaptive Filter (FDAF) can be used to convert the signal to the frequency domain and process the block data using FFT. The update formula in the frequency domain is: W(k+1) = W(k) + μ *conj(X(k)) * E(k) / (P_X(k) + ε), where P_X is the power spectrum.
[0099] Alternatively, the Kalman filter method can be used to model the echo path as a state space and update the weights using Kalman gain.
[0100] The above completes the adaptive echo cancellation process based on mutual inductance mechanism in this application embodiment.
[0101] As described above, this application provides an adaptive echo cancellation method based on a mutual inductance mechanism. By constructing a mutual inductance mechanism between dual-talk detection and nonlinear compensation, the adaptability and robustness of the adaptive echo cancellation system can be improved, and echoes in audio communication systems can be eliminated accurately and efficiently, ensuring the transmission quality of voice signals and thus optimizing the call experience of audio communication.
[0102] To facilitate better implementation of the adaptive echo cancellation method based on mutual inductance mechanism in this application embodiment, this application embodiment also provides an adaptive echo cancellation device based on mutual inductance mechanism. The meanings of the terms used are the same as in the aforementioned adaptive echo cancellation method based on mutual inductance mechanism, and specific implementation details can be found in the description of the method embodiment.
[0103] Please see Figure 2 This application provides an adaptive echo cancellation device based on a mutual inductance mechanism. Specifically, this adaptive echo cancellation device may include a preprocessing module 201, a dual-talk detection module 202, an echo estimation module 203, a compensation module 204, and a signal synthesis module 205, as detailed below: The preprocessing module 201 is used to acquire the far-end reference signal and the near-end microphone signal containing the echo, and to perform synchronous preprocessing on the far-end reference signal and the near-end microphone signal to obtain the preprocessed far-end signal and the preprocessed near-end signal. The dual-talk detection module 202 is used to perform dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, output the current dual-talk status detection result, and dynamically adjust the judgment threshold used for dual-talk status detection based on feedback information. The echo estimation module 203 is used to selectively update the adaptive filter based on the dual-talk status detection result, and use the adaptive filter to filter the preprocessed far-end signal to generate an estimated echo signal. The compensation module 204 is used to receive the dual-talk status detection result and the estimated echo signal, select the corresponding nonlinear compensation strategy according to the dual-talk status detection result to perform nonlinear compensation processing on the estimated echo signal, output the compensated echo signal, and generate feedback information for adjusting the judgment threshold. The signal synthesis module 205 is used to subtract the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
[0104] Optionally, in some embodiments, the preprocessing module 201 is specifically used for: Calculate the time delay between the far-end reference signal and the near-end microphone signal, and perform delay alignment on the far-end reference signal based on the time delay to obtain the delay-aligned far-end reference signal; High-pass filtering is applied to the delayed-aligned far-end reference signal and near-end microphone signal to remove low-frequency noise. The amplitude normalization process is performed on the high-pass filtered far-end reference signal and the near-end microphone signal to obtain the pre-processed far-end signal and the pre-processed near-end signal.
[0105] Optionally, in some embodiments, the dual-talk detection module 202 is specifically used for: Calculate the acoustic features of the preprocessed far-end signal and the preprocessed near-end signal respectively; the acoustic features include at least one of signal power ratio, spectral similarity and temporal envelope correlation; The calculated acoustic features are compared with the current judgment threshold. Based on the comparison result, it is determined whether the current frame is in dual-talk state, and the dual-talk state detection result is output. The decision threshold for dual-talk determination is dynamically updated based on feedback information generated by environmental noise estimation or mutual inductance nonlinear compensation.
[0106] Optionally, in some embodiments, the echo estimation module 203 is specifically used for: The preprocessed far-end signal is input into an adaptive filter to calculate the estimated echo signal. The error signal is obtained by subtracting the preprocessed near-end signal from the estimated echo signal. Based on the error signal and the preprocessed remote signal, the weights of the adaptive filter are updated according to the normalized least mean square algorithm; wherein, when the dual-talk state detection result indicates dual-talk state, the update of the adaptive filter weights is paused or slowed down.
[0107] Optionally, in some embodiments, the compensation module 204 is specifically used for: The distortion of the estimated echo signal is evaluated, and the distortion evaluation result is generated. Based on the distortion assessment results and the dual-talk status detection results, the corresponding nonlinear compensation function is selected, and the selected nonlinear compensation function is applied to the estimated echo signal to generate an intermediate compensation signal. Based on the intermediate compensation signal, feedback information is generated to adjust the judgment threshold, and the intermediate compensation signal is output as the compensated echo signal.
[0108] Optionally, in some embodiments, the compensation module 204 is specifically used for: When the dual-talk status detection result indicates dual-talk status, the compensation strength of the selected nonlinear compensation function is lower than that of the nonlinear compensation function selected when the dual-talk status is not in dual-talk status.
[0109] Optionally, in some embodiments, the signal synthesis module 205 is specifically used for: Subtract the compensated echo signal from the preprocessed near-end signal to obtain the preliminary echo cancellation signal; Detect residual echoes in the initial echo cancellation signal; if residual echoes are present, perform post-filtering processing on the initial echo cancellation signal. The initial echo cancellation signal, after residual echo suppression processing, is injected with comfortable noise in the silent section to generate the final speech signal.
[0110] Optionally, in some embodiments, the adaptive filter is a finite impulse response filter, and the order of the finite impulse response filter is dynamically determined based on the time delay calculated in the signal synchronization preprocessing.
[0111] As described above, this application provides an adaptive echo cancellation device based on a mutual inductance mechanism. By constructing a mutual inductance mechanism between dual-talk detection and nonlinear compensation, the adaptability and robustness of the adaptive echo cancellation system can be improved, and the echo in the audio communication system can be eliminated accurately and efficiently, ensuring the transmission quality of the voice signal and thus optimizing the call experience of audio communication.
[0112] In addition, this application also provides a computer device, such as Figure 3 The diagram illustrates the structure of the computer device involved in this application. Specifically, the computer device may include components such as a processor 301 with one or more processing cores, a memory 302 with one or more computer-readable storage media, a power supply 303, and an input unit 304. Those skilled in the art will understand that... Figure 3 The computer device structure shown does not constitute a limitation on the computer device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 301 is the control center of the computer device. It connects various parts of the computer device via various interfaces and lines, and performs various functions and processes data by running or executing software programs and / or modules stored in the memory 302, and by calling data stored in the memory 302, thereby providing overall monitoring of the computer device. Optionally, the processor 301 may include one or more processing cores; preferably, the processor 301 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 301.
[0113] The memory 302 can be used to store software programs and modules. The processor 301 executes various functional applications and data processing by running the software programs and modules stored in the memory 302. The memory 302 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 302 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 302 may also include a memory controller to provide the processor 301 with access to the memory 302.
[0114] The computer device also includes a power supply 303 that supplies power to various components. Preferably, the power supply 303 is logically connected to the processor 301 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 303 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components. The computer device may also include an input unit 304, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical, or trackball signal inputs related to user settings and function control.
[0115] Although not shown, the computer device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 301 in the computer device loads the executable files corresponding to the processes of one or more applications into the memory 302 according to the following instructions, and the processor 301 runs the applications stored in the memory 302 to realize various functions, as follows: The system acquires a far-end reference signal and a near-end microphone signal containing echo, and performs synchronous preprocessing on both signals to obtain preprocessed far-end and near-end signals. Based on these signals, it performs dual-talk state detection, outputs the current detection result, and dynamically adjusts the decision threshold used for detection based on feedback. According to the detection result, it selectively updates an adaptive filter and uses it to filter the preprocessed far-end signal, outputting an estimated echo signal. It receives the detection result and the estimated echo signal, selects a corresponding nonlinear compensation strategy based on the result, performs nonlinear compensation on the estimated echo signal, outputs the compensated echo signal, and generates feedback to adjust the decision threshold. Finally, it subtracts the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
[0116] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0117] This application embodiment improves the adaptability and robustness of the adaptive echo cancellation system by constructing a mutual inductance mechanism between dual-talk detection and nonlinear compensation, accurately and efficiently eliminating echoes in audio communication systems, ensuring the transmission quality of voice signals, and thus optimizing the call experience of audio communication.
[0118] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0119] To this end, this application provides a storage medium storing a plurality of instructions that can be loaded by a processor to execute steps in any of the adaptive echo cancellation methods based on mutual inductance mechanisms provided in this application. For example, the instructions can execute the following steps: The system acquires a far-end reference signal and a near-end microphone signal containing echo, and performs synchronous preprocessing on both signals to obtain preprocessed far-end and near-end signals. Based on these signals, it performs dual-talk state detection, outputs the current detection result, and dynamically adjusts the decision threshold used for detection based on feedback. According to the detection result, it selectively updates an adaptive filter and uses it to filter the preprocessed far-end signal, outputting an estimated echo signal. It receives the detection result and the estimated echo signal, selects a corresponding nonlinear compensation strategy based on the result, performs nonlinear compensation on the estimated echo signal, outputs the compensated echo signal, and generates feedback to adjust the decision threshold. Finally, it subtracts the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
[0120] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0121] The storage medium may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.
[0122] Since the instructions stored in the storage medium can execute the steps of any of the adaptive echo cancellation methods based on mutual inductance provided in this application, the beneficial effects that any of the adaptive echo cancellation methods based on mutual inductance provided in this application can achieve can be realized, as detailed in the preceding embodiments, and will not be repeated here.
[0123] According to one aspect of this application, a computer program product or computer program is provided, comprising computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various alternative implementations described above.
[0124] The above provides a detailed description of an adaptive echo cancellation method, apparatus, device, and medium based on mutual inductance mechanism provided in this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. An adaptive echo cancellation method based on mutual inductance mechanism, characterized in that, include: Acquire a far-end reference signal and a near-end microphone signal containing echo, and perform synchronous preprocessing on the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal; Dual-talk status detection is performed based on the preprocessed far-end signal and the preprocessed near-end signal, the current dual-talk status detection result is output, and the judgment threshold used for dual-talk status detection is dynamically adjusted based on feedback information. Based on the dual-talk state detection results, the adaptive filter is selectively updated, and the preprocessed far-end signal is filtered using the adaptive filter to output an estimated echo signal. The system receives the dual-talk state detection result and the estimated echo signal, selects a corresponding nonlinear compensation strategy based on the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, outputs the compensated echo signal, and generates feedback information for adjusting the judgment threshold. The compensated echo signal is subtracted from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
2. The adaptive echo cancellation method based on mutual inductance mechanism according to claim 1, characterized in that, The step of synchronously preprocessing the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal includes: Calculate the time delay between the far-end reference signal and the near-end microphone signal, and perform delay alignment on the far-end reference signal based on the time delay to obtain the delay-aligned far-end reference signal; High-pass filtering is applied to the delayed-aligned far-end reference signal and the near-end microphone signal respectively to remove low-frequency noise; The amplitude normalization process is performed on the high-pass filtered far-end reference signal and the near-end microphone signal to obtain the pre-processed far-end signal and the pre-processed near-end signal.
3. The adaptive echo cancellation method based on mutual inductance mechanism according to claim 1, characterized in that, The process of performing dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, outputting the current dual-talk status detection result, and dynamically adjusting the judgment threshold used for dual-talk status detection based on feedback information includes: The acoustic features corresponding to the preprocessed far-end signal and the preprocessed near-end signal are calculated respectively; the acoustic features include at least one of signal power ratio, spectral similarity and temporal envelope correlation; The calculated acoustic features are compared with the current judgment threshold. Based on the comparison result, it is determined whether the current frame is in dual-talk state, and the dual-talk state detection result is output. The decision threshold used for dual-talk determination is dynamically updated based on the feedback information generated by environmental noise estimation or mutual inductance nonlinear compensation.
4. The adaptive echo cancellation method based on mutual inductance mechanism according to claim 1, characterized in that, The step of selectively updating the adaptive filter based on the dual-talk state detection result, and using the adaptive filter to filter the preprocessed far-end signal to output an estimated echo signal includes: The preprocessed far-end signal is input into the adaptive filter to calculate the estimated echo signal; The preprocessed near-end signal and the estimated echo signal are subtracted to obtain the error signal; Based on the error signal and the preprocessed remote signal, the weights of the adaptive filter are updated according to the normalized least mean square algorithm; wherein, when the dual-talk state detection result indicates a dual-talk state, the update of the adaptive filter weights is paused or slowed down.
5. The adaptive echo cancellation method based on mutual inductance mechanism according to claim 1, characterized in that, The process involves receiving the dual-talk state detection result and the estimated echo signal, selecting a corresponding nonlinear compensation strategy based on the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, outputting the compensated echo signal, and generating feedback information for adjusting the judgment threshold, including: The estimated echo signal is subjected to distortion assessment, and distortion assessment results are generated; Based on the distortion assessment results and the dual-talk state detection results, a corresponding nonlinear compensation function is selected, and the selected nonlinear compensation function is applied to the estimated echo signal to generate an intermediate compensation signal. Based on the intermediate compensation signal, feedback information for adjusting the determination threshold is generated, and the intermediate compensation signal is output as the compensated echo signal.
6. The adaptive echo cancellation method based on mutual inductance mechanism according to claim 5, characterized in that, Also includes: When the dual-talk state detection result indicates a dual-talk state, the compensation strength of the selected nonlinear compensation function is lower than that of the nonlinear compensation function selected when the dual-talk state is not in a dual-talk state.
7. The adaptive echo cancellation method based on mutual inductance mechanism according to claim 1, characterized in that, The process of subtracting the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo-cancelled signal, followed by post-processing to output the final speech signal, includes: Subtracting the compensated echo signal from the preprocessed near-end signal yields the preliminary echo cancellation signal. Residual echoes are detected in the preliminary echo cancellation signal. If residual echoes are present, the preliminary echo cancellation signal is subjected to post-filtering processing. The initial echo cancellation signal, after residual echo suppression processing, is injected with comfortable noise in the silent section to generate the final speech signal.
8. An adaptive echo cancellation device based on mutual inductance mechanism, characterized in that, include: The preprocessing module is used to acquire a far-end reference signal and a near-end microphone signal containing echo, and to perform synchronous preprocessing on the far-end reference signal and the near-end microphone signal to obtain preprocessed far-end signal and preprocessed near-end signal. The dual-talk detection module is used to perform dual-talk status detection based on the preprocessed far-end signal and the preprocessed near-end signal, output the current dual-talk status detection result, and dynamically adjust the judgment threshold used for dual-talk status detection based on feedback information. The echo estimation module is used to selectively update the adaptive filter based on the dual-talk state detection result, and use the adaptive filter to filter the preprocessed far-end signal to generate an estimated echo signal. The compensation module is used to receive the dual-talk state detection result and the estimated echo signal, select the corresponding nonlinear compensation strategy according to the dual-talk state detection result to perform nonlinear compensation processing on the estimated echo signal, output the compensated echo signal, and generate feedback information for adjusting the judgment threshold. The signal synthesis module is used to subtract the compensated echo signal from the preprocessed near-end signal to obtain a preliminary echo cancellation signal, which is then post-processed to output the final speech signal.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the adaptive echo cancellation method based on mutual inductance mechanism as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the adaptive echo cancellation method based on mutual inductance mechanism as described in any one of claims 1 to 7.