Digital microphone echo cancellation method
By calculating the time difference and intensity difference to locate the speaking position, adjusting the microphone array beam direction, and combining spectrum analysis and echo path estimation models, the problem of echo cancellation difficulties in multi-person conference scenarios using traditional microphone arrays is solved, achieving high-definition audio output.
Patent Information
- Application Number
- CN202511482796.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-16
- Publication Date
- 2026-01-09
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Traditional microphone array systems struggle to adapt to frequent seating changes and simultaneous speaking by multiple people in dynamic multi-person conference scenarios, leading to difficulties in echo cancellation and affecting audio clarity.
By calculating the time difference and intensity difference of the sound signal received by the microphone, the speaker's position is located, the beam focusing direction of the microphone array is adjusted, and combined with spectral feature analysis and echo path estimation model, echoes are separated and eliminated in real time to generate clear audio signals.
It achieves precise audio separation and echo cancellation for multiple speakers in complex meeting scenarios, improving audio quality and transmission stability.
Smart Images

Figure CN121306161A_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of microphone technology, and particularly relates to a method for echo cancellation of digital microphones. Background Technology
[0002] With the increasing prevalence of remote and hybrid work models, high-quality audio acquisition and processing technology has become crucial for ensuring smooth meetings and improving communication efficiency. However, in real-world meeting scenarios, complex sound environments and dynamically changing participant behaviors place higher demands on microphone networks.
[0003] Traditional microphone array systems, when dealing with complex meeting scenarios, typically rely on fixed microphone layouts and preset signal processing algorithms, making them ill-suited for situations with frequent changes in participant seating or multiple people speaking simultaneously. Furthermore, they struggle to handle echoes when multiple speakers overlap, affecting sound clarity. The system's inability to effectively separate and eliminate echoes results in interfering or overlapping sounds in the meeting audio.
[0004] Therefore, how to accurately locate the speaker's position in real time based on the arrival time difference and intensity difference of multiple microphone signals in a dynamically changing meeting scenario, and accurately separate and eliminate echo components in a complex environment where multiple people are speaking simultaneously, has become a key issue in improving the sound quality of distributed microphone network conferences. Summary of the Invention
[0005] The purpose of this invention is to provide a digital microphone echo cancellation method that addresses the business problems of dynamic changes in speaking positions, sound overlap, and echo interference in multi-person conference scenarios. By fusing time difference, intensity difference, and spectral feature analysis, it achieves high-definition audio output.
[0006] This invention can be achieved through the following technical solutions: This application provides a digital microphone echo cancellation method, including the following steps: The system acquires sound signals from multiple microphones via a distributed microphone network. By calculating the time difference and intensity difference between these signals, a first speaking position is determined. The beam focusing direction of the microphone array is adjusted based on this first speaking position to determine the corresponding audio capture enhancement signal. Frequency domain analysis is performed on the audio capture enhancement signal to extract its spectral features. If these spectral features overlap, it is determined that multiple speakers are present, and the direct sound component and mixed echo component of each speaker are separated. The correlation between the temporal envelope of the mixed echo component and the temporal envelope of the direct sound component is analyzed. An echo path estimation model is constructed; the audio capture enhancement signal is input into the echo path estimation model, echo components are identified and removed, and the purified audio signal is output; the signal intensity change amplitude of the purified audio signal is monitored, and if the signal intensity change amplitude exceeds a preset intensity change threshold, the second speech position estimation data is confirmed, and the beam focusing direction is readjusted; based on the beam focusing direction of the second speech position estimation data, the subsequent incoming audio signals are processed in real time, and if the time difference changes, the intensity difference and temporal envelope information are fused from the real-time processed audio signal to obtain the final clear audio output signal; The final clear audio output signal is used for conference transmission in a distributed microphone network.
[0007] The beneficial effects of this invention are as follows: Firstly, this application estimates the speaking location using time difference and intensity difference, and dynamically adjusts the focusing direction of the microphone array using a delay-summation beamforming algorithm to enhance target audio capture. For overlapping multiple speakers, an independent component analysis algorithm is used to separate direct sound and echo components, and an echo path estimation model is constructed through temporal envelope similarity analysis. This is combined with an adaptive filtering algorithm to eliminate echoes and generate clean audio. When signal intensity changes exceed a threshold, the location estimation and beam direction are updated in real time, and echo cancellation parameters are optimized to ensure the accuracy of subsequent signal processing. This invention, by fusing intensity difference and temporal envelope information, outputs a clear audio signal, significantly improving the audio quality and transmission stability of distributed microphone networks in complex conference scenarios. Attached Figure Description
[0008] To better understand and implement this application, the technical solution is described in detail below with reference to the accompanying drawings.
[0009] Figure 1 This is a flowchart illustrating the steps of a digital microphone echo cancellation method provided in an embodiment of this application. Detailed Implementation
[0010] To further illustrate the technical means and effects adopted by the present invention to achieve its intended purpose, exemplary embodiments will be described in detail below, examples of which are illustrated in the accompanying drawings. In the following description relating to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of methods and systems consistent with some aspects of this application as detailed in the appended claims.
[0011] The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to be limiting of the application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to any and all possible combinations comprising one or more of the associated listed items.
[0012] The following detailed description of the specific implementation methods, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided in detail.
[0013] Please see Figure 1 This application provides a digital microphone echo cancellation method, including: Step S101: Obtain sound signals received by multiple microphones from the distributed microphone network, and obtain the first speaking position by calculating the time difference and intensity difference between each sound signal.
[0014] Multiple audio signals are acquired from a distributed microphone network. A preliminary time difference matrix is obtained by calculating the time difference between each audio signal, where the time difference is based on the difference in the timestamps of the signals arriving at each microphone. Based on this preliminary time difference matrix, intensity difference data is obtained by comparing the peak amplitudes of each audio signal. A coarse coordinate system of the sound source is determined using a triangulation algorithm, which estimates the source location based on known microphone positions and time differences. This estimation is performed by finding the geometric intersection of the distance between microphones and the time difference. Based on the coarse coordinates of the sound source, a cleaned signal is obtained from noise interference filtering. Noise interference filtering removes low-frequency interference using a high-pass filter. For this cleaned signal, intensity difference analysis is used to determine the final location. This analysis compares the relative intensity attenuation of the cleaned signal at each microphone to obtain the first speaking position.
[0015] Specifically, in one implementation, a distributed microphone network is deployed in multiple locations within the meeting room, such as on the walls and ceiling, to cover the sound acquisition needs of the entire space. This network consists of several microphones, each receiving sound signals in real time and transmitting them to a central processing unit via wired or wireless means.
[0016] It's important to note that the distributed microphone network is designed to ensure multi-angle capture of sound signals, thus providing a reliable data foundation for subsequent difference calculations. This arrangement allows for application in scenarios of varying meeting sizes, such as small discussion rooms or large lecture halls, demonstrating the technology's adaptability.
[0017] Specifically, the process of acquiring audio signals from multiple microphones in a distributed microphone network includes signal synchronization and initial filtering. First, the central processing unit synchronizes the audio signals from each microphone in time to avoid transmission delays affecting accuracy. Then, a low-pass filter is applied to the audio signals to remove high-frequency noise, ensuring the purity of the audio signal. This acquisition method is particularly effective in conference environments because it captures the speaker's voice without excessive interference from background noise such as air conditioning. Through these steps, the acquired audio signals have sufficient quality to be used to calculate time differences and intensity variations. Further, the time difference between the audio signals is calculated using a cross-correlation method. Specifically, a pair of microphone audio signals is selected, and their cross-correlation function is calculated to find the time delay corresponding to the peak. This time difference reflects the time difference between the sound source and the different microphones.
[0018] For example, in a microphone array, if sound arrives at microphone A first and then at microphone B, the time difference is positive. In a conference setting, this method helps to quickly locate the first speaker, and even when multiple people are speaking simultaneously, it can isolate the first sound source through signal separation technology.
[0019] In one possible implementation, the calculation of intensity difference focuses on comparing signal amplitudes.
[0020] For example, amplitude spectrum analysis is performed on the sound signals received by each microphone to obtain the average energy level of each sound signal. Then, the differences between these energy levels are compared; for instance, if the signal strength of microphone C is higher than that of microphone D, it indicates that the sound source is closer to microphone C. This intensity difference, combined with a geometric model, can further refine the location estimation.
[0021] Preferably, a distance attenuation model is incorporated into the calculation to account for the decrease in sound intensity with distance, thereby improving the accuracy of positioning. In actual meetings, this method effectively handles interference from echoes or reflections, ensuring the accurate acquisition of the first speaking position.
[0022] For example, in one embodiment, suppose the conference room is equipped with four microphones, located at the four corners of the room. When the first speaker speaks, all sound signals are first acquired, and the time difference between each pair of microphones is calculated. Then, these data are combined with the intensity differences using triangulation to derive the speaker's location coordinates. This embodiment demonstrates the application of this technology in a standard conference layout.
[0023] Understandably, after obtaining the first speaking position, the position information can be output to a display device or a subsequent processing module.
[0024] In one embodiment, if the meeting involves video tracking, the positioning results are used to adjust the camera orientation to focus on the speaker. This extended application maintains versatility within the conferencing domain without introducing other scenarios.
[0025] Furthermore, the fusion of time difference and intensity difference is one of the core aspects of this application. Specifically, a location probability model is constructed, using time difference as the primary constraint and intensity difference as an auxiliary weight, to determine the most probable location through iterative optimization. For example, in the calculation, a hyperbolic trajectory is first generated based on the time difference, and then the intersection points are filtered using intensity difference to pinpoint the first speaking location. This detailed process ensures the robustness of the localization and maintains high accuracy even in noisy environments.
[0026] Preferably, in another embodiment, the microphone network is expanded to eight units, covering a larger conference room. After the sound signal is acquired, multiple pairs of time differences are calculated to form a denser constraint network, while the intensity differences are quantized using logarithmic ratios to reduce nonlinear effects. This approach enhances the flexibility of the technology, making it suitable for sound transmission in different rooms.
[0027] It should be noted that the technical solution of this application enables rapid location of the first speaker, thereby providing real-time support in meeting management, such as automatically switching audio focus or recording the speaking order. These effects stem from the precise integration of difference calculation. In one implementation, the overall process begins with the acquisition of the audio signal and progresses step-by-step to difference calculation and position output, ensuring logical coherence.
[0028] Step S102: Adjust the beam focusing direction of the microphone array according to the first speaking position to determine the audio capture enhancement signal corresponding to the first speaking position.
[0029] The raw audio signal is acquired using a microphone array. A delay-summing algorithm is used to calculate the time delay difference between each microphone to determine the first speaking position. The beam focusing direction is adjusted based on this first speaking position, and the corresponding directional gain parameter is obtained from a preset gain table. The raw audio signal is then processed using the directional gain parameter, and a weighted summation method is used to fuse the signal amplitude to obtain the audio capture enhancement signal corresponding to the first speaking position.
[0030] Specifically, in one implementation, information about the first speaking position is first obtained. This position can be determined using sound source localization technology, such as calculating the speaker's location by using the time difference of the audio signals received by the microphone array.
[0031] Specifically, a microphone array consists of multiple microphone units distributed in space to form an array structure. By analyzing the phase difference of the audio signals received by each microphone, the direction of the sound source can be inferred, thus identifying the first speaking position. This positioning process depends on the array's geometric configuration, such as a linear or circular array. In a conference room environment, linear arrays are often used for desktop placement to cover the speaking area in front. Based on the determined first speaking position, the beam focusing direction of the microphone array is further adjusted. Beam focusing refers to using digital signal processing technology to concentrate the array's receiving sensitivity in a specific direction.
[0032] Specifically, a weighted sum of the signals from each microphone is calculated, and a delay compensation method is used to align the signals from the target direction. For example, in one possible implementation, an adaptive filter is used to adjust the weights so that the signal from the target direction is amplified while interference from other directions is suppressed. This adjustment process is performed in real time to ensure that the beam is always pointed to the first speaking position, thereby improving the clarity of the audio capture.
[0033] It should be noted that the adjusted audio capture enhancement signal corresponds to the first speaking position. This enhancement signal is a composite signal generated by a beamforming algorithm, which has a higher signal-to-noise ratio than the original signal.
[0034] Preferably, in a conference system, this enhanced signal can be directly input into a speech recognition module for further processing into text or commands.
[0035] For example, in a multi-person discussion scenario, when a speaker is detected on the left, the system focuses the beam on the left, generating an enhanced signal that reduces noise interference on the right, achieving more accurate audio capture. In another embodiment, considering dynamic scenarios such as video conferencing, the system uses a camera to assist in locating the first speaker. The image captured by the camera is used to verify the direction of the sound source, and then the beam focus is adjusted.
[0036] Specifically, if the speaking position changes, the system gradually updates the focus direction using a smooth transition algorithm to avoid signal interruption. This method is suitable for large conference halls where microphone arrays are mounted on the ceiling, covering the entire space. Through this enhancement, the system can maintain continuous audio quality, supporting clear hearing for remote participants. Furthermore, determining the enhanced signal may involve post-processing steps such as echo cancellation.
[0037] For example, during a meeting in a conference room, when a user speaks from one side of the room, the system adjusts the beam to focus on that position, generating an enhanced signal to remove background noise, such as television sound, thereby improving the accuracy of voice command recognition. This implementation demonstrates the versatility of the technical solution in the same audio capture field, adapting to different interior layouts without introducing additional complexity.
[0038] Understandably, the logic of this process forms a closed loop from location determination to signal enhancement, ensuring targeted audio capture. In practical applications, this method can optimize the signal for specific speakers and reduce overall system power consumption because it only enhances the signal in the necessary directions, rather than capturing it omnidirectionally.
[0039] Step S103: Perform frequency domain analysis on the audio capture enhancement signal, extract the spectral features in the audio capture enhancement signal. If the spectral features overlap, it is determined that multiple people are speaking, and the direct sound component and mixed echo component of each speaker are separated.
[0040] The audio capture enhancement signal is frequency-domain transformed using Fourier transform, which converts the time-domain signal into a frequency-domain signal to analyze frequency components, resulting in a frequency-domain representation signal. Spectral features, including amplitude and phase information, are extracted from the frequency-domain representation signal. If these spectral features overlap, it is determined that multiple speakers are present. Based on the determination of multiple speakers, a blind source separation algorithm is used to separate the direct sound components of each speaker. The input of the blind source separation algorithm is a mixed signal, and the output is an independent source signal. Based on the direct sound components, the mixed echo components are decomposed from the audio capture enhancement signal. This decomposition involves subtracting the direct sound to obtain the echo, resulting in the direct sound components and mixed echo components of each speaker.
[0041] Specifically, in one implementation, frequency domain analysis is performed on the audio capture enhancement signal by first converting the time-domain signal into a frequency-domain representation.
[0042] Specifically, the enhanced signal is decomposed into multiple frames using a short-time Fourier transform. A window function is applied to each frame to reduce spectral leakage, and then the amplitude and phase spectra are calculated. This transformation helps reveal the frequency components of the signal, providing a foundation for subsequent feature extraction. For example, when multiple participants speak simultaneously, this frequency domain analysis can capture the spectral distribution of the mixed audio, ensuring the analysis process adapts to different acoustic environments.
[0043] Furthermore, spectral features are extracted from the audio capture enhancement signal, including calculating Mel-frequency cepstral coefficients or logarithmic amplitude spectra and using them as feature vectors. These features capture the signal's energy distribution and harmonic structure. For example, Mel-frequency cepstral coefficients are derived by simulating human auditory perception, converting linear frequencies to a Mel scale, and then applying discrete cosine transform to obtain a low-dimensional representation. This extraction process is particularly useful in multi-person dialogue scenarios, highlighting the vocal characteristics of different speakers and avoiding interference from ambient noise.
[0044] For example, if spectral features overlap, it is determined that multiple people are speaking. Overlap detection is based on comparing the similarity of feature vectors or the overlap of spectral peaks. In practical applications, such as online conferencing systems, this method can quickly detect multiple simultaneous speakers when the frequency ranges of two speakers overlap, further triggering a separation mechanism.
[0045] Preferably, the preset threshold can be dynamically adjusted according to the ambient noise level to improve the robustness of the judgment.
[0046] In one possible implementation, once multiple speakers are identified, the direct sound component and the mixed echo component of each speaker are separated. The separation process first estimates the room impulse response to distinguish between direct and reflected sound, and then applies an independent component analysis algorithm to decompose the mixed signal into independent sources.
[0047] Specifically, the independent component analysis algorithm assumes that the signal sources are statistically independent and solves the separation matrix by maximizing the non-Gaussianity.
[0048] For example, a variant of fast independent component analysis can be used to process real-time audio. In virtual conference scenarios, this separation isolates the direct sound path of each speaker while extracting echoes as a mixed component, ensuring a clear output signal.
[0049] Understandably, the direct sound component refers to the sound wave that travels directly from the speaker to the microphone, while the mixed echo component includes multipath signals such as those reflected from walls. The detailed separation process involves frequency-domain blind source separation, which first converts the signal to the frequency domain and then applies a separation filter at each frequency.
[0050] For example, filter coefficients can be optimized by minimizing mutual information. This method can effectively reduce crosstalk when dealing with multiple speakers.
[0051] For example, in a three-part conference call, the direct audio from each participant can be used independently for speech recognition, while the echo component can be used for subsequent echo cancellation processing. This process emphasizes real-time performance and is typically executed on a digital signal processor to support low-latency applications.
[0052] Furthermore, in another embodiment, different microphone array configurations are considered to enhance separation accuracy. For example, a linear array is used to capture the signal, and then beamforming techniques are incorporated into the frequency domain analysis to initially suppress echoes from non-target directions, thereby aiding feature extraction and overlap detection. This variant is suitable for large conference hall scenarios where multiple speakers may be accompanied by more echo interference, optimizing separation performance through array geometry.
[0053] In one embodiment, the determination of overlapping spectral features can be combined with a machine learning model, such as a support vector machine classifier. During training, labeled multi-person and single-person audio datasets are used. The input is the extracted spectral feature vector, and the output is the probability of multiple people speaking. If the probability is higher than a threshold, separation is triggered. This integration improves the accuracy of the determination and is more robust in noisy environments.
[0054] For example, it can accurately distinguish overlapping features in an office meeting with background air conditioning noise.
[0055] Preferably, the separated direct sound component and mixed echo component can be further used for audio enhancement applications.
[0056] For example, the direct sound input is fed into a noise suppression module, while the echo component is used for adaptive filter updates to achieve overall signal purification. This application demonstrates the versatility of the technical solution presented in this application, covering the entire process from acquisition to output within the same audio processing domain.
[0057] Step S104: Analyze the correlation between the temporal envelope of the mixed echo component and the temporal envelope of the direct sound component, and construct an echo path estimation model.
[0058] By acquiring audio signals, time-domain data of the mixed echo components and direct sound components are obtained. Hilbert transform is used to calculate their respective envelopes, whereby the Hilbert transform obtains the amplitude of the analytic signal through convolution, resulting in the envelope curve. Based on the envelope curve, the Pearson correlation coefficient is calculated. The Pearson correlation coefficient is determined by dividing the covariance of the two envelope curves by the product of their respective standard deviations, thus determining the correlation strength between the mixed echo components and the direct sound components. If the correlation strength exceeds a preset threshold, delay features are extracted from the correlation strength, where the delay features are obtained by finding the time offset corresponding to the correlation peak. An initial echo path model is constructed, where the input of the initial echo path model is the delay features, and the output is a preliminary estimate of the path delay and attenuation parameters. By iteratively updating the parameters of the initial echo path model, including the delay and attenuation coefficients, and adjusting by minimizing the residuals, a final echo path estimation model is obtained.
[0059] Specifically, in one implementation, the audio signal is first preprocessed to extract the temporal envelope of the mixed echo component and the direct sound component.
[0060] Specifically, audio signals typically originate from a mixture of sounds captured by a microphone, including the speaker's direct sound and echoes reflected after being played through a speaker. Preprocessing steps include framing the signal and then applying window functions such as the Hamming window to reduce spectral leakage. This ensures signal smoothness for subsequent analysis.
[0061] It should be noted that the time-domain envelope refers to the slowly changing curve of the signal amplitude, which can be obtained by low-pass filtering after absolute value rectification, thereby highlighting the attenuation characteristics of the echo. This extraction process helps to distinguish between the sharp peaks of the direct sound and the gradual tails of the echo, providing basic data for correlation analysis.
[0062] Furthermore, the correlation between the temporal envelope of the mixed echo component and the temporal envelope of the direct sound component is analyzed. This analysis is a key step in model building, as the correlation reflects the temporal similarity between the two components.
[0063] Specifically, the temporal envelope of the direct sound component is first obtained from a reference signal, which can be the original audio played by the speaker. Then, a preliminary echo cancellation estimate is performed on the microphone signal to separate the mixed echo component. The correlation calculation process includes normalizing the two envelope sequences to ensure consistent amplitude ranges. Subsequently, the similarity is calculated using the Pearson correlation coefficient formula, which is obtained by dividing the covariance by the standard deviation; a value closer to 1 indicates a higher correlation.
[0064] For example, in a conference room scenario, if the direct sound envelope shows a rapidly rising peak, while the mixed echo envelope shows a similarly delayed shape, the correlation coefficient may be close to 1, indicating significant reflections along the echo path. This calculation process not only quantifies temporal similarity but also identifies the echo delay time, typically determined by locating the relevant peaks. By iteratively calculating the correlation across different frames, a statistical distribution can be obtained, further revealing the dynamic changes in the path. This detailed analysis helps the model capture the propagation patterns of echoes, effectively reducing residual echoes and improving speech clarity in practical applications such as video call systems.
[0065] Preferably, an echo path estimation model is constructed based on correlation analysis. This model uses the correlation results as input features and employs an adaptive filter structure to simulate the echo path.
[0066] Specifically, the echo path estimation model can be based on the Least Mean Squares (LMS) algorithm, whose core is to minimize the error signal by iteratively updating the filter coefficients. The process involves initializing the filter coefficients as a zero vector, and then calculating the predicted echo in each frame, which is the convolution of the reference signal and the current coefficients. The error signal is the microphone signal minus the predicted echo, and the coefficient update formula involves multiplying the step size parameter by the product of the error and the reference signal.
[0067] For example, in one embodiment, the step size is adaptively adjusted based on signal energy to avoid slow or unstable convergence. The key to this construction process is incorporating temporal envelope correlation as a weighting factor; for instance, if the correlation coefficient is above a threshold, the update step size is increased to accelerate convergence. In this way, the model can better adapt to changes in the room's acoustic environment.
[0068] In one possible implementation, multi-channel input can be introduced to enhance the robustness of the model.
[0069] Specifically, if the system is equipped with multiple microphones, the mixed echo envelope of each channel is extracted separately, and the correlation vector with the direct sound envelope is calculated. The model then fuses these vectors, generating a comprehensive path estimate using a weighted averaging method. This fusion process considers spatial correlations, such as using beamforming techniques to pre-amplify the signal in the direction of the direct sound, thereby improving the accuracy of the correlation calculation.
[0070] For example, in a smart speaker scenario in a conference room, a multi-microphone array can capture echoes from different directions, and the model estimates multipath propagation paths accordingly, ensuring effective cancellation even when the user moves. Furthermore, the model training process involves a supervised learning mechanism.
[0071] Specifically, simulated data with known echo paths are used as the training set, for example, by generating synthetic signals through convolutional room impulse responses. During training, the input is temporal envelope correlation features, and the output is estimated path coefficients. The process includes forward propagation to calculate the prediction error, followed by backward parameter updates. In practical deployments, this method can continuously optimize model parameters through online learning.
[0072] In one embodiment, estimation of nonlinear echo paths is considered. Traditional linear models may be insufficient to handle speaker distortion, therefore nonlinear functions such as polynomial approximations are introduced. Specifically, correlation analysis is extended to higher-order statistics, such as calculating the second moment of the cross-correlation function of the envelope. The model then uses a neural network architecture, such as a multilayer perceptron with hidden layers, taking the correlation features as input and outputting a nonlinear path estimate. In teleconferencing systems, this extension can handle distortion caused by amplifier saturation, improving overall performance.
[0073] Understandably, the model's output is used for echo cancellation applications.
[0074] Specifically, the estimated path is convolved with the reference signal to generate a simulated echo, which is then subtracted from the microphone signal to achieve cleanup. The model's effect is to reduce the level of residual echo. In another implementation, for mobile device scenarios, the model is optimized to a low-computational-complexity version.
[0075] Specifically, the filter order is reduced to 128, and a Fast Fourier Transform is used to accelerate convolution. This optimization process maintains the core of correlation analysis while adapting to processor limitations.
[0076] For example, frequency domain analysis can be combined to enhance the time-domain model. Specifically, the time-domain envelope correlation is converted into a frequency-domain power spectrum correlation, which is then fed back into the path estimation. This combination performs well in noisy environments, such as office background noise, and can more accurately distinguish echo components. Furthermore, model performance is evaluated using metrics such as echo return loss enhancement values.
[0077] Specifically, the power ratio of the input and output signals is calculated after implementation to ensure the accuracy of the estimate.
[0078] Step S105: Input the audio capture enhancement signal into the echo path estimation model, identify and remove the echo components, and output the purified audio signal.
[0079] The audio capture signal is acquired, and its frequency domain representation is obtained through Fourier transform. The spectral energy distribution is calculated from this frequency domain representation to determine the noise distribution characteristics. An adaptive filter is then applied to the audio capture enhancement signal based on these noise distribution characteristics. The adaptive filter takes the audio capture enhancement signal and the noise distribution characteristics as inputs and outputs the filtered signal, resulting in the enhanced signal. Multi-channel echo paths are extracted from the enhanced signal. These echo paths are obtained by separating the time delay signals of different channels. The amplitude differences between the multi-channel echo paths are assessed; if the amplitude difference exceeds a preset threshold, the corresponding echo component is subtracted, resulting in the purified audio signal.
[0080] Specifically, in one implementation, the step of inputting the audio capture enhancement signal into the echo path estimation model begins with the signal acquired by the microphone.
[0081] Specifically, the audio capture enhancement signal is obtained through preprocessing, such as applying noise suppression algorithms to the raw microphone signal to improve the clarity of the direct sound components. This input process ensures that the model receives high-quality data, avoiding the impact of low signal-to-noise ratio on estimation accuracy.
[0082] It should be noted that the echo path estimation model is based on a previously constructed adaptive filter structure designed to simulate the sound propagation path from the speaker to the microphone. Using this input, the model can analyze the mixing components in the signal, providing a basis for subsequent identification. In video call systems, this step helps to handle echoes caused by network latency in real time. Furthermore, the process of identifying echo-subtracted components relies on the model's predictive capabilities.
[0083] Specifically, the model performs a convolution operation on the input enhanced signal and a reference signal to generate a predicted echo signal. The reference signal is typically the raw audio played by the speaker. The recognition step involves comparing the predicted echo with the actual microphone signal and locating the position and intensity of the echo components by calculating the error signal.
[0084] For example, in a conference room environment, if the augmented signal exhibits delayed energy peaks, the model identifies these peaks as echo components. This identification not only quantifies the echo delay but also takes into account the multipath effects of room reflections, ensuring the appropriateness of the subtraction.
[0085] Preferably, when subtracting echo components, an error minimization mechanism is used to update the model parameters. Specifically, the predicted echo is subtracted from the enhanced signal to generate a preliminary cleaned signal. Then, the filter coefficients are iteratively adjusted until the residual error approaches zero. The core of this subtraction mechanism lies in adaptive step size control, such as dynamically adjusting the step size value based on signal energy to adapt to environmental changes.
[0086] In one possible implementation, if sudden noise is detected, the model will pause updates to prevent divergence.
[0087] It should be noted that the internal structure of the echo path estimation model includes a filter coefficient vector and an update algorithm. The model iteratively optimizes the coefficients using the least mean square criterion, a process involving dot-multiplication and accumulation of the input signal and coefficients to output the predicted value.
[0088] For example, in a smart speaker application for conferencing, the model first initializes the coefficients with small random values, and then calculates the gradient descent direction in each frame of the signal to achieve a gradual approximation of the path. This structure explains how the model extracts echo features from the enhanced signal, avoiding the limitations of traditional fixed filters. This explanation clarifies the model's role as an echo simulation engine, providing dynamic adaptability in audio processing.
[0089] In one embodiment, the step of outputting the purified audio signal immediately follows the subtraction process.
[0090] Specifically, the clean signal is the result of subtracting the estimated echo from the microphone signal, and then undergoes further post-processing such as gain adjustment to ensure the naturalness of the output audio.
[0091] Understandably, this output can be directly used in a speech recognition module. For example, in a teleconferencing system, improved signal clarity reduces misunderstandings between participants. This output process not only completes the echo cancellation chain but also provides a scalable interface for the overall system. Furthermore, to enhance the robustness of the model, frequency-domain assisted recognition can be introduced. Specifically, the time-domain enhanced signal is converted into a spectral representation, and then the power distribution of the echo is calculated in the frequency domain, combined with the model's time-domain estimation.
[0092] Preferably, this combination achieves more accurate subtraction by inverse transforming back to the time domain. In one implementation, if the ambient noise is high, frequency domain analysis can highlight the frequency band characteristics of the echo, avoiding time domain aliasing. This extension performs well in office video conferencing scenarios, handling signal cleansing under background conversation interference.
[0093] For example, consider the complete process in a specific scenario: In a hands-free calling device, an enhanced signal is first input to a model, which predicts the path based on historical correlation data. Then, the delayed portion of the echo component is identified, subtracted using a subtraction operation, and finally, anechoic audio is output. This process forms a closed loop from input to output, ensuring real-time performance.
[0094] It should be noted that the model's path estimation acts as a bridge here, connecting signal enhancement and final sanitization. In another implementation, for multi-channel audio capture, the model input fuses multiple enhanced signals.
[0095] Specifically, the signal from each channel is independently input into the model, generating its own echo estimate, which is then synthesized and subtracted using an averaging method. This multi-channel processing explains how the model handles spatial diversity, capturing reflection paths from different directions in large conference rooms and improving recognition accuracy. This approach further enhances the purity of the output signal, making it suitable for collaborative team environments. Furthermore, the quality of the output signal can be evaluated by calculating the residual echo level. This process involves comparing the energy ratio of the input and output signals to ensure the effectiveness of the subtraction.
[0096] In one embodiment, the model is considered successful if the residual is below a threshold. This evaluation supports online optimization of the model, ensuring continuous and clear audio in real-world applications such as distance learning systems.
[0097] Understandably, the technical goal of the entire process is to achieve efficient echo cancellation, thereby providing reliable support in the field of audio communication.
[0098] For example, after constructing the path model, by subtracting steps, the system can handle the challenges posed by complex acoustic environments and output high-quality signals to support seamless interaction.
[0099] Step S106: Monitor the signal strength change of the purified audio signal. If the signal strength change exceeds the preset intensity change threshold, confirm the second speaking position estimation data and readjust the beam focusing direction.
[0100] The original audio signal is acquired through a microphone array. The noise component of the original audio signal is obtained, and Fourier transform is used to separate the noise, resulting in a purified audio signal. For the purified audio signal, the amplitude of signal intensity variation is monitored. If the amplitude of signal intensity variation exceeds a preset amplitude variation threshold, positional offset features are extracted from the amplitude of variation to confirm the estimated second speaking position data and determine the speaker's movement trajectory. Based on the speaker's movement trajectory, the beam focusing direction is readjusted to obtain the adjusted beam focusing direction. The second speaking position refers to the speaker's new position after the position change.
[0101] Specifically, in one implementation, the system first performs purification processing on the acquired audio signal to remove noise interference.
[0102] Specifically, the purification process can be achieved through filtering techniques, such as using adaptive filters to suppress background noise, thereby obtaining a clearer audio signal. This purified signal is then used for subsequent intensity monitoring to ensure accuracy. Furthermore, monitoring the amplitude changes in the intensity of the purified audio signal involves calculating the amplitude fluctuations of the signal within a time window.
[0103] For example, audio signals can be segmented, and the root mean square value of each segment can be calculated as an intensity index. Then, the intensity difference between adjacent segments can be compared to quantify the magnitude of change. When applied in a voice conferencing system, this method can capture the dynamic changes in the speaker's voice in real time.
[0104] It should be noted that the calculation of signal strength variation can be based on a preset time window, such as a window every 5 seconds, calculating the difference between the maximum and minimum intensity within the window. If this difference exceeds a preset intensity variation threshold, it indicates a significant signal change. This threshold can be calibrated using experimental data and adjusted in different conference room environments to accommodate the effects of echo or distance. Through this monitoring, the system can detect potential changes in speaking position, such as a participant moving from one meeting to another, thus providing a basis for position estimation. In practical applications, this mechanism helps improve the accuracy of audio acquisition and avoids signal attenuation caused by positional deviations.
[0105] In one possible implementation, if the signal strength change exceeds a preset threshold, the estimated data for the second speaking position is confirmed.
[0106] Specifically, the second speaker position estimation data can be obtained by calculating the phase difference of the microphone array, for example, by using a time delay estimation algorithm to determine the direction of the sound source. Once the amplitude exceeds a threshold, the system considers the estimation data valid and uses it to update the position information. This confirmation process is particularly useful in multi-speaker conference scenarios, enabling rapid response to speaker changes and ensuring continuous capture of the audio signal.
[0107] Preferably, after confirming the estimated data of the second speaking position, the system readjusts the beam focusing direction.
[0108] For example, the beamforming parameters of a microphone array can be controlled by a digital signal processor to point the main beam to a confirmed location. This adjustment can be implemented in steps: first, the new focus angle is calculated, and then a weighted delay summation algorithm is applied to redirect the beam. In large conference hall applications, this adjustment can significantly improve the voice clarity of distant speakers and reduce blind spots.
[0109] For example, in a small conference room scenario, after the system detects that the signal strength change exceeds a threshold, it confirms the second position estimation data and adjusts the beam to achieve a smooth transition from the stage to the audience seating area. This scenario demonstrates the versatility of the technology, making it suitable for various indoor audio environments.
[0110] Understandably, in another embodiment, the threshold can be dynamically adjusted and automatically optimized based on the ambient noise level.
[0111] For example, if background noise increases during a meeting, the threshold is raised accordingly to avoid false triggering of location confirmation. This flexibility enhances the system's robustness in noisy environments. Furthermore, the beam focusing direction readjustment process includes verifying the quality of the adjusted signal.
[0112] Specifically, after adjustment, the intensity change is monitored again. If it remains stable, the new direction is maintained; otherwise, iterative optimization is performed. This verification loop ensures the reliability of audio processing and delivers more stable voice transmission results in actual business applications.
[0113] In one embodiment, the entire process is integrated into the smart conferencing device, forming a closed-loop control from signal purification to beam adjustment.
[0114] For example, the device collects multi-channel audio in real time, calculates the amplitude after purification, and if it exceeds the threshold, it confirms the location and makes adjustments to ensure that the voices of all speakers in the meeting are picked up evenly.
[0115] Step S107: Based on the beam focusing direction of the estimated data from the second speaking position, the incoming audio signal is processed in real time. If the time difference changes, the intensity difference and temporal envelope information are fused from the processed audio signal to obtain the final clear audio output signal for use in conference transmission of the distributed microphone network.
[0116] Based on the beam focusing direction, the incoming audio signal is acquired, and noise components are separated using Fourier transform to obtain a purified audio signal. For the purified audio signal, if the time difference changes, intensity difference features are extracted from the change to determine intensity difference fusion parameters. Using these intensity difference fusion parameters, time-domain envelope information is fused to obtain fused signal features. Based on these fused signal features, the transmission path of the distributed microphones is adjusted to obtain a final clear audio output signal. The audio transmission and distribution of the conference system are then determined using this output signal.
[0117] Specifically, in one implementation, the system processes the incoming audio signals in real time based on the beam focusing direction of the data estimated from the second speaking position.
[0118] Specifically, this processing is implemented using a digital signal processor. First, the beam focusing direction is taken as a parameter input, and the microphone array's receiving mode is adjusted to enhance the intensity of the sound signal from that direction. This method, applied in a conference environment with a distributed microphone network, can capture participants' speech in real time, ensuring that the signal is not attenuated due to positional deviations.
[0119] For example, in a medium-sized conference room scenario, the system uses pre-confirmed location data to selectively filter background noise, thus providing a basis for subsequent judgments. Furthermore, if the time difference changes, the fusion process is triggered.
[0120] It should be noted that the time difference refers to the time difference in arrival times of audio signals at different microphone units in a distributed microphone network, obtained by calculating phase difference or time delay estimation. The specific process involves acquiring signals from multiple channels, comparing the arrival times of each channel, and considering a change as occurring if the difference exceeds a preset threshold, such as 0.1 seconds. This judgment is used in conference transmissions to detect speaker movement or switching, helping the system respond dynamically. In practical applications, this mechanism avoids signal interruptions, ensuring a continuous audio stream.
[0121] Preferably, intensity differences and temporal envelope information are fused from the real-time processed audio signal to obtain the final clear audio output signal.
[0122] Specifically, intensity differences are obtained by calculating the amplitude differences of the signal across different time periods, such as comparing the mean difference between the current and previous periods. Temporal envelope information involves extracting the signal's amplitude profile, using envelope detection algorithms like the Hilbert transform to obtain the instantaneous amplitude curve. The fusion process integrates this information, for example, by adjusting the envelope curve using a weighted average method that incorporates intensity differences as weights, thereby generating an enhanced output signal. This fusion improves audio clarity in distributed microphone networks and is suitable for online conferencing platforms.
[0123] For example, in one possible implementation, the entire process is applied to a video conferencing system. First, the signal is processed according to the beam direction. Then, the time difference is monitored; if it changes, the intensity differences (e.g., calculating amplitude fluctuations) and the temporal envelope (e.g., extracting the signal peak curve) are fused, ultimately outputting a clear signal for network transmission. This approach demonstrates versatility in multi-participant scenarios, capable of handling dynamic changes in different speaking positions.
[0124] Understandably, the calculation of the time difference requires careful consideration of microphone spacing and sound propagation speed. Specifically, the process involves collecting signals from each microphone, calculating the cross-correlation function to find the time delay difference corresponding to the peak value, and then monitoring the time delay variations across consecutive frames. If the variation is significant, the system immediately enters the fusion phase. This calculation process ensures real-time performance during conference transmission.
[0125] In another embodiment, adaptive weights can be introduced when fusing intensity differences and temporal envelopes. Specifically, the contribution ratio of envelope information is dynamically adjusted based on the magnitude of the intensity differences; for example, its weight is increased when the intensity differences are large, to optimize the clarity of the output signal. This flexibility enhances the robustness of the system in noisy conference environments.
[0126] Furthermore, the generated clear audio output signal is directly used for conference transmission over a distributed microphone network. For example, in a large conference hall, the output signal is distributed via the network to remote participant terminals, ensuring that all sound is balanced and clear. This application demonstrates the scalability of the technology within the same domain.
[0127] In one embodiment, the real-time processing procedure includes determining the time difference after initial filtering. Specifically, this involves extracting time-domain features from the processed signal, comparing the deviation with a reference value, and if a change is found, fusing the information to generate an output. This method is effective in small team meetings, resulting in stable transmission performance.
[0128] It should be noted that the principle of extracting time-domain envelope information is based on tracking the amplitude changes of the signal.
[0129] Specifically, the envelope curve is obtained through rectification and low-pass filtering, and then fused with intensity differences to form the final signal. This explanation of the principle helps improve audio quality in conferencing systems.
[0130] For example, in the actual deployment of a distributed network, after the system confirms the change in time difference, the fusion process can be executed in steps: first, the intensity difference is quantified, then the envelope information is superimposed, and finally, a clear audio output signal is output for transmission.
[0131] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0132] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for echo cancellation of a digital microphone, characterized in that: Includes the following steps: The system acquires sound signals from multiple microphones via a distributed microphone network. By calculating the time difference and intensity difference between these signals, a first speaking position is determined. The beam focusing direction of the microphone array is adjusted based on this first speaking position to determine the corresponding audio capture enhancement signal. Frequency domain analysis is performed on the audio capture enhancement signal to extract its spectral features. If these spectral features overlap, it is determined that multiple speakers are present, and the direct sound component and mixed echo component of each speaker are separated. The correlation between the temporal envelope of the mixed echo component and the temporal envelope of the direct sound component is analyzed. An echo path estimation model is constructed; the audio capture enhancement signal is input into the echo path estimation model, echo components are identified and removed, and the purified audio signal is output; the signal intensity change amplitude of the purified audio signal is monitored, and if the signal intensity change amplitude exceeds a preset intensity change threshold, the second speech position estimation data is confirmed, and the beam focusing direction is readjusted; based on the beam focusing direction of the second speech position estimation data, the subsequent incoming audio signals are processed in real time, and if the time difference changes, the intensity difference and temporal envelope information are fused from the real-time processed audio signal to obtain the final clear audio output signal; The final clear audio output signal is used for conference transmission in a distributed microphone network.
2. The digital microphone echo cancellation method according to claim 1, characterized in that: Obtaining the first speaking position includes: Multiple sound signals are acquired from the distributed microphone network, the time difference between each signal is calculated, and a preliminary time difference matrix is obtained. For the aforementioned preliminary time difference matrix, intensity difference data is obtained; The approximate coordinates of the sound source were determined using a triangulation algorithm. Based on the rough coordinates of the sound source, a purified signal is obtained from noise interference filtering. Based on the purification signal, the final position is determined by intensity difference analysis, and the first speaking position is obtained.
3. The digital microphone echo cancellation method according to claim 1, characterized in that: Determining the audio capture enhancement signal corresponding to the first speaking position includes: The original audio signal is acquired by the microphone array, and the time delay difference between each microphone is calculated by the delay summation algorithm to determine the first speaking position. Adjust the beam focusing direction according to the first speaking position, and obtain the directional gain parameter corresponding to the first speaking position from the preset gain table; The original audio signal is processed according to the directional gain parameter, and the signal amplitude is fused using a weighted summation method to obtain the audio capture enhancement signal corresponding to the first speaking position.
4. The digital microphone echo cancellation method according to claim 1, characterized in that: The frequency domain analysis of the audio capture enhancement signal includes: The audio capture enhancement signal is frequency domain transformed by Fourier transform to obtain a frequency domain representation signal; Spectral features are extracted from the frequency domain representation signal. If the spectral features overlap, it is determined that multiple people are speaking. Based on the determination results of multiple speakers, a blind source separation algorithm is used to separate the direct sound components of each speaker; Based on the direct sound component, the mixed echo component is decomposed from the audio capture enhancement signal to obtain the direct sound component and mixed echo component of each speaker.
5. The digital microphone echo cancellation method according to claim 1, characterized in that: The construction of the echo path estimation model includes: Acquire audio signals to obtain time-domain data of the mixed echo component and the direct sound component; Based on the time-domain data of both, the Hilbert transform is used to calculate their respective envelopes, and the envelope curves are obtained. The Pearson correlation coefficient is calculated based on the envelope curve to determine the correlation strength between the mixed echo component and the direct sound component; If the correlation strength exceeds a preset threshold, then delay features are extracted from the correlation strength to construct an initial echo path model; The parameters of the initial echo path model are iteratively updated to obtain the final echo path estimation model.
6. The digital microphone echo cancellation method according to claim 1, characterized in that: The step of inputting the audio capture enhancement signal into the echo path estimation model, identifying and subtracting echo components, and outputting the cleaned audio signal includes: Acquire the audio capture enhancement signal and obtain its frequency domain representation; The spectral energy distribution is calculated from the frequency domain representation to determine the noise distribution characteristics; An adaptive filter is used to process the audio capture enhancement signal based on the noise distribution characteristics. The multi-channel echo path is extracted from the audio capture enhancement signal, and the amplitude difference of the multi-channel echo path is determined. If the amplitude difference exceeds a preset threshold, the corresponding echo component is deducted to obtain the purified audio signal.
7. The digital microphone echo cancellation method according to claim 1, characterized in that: The monitoring of the signal strength change amplitude of the purified audio signal, if the signal strength change amplitude exceeds a preset intensity change threshold, confirms the second speaking position estimation data and readjusts the beam focusing direction, including: The original audio signal is acquired by a microphone array, the noise component of the original audio signal is obtained, the noise is separated, and a purified audio signal is obtained. For the purified audio signal, the signal intensity change amplitude is monitored. If the signal intensity change amplitude exceeds a preset intensity change threshold, the position offset feature is extracted from the signal intensity change amplitude to confirm the second speaking position estimation data and determine the speaker's movement trajectory. The beam focusing direction is readjusted based on the speaker's movement trajectory to obtain the adjusted beam focusing direction.
8. The digital microphone echo cancellation method according to claim 1, characterized in that: The step of processing subsequent incoming audio signals in real time based on the beam focusing direction estimated from the second speaking position includes: The incoming sound signal is obtained according to the beam focusing direction, the noise component is separated, and the purified sound signal is obtained. For the purified sound signal, if the time difference value changes, the intensity difference feature is extracted from the change, and the intensity difference fusion parameter is determined. The temporal envelope information is fused using the intensity difference fusion parameters to obtain the fused signal features; By adjusting the transmission path of the distributed microphones using the fused signal characteristics, a clear audio output signal is obtained.