Echo cancellation method, device, equipment and storage medium

By processing signals through a dual-microphone array and a deep neural network model, the problems of poor echo cancellation and high hardware costs in access control and intercom systems are solved, achieving efficient echo suppression and low-cost solutions in complex acoustic environments.

CN120496562BActive Publication Date: 2025-09-23SHENZHEN DINSTAR TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510984940.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-09-23
Estimated Expiration
2045-07-17

AI Technical Summary

Technical Problem

The existing echo cancellation algorithm in the field of access control intercom has slow convergence speed, large steady-state error, insufficient environmental adaptability, high hardware cost, and poor echo suppression effect, especially in complex acoustic environments.

Method used

A dual-microphone array is used to collect signals, and a logarithmic Mel spectrum is generated through short-time Fourier transform. A deep neural network model is used to extract the echo spectrum and perform nonlinear mapping. Combined with voice activity detection and reverberation level adjustment, the near-end pure speech signal is determined.

Benefits of technology

It improves the echo suppression effect, reduces hardware costs, and improves voice clarity and system adaptability in various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120496562B_ABST
    Figure CN120496562B_ABST
Patent Text Reader

Abstract

The present invention discloses an echo cancellation method, apparatus, device, and storage medium. The method is applied to a communication device including a dual-microphone array. The method comprises: collecting a near-end mixed signal and a far-end reference signal through the dual-microphone array, extracting spectral features of the near-end mixed signal and the far-end reference signal to obtain feature data; processing the feature data using a deep neural network model to obtain an estimated echo spectrum; and determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal. Because the present invention collects the near-end mixed signal and the far-end reference signal through the dual-microphone array, obtains the estimated echo spectrum using a deep neural network model, and finally determines the near-end clean speech signal, compared to the prior art, the present invention improves the echo suppression effect in various environments and reduces hardware costs.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of audio processing technology, and in particular to an echo cancellation method, device, equipment and storage medium. Background Art

[0002] Currently, in the field of access control and intercoms, terminal devices such as the DP88 are widely used in building security scenarios. However, existing systems have significant shortcomings in echo cancellation: 1. Limitations of traditional algorithms: These rely on G.168-based echo cancellation algorithms (such as NLMS), which simulate the echo path using a fixed-step adaptive filter. This results in slow convergence and large steady-state errors in complex environments (such as metal door frame reflections and multipath propagation). This is particularly true when two parties are speaking simultaneously, resulting in poor echo suppression effectiveness. 2. Inadequate environmental adaptability: In noisy outdoor environments or reverberant indoor environments, the DP88's microphone easily picks up reflected sound from the speaker, causing a noticeable echo to the far-end user, severely impacting call clarity and security system reliability. 3. High hardware dependency: Traditional algorithms rely on high-performance DSP chips for filtering, increasing hardware costs and making it difficult to adapt to dynamic acoustic environments (such as changes in the reflection path caused by movement, doors, and windows).

[0003] Therefore, there is an urgent need for an echo cancellation method that can improve the echo suppression effect in various environments and reduce hardware costs. Summary of the Invention

[0004] The main purpose of the present invention is to provide an echo cancellation method, device, equipment and storage medium, aiming to solve the technical problems of poor echo suppression effect and high hardware cost in the existing technology in double-end calls or strong reverberation scenarios.

[0005] To achieve the above object, the present invention provides an echo cancellation method, which is applied to a communication device including a dual-microphone array, and comprises the following steps:

[0006] Collecting a near-end mixed signal and a far-end reference signal through the dual-microphone array, and performing spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data;

[0007] Processing the feature data using a deep neural network model to obtain an estimated echo spectrum;

[0008] A near-end clean speech signal is determined based on the estimated echo spectrum and the near-end mixed signal.

[0009] Optionally, the step of collecting a near-end mixed signal and a far-end reference signal through the dual-microphone array, and performing spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data includes:

[0010] Collecting a near-end mixed signal and a far-end reference signal by the dual-microphone array, wherein the near-end mixed signal includes an echo signal and a near-end voice signal;

[0011] Performing time-frequency conversion processing on the near-end mixed signal and the far-end reference signal using short-time Fourier transform to generate corresponding logarithmic Mel spectrum;

[0012] Dynamic feature information is extracted from the logarithmic Mel spectrum to obtain feature data.

[0013] Optionally, after the step of collecting the near-end mixed signal and the far-end reference signal by using the dual-microphone array, the method further includes:

[0014] monitoring spectrum entropy values ​​of the near-end mixed signal and the far-end reference signal in real time to obtain monitoring results;

[0015] Determining the reverberation level corresponding to the current environment according to the monitoring results;

[0016] The update step size of the deep neural network model is adjusted based on the reverberation level corresponding to the current environment.

[0017] Optionally, the step of processing the feature data using a deep neural network model to obtain an estimated echo spectrum includes:

[0018] Constructing historical frame sequence data based on the feature data, and inputting the historical frame sequence data into a deep neural network model;

[0019] The deep neural network is used to learn the nonlinear mapping relationship between the echo and the pure speech in the frequency spectrum of the historical frame sequence data, and an estimated echo spectrum is output.

[0020] Optionally, after the step of determining the near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal, the method further includes:

[0021] Performing voice activity detection on the near-end clean voice signal, and determining whether there is near-end voice activity in the current environment according to the voice activity detection result;

[0022] When there is near-end voice activity in the current environment, freezing parameter updates of the deep neural network model;

[0023] When there is no near-end voice activity in the current environment, background noise is generated and inserted into the silent voice signal.

[0024] Optionally, the step of performing voice activity detection on the near-end clean voice signal and determining whether there is near-end voice activity in the current environment according to the voice activity detection result includes:

[0025] Performing voice activity detection on the near-end clean voice signal based on an energy threshold and a double-threshold algorithm to obtain a voice activity detection result;

[0026] Determine whether there is near-end voice activity in the current environment according to the voice activity detection result.

[0027] Optionally, after the step of determining the near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal, the method further includes:

[0028] Encoding the near-end clean speech signal based on a wideband speech coding standard to obtain an encoded near-end clean speech signal;

[0029] The encoded near-end clean voice signal is encrypted to achieve encrypted transmission of the encoded near-end clean voice signal.

[0030] In addition, to achieve the above-mentioned object, the present invention further provides an echo cancellation device, which is applied to a communication device including a dual-microphone array, and includes:

[0031] a signal acquisition module, configured to acquire a near-end mixed signal and a far-end reference signal through the dual-microphone array, and perform spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data;

[0032] An echo extraction module is used to process the feature data using a deep neural network model to obtain an estimated echo spectrum;

[0033] The echo cancellation module is configured to determine a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal.

[0034] In addition, to achieve the above objectives, the present invention also proposes an echo cancellation device, which includes: a memory, a processor, and an echo cancellation program stored in the memory and executable on the processor, wherein the echo cancellation program is configured to implement the steps of the echo cancellation method described above.

[0035] In addition, to achieve the above-mentioned object, the present invention further proposes a storage medium, on which an echo cancellation program is stored. When the echo cancellation program is executed by a processor, the steps of the echo cancellation method described above are implemented.

[0036] The present invention discloses a method for collecting a near-end mixed signal and a far-end reference signal using a dual-microphone array, extracting spectral features from the near-end mixed signal and the far-end reference signal to obtain feature data; processing the feature data using a deep neural network model to obtain an estimated echo spectrum; and determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal. Because the present invention collects the near-end mixed signal and the far-end reference signal using a dual-microphone array, obtains an estimated echo spectrum using a deep neural network model, and finally determines the near-end clean speech signal, compared to existing technologies, the present invention improves echo suppression effectiveness in various environments and reduces hardware costs. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 1 is a flow chart of a first embodiment of an echo cancellation method according to the present invention;

[0038] Figure 2 1. It is a flow chart of a second embodiment of the echo cancellation method of the present invention;

[0039] Figure 3 1. It is a flowchart of a third embodiment of an echo cancellation method according to the present invention;

[0040] Figure 4 is a structural block diagram of a first embodiment of an echo cancellation device according to the present invention;

[0041] Figure 5 It is a structural diagram of an echo cancellation device in a hardware operating environment involved in an embodiment of the present invention.

[0042] The purpose, features and advantages of the present invention will be further described with reference to the accompanying drawings and in conjunction with the embodiments. DETAILED DESCRIPTION

[0043] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0044] The embodiment of the present invention provides an echo cancellation method, referring to Figure 1 , Figure 1 FIG. 1 is a flow chart of the first embodiment of the echo cancellation method of the present invention.

[0045] In this embodiment, the echo cancellation method is applied to a communication device including a dual-microphone array. The method includes steps S10 to S30:

[0046] Step S10: collecting a near-end mixed signal and a far-end reference signal through the dual-microphone array, and performing spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data.

[0047] It should be noted that the execution entity of this embodiment can be a computer server device with data processing, network communication, and program execution functions used in access control intercom scenarios, such as a server, tablet computer, or personal computer, or an electronic device capable of performing the above functions (such as an echo cancellation device or communication device). The following uses an echo cancellation system (hereinafter referred to as the system) that includes a communication device as an example to illustrate this embodiment and the following embodiments.

[0048] It should be understood that a dual-microphone array uses two microphones arranged at a specific distance and in a geometric configuration to collect sound source signals. It then uses information such as the time difference, amplitude difference, and phase difference of the sound waves at the two microphones to locate and enhance the target sound source, as well as suppress noise and interference. This principle is based on the propagation characteristics of acoustic signals. The different distances between the sound source and the two microphones result in differences in the time and phase of the sound waves reaching the two microphones. By analyzing and processing these differences, information related to the target sound source can be extracted and interference from non-target sound sources can be suppressed.

[0049] The aforementioned near-end mixed signal refers to the signal collected by the microphone of a local device (such as an access control intercom or mobile phone). This signal includes the near-end voice signal (i.e., the voice signal generated by the local user speaking), the echo signal (i.e., the voice signal sent from the remote device, which is played through the local device's speaker and then picked up by the microphone. For example, in an access control intercom, the remote user's voice is broadcast from the access control device's speaker. Due to the influence of the acoustic environment (such as reflection and reverberation), part of the sound is received by the access control device's microphone, forming an echo, which will interfere with the transmission of the near-end voice signal), and background noise (i.e., other interfering noises in the local environment, such as fan noise and street noise).

[0050] The far-end reference signal refers to the original speech signal sent from the remote device and serves as a reference signal in the echo cancellation system. Its primary function is to provide the echo cancellation algorithm with characteristic information about the far-end speech, helping the algorithm more accurately identify and estimate the echo components picked up by the local microphone. By comparing and analyzing the far-end reference signal with the signal picked up by the local microphone, the echo cancellation system constructs an echo model, effectively suppressing the echo.

[0051] For example, in an access control intercom system, when an indoor user and an outdoor user speak simultaneously, the access control device's microphone captures a near-end mixed signal (including the indoor user's voice, the echo from the outdoor user's voice played back through the speaker, and indoor background noise), while the far-end reference signal is the outdoor user's original voice signal. The echo cancellation system analyzes these two signals to accurately estimate the echo and remove it from the near-end mixed signal, allowing the indoor user's voice to be clearly transmitted to the outdoor user, and vice versa.

[0052] In a specific implementation, the dual-microphone array can be used to collect a near-end mixed signal and a far-end reference signal, where the near-end mixed signal includes an echo signal and a near-end voice signal; the near-end mixed signal and the far-end reference signal are subjected to time-frequency conversion processing using short-time Fourier transform to generate corresponding logarithmic Mel spectrums; dynamic feature information is extracted from the logarithmic Mel spectrum to obtain feature data.

[0053] It should be explained that log-Mel Spectrogram is a feature representation method widely used in the field of speech signal processing. It combines the simulation of human auditory perception by the Mel scale and the compression of the signal dynamic range by the logarithmic transformation, and can effectively capture important information in the speech signal.

[0054] It should be noted that through the logarithmic Mel spectrum, the speech signal can be converted from the time domain to the frequency domain, and the features can be compressed and represented on the Mel frequency scale, so that the subsequent deep neural network can better learn the echo features in the speech signal and thus achieve effective echo cancellation.

[0055] It should be understood that dynamic feature information is extracted from the logarithmic Mel spectrum to obtain feature data. The feature data can capture information about how the speech signal changes over time and provide richer feature representations for tasks such as speech recognition and echo cancellation.

[0056] Step S20: Process the feature data using a deep neural network model to obtain an estimated echo spectrum.

[0057] It should be understood that the deep neural network (DNN) model is a complex computational model composed of a large number of interconnected artificial neurons, which can efficiently represent and classify input data through hierarchical feature learning and nonlinear transformation.

[0058] In a specific implementation, historical frame sequence data can be constructed based on the feature data, and the historical frame sequence data can be input into a deep neural network model; the deep neural network is used to learn the nonlinear mapping relationship between echo and pure speech in the spectrum of the historical frame sequence data, and output an estimated echo spectrum.

[0059] It should be noted that in order to improve the reliability of the deep neural network model, before using the deep neural network model to process the feature data to obtain the estimated echo spectrum, it can also include: training the deep neural network model, taking the sample input feature data and the corresponding sample echo spectrum as the target during the training process, and adjusting the network parameters so that the error between the estimated echo spectrum output by the model and the sample echo spectrum is minimized.

[0060] Step S30: determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal.

[0061] It should be noted that the near-end mixed signal includes an echo signal and a near-end speech signal. After obtaining the estimated echo spectrum, the estimated echo spectrum is subtracted from the near-end mixed signal to obtain the near-end pure speech signal.

[0062] In a specific implementation, to achieve encrypted transmission of near-end clean voice signals and ensure security during voice communication, after step S30, the process further includes: encoding the near-end clean voice signals based on a wideband voice coding standard to obtain an encoded near-end clean voice signal; and encrypting the encoded near-end clean voice signal to achieve encrypted transmission of the encoded near-end clean voice signal. For example, G.722 wideband coding is supported to preserve high-frequency voice details, combined with SRTP encrypted transmission.

[0063] This embodiment discloses a method for collecting a near-end mixed signal and a far-end reference signal using a dual-microphone array, extracting spectral features from the near-end mixed signal and the far-end reference signal to obtain feature data; processing the feature data using a deep neural network model to obtain an estimated echo spectrum; and determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal. Because this embodiment uses a dual-microphone array to collect the near-end mixed signal and the far-end reference signal, and uses a deep neural network model to obtain an estimated echo spectrum before finally determining the near-end clean speech signal, compared to existing technologies, this embodiment improves echo suppression effectiveness in various environments and reduces hardware costs.

[0064] refer to Figure 2 , Figure 2 FIG. 1 is a flow chart of a second embodiment of an echo cancellation method according to the present invention.

[0065] Based on the first embodiment described above, in this embodiment, after step S10, steps S101 to S103 are further included:

[0066] Step S101: monitor the spectrum entropy values ​​of the near-end mixed signal and the far-end reference signal in real time to obtain monitoring results.

[0067] Step S102: determining the reverberation level corresponding to the current environment according to the monitoring result.

[0068] Step S103: adjusting the update step size of the deep neural network model based on the reverberation level corresponding to the current environment.

[0069] It should be understood that reverberation level is a key parameter for evaluating speech signal processing quality, acoustic design, sound quality control, and other fields. The reverberation level can be determined by calculating the theoretical reverberation time based on the room's volume, surface area, sound absorption coefficient, and other parameters using theoretical formulas such as the Sabine formula and the Eeling formula. This is then compared with the actual measured reverberation time.

[0070] For example, by monitoring the spectral entropy value of the input signal (i.e., the near-end mixed signal and the far-end reference signal) in real time, the reverberation level of the current environment can be judged, and the update step size of the deep neural network model (CNN model) can be automatically adjusted (such as increasing the learning rate when the reverberation is strong to accelerate convergence).

[0071] It should be noted that by real-time monitoring of spectral entropy values ​​and dynamically adjusting the update step size, the DNN model can process speech signals more efficiently, ensuring that the system can maintain real-time speech processing capabilities in different reverberation environments.

[0072] For example, in environments with strong reverberation, where spectral entropy is typically high, increasing the DNN model's update step size can accelerate the model's adaptation to environmental changes, allowing it to more quickly learn and adapt to the reverberant environment's characteristic patterns. In environments with weak reverberation, reducing the update step size allows the model to more precisely learn the detailed features of the speech signal, preventing model instability or missed important features caused by excessively large update step sizes.

[0073] Furthermore, in environments with strong reverberation, increasing the update step size allows the DNN model to more quickly learn and suppress echo characteristics, thereby improving echo cancellation and speech signal clarity. Therefore, by automatically adjusting the update step size based on the reverberation level, the DNN model can better adapt to different acoustic environments, accelerating learning in strong reverberation environments to identify useful information in speech, while finely learning speech details in weak reverberation environments, thereby improving speech recognition accuracy.

[0074] It should be added that in order to meet the needs of real-time intercom, FPGA accelerators can be used to implement parallel computing of the forward propagation of the DNN model, controlling the single-frame processing delay to within 10ms to meet the needs of real-time intercom.

[0075] This embodiment discloses collecting a near-end mixed signal and a far-end reference signal through a dual-microphone array, monitoring the spectral entropy values ​​of the near-end mixed signal and the far-end reference signal in real time, and obtaining monitoring results; determining the reverberation level corresponding to the current environment based on the monitoring results; adjusting the update step size of a deep neural network model based on the reverberation level corresponding to the current environment; extracting spectral features from the near-end mixed signal and the far-end reference signal to obtain feature data; processing the feature data using a deep neural network model to obtain an estimated echo spectrum; and determining a near-end pure speech signal based on the estimated echo spectrum and the near-end mixed signal. Because this embodiment determines the reverberation level corresponding to the current environment by monitoring the spectral entropy values ​​of the near-end mixed signal and the far-end reference signal in real time, and adjusts the update step size of the deep neural network model based on the reverberation level, compared to the existing technology, this embodiment can effectively improve the adaptability of the deep neural network model, and improve the speech processing effect and system real-time performance.

[0076] refer to Figure 3 , Figure 3 FIG. 4 is a flow chart of a third embodiment of an echo cancellation method according to the present invention.

[0077] Based on the above embodiments, in this embodiment, after step S30, steps S40 to S60 are further included:

[0078] Step S40: performing voice activity detection on the near-end clean voice signal, and determining whether there is near-end voice activity in the current environment according to the voice activity detection result.

[0079] Step S50: When there is near-end voice activity in the current environment, the parameter update of the deep neural network model is frozen.

[0080] Step S60: When there is no near-end voice activity in the current environment, background noise is generated and inserted into the silent voice signal.

[0081] In a specific implementation, voice activity detection can be performed on the near-end pure voice signal based on an energy threshold and a double-threshold algorithm to obtain a voice activity detection result; and whether there is near-end voice activity in the current environment is determined based on the voice activity detection result.

[0082] It should be noted that based on the energy threshold and dual-threshold algorithm, when near-end voice activity is detected in the current environment, freezing the parameter update of the DNN model can effectively avoid filter divergence during double-end speech.

[0083] In addition, when there is no near-end voice activity in the current environment, that is, the near-end pure voice signal is a silent voice signal, background noise may be generated and inserted into the silent voice signal in order to improve the naturalness of the call.

[0084] This embodiment discloses collecting a near-end mixed signal and a far-end reference signal using a dual-microphone array, extracting spectral features from the near-end mixed signal and the far-end reference signal to obtain feature data; processing the feature data using a deep neural network model to obtain an estimated echo spectrum; determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal; performing voice activity detection on the near-end clean speech signal and determining whether near-end voice activity exists in the current environment based on the voice activity detection result; freezing parameter updates of the deep neural network model when near-end voice activity exists in the current environment; and generating background noise to be inserted into the silent speech signal when near-end voice activity does not exist in the current environment. Because this embodiment performs voice activity detection on the near-end clean speech signal and freezes parameter updates of the deep neural network model when near-end voice activity exists in the current environment, and generates background noise to be inserted into the silent speech signal when near-end voice activity does not exist, compared to the prior art, the present invention not only improves the echo cancellation effect but also enhances the user experience.

[0085] In addition, an embodiment of the present invention further provides a storage medium, on which an echo cancellation program is stored. When the echo cancellation program is executed by a processor, the steps of the echo cancellation method described above are implemented.

[0086] Reference Figure 4 , Figure 4 This is a structural block diagram of the first embodiment of the echo cancellation device of the present invention.

[0087] like Figure 4 As shown, the echo cancellation device proposed in the embodiment of the present invention is applied to a communication device, which includes a dual-microphone array. The device includes: a signal acquisition module 501 , an echo extraction module 502 and an echo cancellation module 503 .

[0088] The signal acquisition module 501 is configured to acquire a near-end mixed signal and a far-end reference signal through the dual-microphone array, and perform spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data.

[0089] The echo extraction module 502 is used to process the feature data using a deep neural network model to obtain an estimated echo spectrum.

[0090] The echo cancellation module 503 is configured to determine a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal.

[0091] The signal acquisition module 501 is further configured to collect a near-end mixed signal and a far-end reference signal through the dual-microphone array, where the near-end mixed signal includes an echo signal and a near-end speech signal; perform time-frequency conversion on the near-end mixed signal and the far-end reference signal using a short-time Fourier transform to generate corresponding logarithmic Mel-spectra; and extract dynamic feature information from the logarithmic Mel-spectra to obtain feature data.

[0092] The echo extraction module 502 is further configured to construct historical frame sequence data based on the feature data, and input the historical frame sequence data into a deep neural network model; utilize the deep neural network to learn the nonlinear mapping relationship between the echo and the pure speech in the frequency spectrum of the historical frame sequence data, and output an estimated echo spectrum.

[0093] The echo cancellation module 503 is further configured to encode the near-end clean voice signal based on a wideband voice coding standard to obtain an encoded near-end clean voice signal; and encrypt the encoded near-end clean voice signal to achieve encrypted transmission of the encoded near-end clean voice signal.

[0094] This device embodiment discloses collecting a near-end mixed signal and a far-end reference signal using a dual-microphone array, extracting spectral features from the near-end mixed signal and the far-end reference signal to obtain feature data; processing the feature data using a deep neural network model to obtain an estimated echo spectrum; and determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal. Because this device embodiment collects the near-end mixed signal and the far-end reference signal using a dual-microphone array, obtains an estimated echo spectrum using a deep neural network model, and finally determines the near-end clean speech signal, compared to existing technologies, this device embodiment improves echo suppression effectiveness in various environments and reduces hardware costs.

[0095] Based on the above-mentioned first embodiment of the echo cancellation device of the present invention, a second embodiment of the echo cancellation device of the present invention is proposed.

[0096] In this embodiment, the signal acquisition module 501 is also used to monitor the spectral entropy values ​​of the near-end mixed signal and the far-end reference signal in real time to obtain monitoring results; determine the reverberation level corresponding to the current environment based on the monitoring results; and adjust the update step size of the deep neural network model based on the reverberation level corresponding to the current environment.

[0097] Other embodiments or specific implementations of the echo cancellation device of the present invention can refer to the above-mentioned method embodiments and will not be described in detail here.

[0098] The present application provides an echo cancellation device, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the echo cancellation method in the above-mentioned embodiment 1.

[0099] Reference below Figure 5 , which shows a schematic diagram of the structure of an echo cancellation device suitable for implementing an embodiment of the present application. The echo cancellation device in the embodiment of the present application may include, but is not limited to, mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (Portable Application Descriptions), PMPs (Portable Media Players), and in-vehicle terminals (e.g., in-vehicle navigation terminals), as well as fixed terminals such as digital TVs and desktop computers. Figure 5 The echo cancellation device shown is only an example and should not limit the functions and scope of use of the embodiments of the present application.

[0100] like Figure 5 As shown, the echo cancellation device may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in a read-only memory 1002 or programs loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the echo cancellation device. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, a magnetic tape or hard disk; and a communication device 1009. Communication device 1009 can allow the echo cancellation device to communicate with other devices wirelessly or by wire to exchange data. Although the figure shows an echo cancellation device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems can be implemented or have instead.

[0101] In particular, according to the embodiments disclosed in the present application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, the embodiments disclosed in the present application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program comprising program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device 1003, or installed from a read-only memory 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are performed.

[0102] The echo cancellation device provided in this application, utilizing the echo cancellation method described in the aforementioned embodiments, can address the technical issues of prior art, such as poor echo suppression and high hardware costs in double-ended conversations or strong reverberation scenarios. Compared to prior art, the echo cancellation device provided in this application achieves the same beneficial effects as the echo cancellation method described in the aforementioned embodiments. Other technical features of this echo cancellation device are the same as those disclosed in the aforementioned embodiments and are not further elaborated here.

[0103] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0104] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0105] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or system comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or system. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or system comprising the element.

[0106] The serial numbers of the above embodiments of the present invention are for description only and do not represent the advantages or disadvantages of the embodiments.

[0107] Through the above description of the embodiments, those skilled in the art will clearly understand that the above-mentioned embodiments and methods can be implemented by means of software plus the necessary general-purpose hardware platform. Of course, hardware can also be used, but in many cases the former is a more preferred embodiment. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, a magnetic disk, or an optical disk) and includes a number of instructions for enabling a terminal device (which can be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0108] The above are only preferred embodiments of the present invention and are not intended to limit the patent scope of the present invention. Any equivalent structure or equivalent process transformation made using the contents of the present invention description and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present invention.

Claims

1. An echo cancellation method, characterized in that: The method is applied to a communication device including a dual-microphone array, and the method includes: Collecting a near-end mixed signal and a far-end reference signal through the dual-microphone array, and performing spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data; Processing the feature data using a deep neural network model to obtain an estimated echo spectrum; determining a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal; After the step of collecting the near-end mixed signal and the far-end reference signal by the dual-microphone array, the method further includes: monitoring spectrum entropy values ​​of the near-end mixed signal and the far-end reference signal in real time to obtain monitoring results; Determining the reverberation level corresponding to the current environment according to the monitoring results; Adjusting the update step size of the deep neural network model based on the reverberation level corresponding to the current environment; The step of processing the feature data using a deep neural network model to obtain an estimated echo spectrum includes: Constructing historical frame sequence data based on the feature data, and inputting the historical frame sequence data into a deep neural network model; The deep neural network is used to learn the nonlinear mapping relationship between the echo and the pure speech in the frequency spectrum of the historical frame sequence data, and an estimated echo spectrum is output.

2. The echo cancellation method according to claim 1, wherein: The step of collecting a near-end mixed signal and a far-end reference signal through the dual-microphone array, and extracting spectrum features of the near-end mixed signal and the far-end reference signal to obtain feature data includes: Collecting a near-end mixed signal and a far-end reference signal by the dual-microphone array, wherein the near-end mixed signal includes an echo signal and a near-end voice signal; Performing time-frequency conversion processing on the near-end mixed signal and the far-end reference signal using short-time Fourier transform to generate corresponding logarithmic Mel spectrum; Dynamic feature information is extracted from the logarithmic Mel spectrum to obtain feature data.

3. The echo cancellation method according to claim 1, wherein: After the step of determining the near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal, the method further includes: Performing voice activity detection on the near-end clean voice signal, and determining whether there is near-end voice activity in the current environment according to the voice activity detection result; When there is near-end voice activity in the current environment, freezing parameter updates of the deep neural network model; When there is no near-end voice activity in the current environment, background noise is generated and inserted into the silent voice signal.

4. The echo cancellation method according to claim 3, wherein: The step of performing voice activity detection on the near-end clean voice signal and determining whether there is near-end voice activity in the current environment according to the voice activity detection result includes: Performing voice activity detection on the near-end clean voice signal based on an energy threshold and a double-threshold algorithm to obtain a voice activity detection result; Determine whether there is near-end voice activity in the current environment according to the voice activity detection result.

5. The echo cancellation method according to claim 1, wherein: After the step of determining the near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal, the method further includes: Encoding the near-end clean speech signal based on a wideband speech coding standard to obtain an encoded near-end clean speech signal; The encoded near-end clean voice signal is encrypted to achieve encrypted transmission of the encoded near-end clean voice signal.

6. An echo cancellation device, characterized in that: The apparatus is applied to a communication device including a dual-microphone array, and the apparatus includes: a signal acquisition module, configured to acquire a near-end mixed signal and a far-end reference signal through the dual-microphone array, and perform spectrum feature extraction on the near-end mixed signal and the far-end reference signal to obtain feature data; An echo extraction module is used to process the feature data using a deep neural network model to obtain an estimated echo spectrum; an echo cancellation module, configured to determine a near-end clean speech signal based on the estimated echo spectrum and the near-end mixed signal; The signal acquisition module is further configured to monitor the spectrum entropy values ​​of the near-end mixed signal and the far-end reference signal in real time to obtain monitoring results; determine the reverberation level corresponding to the current environment based on the monitoring results; and adjust the update step size of the deep neural network model based on the reverberation level corresponding to the current environment; The echo extraction module is further used to construct historical frame sequence data based on the feature data, and input the historical frame sequence data into a deep neural network model; use the deep neural network to learn the nonlinear mapping relationship between the echo and the pure speech in the frequency spectrum of the historical frame sequence data, and output an estimated echo spectrum.

7. An echo cancellation device, characterized in that: The device includes: a memory, a processor, and an echo cancellation program stored in the memory and executable on the processor, wherein the echo cancellation program is configured to implement the steps of the echo cancellation method according to any one of claims 1 to 5.

8. A storage medium, characterized in that: The storage medium stores an echo cancellation program, which, when executed by a processor, implements the steps of the echo cancellation method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Echo cancellation method and device, audio equipment and storage medium

    CN117727317A

  • Method for eliminating echo of intercom system

    CN119342151A