Echo cancellation method and apparatus, storage medium and electronic device

By acquiring audio and reference signal features in the near-end device, performing feature extraction and frequency division decoding, the problem of poor echo cancellation caused by inconsistent high and low frequency attenuation in the far field is solved, achieving a highly efficient echo cancellation effect.

WO2025260857A1PCT designated stage Publication Date: 2025-12-26SHENZHEN TCL DIGITAL TECH CO LTD
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2025/082467
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-06-19
Filing Date
2025-03-13
Publication Date
2025-12-26

AI Technical Summary

Technical Problem

In the far-field scenario of near-end equipment, existing echo cancellation methods cannot effectively address the problem of inconsistent high and low frequency attenuation, resulting in poor echo cancellation performance.

Method used

By acquiring the features of the audio to be processed and the reference signal, feature extraction and encoding are performed, and the audio is divided into multiple channels for frequency division decoding to obtain decoded audio in different frequency ranges. These decoded audios are then combined to eliminate echo.

Benefits of technology

It improves the echo cancellation effect of near-end devices in the far field under high and low frequency attenuation conditions, and achieves effective echo cancellation, which is suitable for terminal devices such as televisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2025082467_26122025_PF_FP_ABST
    Figure CN2025082467_26122025_PF_FP_ABST
Patent Text Reader

Abstract

An echo cancellation method and apparatus, a storage medium and an electronic device, relating to the technical field of audio processing. The method comprises: acquiring an audio signal feature of an audio to be processed and a reference signal feature of a reference signal corresponding to the audio to be processed (S110); performing feature extraction processing on the basis of the audio signal feature and the reference signal feature, so as to obtain an effective near-end speech feature (S120); dividing the effective near-end speech feature into multiple channels and performing sub-band decoding processing, so as to obtain multiple decoded audios in different frequency ranges (S130); and combining the multiple decoded audios in different frequency ranges, so as to obtain an echo-cancelled audio corresponding to the audio to be processed (S140). The method improves the echo cancellation effect of near-end devices in far fields under high- and low-frequency attenuation.
Need to check novelty before this filing date? Find Prior Art

Description

Echo cancellation methods, apparatus, storage media and electronic equipment

[0001] This application claims priority to Chinese Patent Application No. 202410798599.1, filed on June 19, 2024, entitled "Echo Cancellation Method, Apparatus, Storage Medium and Electronic Device", the entire contents of which are incorporated herein by reference. Technical Field

[0002] This application relates to the field of audio processing technology, specifically to an echo cancellation method, apparatus, storage medium, and electronic device. Background Technology

[0003] Many near-end devices, such as conference machines and televisions, often face the problem of echo cancellation during use. For example, the sound signal played by the external speakers of a television is picked up by the television's microphone after a series of acoustic reflections, and then mixed and played again, resulting in an echo in the sound played by the speakers. Technical issues

[0004] Currently, there are techniques for eliminating echoes in audio by using echo cancellation models. However, current methods estimate the characteristics of the entire frequency band in the audio by end-to-end estimation. But in the case of near-end devices in the far field (e.g., when the microphone of a TV is far from the sound source played by the speaker), the current methods cannot effectively process the echoes because the high and low frequencies attenuate at different rates during the propagation of the sound source. This results in poor echo cancellation performance. Technical solutions

[0005] This application provides a solution that can effectively improve the echo cancellation effect of near-end devices in the far field under high and low frequency attenuation conditions.

[0006] The embodiments of this application provide the following technical solutions:

[0007] According to one embodiment of this application, an echo cancellation method includes: acquiring audio signal features of an audio to be processed and reference signal features of a reference signal corresponding to the audio to be processed; performing feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features; performing frequency division decoding processing on the effective near-end speech features into multiple channels to obtain multiple decoded audios of different frequency ranges; and combining the multiple decoded audios of different frequency ranges to obtain echo-cancelled audio corresponding to the audio to be processed.

[0008] In some embodiments of this application, the step of performing feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features includes: encoding the audio signal features and the reference signal features to obtain audio encoded features; and performing feature extraction processing on the audio encoded features to obtain the effective near-end speech features.

[0009] In some embodiments of this application, the step of performing feature extraction processing on the audio coding features to obtain the effective near-end speech features includes: performing feature analysis and extraction on the audio coding features in multiple paths to obtain multiple effective path speech features; and performing feature interaction rearrangement on the multiple effective path speech features to obtain the effective near-end speech features.

[0010] In some embodiments of this application, the step of encoding the audio signal features and the reference signal features to obtain audio encoded features includes: sequentially encoding the audio signal features and the reference signal features using multiple encoders that are connected in series and have the same structure and parameters to obtain the audio encoded features.

[0011] In some embodiments of this application, the step of dividing the effective near-end speech features into multiple paths for frequency division decoding to obtain multiple decoded audios of different frequency ranges includes: dividing the effective near-end speech features into two paths for frequency division decoding to obtain decoded audios of a first frequency range and a second frequency range, wherein the first frequency range is higher than the second frequency range.

[0012] In some embodiments of this application, the step of dividing the effective near-end speech features into multiple paths for frequency division decoding to obtain multiple decoded audios with different frequency ranges includes: acquiring audio playback scene information; determining the target number of decoding paths based on the audio playback scene information; and dividing the effective near-end speech features into multiple paths for frequency division decoding according to the decoding paths corresponding to the target number of decoding paths to obtain multiple decoded audios with different frequency ranges.

[0013] In some embodiments of this application, the step of obtaining the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed includes: obtaining the audio to be processed and the reference signal corresponding to the audio to be processed; performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0014] In some embodiments of this application, the step of performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal includes: performing linear echo cancellation processing on the audio to be processed and the reference signal using a linear filter to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0015] According to one embodiment of this application, an echo cancellation device includes: a feature acquisition module for acquiring audio signal features of an audio to be processed and reference signal features of a reference signal corresponding to the audio to be processed; a feature extraction module for performing feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features; a frequency division decoding module for dividing the effective near-end speech features into multiple channels for frequency division decoding processing to obtain multiple decoded audios of different frequency ranges; and a combination module for combining the multiple decoded audios of different frequency ranges to obtain echo-cancelled audio corresponding to the audio to be processed.

[0016] In some embodiments of this application, the feature extraction module is configured to: encode the audio signal features and the reference signal features to obtain audio encoded features; and perform feature extraction processing on the audio encoded features to obtain the effective near-end speech features.

[0017] In some embodiments of this application, the feature extraction module is used to: perform feature analysis and extraction on the audio coding features in multiple paths to obtain multiple effective path speech features; and perform feature interaction rearrangement on the multiple effective path speech features to obtain the effective near-end speech features.

[0018] In some embodiments of this application, the feature extraction module is used to: sequentially encode the audio signal features and the reference signal features using multiple encoders connected in series and having the same structure and parameters to obtain the audio encoded features.

[0019] In some embodiments of this application, the frequency division decoding module is used to: divide the effective near-end speech features into two paths for frequency division decoding processing to obtain decoded audio in a first frequency range and decoded audio in a second frequency range, wherein the first frequency range is higher than the second frequency range.

[0020] In some embodiments of this application, the frequency division decoding module is used to: acquire audio playback scene information; determine the target number of decoding paths based on the audio playback scene information; and divide the effective near-end speech features into multiple paths for frequency division decoding processing according to the decoding paths corresponding to the target number of decoding paths, thereby obtaining multiple decoded audio paths with different frequency ranges.

[0021] In some embodiments of this application, the feature acquisition module is used to: acquire the audio to be processed and the reference signal corresponding to the audio to be processed; perform feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0022] In some embodiments of this application, the feature acquisition module is used to: perform linear echo cancellation processing on the audio to be processed and the reference signal using a linear filter to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0023] According to another embodiment of this application, a storage medium stores a computer program thereon, which, when executed by a computer's processor, causes the computer to perform the methods described in the embodiments of this application.

[0024] According to another embodiment of this application, an electronic device may include: a memory storing a computer program; and a processor reading the computer program stored in the memory to execute the methods described in the embodiments of this application.

[0025] According to another embodiment of this application, a computer program product or computer program includes computer instructions stored in a computer-readable storage medium. A processor of a computer device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the computer device to perform the methods provided in the various optional implementations described in the embodiments of this application. Beneficial effects

[0026] In this embodiment, the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed are obtained; feature extraction processing is performed based on the audio signal features and the reference signal features to obtain effective near-end speech features; the effective near-end speech features are divided into multiple paths for frequency division decoding processing to obtain multiple decoded audios of different frequency ranges; the multiple decoded audios of different frequency ranges are combined to obtain the echo cancellation audio corresponding to the audio to be processed.

[0027] In this way, for the audio to be processed, effective near-end speech features are extracted based on the audio signal features of the audio to be processed and the reference signal features of the reference signal. Then, the effective near-end speech features are divided into multiple paths for frequency division decoding. This allows the feature decoding to be divided into multiple branches of different frequencies and decoded separately, resulting in multiple decoded audios of different frequency ranges, which are then used to synthesize the echo-cancelled audio corresponding to the audio to be processed. This can effectively improve the echo cancellation effect of near-end devices in the far field under high and low frequency attenuation conditions. Moreover, this echo cancellation scheme can be easily implemented through a small-scale echo cancellation model, making the echo cancellation model effectively applicable to terminal devices such as televisions. Attached Figure Description

[0028] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0029] Figure 1 shows a flowchart of an echo cancellation method according to an embodiment of this application.

[0030] Figure 2 shows a structural diagram of an echo cancellation model according to an embodiment of this application.

[0031] Figure 3 shows a flowchart of feature interaction rearrangement according to an embodiment of this application.

[0032] Figure 4 shows a block diagram of an echo cancellation device according to an embodiment of this application.

[0033] Figure 5 shows a block diagram of an electronic device according to an embodiment of this application.

[0034] Implementation methods of this application

[0035] The present disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the embodiments provided herein are merely illustrative of the present disclosure and are not intended to limit the present disclosure. Furthermore, the embodiments provided below are some embodiments for implementing the present disclosure, and not all embodiments for implementing the present disclosure. Unless otherwise specified, the technical solutions described in the embodiments of the present disclosure can be implemented in any combination.

[0036] It should be noted that, in the embodiments of this disclosure, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a method or apparatus that includes a list of elements includes not only the elements expressly described, but also other elements not expressly listed, or elements inherent to implementing the method or apparatus. Without further limitations, an element defined by the phrase "comprising a..." does not exclude the presence of other related elements (e.g., steps in the method or units in the apparatus, such as portions of circuitry, processors, programs, or software, etc.) in the method or apparatus that includes that element.

[0037] For example, the echo cancellation method provided in this embodiment includes a series of steps, but the echo cancellation method provided in this embodiment is not limited to the steps described. Similarly, the echo cancellation device provided in this embodiment includes a series of units, but the device provided in this embodiment is not limited to the units explicitly described, but may also include units that need to be set up for obtaining relevant information or processing based on the information.

[0038] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure belongs. The terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting of this disclosure.

[0039] Figure 1 schematically illustrates a flowchart of an echo cancellation method according to an embodiment of this application. The entity performing this echo cancellation method can be any near-end device with processing capabilities, such as a television, computer, and home appliances, etc.

[0040] As shown in Figure 1, the echo cancellation method may include steps S110 to S140.

[0041] Step S110: Obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed;

[0042] Step S120: Perform feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features;

[0043] Step S130: Divide the effective near-end speech features into multiple channels and perform frequency division decoding to obtain multiple decoded audios with different frequency ranges;

[0044] Step S140: Combine the decoded audio from multiple channels with different frequency ranges to obtain the echo-cancelled audio corresponding to the audio to be processed.

[0045] The audio to be processed is the audio containing the echo signal. For example, the audio to be processed could be the audio to be transmitted to a speaker in a near-end device for playback. The reference signal is the signal used to reference the cancellation of the echo. For example, the reference signal corresponding to the audio to be processed could be the audio signal played by the speaker some time before the audio to be processed is played. The audio to be processed and the reference signal can be obtained from the audio playback system of the near-end device.

[0046] The audio signal characteristics of the audio to be processed are the signal features used to characterize the audio to be processed. These audio signal characteristics can be amplitude spectrum or other features. The reference signal characteristics of the reference signal are the signal features used to characterize the reference signal. These reference signal characteristics can also be amplitude spectrum or other features.

[0047] By employing a neural network module for feature extraction, feature extraction processing can be performed based on audio signal features and reference signal features to extract effective near-end speech features, which are the audio features of the extracted audio that do not contain echo signals.

[0048] Furthermore, the effective near-end speech features are divided into multiple paths for frequency-division decoding, resulting in multiple decoded audio streams with different frequency ranges. The combined frequency range of these multiple streams constitutes a complete frequency range. For example, the effective near-end speech features are decoded in the first path to obtain the decoded audio stream of the first frequency range, and the effective near-end speech features are decoded in the second path to obtain the decoded audio stream of the second frequency range. Here, the first frequency range can be a high-frequency range, and the second frequency range can be a low-frequency range.

[0049] Then, by combining the decoded audio from multiple channels with different frequency ranges, a complete frequency range of audio can be obtained, which is the echo-cancelled audio corresponding to the audio to be processed. For example, by combining the decoded audio from the first frequency range and the decoded audio from the second frequency range, the echo-cancelled audio corresponding to the audio to be processed can be obtained. The echo-cancelled audio corresponding to the audio to be processed is the audio after the echo signal in the audio to be processed has been eliminated.

[0050] In this way, based on steps S110 to S140, for the audio to be processed, effective near-end speech features are extracted based on the audio signal features of the audio to be processed and the reference signal features of the reference signal. Then, the effective near-end speech features are divided into multiple paths for frequency division decoding. This allows the feature decoding to be divided into multiple branches of different frequencies and decoded separately, resulting in multiple decoded audios of different frequency ranges to synthesize the echo-cancelled audio corresponding to the audio to be processed. This can effectively improve the echo cancellation effect under high and low frequency attenuation in the far field of near-end devices. Moreover, this echo cancellation scheme can be easily implemented through a small-scale echo cancellation model, making the echo cancellation model effectively implemented in terminal devices such as televisions.

[0051] The following describes further optional embodiments of the steps performed during echo cancellation in the embodiment shown in Figure 1.

[0052] In one embodiment, acquiring the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed may include: acquiring the audio to be processed and the reference signal corresponding to the audio to be processed; performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0053] After acquiring the audio to be processed and its corresponding reference signal, feature extraction is performed on both to obtain audio signal features characterizing the audio to be processed and reference signal features characterizing the reference signal. These audio signal features can be amplitude spectrum or other features, and the reference signal features can also be amplitude spectrum or other features.

[0054] In one embodiment, feature extraction processing is performed on the audio to be processed and the reference signal to obtain audio signal features of the audio to be processed and reference signal features of the reference signal. Specifically, this may involve: preprocessing the audio to be processed to obtain preprocessed audio; performing a Fourier transform on the preprocessed audio to obtain first spectral data; calculating the amplitude spectrum of the first spectral data to obtain a first amplitude spectrum; extracting features from the first amplitude spectrum to obtain audio signal features of the audio to be processed; preprocessing the reference signal to obtain a preprocessed signal; performing a Fourier transform on the preprocessed signal to obtain second spectral data; calculating the amplitude spectrum of the second spectral data to obtain a second amplitude spectrum; and extracting features from the second amplitude spectrum to obtain reference signal features of the reference signal.

[0055] Preprocessing can include downsampling, windowing, and framing. Fourier transform can convert the preprocessed audio or signal from the time domain to the frequency domain, yielding first and second frequency spectrum data. Since the first and second frequency spectrum data are complex arrays, calculating their absolute values ​​or moduli yields the first and second amplitude spectra. Then, features such as the first moment, second moment, dominant bandwidth, and dominant frequency are extracted from the amplitude spectrum to serve as effective audio signal features characterizing the audio to be processed and reference signal features characterizing the reference signal.

[0056] Furthermore, in one embodiment, the step of performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal may include: performing linear echo cancellation processing on the audio to be processed and the reference signal using a linear filter to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0057] A linear filter is a filter that eliminates linear echo signals through adaptive filtering. The linear filter can be a block frequency domain adaptive filter or a dual filter based on the proportionally normalized least mean square algorithm, etc.

[0058] Linear filters can be used to perform linear echo cancellation on the audio to be processed and the reference signal, eliminating the linear echo signal in the audio to be processed. After linear echo cancellation, the linear filter can output the "audio signal characteristics of the audio to be processed after linear echo cancellation" and the "reference signal characteristics of the reference signal".

[0059] The "audio signal characteristics of the audio to be processed after linear echo cancellation processing" and the "reference signal characteristics of the reference signal" have already eliminated the linear echo signal. They can be further used to perform echo cancellation through the nonlinear echo cancellation method in this application embodiment, which can effectively eliminate the residual nonlinear echo signal, thereby further improving the echo cancellation effect.

[0060] In one embodiment, the step of performing feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features may include: encoding the audio signal features and the reference signal features to obtain audio encoded features; and performing feature extraction processing on the audio encoded features to obtain the effective near-end speech features.

[0061] An encoder can encode the audio signal features and the reference signal features. This encoding process can compress and reduce the dimensions of the audio signal features and the reference signal features, resulting in compressed and reduced audio coded features. These compressed and reduced audio coded features include both the compressed and reduced audio signal features and the reference signal features.

[0062] Furthermore, through the effective near-end audio feature extraction module, feature extraction processing can be performed on the audio coding features to extract effective near-end speech features, which are the audio coding features after eliminating echo signal features (such as the signal features of nonlinear echo signals).

[0063] In one embodiment, encoding the audio signal features and the reference signal features to obtain audio encoded features may include: sequentially encoding the audio signal features and the reference signal features using multiple encoders connected in series and having the same structure and parameters to obtain the audio encoded features.

[0064] Each encoder can be a neural network module for feature encoding, and each encoder can include multiple layers, such as convolutional layers, fully connected layers, pooling layers, etc. Multiple encoders are connected in series to form an encoding module, and multiple encoders have the same structure and parameters. For example, as shown in Figure 2, in one example, encoding module 210 can include two encoders 211 and 212 connected in series, and encoders 211 and 212 have the same structure and parameters.

[0065] In the audio signal feature and reference signal feature input encoding module, multiple encoders can sequentially encode the audio signal features and reference signal features to obtain audio coded features. Since multiple encoders have the same structure and parameters, cascading encoding can effectively compress and reduce the dimensionality of features with less computation.

[0066] It is understood that in some other embodiments, the process of encoding the audio signal features and the reference signal features to obtain audio encoded features may include: encoding the audio signal features and the reference signal features sequentially using multiple encoders that are connected in series and do not have the same structure and parameters to obtain the audio encoded features.

[0067] In one embodiment, referring to FIG3, the feature extraction processing of the audio coding features to obtain the effective near-end speech features may include:

[0068] Step S310: Perform feature analysis and extraction on the audio coding features in multiple channels to obtain effective speech features in multiple channels; Step S320: Perform feature interaction rearrangement on the effective speech features in multiple channels to obtain the effective near-end speech features.

[0069] The effective near-end audio feature extraction module can include multiple analyzers arranged in parallel and a feature rearrangement unit. Each analyzer is a neural network module, which can also contain multiple layers. Each analyzer can perform feature analysis and extraction on the audio encoded features, extracting the effective segmented speech features after eliminating echo signal features. The feature rearrangement unit can also be a neural network module, which can perform feature rearrangement on multiple effective segmented speech features to obtain effective near-end speech features.

[0070] For example, referring to Figure 2, in one example, the effective near-end audio feature extraction module 220 may include two analyzers 221 and 222 arranged in parallel, and a feature interaction rearranger 223 connected to the analyzers 221 and 222. The effective split-path speech features extracted by the analyzers 221 and 222 are input to the feature interaction rearranger 223. The feature interaction rearranger 223 can perform feature interaction rearrangement on the two effective split-path speech features to obtain effective near-end speech features.

[0071] In one embodiment, the step of dividing the effective near-end speech features into multiple paths for frequency division decoding to obtain multiple decoded audio streams with different frequency ranges may include:

[0072] The effective near-end speech features are divided into two paths for frequency division decoding to obtain decoded audio in a first frequency range and decoded audio in a second frequency range, wherein the first frequency range is higher than the second frequency range.

[0073] A decoder can be used to decode effective near-end speech features. By setting two decoders with different frequency ranges, effective near-end speech features can be decoded separately (i.e., frequency division decoding). The decoder can be a neural network module.

[0074] Specifically, as shown in Figure 2, the frequency division decoding module 230 can be equipped with two decoders 231 and 232. The effective near-end speech features are decoded in the first decoder 231 to obtain the decoded audio of the first frequency range, and the effective near-end speech features are decoded in the second decoder 232 to obtain the decoded audio of the second frequency range. The first frequency range is higher than the second frequency range. The first frequency range can be a high frequency range and the second frequency range can be a low frequency range, thereby realizing high and low frequency group decoding.

[0075] By using a two-way high- and low-frequency group decoding method, the inconsistent attenuation of high and low frequencies during sound source propagation can be effectively addressed, thus improving the echo cancellation effect.

[0076] Furthermore, in one embodiment, the step of dividing the effective near-end speech features into multiple paths for frequency division decoding to obtain multiple decoded audio streams with different frequency ranges may include:

[0077] Obtain audio playback scene information; determine the target number of decoding paths based on the audio playback scene information; divide the effective near-end speech features into multiple paths for frequency division decoding according to the decoding paths corresponding to the target number of decoding paths, and obtain multiple decoded audios with different frequency ranges.

[0078] The audio playback scene information may specifically include the size of the audio playback space, the distance between the sound source and the microphone (e.g., the distance between the external speakers of the TV and the microphone in the TV), and other scene information. The audio playback scene information may be pre-input into the near-end device or dynamically input into the near-end device by the user.

[0079] The preset path lookup table can be used to retrieve the predetermined path corresponding to the audio playback scenario information as the target decoding path. Therefore, the effective near-end speech features can be decoded separately using decoders with different frequency ranges corresponding to the target decoding path, resulting in multiple decoded audio streams with different frequency ranges. Frequency division decoding based on the decoding path matched to the audio playback scenario information makes the decoding process more consistent with the audio playback scenario, further improving the overall echo cancellation effect.

[0080] Furthermore, in one example of this application, echo cancellation can be performed based on the aforementioned embodiments of this application using the overall echo cancellation model shown in Figure 2. The echo cancellation model shown in Figure 2 can be implemented as a small-scale echo cancellation model, which can be effectively implemented in terminal devices such as televisions and effectively improve the echo cancellation effect.

[0081] To facilitate better implementation of the echo cancellation method provided in the embodiments of this application, the embodiments of this application also provide an echo cancellation device based on the above-described echo cancellation method. The meanings of the terms used are the same as in the echo cancellation method described above, and specific implementation details can be found in the description of the method embodiments. Figure 4 shows a block diagram of an echo cancellation device according to an embodiment of this application.

[0082] As shown in Figure 4, the echo cancellation device 400 may include: a feature acquisition module 410, used to acquire the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed; a feature extraction module 420, used to perform feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features; a frequency division decoding module 430, used to divide the effective near-end speech features into multiple channels for frequency division decoding processing to obtain multiple decoded audios of different frequency ranges; and a combination module 440, used to combine the multiple decoded audios of different frequency ranges to obtain the echo-cancelled audio corresponding to the audio to be processed.

[0083] In some embodiments of this application, the feature extraction module is configured to: encode the audio signal features and the reference signal features to obtain audio encoded features; and perform feature extraction processing on the audio encoded features to obtain the effective near-end speech features.

[0084] In some embodiments of this application, the feature extraction module is used to: perform feature analysis and extraction on the audio coding features in multiple paths to obtain multiple effective path speech features; and perform feature interaction rearrangement on the multiple effective path speech features to obtain the effective near-end speech features.

[0085] In some embodiments of this application, the feature extraction module is used to: sequentially encode the audio signal features and the reference signal features using multiple encoders connected in series and having the same structure and parameters to obtain the audio encoded features.

[0086] In some embodiments of this application, the frequency division decoding module is used to: divide the effective near-end speech features into two paths for frequency division decoding processing to obtain decoded audio in a first frequency range and decoded audio in a second frequency range, wherein the first frequency range is higher than the second frequency range.

[0087] In some embodiments of this application, the frequency division decoding module is used to: acquire audio playback scene information; determine the target number of decoding paths based on the audio playback scene information; and divide the effective near-end speech features into multiple paths for frequency division decoding processing according to the decoding paths corresponding to the target number of decoding paths, thereby obtaining multiple decoded audio paths with different frequency ranges.

[0088] In some embodiments of this application, the feature acquisition module is used to: acquire the audio to be processed and the reference signal corresponding to the audio to be processed; perform feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0089] In some embodiments of this application, the feature acquisition module is used to: perform linear echo cancellation processing on the audio to be processed and the reference signal using a linear filter to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0090] It should be noted that although several modules or units for the device used to perform actions have been mentioned in the detailed description above, this division is not mandatory. In fact, according to the embodiments of this application, the features and functions of two or more modules or units described above can be embodied in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided and embodied by multiple modules or units.

[0091] Furthermore, this application also provides an electronic device, as shown in FIG5. FIG5 shows a block diagram of an electronic device according to an embodiment of this application, specifically:

[0092] The electronic device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that the electronic device structure shown in FIG. 5 does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0093] The processor 501 is the control center of the electronic device. It connects to various parts of the computer device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, it performs various functions of the computer device and processes data, thereby providing overall monitoring of the electronic device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user page, and application programs, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 501.

[0094] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and data processing by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the computer device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.

[0095] The electronic device also includes a power supply 503 that supplies power to various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0096] The electronic device may also include an input unit 504, which can be used to receive input digital or character information and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0097] Although not shown, the electronic device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the electronic device loads the executable files corresponding to the processes of one or more computer programs into the memory 502 according to the following instructions, and the processor 501 runs the computer programs stored in the memory 502, thereby realizing the various functions in the foregoing embodiments of this application. For example, the processor 501 can perform the following steps:

[0098] The process involves: acquiring the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed; performing feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features; dividing the effective near-end speech features into multiple channels for frequency division decoding processing to obtain multiple decoded audios of different frequency ranges; and combining the multiple decoded audios of different frequency ranges to obtain the echo-cancelled audio corresponding to the audio to be processed.

[0099] In some embodiments of this application, the step of performing feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features includes: encoding the audio signal features and the reference signal features to obtain audio encoded features; and performing feature extraction processing on the audio encoded features to obtain the effective near-end speech features.

[0100] In some embodiments of this application, the step of performing feature extraction processing on the audio coding features to obtain the effective near-end speech features includes: performing feature analysis and extraction on the audio coding features in multiple paths to obtain multiple effective path speech features; and performing feature interaction rearrangement on the multiple effective path speech features to obtain the effective near-end speech features.

[0101] In some embodiments of this application, the step of encoding the audio signal features and the reference signal features to obtain audio encoded features includes: sequentially encoding the audio signal features and the reference signal features using multiple encoders that are connected in series and have the same structure and parameters to obtain the audio encoded features.

[0102] In some embodiments of this application, the step of dividing the effective near-end speech features into multiple paths for frequency division decoding to obtain multiple decoded audios of different frequency ranges includes: dividing the effective near-end speech features into two paths for frequency division decoding to obtain decoded audios of a first frequency range and a second frequency range, wherein the first frequency range is higher than the second frequency range.

[0103] In some embodiments of this application, the step of dividing the effective near-end speech features into multiple paths for frequency division decoding to obtain multiple decoded audios with different frequency ranges includes: acquiring audio playback scene information; determining the target number of decoding paths based on the audio playback scene information; and dividing the effective near-end speech features into multiple paths for frequency division decoding according to the decoding paths corresponding to the target number of decoding paths to obtain multiple decoded audios with different frequency ranges.

[0104] In some embodiments of this application, the step of obtaining the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed includes: obtaining the audio to be processed and the reference signal corresponding to the audio to be processed; performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0105] In some embodiments of this application, the step of performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal includes: performing linear echo cancellation processing on the audio to be processed and the reference signal using a linear filter to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

[0106] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by a computer program, or by a computer program controlling related hardware. The computer program can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0107] Therefore, embodiments of this application also provide a storage medium storing a computer program that can be loaded by a processor to execute the steps in any of the methods provided in embodiments of this application.

[0108] The storage medium can be a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), disk or optical disk, etc.

[0109] Since the computer program stored in the storage medium can execute the steps of any of the methods provided in the embodiments of this application, the beneficial effects that the methods provided in the embodiments of this application can achieve can be realized. For details, please refer to the previous embodiments, which will not be repeated here.

[0110] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the embodiments disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein.

[0111] It should be understood that this application is not limited to the embodiments described above and shown in the accompanying drawings, but various modifications and changes can be made without departing from its scope.

Claims

1. An echo cancellation method, wherein, include: Acquire the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed; Based on the audio signal features and the reference signal features, feature extraction processing is performed to obtain effective near-end speech features; The effective near-end speech features are divided into multiple paths and frequency-division decoding processes to obtain multiple decoded audios with different frequency ranges; The decoded audio from multiple channels with different frequency ranges is combined to obtain the echo-cancelled audio corresponding to the audio to be processed.

2. The method according to claim 1, wherein, The feature extraction process based on the audio signal features and the reference signal features to obtain effective near-end speech features includes: The audio signal features and the reference signal features are encoded to obtain audio encoded features; The audio coding features are subjected to feature extraction processing to obtain the effective near-end speech features.

3. The method according to claim 2, wherein, The step of performing feature extraction processing on the audio coding features to obtain the effective near-end speech features includes: The audio coding features are analyzed and extracted from multiple paths to obtain effective multi-path speech features; The effective near-end speech features are obtained by performing feature interaction rearrangement on the effective split-path speech features of the multiple paths.

4. The method according to claim 2, wherein, The step of encoding the audio signal features and the reference signal features to obtain audio encoded features includes: The audio signal features and the reference signal features are sequentially encoded using multiple encoders connected in series with the same structure and parameters to obtain the audio encoded features.

5. The method according to claim 1, wherein, The step of dividing the effective near-end speech features into multiple paths and performing frequency division decoding to obtain multiple decoded audio streams with different frequency ranges includes: The effective near-end speech features are divided into two paths for frequency division decoding to obtain decoded audio in a first frequency range and decoded audio in a second frequency range, wherein the first frequency range is higher than the second frequency range.

6. The method according to claim 1, wherein, The step of dividing the effective near-end speech features into multiple paths and performing frequency division decoding to obtain multiple decoded audio streams with different frequency ranges includes: Obtain audio playback scene information; Based on the audio playback scenario information, determine the target number of decoding paths; According to the decoding path corresponding to the target decoding path, the effective near-end speech features are divided into multiple paths for frequency division decoding processing to obtain multiple decoded audios with different frequency ranges.

7. The method according to claim 1, wherein, The acquisition of the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed includes: Acquire the audio to be processed and the reference signal corresponding to the audio to be processed; Feature extraction processing is performed on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

8. The method according to claim 7, wherein, The step of performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal includes: A linear filter is used to perform linear echo cancellation processing on the audio to be processed and the reference signal to obtain the audio signal characteristics of the audio to be processed and the reference signal characteristics of the reference signal.

9. The method according to claim 7, wherein, The step of performing feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal includes: The audio to be processed is preprocessed to obtain preprocessed audio; The preprocessed audio is subjected to Fourier transform to obtain the first spectrum data; The amplitude spectrum is calculated from the first spectral data to obtain the first amplitude spectrum; Extract the features of the first amplitude spectrum to obtain the audio signal features of the audio to be processed; The reference signal is preprocessed to obtain a preprocessed signal; The preprocessed signal is subjected to Fourier transform to obtain the second spectrum data; The amplitude spectrum is calculated from the second spectral data to obtain the second amplitude spectrum; The features of the second amplitude spectrum are extracted to obtain the reference signal features of the reference signal.

10. The method according to claim 2, wherein, The step of encoding the audio signal features and the reference signal features to obtain audio encoded features includes: The audio signal features and the reference signal features are sequentially encoded using multiple encoders that are cascaded together and have different structures and parameters to obtain the audio encoded features.

11. The method according to claim 6, wherein, Determining the target decoding path based on the audio playback scene information includes: The predetermined path corresponding to the audio playback scene information is retrieved from the preset path lookup table as the target decoding path.

12. An echo cancellation device, wherein, include: The feature acquisition module is used to acquire the audio signal features of the audio to be processed and the reference signal features of the reference signal corresponding to the audio to be processed; The feature extraction module is used to perform feature extraction processing based on the audio signal features and the reference signal features to obtain effective near-end speech features; The frequency division decoding module is used to divide the effective near-end speech features into multiple paths for frequency division decoding processing to obtain multiple decoded audios with different frequency ranges. The combination module is used to combine the decoded audio from multiple channels with different frequency ranges to obtain the echo-cancelled audio corresponding to the audio to be processed.

13. The apparatus according to claim 12, wherein, The feature extraction module is used to: encode the audio signal features and the reference signal features to obtain audio encoded features; and perform feature extraction processing on the audio encoded features to obtain the effective near-end speech features.

14. The apparatus according to claim 13, wherein, The feature extraction module is used to: perform feature analysis and extraction on the audio coding features in multiple channels to obtain multiple effective channel speech features; and perform feature interaction rearrangement on the multiple effective channel speech features to obtain the effective near-end speech features.

15. The apparatus according to claim 13, wherein, The feature extraction module is used to: sequentially encode the audio signal features and the reference signal features using multiple encoders connected in series and having the same structure and parameters, to obtain the audio encoded features.

16. The apparatus according to claim 12, wherein, The frequency division decoding module is used to: divide the effective near-end speech features into two paths for frequency division decoding processing to obtain decoded audio in a first frequency range and decoded audio in a second frequency range, wherein the first frequency range is higher than the second frequency range.

17. The apparatus according to claim 12, wherein, The frequency division decoding module is used to: acquire audio playback scene information; determine the target number of decoding paths based on the audio playback scene information; and divide the effective near-end speech features into multiple paths for frequency division decoding processing according to the decoding paths corresponding to the target number of decoding paths, thereby obtaining multiple decoded audio paths with different frequency ranges.

18. The apparatus according to claim 12, wherein, The feature acquisition module is used to: acquire the audio to be processed and the reference signal corresponding to the audio to be processed; perform feature extraction processing on the audio to be processed and the reference signal to obtain the audio signal features of the audio to be processed and the reference signal features of the reference signal.

19. A storage medium, wherein, It stores a computer program that, when executed by the computer's processor, causes the computer to perform the method described in any one of claims 1 to 11.

20. An electronic device, wherein, include: Memory, which stores computer programs; A processor reads a computer program stored in memory to perform the method described in any one of claims 1 to 11.

Citation Information

Patent Citations

  • Echo cancellation method and device, storage medium and computer equipment

    CN110176244A

  • Voice processing method and device and voice processing model generation method and device

    CN112466318A

  • Single-channel voice echo cancellation method and device based on deep neural network

    CN115565543A

  • Echo cancellation method and device, storage medium and electronic equipment

    CN118737177A

  • System and method for canceling acoustic echo in an audio conferencing communication system

    JP2010507105A