Echo cancellation apparatus, method, computer device and storage medium

CN115631761BActive Publication Date: 2026-09-15BEIJING ESWIN COMPUTING TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202211275333.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-18
Publication Date
2026-09-15
Estimated Expiration
2042-10-18

AI Technical Summary

Technical Problem

然而,现有的回声消除装置在进行回声消除时并未考虑房间内的混响程度,从而导致回声装置对房间的回声抑制效果不佳,产生回声残余较大等副作用

Benefits of technology

[0017] The echo cancellation device provided in this application, through a filter configuration module, pre-configures the length of an adaptive filter to a pre-configured length that matches the reverberation level of the target room. After acquiring the original speech signal from the target room through a speech acquisition module, the echo cancellation module uses the pre-configured length adaptive filter to filter the original speech signal, removing echoes to obtain the target speech signal. Using this method, the adaptive filter can be tuned to a length matching the reverberation time of the target room. Therefore, when the tuned adaptive filter is used in the target room, it can accurately simulate the echo transmission characteristics of the target room, filtering out echoes from the original speech signal, achieving better echo suppression performance, avoiding side effects, and improving the accuracy of echo cancellation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115631761B_ABST
    Figure CN115631761B_ABST
Patent Text Reader

Abstract

The application provides an echo cancellation device and method, a computer device and a storage medium, and relates to the technical field of audio processing. By pre-configuring the length of an adaptive filter as a pre-configured length matching the reverberation degree of a target room, when an original voice signal of the target room is collected, the original voice signal can be filtered by the adaptive filter with the pre-configured length to filter out echoes to obtain a target voice signal. By the method, the adaptive filter can be debugged to a length matching the reverberation time of the target room, so that when the adaptive filter is used in the target room, the adaptive filter can accurately simulate the echo transmission characteristics of the target room, filter out echoes in the original voice signal, achieve better echo suppression performance, avoid the generation of side effects, and improve the accuracy of echo cancellation by the adaptive filter.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and specifically to an echo cancellation device, method, computer equipment, and storage medium. Background Technology

[0002] Acoustic echo refers to the sound emitted by a speaker being picked up by a microphone at the same time that the receiver hears it. For example, in a multimedia classroom setting, the signal collected by the indoor microphones sometimes includes not only the teacher's voice but also the sound emitted by the speaker. Therefore, echo cancellation technology is needed to remove the sound from the speaker.

[0003] In related technologies, adaptive linear filtering is mainly used for echo cancellation. For example, an adaptive filter estimates the transfer function from the speaker to the microphone based on the audio signal collected by the microphone and the speaker reference signal, thereby canceling the echo generated by the speaker and achieving the purpose of eliminating the echo. However, existing echo cancellation devices do not consider the reverberation level in the room when performing echo cancellation, resulting in poor echo suppression effect and side effects such as large echo residue. Summary of the Invention

[0004] This application provides an echo cancellation device, method, computer equipment, and storage medium. The technical solution is as follows:

[0005] On one hand, an echo cancellation device is provided, the device comprising:

[0006] A filter configuration module is used to pre-configure the length of the adaptive filter to a pre-configured length that matches the reverberation level of the target room;

[0007] The voice acquisition module is used to acquire voice data from target sound sources in the target room and obtain the original voice signal.

[0008] An echo cancellation module is used to filter the original speech signal using an adaptive filter of a pre-configured length to obtain the target speech signal of the target sound source. The adaptive filter of the pre-configured length is used to simulate the echo transmission characteristics of the target room to filter out the echo in the original speech signal.

[0009] On the other hand, an echo cancellation method is provided, the method comprising:

[0010] The length of the adaptive filter is pre-configured to match the reverberation level of the target room;

[0011] Speech recordings are performed on the target sound source in the target room to obtain the raw speech signal;

[0012] The original speech signal is filtered by an adaptive filter of a pre-configured length to obtain the target speech signal of the target sound source. The adaptive filter of the pre-configured length is used to simulate the echo transmission characteristics of the target room to filter out the echo in the original speech signal.

[0013] On the other hand, a computer device is provided, including a memory, a processor, and a computer program stored in the memory, wherein the processor executes the computer program to implement the echo cancellation method described above.

[0014] On the other hand, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, implements the above-described echo cancellation method.

[0015] On the other hand, a computer program product is provided, including a computer program that, when executed by a processor, implements the above-described echo cancellation method.

[0016] The beneficial effects of the technical solutions provided in this application are:

[0017] The echo cancellation device provided in this application, through a filter configuration module, pre-configures the length of an adaptive filter to a pre-configured length that matches the reverberation level of the target room. After acquiring the original speech signal from the target room through a speech acquisition module, the echo cancellation module uses the pre-configured length adaptive filter to filter the original speech signal, removing echoes to obtain the target speech signal. Using this method, the adaptive filter can be tuned to a length matching the reverberation time of the target room. Therefore, when the tuned adaptive filter is used in the target room, it can accurately simulate the echo transmission characteristics of the target room, filtering out echoes from the original speech signal, achieving better echo suppression performance, avoiding side effects, and improving the accuracy of echo cancellation. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments of this application will be briefly introduced below.

[0019] Figure 1 This is a schematic diagram of the structure of an echo cancellation device provided in an embodiment of this application;

[0020] Figure 2a This is a schematic diagram of the structure of an echo cancellation device provided in an embodiment of this application;

[0021] Figure 2b A flowchart illustrating a method for obtaining the pre-configuration length of an adaptive filter, provided in an embodiment of this application;

[0022] Figure 3 A schematic diagram of an echo cancellation principle provided in an embodiment of this application;

[0023] Figure 4 A flowchart illustrating an echo cancellation method provided in an embodiment of this application;

[0024] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation

[0025] The embodiments of this application are described below with reference to the accompanying drawings. It should be understood that the embodiments described below with reference to the accompanying drawings are exemplary descriptions for explaining the technical solutions of the embodiments of this application, and do not constitute a limitation on the technical solutions of the embodiments of this application.

[0026] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. It should be further understood that the terms “comprising” and “including” as used in embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, operation, element, and / or component, but do not exclude implementation as other features, information, data, step, operation, element, component, and / or combinations thereof supported by the art. It should be understood that when we say that an element is “connected” or “coupled” to another element, the one element can be directly connected or coupled to the other element, or it can mean that the one element and the other element establish a connection relationship through an intermediate element. Furthermore, “connected” or “coupled” as used herein can include wireless connection or wireless coupling. The term “and / or” as used herein indicates at least one of the items defined by the term; for example, “A and / or B” can be implemented as “A,” or as “B,” or as “A and B.”

[0027] Figure 1 This is a schematic diagram of the structure of an echo cancellation device provided in an embodiment of this application. Figure 1 As shown, the device includes: a filter configuration module 101, a voice acquisition module 102, and an echo cancellation module 103.

[0028] The filter configuration module 101 is used to pre-configure the length of the adaptive filter to a pre-configured length that matches the reverberation level of the target room.

[0029] The voice acquisition module 102 is used to acquire voice data from the target sound source in the target room and obtain the original voice signal.

[0030] The echo cancellation module 103 is used to filter the original speech signal through an adaptive filter of a pre-configured length to obtain the target speech signal of the target sound source. The adaptive filter is used to simulate the echo transmission characteristics of the target room to filter out the echo in the original speech signal.

[0031] The echo transmission characteristic refers to the characteristics of the echo transmission path from the speaker to the microphone in the target room. For example, the voice acquisition module 102 is specifically used to acquire voice data from the target room using the microphone to obtain the original voice signal; the echo cancellation module 103 is specifically used to determine the analog echo signal using an adaptive filter of pre-configured length, and to cancel the analog echo signal from the original voice signal to obtain the target voice signal.

[0032] A microphone can be used to collect audio signals from a target room and then play them through a speaker. The microphone is used to collect audio signals from the target room, and the speaker is used to play the collected audio signals. For example, the microphone and speaker can be located inside the target room; for example, in a smart classroom scenario, a multimedia classroom can be equipped with multimedia equipment, including microphones and speakers; of course, multimedia equipment can also include computers, projectors, etc.; the microphone in the room collects the teacher's voice and plays it through the speaker; however, the audio signal collected by the microphone sometimes includes not only the teacher's voice but also the sound played by the speaker, therefore, the echo cancellation method of this application can be used to remove the sound from the speaker.

[0033] The computer device housing the adaptive filter can be a device that is not moved during use; after the adaptive filter is properly adjusted according to the method for obtaining the pre-configured length of the adaptive filter in this application, the position of the computer device does not need to be moved. The computer device can be a device equipped with a microphone and speakers. For example, the computer device can be a smart microphone, or a multimedia teaching device in a multimedia classroom, an online conferencing system, a digital broadcast receiver, a desktop computer, or a vehicle-mounted terminal (such as a vehicle navigation terminal, vehicle-mounted computer, etc.) with audio acquisition and playback functions; of course, the computer device can also be a server with audio acquisition and playback functions.

[0034] For example, the target speech signal can be a speech signal emitted by a target sound source, such as the teacher's voice in a multimedia classroom setting. The simulated echo signal is used to simulate an echo signal propagating along the echo transmission path of the target room. This simulated echo signal can be an echo signal transmitted from a speaker to a microphone, with the speaker as the sound source. The original speech signal is the speech signal actually captured by the microphone, including the speech signal from the target sound source and the echo signal transmitted from the speaker to the microphone. It should be noted that the microphone's intended goal is to capture the sound from the target sound source (such as the teacher). However, since the speaker also emits sound, the microphone easily captures sound from the speaker, resulting in the teacher's and speaker's sounds being superimposed in the actual audio signal captured by the microphone.

[0035] In this application, an adaptive filter can be used to simulate the acoustic transmission characteristics from the speaker to the microphone in the target room, generating a simulated echo signal from the speaker to the microphone, which is the echo to be eliminated. In one possible example, the simulated echo signal can be generated using the adaptive filter and the far-end speech signal input to the speaker, and the simulated echo signal can be subtracted from the original speech signal to obtain the target speech signal. Here, the far-end speech signal is the signal that needs to be played through the speaker. For example, in a scenario where user A and user B are making a phone call, user A's voice is transmitted to user B's mobile phone and played through the speaker on user B's mobile phone; another example is a multimedia classroom scenario where the microphone captures the teacher's voice and plays it through the speaker.

[0036] The distant speech signal can include speech signals from multiple moments. The analog echo signal for each moment can be calculated using multiple frames of speech signal corresponding to each moment. These multiple frames of speech signal include the speech signal at that moment and the speech signals from all moments preceding it. For example, the process of generating an analog echo signal and obtaining the target speech signal using an adaptive filter and the distant speech signal can include: when echo cancellation is required on the speech signal at the current moment, the adaptive filter can combine a single frame of speech signal at the current moment with multiple frames of speech signal preceding it to calculate the analog echo signal at the current moment; then, the analog echo signal at the current moment is subtracted from the single frame of speech signal at the current moment in the original speech signal to obtain the target speech signal at the current moment. This process is repeated to obtain the target speech signals at each moment.

[0037] It should be noted that because the room impulse response (RIR) is generally quite long, adaptive filters often need to be of the order of thousands to achieve good results, leading to excessive processing delays. To reduce processing delays, a block-based frequency domain adaptive algorithm can be used for echo cancellation. Block-based processing involves dividing the adaptive filter into multiple blocks, each block processing one frame of the speech signal. For example, each time step corresponds to multiple frames of speech signal, and a certain block processes a past frame from those multiple frames.

[0038] However, the specific number of blocks to divide the adaptive filter into, i.e., the length of the adaptive filter, needs to be determined based on the actual situation. Different rooms have different reverberation levels. Reverberation level is usually measured by reverberation time T60, defined as the time required to reduce reverberation by 60 dB. For example, in a multimedia classroom, such as an auditorium, with a large reverberation time (T60 reaching 0.8–2 seconds), using a shorter adaptive filter will result in a larger echo residue. Conversely, in a room with low reverberation, using a longer adaptive filter will also lead to a decrease in echo suppression performance, or even side effects. Therefore, how to configure the length of the adaptive filter is one of the urgent problems to be solved, and there is relatively little research on this aspect in this field.

[0039] This application provides a method for obtaining the pre-configured length of an adaptive filter. This method can be executed through a filter configuration module to pre-configure the length of the adaptive filter to a pre-configured length that matches the reverberation level of the target room. The following describes the method in conjunction with... Figure 2a The device structure shown and Figure 2b The following flowchart illustrates the method:

[0040] Figure 2a This is a schematic diagram of the structure of an echo cancellation device provided in an embodiment of this application. Figure 2a As shown, the filter configuration module in the echo cancellation device may include: an audio signal acquisition unit 1011, a filtering unit 1012, a fitting unit 1013, a reverberation time determination unit 1014, and a length mapping unit 1015. Figure 2b This is a schematic diagram illustrating the process of obtaining the pre-configured length of an adaptive filter, as provided in an embodiment of this application. Figure 2b As shown, the method includes steps 201 to 205. Steps 201 to 205 can be implemented by various units in the filter configuration module.

[0041] The filter configuration module includes an audio signal acquisition unit 1011, which is used to implement the following step 201:

[0042] Step 201: Play the debugging audio in the target room and collect the audio signal in the target room during the playback process to obtain the collected audio signal.

[0043] The acquired audio signal includes frequency domain acquisition signals at multiple times. Each time-domain acquisition signal includes multiple frequency points. The target audio can be a debugging audio segment used to adjust the adaptive filter; for example, the target audio is white noise debugging audio.

[0044] By performing steps 201-205, debugging audio is played in the target room to adjust the adaptive filter to a suitable length that matches the target room based on the playback process of the debugging audio. In this step, the target audio can be played in the target room through the speaker, and the audio signal of the target room can be acquired through the microphone.

[0045] In one possible implementation, a reference signal of the target audio can be input to a speaker for playback; and the acquired audio signal can be a frequency domain signal. A time domain signal is acquired via a microphone, and the acquired audio signal in the frequency domain is obtained through time-frequency transformation. For example, step 201 can be implemented via steps 2011 and 2012.

[0046] Step 2011: Input the reference signal of the target audio into the speaker so that the target audio can be played in the target room through the speaker.

[0047] Step 2012: During the playback of the target audio through the speaker, the time-domain audio signal in the target room is acquired through the microphone, and the time-domain audio signal is transformed by time-frequency conversion to obtain multi-frame frequency-domain acquisition signals corresponding to multiple moments.

[0048] The acquired audio signal includes multiple frames of frequency domain acquired signals corresponding to these multiple moments.

[0049] For example, the time-domain audio signal can be a short-time stationary signal. After performing a short-time Fourier transform (STFT) on the time-domain audio signal, processing can be done using an adaptive filter to process each audio signal in the frequency domain. For example, the STFT process includes: first, dividing the time-domain audio signal into frames to obtain multiple frames; for example, the length of a frame is typically between 10ms and 30ms, and the inter-frame overlap rate can be set to 50%; second, using a pre-configured window function to window the time-domain audio signal frame by frame; for example, the pre-configured window function can be a time-domain window function, such as a Kaiser window or a moving window function, windowing the time-domain audio signal frame by frame; finally, performing a Fast Fourier Transform (FFT) on each frame to transform the time-domain audio signal to the frequency domain, obtaining the acquired audio signal, which includes frequency-domain acquired signals from multiple time points.

[0050] For example, Figure 3 A schematic diagram illustrating the principle of echo cancellation provided in this application is shown below. Figure 3 As shown, k represents time; S(k) represents the unknown near-end speech signal; Y(k) represents the acquired audio signal collected by the microphone; X(k) represents the reference signal of the target audio, which is the far-end speech signal (also known as the reference signal); E(k) represents the output residual signal; and W(k) represents the coefficients of the adaptive filter, which can be the coefficients of a linear adaptive filter. Figure 3 The curve in the figure represents the acoustic transmission path from the loudspeaker to the microphone, and R(k) represents the echo signal to be eliminated.

[0051] The near-end speech signal S(k) can be represented by a vector, and its corresponding speech vector S(k) is defined as:

[0052]

[0053] Where S(k) represents a frame of audio signal at the current time k, and n represents the number of frequency points. Each frame has N frequency points, n = 0, 1, ..., N-1. Correspondingly, the acquisition audio signal Y(k), echo signal R(k), and residual signal E(k) can be defined in the same way as S(k), which will not be listed here.

[0054] The filter configuration module also includes a filter unit 1012, which is used to implement the following step 202:

[0055] Step 202: Filter the acquired audio signal frame by frame using an adaptive filter of initial length, and determine the convergence coefficient of the adaptive filter based on the filtering process.

[0056] The acquired audio signal includes frequency domain acquisition signals at various time points. The frequency domain acquisition signals at each time point can be filtered using the adaptive filter coefficients corresponding to that time point. The convergence coefficient is the convergence coefficient at which the adaptive filter coefficients at each time point reach convergence. For example, when the acquired audio signal tends to stabilize, the adaptive filter coefficients at the corresponding time point also tend to converge, and the convergence coefficient can then be obtained.

[0057] In one possible implementation, the filtering unit 1012 includes a filtering subunit and a convergence coefficient determination subunit, wherein the filtering subunit is used to implement the following step 2021 included in step 202, and the convergence coefficient determination subunit is used to implement the following step 2022 included in step 202:

[0058] Step 2021: Filter the frequency domain acquisition signal at each time step using an adaptive filter of initial length to obtain the residual signal corresponding to each time step.

[0059] In this step, the adaptive filter coefficients can be applied to the reference signal to simulate the echo transmission characteristics of the target room, and the frequency domain acquisition signal can be filtered based on the simulated echo.

[0060] In one possible implementation, the filter subunit, when implementing step 2021, can specifically be used to perform the following steps A1-A2:

[0061] Step A1: Using an adaptive filter of initial length, based on the frame set corresponding to the reference signal of the target audio at each time moment and the adaptive filter coefficients corresponding to each time moment, estimate the echo signal corresponding to each time moment.

[0062] In this application, the length of the adaptive filter refers to the number of blocks included in the adaptive filter, and the adaptive filter of the initial length includes an initial number of blocks.

[0063] The reference signal for the target audio includes frequency domain reference signals at multiple times. A time point can be the sampling time corresponding to a frame sampled from the target audio. For example, the current time k=1 corresponds to the first frame from 0 to 10ms; the current time k=2 corresponds to the second frame from 5 to 15ms. For each time point, the adaptive filter can combine one frame from that time point with multiple frames preceding that time point to filter the frequency domain sampled signal at that time point. In other words, the frame set corresponding to each time point is a set of multiple frames of the reference signal used by the adaptive filter when calculating the residual signal corresponding to the current time point. Specifically, the frame set corresponding to each time point includes one frame of frequency domain reference signal at that time point and the frequency domain reference signals of each frame from each preceding time point.

[0064] For each block of the adaptive filter, each block is used to process the corresponding frame frequency domain reference signal in the corresponding frame set at each time step.

[0065] Step A2: Eliminate the echo signal corresponding to each time moment from the frequency domain acquired signal at each time moment to obtain the residual signal corresponding to each time moment.

[0066] For example, assuming the adaptive filter comprises B blocks, such as the NLMS block adaptive filter having B blocks, the adaptive filter coefficients can be defined as follows:

[0067] W(k)=(W0(k),W1(k),…,W b (k),…W B-1 (k));

[0068] Where W0(k), W1(k), ..., W b (k),…W B-1 (k) represent the block coefficients of the 0th block, the 1st block, ..., the bth block ... the B-1th block in the adaptive filter.

[0069] X(k) represents the reference signal of the target audio, and the frame set X(k) corresponding to time k can be expressed as: X(k) = (X0(k), X1(k), ..., X b (k),…X B-1 (k)); where k is the current time, then X0(k) represents the frequency domain reference signal that is 0 times away from the current time, that is, the frequency domain reference signal at the current time k; X b X1(k) represents the frequency domain reference signal at time b, that is, the frequency domain reference signal at time b before the current time. For example, if the current k = 8, then X1(k) represents the frequency domain reference signal at time 7.

[0070] Each block is used to process the frequency domain reference signal of the corresponding frame in the frame set at each time step. For example, the 0th block, the 1st block, ... the bth block ... the (B-1)th block respectively process the frequency domain reference signals at times 0, 1, ... b, ... B-1 from the current k-time. b (k),…W B-1 (k) correspond to X0(k), X1(k), ..., X b (k),…X B-1 The coefficient of (k).

[0071] The b-th filter is defined as follows:

[0072]

[0073] For example, at time k, the frame set corresponding to time k includes B frames of frequency-domain audio signals from B time points; the b-th block is used to process the b-th frame signal in the frame set at time k. A single frame of frequency-domain audio signal may include multiple frequency points, and each block coefficient may include multiple frequency point coefficients, each corresponding to a specific frequency point; for example, a single frame of frequency-domain audio signal may include N frequency points, W... b (k) includes N frequency coefficients, W b,0 (k), W b,1 (k), ...W b,N-1 (k) are the frequency coefficients corresponding to the 0th frequency point, the 1st frequency point, ..., the Nth frequency point, respectively.

[0074] For example, the filtering process using the NLMS block adaptive filter can be represented as shown in the following formula 1. That is, the residual signal corresponding to each time step can be obtained by performing a passivation on the frequency domain acquisition signal at each time step using the following formula 1:

[0075] Formula 1:

[0076] Where X(k) represents the reference signal of the target audio, specifically the frame set corresponding to the k-th time in the reference signal. Y(k) represents the acquired audio signal collected by the microphone, specifically the acquired audio signal at the k-th time. E(k) represents the output residual signal, specifically the residual signal corresponding to the k-th time.

[0077] B represents the length of the adaptive filter W(k). This represents the echo signal corresponding to the k-th time step in the adaptive filter simulation. Initially, W(k) can be set to all zeros. In Equation 1, the adaptive filter of this initial length includes an initial number of blocks; for example, the initial value of B can be 64.

[0078] In this context, if there is an audio signal from a target sound source nearby, such as when the teacher is speaking, E(k) represents the audio signal from the target sound source, which is the target speech signal. During the tuning phase of the adaptive filter using the tuning audio, only the speaker plays the target audio, and the target sound source is not set. Therefore, the expected value of E(k) should be 0.

[0079] Step 2022: Based on the residual signals at each time point, iterate the adaptive filter coefficients at each time point until the adaptive filter coefficients at each time point meet the convergence condition, then stop the iteration and obtain the convergence coefficients.

[0080] For each time step k, the adaptive filter coefficients for time step k+1 can be calculated iteratively using the adaptive filter coefficients corresponding to time step k. In one possible implementation, the initial length of the adaptive filter includes an initial number of blocks, each block being used to process the corresponding frame frequency domain reference signal in the frame set corresponding to each time step. The adaptive filter coefficients include the block coefficients of each block. The coefficients of the corresponding block for the next time step can be updated using the block coefficients, update step size, etc. Accordingly, the convergence coefficient determination subunit, when iterating the adaptive filter coefficients corresponding to each time step based on the residual signal corresponding to each time step, the iteration process may include the following step B1:

[0081] Step B1: For each block included in the adaptive filter, iteratively execute the following steps: Based on the block coefficients of the block at each time step, the corresponding frame frequency domain reference signal, the residual signal corresponding to each time step, and the update step size, determine the block coefficients of the block at the next time step.

[0082] In each iteration, the difference in filter coefficients corresponding to that iteration is calculated. A convergence condition can be that the difference in filter coefficients is less than a target convergence threshold. For example, this difference can be the difference between the adaptive filter coefficients at each time step of the current iteration and the next time step. If this difference in filter coefficients is less than the target convergence threshold, then the adaptive filter coefficients at each time step meet the convergence condition. For instance, it can also be determined that the difference in filter coefficients for multiple consecutive iterations is less than the target convergence threshold before confirming that the adaptive filter coefficients at each time step meet the convergence condition. When the convergence condition is met for multiple consecutive frames, it indicates that the filter coefficients have changed little and have entered a convergence state.

[0083] For example, taking the b-th block as an example, the coefficients of each block in the adaptive filter coefficients at each time step can be iterated using the following formula:

[0084] Formula 2:

[0085] Where, μ b (k) represents the update step size, also known as the step size factor, such as μ. b (k) = 0.1. Represents vector X b The transpose of (k). The estimated W b (k+1) is used to substitute into Formula 1 for filtering calculation at the next time step.

[0086] For example, taking the b-th block as an example, the convergence condition of the adaptive filter coefficients at each time step can be determined according to the following formula three:

[0087] Formula 3: ||W b (k+1)-W b (k)|| <Th;

[0088] Where Th represents the target convergence threshold, the value of which can be set as needed. When multiple consecutive frames satisfy the convergence condition corresponding to Formula 3, it indicates that the filter coefficients change little and have entered the convergence state. For example, if the convergence condition is reached at the k-th time, then time k is a stationary time, and the adaptive filter coefficient corresponding to the k-th time is the convergence coefficient.

[0089] The filter configuration module also includes a fitting unit 1013, which is used to implement the following step 203:

[0090] Step 203: Based on the convergence coefficient, obtain the energy attenuation line of the debug audio.

[0091] The energy attenuation line is used to simulate the attenuation trend of target audio as it travels along the echo transmission path, which is the acoustic transmission path from the speaker to the microphone in the target room.

[0092] The reverberation time represents the time required for the reverberation to decrease by 60 dB. It is the time elapsed from when the sound source stops emitting sound after the indoor sound field reaches a steady state until the sound pressure level decays by 60 dB, denoted as T60 or RT, in seconds (s). Different rooms have different reverberation levels. Reverberation level is usually measured by the reverberation time T60. In larger multimedia classrooms, such as lecture halls, T60 can reach 0.8–2 seconds, and using a shorter adaptive filter will result in a larger echo residue. Conversely, in rooms with lower reverberation, using a longer adaptive filter will also lead to a decrease in echo suppression performance, or even side effects. Therefore, this application allows for determining the reverberation time of the target room during the commissioning phase by fitting an energy attenuation line, and then using the reverberation time to adjust the adaptive filter to a suitable length.

[0093] In this step, multiple attenuation values ​​during the transmission of the target audio in the target room can be determined based on the convergence coefficient. An energy attenuation line is then fitted based on these multiple attenuation values, which reflect the rate of energy attenuation over time during the transmission of the target audio. Then, the reverberation time is estimated based on the rate of attenuation represented by the energy attenuation line.

[0094] In one possible implementation, the convergence coefficient includes the block convergence coefficients of each block; each decay time can be located based on the block convergence coefficients, and curve fitting can be performed based on the decay amount corresponding to each decay time. For example, when implementing step 203, the fitting unit 1013 can specifically execute the following steps 2031 to 2033 included in step 203:

[0095] Step 2031: For each block, determine the average value of the frequency coefficients included in the convergence coefficient of the corresponding block, and obtain the average value of the frequency coefficients corresponding to each block at the stationary time.

[0096] The stationary moment refers to the moment corresponding to the convergence coefficient. The adaptive filter of the initial length includes an initial number of blocks, and the convergence coefficient includes the block convergence coefficients corresponding to each block. In each block convergence coefficient, each frequency point coefficient corresponds to each frequency point included in the corresponding frame frequency domain reference signal of that block. In the vector corresponding to the adaptive filter coefficients, each column can correspond to a block convergence coefficient, and each column includes the frequency point coefficients in the corresponding block convergence coefficient. For example, the average value of the frequency point coefficients of each block convergence coefficient in W(k) can be calculated along the column vector direction using the following formula four:

[0097] Formula 4:

[0098] in, The value of the adaptive filter coefficients W(k) is represented along the column vector, and N represents the total number of frequency points. This represents the average value of the frequency coefficients corresponding to the block.

[0099] Step 2032: Based on the average frequency coefficient of each block at the stationary moment, determine the attenuation amount of the reference signal of the target audio at each attenuation moment.

[0100] In this step, the maximum value among the average values ​​of each frequency point corresponding to each block is determined, and the average frequency point corresponding to each target block is selected from the average values ​​based on the maximum value. The corresponding time of each target block in the current frame set is taken as the attenuation time. Each target block includes the block corresponding to the maximum value in the adaptive filter, and the blocks following the block corresponding to the maximum value. The current frame set is the frame set corresponding to the stationary time in the reference signal.

[0101] For example, the block corresponding to the maximum value is the b-th block, denoted as b. max That is to say, The position of the maximum value is b. max The target blocks sequentially include the b-th... max The b+1, b+2, ..., B-1, B blocks; correspondingly, each decay time may include the b-th block. maxThe corresponding times of the b+1, b+2, ... B-1, B blocks in the current frame set. For example, suppose the effective length of a frame is 8ms, and there are 10 target blocks. If the stationary time is 80ms, the frame set corresponding to time 80ms includes one frame each at 0ms (0~8ms), 8ms (8~16ms), 16ms (16~24ms)..., 64ms (64~72ms), and 72ms (72~80ms). If b max If it is a block used in adaptive filtering to process the 64ms frame, then the decay times include 0ms, 8ms, 16ms, ... 64ms.

[0102] For example, the maximum value among the various average values ​​is used to filter the frequency average value corresponding to the target block from the average values ​​corresponding to each block. For example, Can be adjusted to filter

[0103] For example, the average frequency points corresponding to the selected target blocks can also be calculated using the following formula five. Normalization is performed:

[0104] Formula 5:

[0105] in, This represents the average frequency value of the target block after normalization.

[0106] The magnitude of each block coefficient represents the influence of the frequency domain reference signal corresponding to each block at the current time, such as the echo at the current stationary time. The larger the average value of the frequency coefficients corresponding to each block, the greater the influence. In this step, the average value of the frequency coefficients corresponding to each target block can be converted into the attenuation amount at each attenuation time.

[0107] For example, define the expression for the energy decay curve: Here, "smooth" indicates a smooth calculation, which can be achieved using a moving average window with a width of 7. Smoothing is then performed. The average value of the frequency coefficients corresponding to each target block can be substituted into this expression to obtain the attenuation amount at each attenuation time. For example, if the attenuation times include 0ms, 8ms, 16ms, ... 64ms, then the attenuation amounts at 0ms, 8ms, 16ms, ... 64ms are 20dB, 15dB, 11dB, ... 0.1dB, respectively.

[0108] Step 2033: Based on the decay amount at each decay time, fit the energy decay line.

[0109] This energy attenuation line is used to simulate the attenuation trend of audio transmitted along an echo transmission path, which can be the acoustic transmission path from the speaker to the microphone in the target room. For example, a straight line can be fitted based on each attenuation moment and the corresponding attenuation amount at each moment to obtain the energy attenuation line. The horizontal axis represents the time, and the vertical axis represents the attenuation amount. The fitted line can be expressed as: y = p0(k) + p1(k)x, where x represents the horizontal axis and y represents the vertical axis. For example, the segment from -5dB to -25dB of the attenuation amount at each attenuation moment can be used for straight line fitting.

[0110] The filter configuration module also includes a reverberation time determination unit 1014, which is used to implement the following step 204:

[0111] Step 204: Determine the reverberation time of the target room based on the energy decay curve.

[0112] In this step, the reverberation time of the target room can be determined based on the rate of change of attenuation corresponding to the energy attenuation line. For example, the energy attenuation line is calculated using the adaptive filter coefficients at a stationary moment. In this step, the adaptive filter coefficients at each time point after the stationary moment can also be combined to estimate various reference reverberation times; combining multiple estimated reverberation times, the final reverberation time is obtained. In one possible implementation, when implementing step 204, the reverberation time determination unit can specifically be used to execute the following steps 2041 to 2042 included in step 204:

[0113] Step 2041 determines the initial reverberation time of the target room at a steady moment based on the rate of change of attenuation corresponding to the energy attenuation curve.

[0114] Step 2042: Based on the initial reverberation time and the reverberation time after at least one stationary moment, determine the reverberation time of the target room, wherein the stationary moment is the moment after the stationary moment.

[0115] For example, the initial reverberation time can be calculated using the following formula six:

[0116] Formula Six:

[0117] The initial reverberation time can be the reverberation time corresponding to a stationary moment. For example, if the stationary moment is the k-th moment, then the initial reverberation time can be expressed as T. 60 (k); p1(k) represents the rate of change of decay, which can be the slope of the energy decay curve; T frameThis represents the time corresponding to the effective length of a frame of frequency domain reference signal. For example, the first frame is 0 to 10 ms, the second frame is 5 to 15 ms, and the effective length can correspond to a time of 5 ms.

[0118] For example, step 2042 can be implemented in either of the following two ways: Method 1 and Method 2. Accordingly, when the reverberation time determination unit 1014 determines the reverberation time of the target room based on the initial reverberation time and the reverberation time after at least one stationary moment, it can specifically execute either the step of Method 1 or the step of Method 2.

[0119] Method 1: Based on the initial reverberation time and the rate of change of attenuation corresponding to each stable reverberation time, iterate the stable reverberation time corresponding to each stable time until each stable reverberation time meets the target condition, then stop the iteration to obtain the reverberation time of the target room.

[0120] The target conditions may include, but are not limited to: iterating to the last frame of the target video, or exceeding the first target threshold. For example, if the debugging audio length is 50 seconds, and the plateau time is 80 ms, assuming an effective frame length of 8 ms, then starting from the initial reverberation time corresponding to 80 ms, iteratively calculate the reverberation time corresponding to 88 ms, 96 ms, ... up to 50 ms, and use the reverberation time corresponding to 50 ms as the final reverberation time. As another example, if the first target threshold can be 20, then the 20th plateau reverberation time after the plateau time is obtained through iteration and used as the final reverberation time.

[0121] For example, taking the next moment after a stationary period as an example, in one iteration, the reverberation time after stationarity can be obtained in the following ways: For the next moment after a stationary period, the reverberation time corresponding to the next moment can be calculated using the following formula seven, based on the initial reverberation time, the rate of change of attenuation corresponding to the next moment, and the block number corresponding to the maximum value. Here, the block number corresponding to the maximum value is the sequence number of the block corresponding to the maximum value among the average values ​​of each frequency point corresponding to each block.

[0122] Formula 7:

[0123] Among them, T 60 (k+1) represents the reverberation time corresponding to the (k+1)th time; if the stationary time is the kth time, then T 60 (k+1) represents the reverberation time corresponding to the next time step after the stationary time. p1(k+1) represents the rate of change of decay at time k+1, which is also the rate of change of decay corresponding to the next time step after the stationary time. T 60 (k) represents the initial reverberation time. b maxThis indicates the block number corresponding to the maximum value. α represents the smoothing coefficient, which can be configured as needed. T frame This represents the time corresponding to the effective length of a frame of frequency domain reference signal.

[0124] For example, for each stationary time point, the energy attenuation line of the tuned audio at that stationary time point can be fitted using the adaptive filter coefficients corresponding to each stationary time point, and the corresponding rate of change of attenuation can be obtained based on the energy attenuation line at that stationary time point. For example, for the (k+1)th time point, the energy attenuation line at the (k+1)th time point can be fitted using the coefficients of the adaptive filter at the (k+1)th time point through the same process in step 203, so as to obtain p1(k+1) corresponding to the (k+1)th time point.

[0125] Method 2: The average of the initial reverberation time and at least one stable reverberation time is determined as the reverberation time of the target room.

[0126] For example, in Method 2, the corresponding post-stationary reverberation time can be calculated using the adaptive filter coefficients corresponding to the stationary time, thus obtaining multiple post-stationary reverberation times. The average of the post-stationary reverberation times corresponding to the second target threshold can be used as the final reverberation time. The calculation method for each post-stationary reverberation time can refer to the reverberation time corresponding to the (k+1)th time in Method 1, i.e., T. 60 The calculation method for (k+1) will not be elaborated here.

[0127] The filter configuration module also includes a length mapping unit 1015, which is used to implement the following step 205:

[0128] Step 205: Map the reverberation time to the pre-configured length of the adaptive filter, and configure the length of the adaptive filter to the pre-configured length.

[0129] For example, a mapping relationship between filter length and reverberation time can be obtained; based on the mapping relationship between filter length and reverberation time, the reverberation time of the target room is mapped to the pre-configured length of the adaptive filter.

[0130] For example, in this step, the pre-configured length can be calculated based on the reverberation time using the following formula eight:

[0131] Formula 8:

[0132] Among them, T 60 The reverberation time can be the final reverberation time of the target room obtained through step 204. frameThis represents the time corresponding to the effective length of a frame of frequency domain reference signal. In this application, the mapping relationship shown in Formula 8 can be used to determine the optimal filter length. For example, the effective length T of a frame of speech... frame It is 16ms, if T 60 =500ms, then substituting into the above formula eight, we can get B=32, that is, the length of the adaptive filter in the target room can be adjusted to 32, that is, it includes 32 blocks.

[0133] It should be noted that, depending on different T 60 Experiments were conducted using the optimal filter length to determine the mapping relationship between filter length and reverberation time. For block frequency domain filters, this application only uses T... 60 Let's take the number of filter blocks as an example, rounded up to an integer multiple of 8. However, there are no specific limitations on the mapping relationship between filter length and reverberation time.

[0134] The following example flow further illustrates the length adjustment process of the adaptive filter implemented in the filter configuration module of this application: First, a debugging audio clip is played through a speaker, while a microphone captures the signal. The captured time-domain signal is then converted to a frequency-domain signal via STFT, yielding the captured audio signal. Next, adaptive filter bank alignment is used for filtering, such as a frequency-domain NLMS filter bank. Then, the NLMS filter coefficients are updated based on the filtering results. After the filter coefficients have stabilized and converged, the reverberation time T is estimated. 60 Finally, the optimal filter length is calculated based on the reverberation time. It should be noted that the method for estimating the optimal filter length during the initial debugging phase in this application ensures that the length of the adaptive filter matches the target room. This allows the adaptive filter to achieve better echo suppression performance when officially used in the target room, while avoiding side effects. Furthermore, the adaptive filter length estimation method provided in this application utilizes the coefficients after convergence of the frequency-domain block adaptive filter, and is independent of specific algorithms. This method can be applied to block adaptive NLMS, RLS, Kalman, other variable step-size methods, and dual-filter structures, demonstrating broad applicability. Moreover, the reverberation time calculation method in this application is simple and inexpensive, making it suitable for various hardware platforms.

[0135] The echo cancellation device provided in this application, through a filter configuration module, pre-configures the length of an adaptive filter to a pre-configured length that matches the reverberation level of the target room. After acquiring the original speech signal from the target room through a speech acquisition module, the echo cancellation module uses the pre-configured length adaptive filter to filter the original speech signal, removing echoes to obtain the target speech signal. Using this method, the adaptive filter can be tuned to a length matching the reverberation time of the target room. Therefore, when the tuned adaptive filter is used in the target room, it accurately simulates the echo transmission characteristics of the target room, filtering out echoes from the original speech signal, achieving better echo suppression performance, avoiding side effects, and improving the accuracy of echo cancellation.

[0136] Figure 4 This is a schematic diagram of the structure of an echo cancellation method provided in an embodiment of this application. Figure 4 As shown, the method includes the following steps:

[0137] Step 401: Pre-configure the length of the adaptive filter to a pre-configured length that matches the reverberation level of the target room;

[0138] Step 402: Collect speech from the target sound source in the target room to obtain the original speech signal;

[0139] Step 403: Filter the original speech signal using an adaptive filter of pre-configured length to obtain the target speech signal of the target sound source. The adaptive filter of pre-configured length is used to simulate the echo transmission characteristics of the target room to filter out the echo in the original speech signal.

[0140] In one possible implementation, the length of the adaptive filter is pre-configured to a pre-configured length that matches the reverberation level of the target room, including:

[0141] Play the debugging audio in the target room and collect the audio signal in the target room during the playback to obtain the collected audio signal;

[0142] The acquired audio signal is filtered frame by frame using an adaptive filter of initial length, and the convergence coefficient of the adaptive filter is determined based on the filtering process.

[0143] The energy attenuation line of the debug audio was obtained by fitting the convergence coefficient.

[0144] The reverberation time of the target room is determined based on this energy decay line;

[0145] The reverberation time is mapped to the pre-configured length of the adaptive filter, and the length of the adaptive filter is configured to the pre-configured length.

[0146] In one possible implementation, the acquired audio signal includes multi-frame frequency domain acquired signals corresponding to multiple time points;

[0147] The acquired audio signal is filtered frame by frame using an adaptive filter of initial length, and the convergence coefficient of the adaptive filter is determined based on the filtering process, including:

[0148] The frequency domain acquisition signal at each time step is filtered using an adaptive filter of this initial length to obtain the residual signal corresponding to each time step.

[0149] Based on the residual signals at each time step, the adaptive filter coefficients at each time step are iterated until the adaptive filter coefficients at each time step meet the convergence condition. Then the iteration stops and the convergence coefficients are obtained.

[0150] In one possible implementation, the frequency domain acquisition signal at each time step is filtered using an adaptive filter of initial length to obtain the residual signal corresponding to each time step, including:

[0151] Using an adaptive filter of initial length, the echo signal corresponding to each moment is estimated based on the frame set corresponding to the reference signal of the target audio at each moment and the adaptive filter coefficients corresponding to each moment.

[0152] The frame set corresponding to each time moment includes a frame of frequency domain reference signal at each time moment and a frame of frequency domain reference signal at each preceding time moment, wherein each preceding time moment is a time moment before each time moment.

[0153] The echo signal corresponding to each time moment is eliminated from the frequency domain acquired signal at each time moment to obtain the residual signal corresponding to each time moment.

[0154] In one possible implementation, the adaptive filter of the initial length includes an initial number of blocks, each block being used to process the frequency domain reference signal of the corresponding frame at the corresponding time in the frame set at each time step;

[0155] Based on the residual signal at each time step, the adaptive filter coefficients at each time step are iterated, including:

[0156] For each block included in the adaptive filter, the following steps are performed iteratively:

[0157] Based on the block coefficients of the block at each time step, the corresponding frame frequency domain reference signal, the residual signal corresponding to each time step, and the update step size, the block coefficients of the block at the next time step are determined.

[0158] The adaptive filter coefficients include the block coefficients of each block.

[0159] In one possible implementation, the adaptive filter of the initial length includes an initial number of blocks, and the convergence coefficients include the block convergence coefficients corresponding to each block.

[0160] The energy attenuation line of the tuned audio is obtained by fitting based on the convergence coefficient, including:

[0161] For each block, determine the average value of the frequency coefficients included in the convergence coefficient of the corresponding block, and obtain the average value of the frequency coefficients of each block at the stationary time, where the stationary time refers to the time corresponding to the convergence coefficient.

[0162] Based on the average frequency coefficient of each block at a steady moment, the attenuation of the reference signal of the target audio at each attenuation moment is determined.

[0163] Based on the attenuation amount at each attenuation moment, the energy attenuation line is fitted. This energy attenuation line is used to simulate and debug the attenuation trend of audio transmission along the echo transmission path, which is the acoustic transmission path from the speaker to the microphone in the target room.

[0164] In one possible implementation, determining the reverberation time of the target room based on the energy decay line includes:

[0165] Based on the rate of change of attenuation corresponding to the energy attenuation line, the initial reverberation time of the target room at a steady moment is determined.

[0166] The reverberation time of the target room is determined based on the initial reverberation time and the post-stable reverberation time corresponding to at least one post-stable moment, wherein the post-stable moment is the moment after the current stable moment.

[0167] In one possible implementation, determining the reverberation time of the target room based on the initial reverberation time and the post-stationary reverberation time corresponding to at least one stationary time step includes:

[0168] The average of the initial reverberation time and at least one stable reverberation time is determined as the reverberation time of the target room.

[0169] In one possible implementation, determining the reverberation time of the target room based on the initial reverberation time and the post-stationary reverberation time corresponding to at least one stationary time step includes:

[0170] Based on the initial reverberation time and the rate of change of attenuation corresponding to each stationary reverberation time, the stationary reverberation time corresponding to each stationary time is iterated until each stationary reverberation time meets the target condition, at which point the iteration stops and the reverberation time of the target room is obtained.

[0171] In one possible implementation, the echo transmission characteristic is a characteristic of the echo transmission path along the speaker to the microphone in the target room;

[0172] The program plays debugging audio in the target room and collects audio signals from the target room during playback. The collected audio signals include:

[0173] The reference signal of the target audio is input to a speaker so that the target audio can be played in the target room through the speaker;

[0174] During the playback of target audio through a speaker, a time-domain audio signal in the target room is acquired through a microphone, and the time-domain audio signal is transformed by time-frequency conversion to obtain multi-frame frequency-domain acquisition signals corresponding to multiple moments. The acquired audio signal includes the multi-frame frequency-domain acquisition signals corresponding to the multiple moments.

[0175] The echo cancellation method provided in this application involves pre-configuring the length of an adaptive filter to match the reverberation level of the target room. After acquiring the original speech signal from the target room, the original speech signal is filtered using the pre-configured length adaptive filter to remove echoes and obtain the target speech signal. This method allows the adaptive filter to be tuned to match the reverberation time of the target room. When the tuned adaptive filter is used in the target room, it accurately simulates the echo transmission characteristics of the target room, filtering out echoes from the original speech signal. This achieves better echo suppression performance while avoiding side effects, thus improving the accuracy of echo cancellation using the adaptive filter.

[0176] The apparatus in this application embodiment can execute the method provided in this application embodiment, and the implementation principle is similar. The actions performed by each module in the apparatus of each embodiment of this application correspond to the steps in the method of each embodiment of this application. For detailed functional descriptions of each module of the apparatus, please refer to the descriptions in the corresponding methods shown above, which will not be repeated here.

[0177] Figure 5 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. For example... Figure 5 As shown, the computer device includes: a memory, a processor, and a computer program stored in the memory. The processor executes the computer program to implement the steps of the echo cancellation method, which, compared with related technologies, can achieve:

[0178] The echo cancellation method provided in this application involves pre-configuring the length of an adaptive filter to match the reverberation level of the target room. After acquiring the original speech signal from the target room, the original speech signal is filtered using the pre-configured length adaptive filter to remove echoes and obtain the target speech signal. This method allows the adaptive filter to be tuned to match the reverberation time of the target room. When the tuned adaptive filter is used in the target room, it accurately simulates the echo transmission characteristics of the target room, filtering out echoes from the original speech signal. This achieves better echo suppression performance while avoiding side effects, thus improving the accuracy of echo cancellation using the adaptive filter.

[0179] In one alternative embodiment, a computer device is provided, such as Figure 5 As shown, Figure 5 The computer device 500 shown includes a processor 501 and a memory 503. The processor 501 and the memory 503 are connected, for example, via a bus 502. Optionally, the computer device 500 may further include a transceiver 504, which can be used for data interaction between the computer device and other computer devices, such as sending and / or receiving data. It should be noted that in practical applications, the transceiver 504 is not limited to one type, and the structure of the computer device 500 does not constitute a limitation on the embodiments of this application.

[0180] Processor 501 may be a CPU (Central Processing Unit), a general-purpose processor, a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), an FPGA (Field Programmable Gate Array), or other programmable logic devices, transistor logic devices, hardware components, or any combination thereof. It can implement or execute the various exemplary logic blocks, modules, and circuits described in conjunction with the disclosure of this application. Processor 501 may also be a combination that implements computational functions, such as including one or more microprocessor combinations, a combination of a DSP and a microprocessor, etc.

[0181] Bus 502 may include a pathway for transmitting information between the aforementioned components. Bus 502 may be a PCI (Peripheral Component Interconnect) bus or an EISA (Extended Industry Standard Architecture) bus, etc. Bus 502 can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 5 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.

[0182] The memory 503 may be ROM (Read Only Memory) or other types of static storage devices capable of storing static information and instructions, RAM (Random Access Memory) or other types of dynamic storage devices capable of storing information and instructions, or EEPROM (Electrically Erasable Programmable Read Only Memory), CD-ROM (Compact Disc Read Only Memory) or other optical disc storage, optical disc storage (including compressed optical discs, laser discs, optical discs, digital universal optical discs, Blu-ray discs, etc.), magnetic disk storage media, other magnetic storage devices, or any other medium capable of carrying or storing computer programs and capable of being read by a computer, without limitation herein.

[0183] The memory 503 is used to store computer programs that execute the embodiments of this application, and the execution is controlled by the processor 501. The processor 501 is used to execute the computer programs stored in the memory 503 to implement the steps shown in the foregoing method embodiments.

[0184] Electronic devices include, but are not limited to, servers, terminals, or cloud computing center equipment.

[0185] This application provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, it can implement the steps and corresponding content of the aforementioned method embodiments.

[0186] This application also provides a computer program product, including a computer program that, when executed by a processor, can implement the steps and corresponding content of the aforementioned method embodiments.

[0187] Those skilled in the art will understand that, unless otherwise stated, the singular forms “a,” “an,” “the,” and “the” used herein may also include the plural forms. The terms “comprising” and “including” as used in the embodiments of this application mean that the corresponding feature can be implemented as the presented feature, information, data, step, or operation, but do not exclude implementation as other features, information, data, steps, or operations supported by this art.

[0188] The terms "first," "second," "third," "fourth," "1," "2," etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented in a sequence other than that shown in the figures or text.

[0189] It should be understood that although arrows indicate various operation steps in the flowcharts of this application's embodiments, the order in which these steps are implemented is not limited to the order indicated by the arrows. Unless explicitly stated herein, in some implementation scenarios of this application's embodiments, the implementation steps in each flowchart can be executed in other orders as required. Furthermore, some or all steps in each flowchart, based on the actual implementation scenario, may include multiple sub-steps or multiple stages. Some or all of these sub-steps or stages can be executed at the same time, and each sub-step or stage can also be executed at different times. In scenarios where execution times differ, the execution order of these sub-steps or stages can be flexibly configured according to requirements, and this application's embodiments do not limit this.

[0190] The above description is only an optional implementation method for some implementation scenarios of this application. It should be noted that for those skilled in the art, other similar implementation methods based on the technical concept of this application without departing from the technical concept of this application also fall within the protection scope of the embodiments of this application.

Claims

1. An echo cancellation device, characterized in that, The device includes: A filter configuration module is used to pre-configure the length of the adaptive filter to a pre-configured length that matches the reverberation level of the target room; The voice acquisition module is used to acquire voice data from target sound sources in the target room and obtain the original voice signal. An echo cancellation module is used to filter the original speech signal through an adaptive filter of a pre-configured length to obtain the target speech signal of the target sound source. The adaptive filter of the pre-configured length is used to simulate the echo transmission characteristics of the target room to filter out the echo in the original speech signal. The filter configuration module includes: An audio signal acquisition unit is used to play debugging audio in a target room and acquire audio signals in the target room during the playback process to obtain the acquired audio signals. A filtering unit is used to filter the acquired audio signal frame by frame using an adaptive filter of initial length, and to determine the convergence coefficient of the adaptive filter based on the filtering process. A fitting unit is used to fit the energy attenuation line of the debug audio based on the convergence coefficient; A reverberation time determination unit is used to determine the reverberation time of the target room based on the energy decay line; A length mapping unit is used to map the reverberation time to a pre-configured length of the adaptive filter, and to configure the length of the adaptive filter to the pre-configured length.

2. The apparatus according to claim 1, characterized in that, The acquired audio signals include multi-frame frequency domain acquired signals corresponding to multiple time points; The filtering unit includes: The filtering subunit is used to filter the frequency domain acquisition signal at each time step through the adaptive filter of the initial length to obtain the residual signal corresponding to each time step. The convergence coefficient determination subunit is used to iterate the adaptive filter coefficients at each time step based on the residual signal at each time step until the adaptive filter coefficients at each time step meet the convergence condition, at which point the iteration stops and the convergence coefficients are obtained.

3. The apparatus according to claim 2, characterized in that, The filtering subunit is used for: Using an adaptive filter of initial length, the echo signal corresponding to each moment is estimated based on the frame set corresponding to the reference signal of the debug audio at each moment and the adaptive filter coefficients corresponding to each moment. The frame set corresponding to each time moment includes a frame of frequency domain reference signal at each time moment and a frame of frequency domain reference signal at each preceding time moment, wherein each preceding time moment is a time moment before each time moment. The echo signal corresponding to each time moment is eliminated from the frequency domain acquired signal at each time moment to obtain the residual signal corresponding to each time moment.

4. The apparatus according to claim 2, characterized in that, The initial length of the adaptive filter includes an initial number of blocks, each block being used to process the corresponding frame frequency domain reference signal at the corresponding time in the frame set at each time. The convergence coefficient determination subunit is used for: For each block included in the adaptive filter, the following steps are performed iteratively: Based on the block coefficients of the block at each time step, the corresponding frame frequency domain reference signal, the residual signal corresponding to each time step, and the update step size, the block coefficients of the block at the next time step are determined. The adaptive filter coefficients include the block coefficients of each block.

5. The apparatus according to claim 1, characterized in that, The initial length of the adaptive filter includes an initial number of blocks, and the convergence coefficients include the block convergence coefficients corresponding to each block. The fitting unit is used for: For each block, determine the average value of each frequency point coefficient included in the convergence coefficient of the corresponding block, and obtain the average value of the frequency point coefficients corresponding to each block at the stationary time, where the stationary time refers to the time corresponding to the convergence coefficient. Based on the average frequency coefficient of each block at a steady moment, the attenuation of the reference signal of the debug audio at each attenuation moment is determined. Based on the attenuation amount at each attenuation moment, the energy attenuation line is fitted. The energy attenuation line is used to simulate and debug the attenuation trend of audio transmission along the echo transmission path, which is the acoustic transmission path from the speaker to the microphone in the target room.

6. The apparatus according to claim 1, characterized in that, The reverberation time determination unit is used for: Based on the rate of change of attenuation corresponding to the energy attenuation line, the initial reverberation time of the target room at a steady moment is determined. The reverberation time of the target room is determined based on the initial reverberation time and the post-stable reverberation time corresponding to at least one post-stable moment, wherein the post-stable moment is the moment after the stable moment.

7. The apparatus according to claim 6, characterized in that, When determining the reverberation time of the target room based on the initial reverberation time and the reverberation time after at least one stationary moment, the reverberation time determination unit is specifically used for: The average of the initial reverberation time and at least one stable reverberation time is determined as the reverberation time of the target room.

8. The apparatus according to claim 6, characterized in that, When determining the reverberation time of the target room based on the initial reverberation time and the reverberation time after at least one stationary moment, the reverberation time determination unit is specifically used for: Based on the initial reverberation time and the rate of change of attenuation corresponding to each stable reverberation time, the stable reverberation time corresponding to each stable time point is iterated until each stable reverberation time meets the target condition, at which point the iteration stops, and the reverberation time of the target room is obtained.

9. The apparatus according to claim 1, characterized in that, The echo transmission characteristic is the characteristic along the echo transmission path from the speaker to the microphone in the target room; The audio signal acquisition unit is used for: A reference signal for the test audio is input to a speaker to play the test audio in the target room through the speaker; During the process of playing and debugging audio through a speaker, the time-domain audio signal in the target room is acquired through a microphone, and the time-domain audio signal is transformed by time-frequency conversion to obtain multi-frame frequency-domain acquisition signals corresponding to multiple times. The acquired audio signal includes the multi-frame frequency-domain acquisition signals corresponding to the multiple times.

10. An echo cancellation method, characterized in that, The method includes: The length of the adaptive filter is pre-configured to match the reverberation level of the target room; Speech recordings are performed on the target sound source in the target room to obtain the raw speech signal; The original speech signal is filtered by an adaptive filter of a pre-configured length to obtain the target speech signal of the target sound source. The adaptive filter is used to simulate the echo transmission characteristics of the target room to filter out the echo in the original speech signal. The length of the adaptive filter is pre-configured to match the reverberation level of the target room, including: Play the debugging audio in the target room and collect the audio signal in the target room during the playback to obtain the collected audio signal; The acquired audio signal is filtered frame by frame using an adaptive filter of initial length, and the convergence coefficient of the adaptive filter is determined based on the filtering process. The energy attenuation line of the debug audio was obtained by fitting the convergence coefficient. The reverberation time of the target room is determined based on this energy decay line; The reverberation time is mapped to the pre-configured length of the adaptive filter, and the length of the adaptive filter is configured to the pre-configured length.

11. A computer device, comprising a memory, a processor, and a computer program stored in the memory, characterized in that, The processor executes the computer program to implement the echo cancellation method of claim 10.

12. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the echo cancellation method of claim 10.

Citation Information

Patent Citations

  • Sound echo canceller, hands free telephone using the same, and sound echo canceling method

    JP2006157498A