Echo cancellation method and apparatus, and device and storage medium
By optimizing the Kalman filter function and transforming the audio signal, the problem of near-end speech damage in echo cancellation technology is solved, and effective protection and robustness improvement of near-end speech are achieved during the echo cancellation process.
Patent Information
- Application Number
- PCT/CN2025/072843
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-07
- Filing Date
- 2025-01-16
- Publication Date
- 2025-10-16
AI Technical Summary
Existing echo cancellation technology will damage the near-end voice while suppressing the echo signal, and cannot take into account both echo suppression and protection of the near-end voice.
A Kalman filter is used for echo cancellation. By transforming and inversely transforming the audio signal and combining the posterior state misalignment covariance, the change in the echo path transfer function, and the constraint weights, the optimization function of the Kalman filter is optimized to enhance the protection of near-end speech.
While ensuring the echo cancellation effect, it effectively protects the near-end speech and improves the robustness of the Kalman filter, especially reducing the near-end speech damage in an environment with a relatively low signal-to-noise ratio.
Smart Images

Figure CN2025072843_16102025_PF_FP_ABST
Abstract
Description
Echo cancellation method, device, apparatus and storage medium
[0001] This application claims priority to Chinese Patent Application No. 202410411490.8, filed on April 7, 2024, the disclosure of which is incorporated herein in its entirety as part of the present application. TECHNICAL FIELD
[0002] Embodiments of the present disclosure relate to an echo cancellation method, device, apparatus and storage medium. BACKGROUND
[0003] In the field of real-time audio-video communication, Acoustic Echo Cancellation (AEC) is an important part of audio signal processing. The cause of acoustic echo is the acoustic coupling between the audio output device and the audio acquisition device, which makes the sound emitted by the audio output device also be collected by the audio acquisition device to form an echo. The existence of echo seriously affects the intelligibility of speech and the quality of communication. At present, the echo cancellation method will damage the near-end speech while suppressing the echo signal, and cannot balance the echo suppression and the protection of the near-end speech. SUMMARY
[0004] The present disclosure provides an echo cancellation method, device, apparatus and storage medium to improve the robustness of Kalman filter and ensure the echo cancellation effect on the basis of protecting the near-end speech.
[0005] In a first aspect, the embodiments of the present disclosure provide an echo cancellation method, comprising:
[0006] obtaining a first audio signal collected by an audio acquisition device and a second audio signal output by an audio output device, wherein the first audio signal comprises a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal transmitted by the second audio signal through an echo path;
[0007] transforming the first audio signal to obtain a collection signal corresponding to each frequency in each time window, and transforming the second audio signal to obtain a reference signal corresponding to each frequency in each time window;
[0008] The echo cancellation module is configured to perform echo cancellation on each frequency corresponding to the acquisition signal and the reference information in each time window based on a Kalman filter to determine a target signal corresponding to each frequency in each time window after echo cancellation, wherein an optimization function in the Kalman filter is determined based on a posterior state misadjustment covariance, an echo path transfer function variation, and a constraint weight corresponding to the echo path transfer function variation, the echo path transfer function variation is determined based on an error signal variance and a Kalman gain, and the constraint weight is determined based on a near-end signal variance and a frequency.
[0009] The target signal inverse transformation module is configured to perform inverse transformation on the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal.
[0010] In a second aspect, the embodiments of the present disclosure further provide an echo cancellation device, including:
[0011] The audio signal acquisition module is configured to acquire a first audio signal collected by an audio collection device and a second audio signal output by an audio output device, wherein the first audio signal includes a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal of the second audio signal transmitted through an echo path.
[0012] The audio signal transformation module is configured to transform the first audio signal to obtain an acquisition signal corresponding to each frequency in each time window, and transform the second audio signal to obtain a reference signal corresponding to each frequency in each time window.
[0013] The echo cancellation module is configured to perform echo cancellation on each frequency corresponding to the acquisition signal and the reference information in each time window based on a Kalman filter to determine a target signal corresponding to each frequency in each time window after echo cancellation, wherein an optimization function in the Kalman filter is determined based on a posterior state misadjustment covariance, an echo path transfer function variation, and a constraint weight corresponding to the echo path transfer function variation, the echo path transfer function variation is determined based on an error signal variance and a Kalman gain, and the constraint weight is determined based on a near-end signal variance and a frequency.
[0014] The target signal inverse transformation module is configured to perform inverse transformation on the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal.
[0015] In a third aspect, the embodiments of the present disclosure further provide an electronic device, including:
[0016] one or more processors;
[0017] a storage device configured to store one or more programs,
[0018] When the one or more programs are executed by the one or more processors, the one or more processors implement the echo cancellation method according to any of the embodiments of the present disclosure.
[0019] In a fourth aspect, the embodiments of the present disclosure further provide a storage medium containing computer executable instructions for executing the echo cancellation method according to any of the embodiments of the present disclosure when executed by a computer processor. BRIEF DESCRIPTION OF DRAWINGS
[0020] The above and other features, aspects and advantages of the embodiments of the present disclosure will become more apparent from the following detailed description when taken in conjunction with the accompanying drawings. In the drawings similar or common elements of the drawings are denoted by like reference numbers. It is to be understood that the drawings are schematic, and elements and features are not necessarily to scale.
[0021] Fig. 1 is a flow diagram of an echo cancellation method according to an embodiment of the present disclosure;
[0022] Fig. 2 is an example of an echo cancellation process according to an embodiment of the present disclosure;
[0023] Fig. 3 is a flow diagram of another echo cancellation method according to an embodiment of the present disclosure;
[0024] Fig. 4 is a flow diagram of yet another echo cancellation method according to an embodiment of the present disclosure;
[0025] Fig. 5 is a structural diagram of an echo cancellation apparatus according to an embodiment of the present disclosure; and
[0026] Fig. 6 is a structural diagram of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0027] Embodiments of the present disclosure will be described more fully hereinafter with reference to the accompanying drawings. While several embodiments of the present disclosure are shown in the drawings, it is understood that the present disclosure can be embodied in various forms and should not be construed as being limited to the embodiments set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete, and fully convey the scope of the present disclosure to those skilled in the art. It should be understood that the drawings and embodiments are only for illustrative purposes and are not intended to limit the scope of the present disclosure.
[0028] It should be understood that the various steps in the method embodiments of the present disclosure can be performed in different orders and / or in parallel. In addition, the method embodiments can include additional steps and / or omit performing the steps shown. The scope of the present disclosure is not limited in this respect.
[0029] The term "include," and derivations thereof, is an open term that means "including, but not limited to." The term "based on" means "based, at least in part, on." The term "one embodiment" means "at least one embodiment"; the term "another embodiment" means "at least one additional embodiment"; the term "some embodiments" means "at least some embodiments." Related terms shall be construed accordingly.
[0030] It should be noted that the terms "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules or units, and are not intended to limit the order or interdependence of the functions performed by these devices, modules or units.
[0031] It should be noted that the modification of "one" or "multiple" mentioned in the present disclosure is illustrative rather than limiting, and those skilled in the art should understand that unless otherwise explicitly indicated in the context, it should be understood as "one or more".
[0032] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0033] FIG. 1 is a flow diagram of an echo cancellation method according to an embodiment of the present disclosure. The present embodiment is applicable to the case of performing echo cancellation on the audio signal collected by an audio collection device, and is particularly applicable to the application scenario in which the signal-to-echo ratio is relatively low and there is environmental noise. The method can be performed by an echo cancellation device, which can be implemented in the form of software and / or hardware, and can be implemented by an electronic device, which can be a mobile terminal, a PC terminal or a server, etc.
[0034] As shown in FIG. 1, the echo cancellation method specifically includes the following steps:
[0035] S110, obtaining a first audio signal collected by an audio collection device and a second audio signal output by an audio output device, wherein the first audio signal includes a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal of the second audio signal transmitted through an echo path.
[0036] The audio acquisition device can be any device capable of acquiring an audio signal, such as a microphone or the like. The audio output device can be any device capable of outputting an audio signal, such as a loudspeaker or the like. The first audio signal refers to a mixed audio signal acquired by the audio acquisition device. The second audio signal refers to a far-end voice emitted by the audio output device. The near-end audio signal can refer to a near-end voice emitted near the audio acquisition device. For example, the near-end audio signal can refer to a voice of a user speaking. The echo path refers to an audio transmission path between the audio output device and the audio acquisition device. The far-end audio signal refers to an echo signal of the second audio signal emitted by the audio output device and transmitted to the audio acquisition device through the echo path. The first audio signal can include, in addition to the near-end audio signal and the far-end audio signal, an environmental noise signal acquired by the audio acquisition device. For example, the first audio signal is a mixed audio signal obtained by superimposing the near-end audio signal, the far-end audio signal, and the environmental noise signal acquired at the same time.
[0037] Specifically, the first audio signal currently acquired by the audio acquisition device is acquired so as to eliminate the echo signal in the first audio signal and minimize the loss of the near-end harmonic when eliminating the echo signal between the near-end harmonics, thereby balancing the echo suppression and the protection of the near-end voice. The second audio signal currently output by the audio output device is acquired so as to estimate the echo signal by taking the second audio signal as a reference signal.
[0038] S120, transforming the first audio signal to obtain an acquisition signal corresponding to each frequency in each time window, and transforming the second audio signal to obtain a reference signal corresponding to each frequency in each time window.
[0039] Specifically, the first audio signal d(n) is subjected to a short-time Fourier transform (STFT) to obtain an acquisition signal D(n, k) corresponding to each frequency k in each time window n, thereby converting the first audio signal in the time domain into an acquisition signal in the frequency domain. The second audio signal x(n) is subjected to a short-time Fourier transform STFT to obtain a reference signal X(n, k) corresponding to each frequency k in each time window n, thereby converting the second audio signal in the time domain into a reference signal in the frequency domain. Wherein n represents a time window (i.e. a time frame) index, and k represents a frequency index.
[0040] Exemplarily, referring to FIG. 2, is the echo path transfer function estimated by the Kalman filter. y(n) is the echo signal estimated by the Kalman filter. e(n) is the error signal obtained after filtering by the adaptive filter, that is, the audio signal obtained by subtracting the echo signal y(n) from the first audio signal d(n). As shown in FIG. 2, the state equation and the observation equation can be defined as follows: D(n, k) ≈ X T (n, k) H * (n, k) + V(n, k) H(n, k) = H(n-1, k) + W(n, k) X(n, k) = [X(n, k), X(n-1, k), …, X(n-L+1, k)] T H(n, k) = [h(0, k), h(1, k), …, h(L-1, k)] T
[0041] wherein V(n, k) refers to the STFT transform of the collected near-end audio signal v(n). H(n, k) is the STFT transform vector of the echo path transfer function. W(n, k) is the echo path transfer function variation of a single time window of the filter, and the meaning is that H(n, k) is a first-order Markov model. The bold variable represents a vector, and the non-bold variable represents a scalar. L is the filter length. The superscript T represents vector transposition, and the superscript * represents complex conjugate.
[0042] S130, based on the Kalman filter, the collected signal corresponding to each frequency in each time window, and the reference information, performing echo cancellation to determine the target signal corresponding to each frequency in each time window after echo cancellation, wherein the optimization function in the Kalman filter is determined based on the posterior state misalignment covariance, the echo path transfer function variation, and the constraint weight corresponding to the echo path transfer function variation, the echo path transfer function variation is determined based on the error signal variance and the Kalman gain, and the constraint weight is determined based on the near-end signal variance and the frequency.
[0043] wherein the Kalman filter is an algorithm for optimally estimating the state of a system using a set of measurements observed over time. Since the observed data includes the influence of noise and interference in the system, the optimal estimation can also be regarded as a filtering process. The target signal E(n, k) refers to the STFT transform of the error signal e(n) obtained after echo cancellation. The optimization function refers to the cost function in the Kalman filter, that is, the optimization object. The posterior state misalignment covariance R μ (n, k) is the echo path transfer function estimated by the Kalman filter The echo path transfer function variation can be used to characterize the optimization step of the Kalman filter. The constraint weight can be an adaptive weighting factor of the echo path transfer function variation. The Kalman gain K(n, k) is used to describe the degree of uncertainty in the Kalman filter. The error signal variance The error signal variance The near-end signal variance
[0044] Specifically, by using the Kalman filter, the collected signal D(n, k) corresponding to each frequency in each time window is subjected to echo cancellation based on the reference signal X(n, k), and the target signal E(n, k) corresponding to each frequency in each time window after echo cancellation is determined. In the process of echo cancellation of the collected signal D(n, k) corresponding to each frequency in each time window, the Kalman filter determines the constraint weight β(n, k) based on the near-end signal variance The error signal variance The echo path transfer function variation is determined based on the Kalman gain K(n, k) and the posterior state misalignment covariance R μ (n, k), the echo path transfer function variation, and the constraint weight β(n, k) to determine the optimization function, and by minimizing the optimization function and iteratively updating the echo path transfer function, the target signal E(n, k) after echo cancellation is obtained. Since the penalty term (i.e., the echo path transfer function variation) related to the Kalman gain is added in the optimization function, the gain can be weighted and constrained, which improves the protection effect of the filter on the near-end speech, and the constraint factor can be adaptively adjusted based on the near-end signal variance And the frequency k, thereby ensuring the echo cancellation effect on the basis of protecting the near-end speech.
[0045] Exemplarily, the optimization function J proposed (n, k) in the Kalman filter can be represented as follows:
[0046] Wherein, the posterior state misalignment covariance R μ (n, k) is obtained by subtracting the echo path transfer function Estimated by the Kalman filter from the actual echo path transfer function H(n, k) and performing mean value processing on the subtraction result, that is, R μ (n, k) = E[μ(n, k)μ T (n, k)].
[0047] wherein the error signal variance the transpose K of the Kalman gain K(n, k) T (n, k) and the Kalman gain K(n, k) to obtain a multiplication result, and the multiplication result is determined as the echo path transfer function variation. The Kalman gain K(n, k) is negatively correlated with the posterior state misadjustment covariance R μ (n, k), that is, the greater K(n, k) is, the smaller R μ (n, k) is. The error signal variance (n, k) in the last time window and the error signal E(n, k) corresponding to the current frequency in the current time window are determined, for example, the error signal variance (n, k) in the last time window can be determined by the following formula:
[0048] For example, the determination process of the current constraint weight β(n, k) corresponding to the current frequency in the current time window includes: based on the filter energy upper limit value, normalizing the near-end signal variance to obtain a normalization result; comparing the current frequency with a preset frequency, and determining a normalization coefficient based on the comparison result; and determining the current constraint weight corresponding to the current frequency in the current time window based on the normalization result and the normalization coefficient.
[0049] Specifically, the near-end signal variance may be divided by the filter energy upper limit value P max , and the multiplication result is mapped between m and 1 to obtain the normalization result, where m is a value greater than or equal to 0 and less than 1. For example, m is 0.1 or 0.01 to avoid the case that the normalization result is too small. The current frequency k is compared with a preset frequency k', and if the current frequency k is less than the preset frequency k', the normalization coefficient is determined as a preset maximum coefficient r1, otherwise the normalization coefficient is determined as a preset minimum coefficient r2. For example, the preset maximum coefficient r1 is 1 and the preset minimum coefficient r2 is 0.05. The preset frequency k' can be determined based on the harmonic frequency of the near-end speech that needs to be protected. For example, in a real-time chorus scene, the first three harmonics of the near-end speech are generally within 1200Hz, and the harmonics in this frequency band need to be protected, so the preset frequency k' can be set to 1200Hz.
[0050] For example, the current constraint weight β(n, k) corresponding to the current frequency in the current time window can be determined by the following formula:
[0051] It should be noted that if the normalized result decreases overall within the target frequency band, the period is likely pure echo, and the value of c(k) can be appropriately reduced. By using the constraint weight β(n,k), the change in the echo path transfer function can be adaptively adjusted, thereby minimizing the protection of near-end speech harmonics without affecting convergence speed.
[0052] Among them, the near-end signal variance corresponding to the current frequency in the current time window is is the error signal E(n,k) corresponding to the current frequency in the current time window, the reference signal X(n,k), and the correlation coefficient r between the error signal E(n,k) and the reference signal X(n,k) ex (n,k), error signal variance and the reference signal variance Sure.
[0053] Specifically, based on the error signal E(n,k) corresponding to the current frequency in the current time window and the error signal variance Error signal variance Update and obtain the updated error signal variance Right now: Based on the reference signal X(n,k) and the reference signal variance Variance of the reference signal Update and obtain the updated reference signal variance Right now: Based on the reference signal X(n,k) corresponding to the current frequency in the current time window, the error signal E(n,k), and the correlation coefficient r between the error signal E(n,k) and the reference signal X(n,k) ex (n,k), correlation coefficient r ex (n,k) is updated to obtain the updated correlation coefficient, namely: r ex (n,k)=(1-α)r ex (n,k)+αX(n,k)E * (n,k). Based on the updated error signal variance Reference signal variance and correlation coefficient r ex (n,k), determines the near-end signal variance corresponding to the current frequency in the current time window Right now: The superscript H indicates the conjugate transpose.
[0054] Exemplarily, normalizing the near-end signal variance based on the filter energy upper limit value to obtain a normalized result may include: calculating a minimum estimated value of the near-end signal variance based on the filter energy upper limit value. The normalization is performed to obtain a normalized result. The minimum estimation value of the near-end signal variance The updated error signal variance The reference signal variance And the correlation coefficient r ex (n, k) is determined. For example, The minimum estimation value of the near-end signal variance is replaced by the near-end signal variance The normalization is performed, which can respond more slowly and further protect the near-end speech.
[0055] S140, inverse transform the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal.
[0056] Specifically, after obtaining the target signal corresponding to each frequency in each time window through the Kalman filter, the STFT inverse transform is performed on the target signal to convert the target signal E(n, k) in the frequency domain into the target audio signal e(n) in the time domain, thereby effectively eliminating the echo signal in the first audio signal, and giving consideration to echo suppression and protection of the near-end speech, especially in an environment with a relatively low signal-to-echo ratio and noise, the near-end speech can be effectively protected from damage, and the echo signal between the near-end harmonics is eliminated, thereby reducing filter misadjustment and improving the robustness of the Kalman filter.
[0057] The technical scheme of the embodiment of the present disclosure transforms the first audio signal collected by the audio collection device to obtain the collected signal corresponding to each frequency in each time window, and transforms the second audio signal output by the audio output device to obtain the reference signal corresponding to each frequency in each time window; based on the Kalman filter, the collected signal corresponding to each frequency in each time window and the reference information, the target signal corresponding to each frequency in each time window after echo cancellation is determined, and the inverse transform is performed on the target signal corresponding to each frequency in each time window to obtain the target audio signal after echo cancellation of the first audio signal. Since the penalty term of the echo path transfer function variation is added to the optimization function of the Kalman filter, and the echo path transfer function variation is determined based on the error signal variance and the Kalman gain, the Kalman gain can be weighted and constrained, the protection effect of the Kalman filter on the near-end speech is improved, and the constraint weight is determined based on the near-end signal variance and the frequency, thereby realizing adaptive constraint of the Kalman gain, ensuring the echo cancellation effect on the basis of protecting the near-end speech, and improving the robustness of the Kalman filter.
[0058] FIG. 3 is a flow diagram of another echo cancellation method provided by an embodiment of the present disclosure. The embodiment of the present disclosure optimizes the step of "performing echo cancellation on each frequency corresponding to the collected signal and the reference information in each time window based on the Kalman filter, and determining the target signal of each frequency corresponding to each time window after echo cancellation" based on the above disclosed embodiments. The explanations of the same or corresponding terms in the above disclosed embodiments are not repeated here.
[0059] As shown in FIG. 3, the echo cancellation method specifically includes the following steps:
[0060] S310, obtaining a first audio signal collected by an audio collection device and a second audio signal output by an audio output device.
[0061] S320, transforming the first audio signal to obtain a collected signal corresponding to each frequency in each time window, and transforming the second audio signal to obtain a reference signal corresponding to each frequency in each time window.
[0062] It should be noted that the improved optimization function J proposed (n, k) is derived, and the Kalman filter can perform echo cancellation by performing the following steps S330-S390. In order to simplify the filtering, it is assumed that the autocorrelation matrix R w (n, k) and the prior state disturbance covariance matrix R m (n, k) are derived based on the unit matrix, that is, R m (n, k) = R m (n, k)I. Wherein, I refers to the unit matrix. The prior state disturbance covariance matrix R m (n, k) is determined based on the actual echo path transfer function H(n, k) in the current time window and the estimated echo path transfer function H R m (n, k) = E[m(n, k)m T (n, k)]. By simplifying R w (n, k) and R m (n, k), a simplified Kalman filtering process can be obtained, thereby improving the filtering efficiency.
[0063] S330, determining a current error signal corresponding to a current frequency in a current time window based on a current collected signal and a current reference signal corresponding to the current frequency in the current time window and a last echo path transfer function corresponding to the current frequency in the last time window estimated by the Kalman filter.
[0064] wherein each frequency in the transformed each time window is echo cancelled as a current frequency in a current time window. The Kalman filter is required to update the filtering iteratively, so that the initial value of the last update result can be set when filtering for the first time, so as to update iteratively.
[0065] Specifically, based on the current reference signal X(n, k) corresponding to the current frequency in the current time window and the last echo path transfer function estimated by the Kalman filter corresponding to the current frequency in the last time window determining the echo signal estimated in the last time window, i.e. and subtracting the echo signal estimated in the last time window from the current acquisition signal D(n, k) to obtain the current error signal E(n, k), i.e.
[0066] S340, based on the current error signal corresponding to the current frequency in the current time window, the current reference signal, the current correlation coefficient between the current error signal and the current reference signal, the current error signal variance and the current reference signal variance, determining the current near-end signal variance corresponding to the current frequency in the current time window.
[0067] Specifically, first, based on the current error signal E(n, k) corresponding to the current frequency in the current time window, the current reference signal X(n, k), the current correlation coefficient r ex (n, k) between the current error signal E(n, k) and the current reference signal X(n, k), the current error signal variance the current reference signal variance and the current correlation coefficient r ex (n, k) are updated, and then based on the updated current error signal variance the current reference signal variance and the current correlation coefficient r ex (n, k), the current near-end signal variance
[0068] Exemplarily, step S340 can include: updating the current error signal variance, the current reference signal variance and the current correlation coefficient between the current error signal and the current reference signal based on the current error signal corresponding to the current frequency in the current time window and the current reference signal; and determining the current near-end signal variance corresponding to the current frequency in the current time window based on the updated current error signal variance, the current reference signal variance and the current correlation coefficient.
[0069] Specifically, based on the current error signal E(n, k) corresponding to the current frequency in the current time window and the current error signal variance the current error signal variance updating the current error signal variance to obtain an updated current error signal variance that is, based on the current reference signal X(n, k) and the current reference signal variance of the current reference signal updating the current reference signal variance to obtain an updated current reference signal variance that is, based on the current reference signal X(n, k), the current error signal E(n, k) and the current correlation coefficient r ex (n, k) between the current error signal E(n, k) and the current reference signal X(n, k), updating the current correlation coefficient r ex (n, k) to obtain an updated current correlation coefficient, that is, r ex (n, k) = (1 - a) r ex (n, k) + aX(n, k)E * (n, k). Based on the updated current error signal variance the current reference signal variance and the current correlation coefficient r ex (n, k), the current near-end signal variance corresponding to the current frequency in the current time window is determined that is, the superscript H represents conjugate transpose, a is a real number greater than 0 and less than 1. The estimates of the current error signal variance, the current reference signal variance and the current near-end signal variance need to be updated at each filtering, and the initial values of these variances can be set to 0.
[0070] S350, based on the current near-end signal variance and the current frequency, determining a current constraint weight corresponding to the current frequency in the current time window.
[0071] wherein the current constraint weight refers to the weight value of the change amount of the echo path transfer function in the optimization function with respect to the current frequency in the current time window, which can be used to adaptively constrain the Kalman gain.
[0072] Specifically, the optimal constraint weight for protecting the near-end speech harmonics is determined based on different near-end signal variances and different frequencies k, so as to weight and constrain the Kalman gain, and ensure the echo cancellation effect on the basis of protecting the near-end speech.
[0073] Exemplarily, the step S350 can comprise: normalizing the current near-end signal variance based on the filter energy upper limit value to obtain a normalization result; comparing the current frequency with a preset frequency, and determining a normalization coefficient based on a comparison result; and determining the current constraint weight corresponding to the current frequency within the current time window based on the normalization result and the normalization coefficient.
[0074] Specifically, the near-end signal variance may be divided by the filter energy upper limit value P max , and a multiplication result is mapped between m and 1 to obtain the normalization result, where m is a value greater than or equal to 0 and less than 1. For example, m is 0.1 or 0.01 to avoid the case that the normalization result is too small. The current frequency k is compared with a preset frequency k', and if the current frequency k is less than the preset frequency k', it is determined that the normalization coefficient is a preset maximum coefficient r1, otherwise it is determined that the normalization coefficient is a preset minimum coefficient r2. For example, the preset maximum coefficient r1 is 1, and the preset minimum coefficient r2 is 0.05. The preset frequency k' can be determined based on the harmonic frequency of the near-end speech that needs to be protected. For example, in a real-time chorus scene, the first three harmonics of the near-end speech are generally within 1200Hz, and the harmonics within this frequency band need to be protected, so the preset frequency k' can be set to 1200Hz.
[0075] For example, the current constraint weight β(n, k) corresponding to the current frequency within the current time window can be determined by the following formula:
[0076] It should be noted that by using the constraint weight β(n, k), the echo path transfer function variation can be adaptively adjusted, that is, the Kalman gain is adaptively adjusted, so as to protect the near-end speech harmonics as much as possible without affecting the convergence speed.
[0077] Exemplarily, the normalization of the current near-end signal variance based on the filter energy upper limit value to obtain a normalization result can comprise: normalizing a minimum estimation value of the current near-end signal variance based on the filter energy upper limit value to obtain a normalization result. Wherein, the minimum estimation value of the current near-end signal variance may also be determined based on the updated current error signal variance , the current reference signal variance , and the current correlation coefficient r ex (n, k). For example, By replacing the near-end signal variance with the minimum estimation value of the near-end signal variance for normalization, a response can be slower, and the near-end speech can be further protected.
[0078] S360 : Determine a current a priori state imbalance covariance corresponding to the current frequency in the current time window based on a previous filter disturbance variance and a previous a posteriori state imbalance covariance corresponding to the current frequency in the previous time window.
[0079] Specifically, the previous filter perturbation variance corresponding to the current frequency in the previous time window is and the previous posterior state imbalance covariance R μ (n-1, k) are added, and the obtained addition result is determined as the current prior state imbalance covariance R m (n,k), that is
[0080] S370 : Determine a current Kalman gain corresponding to a current frequency in a current time window based on a current priori state misalignment covariance, a current reference signal, a current proximal signal variance, a current constraint weight, and a current error signal variance.
[0081] Specifically, based on the current prior state misalignment covariance R m (n,k), current reference signal X(n,k), current near-end signal variance The current constraint weight β(n,k) and the updated current error signal variance Determine the current Kalman gain K(n,k) corresponding to the current frequency in the current time window.
[0082] Exemplarily, step S370 may include: determining the current residual echo signal variance based on the current prior state imbalance covariance and the current reference signal; determining the current residual signal variance based on the current residual echo signal variance, the current near-end signal variance, the current constraint weight and the current error signal variance; and determining the current Kalman gain corresponding to the current frequency in the current time window based on the current prior state imbalance covariance, the current reference signal and the current residual signal variance.
[0083] Specifically, the current prior state misalignment covariance R m (n, k), the conjugate transpose X of the current reference signal H (n, k) is multiplied by the current reference signal X(n, k) to obtain the current residual echo signal variance corresponding to the current frequency in the current time window, that is: R m (n,k)X H (n,k)X(n,k). The current constraint weight β(n,k) is updated with the current error signal variance Multiply them and calculate the multiplication result, the current residual echo signal variance and the current near-end signal variance. Add and obtain the current residual signal variance, that is, The current prior state imbalance covariance Rm (n, k) and the current reference signal X(n, k), and divides the multiplication result by the current residual signal variance, to obtain the current Kalman gain K(n, k) corresponding to the current frequency in the current time window. For example, K(n, k) can be determined by the following formula:
[0084] It should be noted that, since there is no in the formula for determining the Kalman gain K(n, k), the filtering effect only depends on the estimation of and , wherein is positively correlated with the convergence speed and negatively correlated with the steady-state misadjustment. When it is overestimated, the echo suppression is better, but the near-end speech is damaged. is negatively correlated with the convergence speed. When it is overestimated, the near-end speech protection is better but the echo suppression is poor, so it is difficult to balance the echo suppression and the near-end speech protection, especially in an environment with a relatively low signal-to-echo ratio and noise. In view of this, the embodiment performs adaptive constraint on the Kalman gain K(n, k) by using , thereby protecting the near-end speech and ensuring the echo cancellation effect.
[0085] S380, based on the last echo path transfer function, the current Kalman gain and the current error signal, determining the current echo path transfer function estimated by the Kalman filter corresponding to the current frequency in the current time window.
[0086] Specifically, the current Kalman gain K(n, k) and the complex conjugate E * (n, k) are multiplied, and the multiplication result is added to the last echo path transfer function , to obtain the current echo path transfer function estimated by the Kalman filter corresponding to the current frequency in the current time window , that is The echo path transfer function of the adaptive filter is dynamically updated by using the Kalman gain.
[0087] S390, performing echo cancellation based on the current collected signal, the current reference signal and the current echo path transfer function, to determine the target signal corresponding to the current frequency in the current time window after echo cancellation.
[0088] Specifically, the transpose X T (n, k) of the current collected signal is multiplied by the complex conjugate of the current echo path transfer function, and the current collected signal D(n, k) is subtracted from the multiplication result, to obtain the target signal E(n, k) after echo cancellation, that is: The target signal E(n, k) corresponding to the current frequency in the current time window is output to push into the next frame of data, and the operations of steps S330-S390 are repeatedly executed until the target signals corresponding to all frequencies in all time windows are determined.
[0089] S391, the target signal corresponding to each frequency in each time window is inversely transformed to obtain a target audio signal after echo cancellation of the first audio signal.
[0090] The technical scheme of the embodiment of the present disclosure determines the current Kalman gain corresponding to the current frequency in the current time window based on the current prior state misadjustment covariance, the current reference signal, the current near-end signal variance, the current constraint weight, and the current error signal variance, thereby introducing the current constraint weight into the current Kalman gain, realizing adaptive constraint on the Kalman gain, and then on the basis of protecting the near-end voice, trying not to affect the suppression amount of the echo, thereby improving the robustness of the Kalman filter.
[0091] On the basis of each of the above technical schemes, the method further includes steps S392 and S393 as follows:
[0092] S392, based on the current echo path transfer function, the last echo path transfer function, and the filter length, determining a current filter disturbance variance corresponding to the current frequency in the current time window.
[0093] Specifically, after the current echo path transfer function estimated by the Kalman filter is determined Then, the current filter disturbance variance needs to be determined in order to perform filtering processing in the next time window based on the current filter disturbance variance. For example, the current echo path transfer function is subtracted from the last echo path transfer function and the square of the 2-norm of the subtraction result is calculated, and the square result is divided by the filter length L to obtain the current filter disturbance variance That is:
[0094] S393, based on the current prior state misadjustment covariance, the current reference signal, the current Kalman gain, the current constraint weight, and the current error signal variance, determining a current posterior state misadjustment covariance corresponding to the current frequency in the current time window.
[0095] Specifically, after the current Kalman gain K(n, k) is determined, the current filter disturbance variance also needs to be determined, so as to perform filtering in the next time window based on the current filter disturbance variance. Since the current echo path transfer function variation and the current constraint weight are added in the optimization function, the current echo path transfer function variation and the current constraint weight are also added in the formula for determining the current filter disturbance variance, and the current echo path transfer function variation is determined based on the current Kalman gain and the updated current error signal variance.
[0096] For example, step S393 can include: determining a candidate posterior state misadjustment covariance based on the current Kalman gain, the current reference signal and the current prior state misadjustment covariance; determining a current echo path transfer function variation corresponding to the current frequency in the current time window based on the current Kalman gain and the current error signal variance, and multiplying the current echo path transfer function variation by the current constraint weight to obtain a target echo path transfer function variation; determining a current posterior state misadjustment covariance corresponding to the current frequency in the current time window based on the candidate posterior state misadjustment covariance and the target echo path transfer function variation.
[0097] Specifically, the transpose K T (n, k) of the current Kalman gain is multiplied by the conjugate X * (n, k) of the current reference signal, and a ratio of a multiplication result to the filter length L is determined, and a subtraction result obtained by subtracting 1 from the ratio is multiplied by the current prior state misadjustment covariance R m (n, k) to obtain the candidate posterior state misadjustment covariance, that is, The latest current error signal variance The conjugate transpose K H (n, k) of the current Kalman gain is multiplied by the current Kalman gain K(n, k) to obtain the current echo path transfer function variation, and the current echo path transfer function variation is multiplied by the current constraint weight to obtain the target echo path transfer function variation, that is: The candidate posterior state misadjustment covariance is subtracted by the target echo path transfer function variation to obtain the current posterior state misadjustment covariance R μ (n, k), that is:
[0098] It should be noted that in some technical solutions, the filtering manner is to directly determine the candidate posterior state misadjustment covariance as the current posterior state misadjustment covariance R μ (n, k). The current filter disturbance variance is determined by the current echo path transfer function variation and the current constraint weight in the embodiment, so that the echo cancellation effect is ensured on the basis of protecting the near-end speech.
[0099] FIG. 4 is a flowchart of another echo cancellation method provided by an embodiment of the present disclosure. The embodiment of the present disclosure optimizes the step of "determining the current prior state misalignment covariance corresponding to the current frequency in the current time window based on the last filter disturbance variance and the last posterior state misalignment covariance corresponding to the current frequency in the last time window" in the above disclosed embodiments. The explanations of the same or corresponding terms in the above disclosed embodiments are not repeated here.
[0100] As shown in FIG. 4, the echo cancellation method specifically includes the following steps:
[0101] S410, obtaining a first audio signal collected by an audio collection device and a second audio signal output by an audio output device.
[0102] S420, transforming the first audio signal to obtain a collection signal corresponding to each frequency in each time window, and transforming the second audio signal to obtain a reference signal corresponding to each frequency in each time window.
[0103] S430, determining a current error signal corresponding to the current frequency in the current time window based on the current collection signal and the current reference signal corresponding to the current frequency in the current time window, and the last echo path transfer function corresponding to the current frequency in the last time window estimated by the Kalman filter.
[0104] S440, determining a current near-end signal variance corresponding to the current frequency in the current time window based on the current error signal, the current reference signal, a current correlation coefficient between the current error signal and the current reference signal, a current error signal variance, and a current reference signal variance corresponding to the current frequency in the current time window.
[0105] S450, determining a current constraint weight corresponding to the current frequency in the current time window based on the current near-end signal variance and the current frequency.
[0106] It should be noted that, by means of adaptive adjustment based on the adaptive weighted constraint weight, the echo suppression amount can be minimized without affecting the near-end harmonics as much as possible, but if the traditional method is still used to calculate the prior state misalignment covariance, the method will have the problem of insufficient echo suppression amount when there is a path mutation or other transient change. To solve this problem, the embodiment determines the current prior state misalignment covariance corresponding to the current frequency in the current time window by performing the following steps S460-S492.
[0107] S460, determine the first prior state misalignment covariance corresponding to the current frequency in the current time window based on the last filter disturbance variance corresponding to the current frequency in the last time window and the last posterior state misalignment covariance.
[0108] Specifically, the first prior state misalignment covariance can be determined in a conventional manner. For example, the last filter disturbance variance corresponding to the current frequency in the last time window and the last posterior state misalignment covariance R μ (n-1, k) are added to obtain an addition result, and the addition result is determined as the first prior state misalignment covariance R m1 (n, k), that is,
[0109] S470, determine the estimated residual echo signal variance based on the last prior state misalignment covariance corresponding to the current frequency in the last time window and the current reference signal, and determine the estimated residual error signal variance based on the estimated residual echo signal variance and the current near-end signal variance.
[0110] Specifically, the last prior state misalignment covariance R m (n-1, k), the conjugate transpose X H (n, k) of the current reference signal, and the current reference signal X(n, k) are multiplied to obtain the estimated residual echo signal variance, that is, R m (n-1, k)X H (n, k)X(n, k). The estimated residual echo signal variance is added to the current near-end signal variance to obtain an addition result, and the addition result is determined as the estimated residual error signal variance R e (n, k), that is:
[0111] Exemplarily, the “determining the estimated residual error signal variance based on the estimated residual echo signal variance and the current near-end signal variance” in step S470 can include: determining the estimated residual error signal variance based on the estimated residual echo signal variance, the current near-end signal variance, the current constraint weight, and the current error signal variance.
[0112] Specifically, if the current constraint weight is considered, the current constraint weight β(n, k) and the current error signal variance can be multiplied and the multiplication result, the estimated residual echo signal variance, and the current near-end signal variance are added to obtain an addition result, and the addition result is determined as the estimated residual error signal variance R e (n, k), that is,
[0113] S480, determining a current filter steady-state quantity and a current filter transient quantity based on the estimated residual error signal variance, the last prior state misadjustment covariance and the current error signal.
[0114] The current filter steady-state quantity refers to a steady-state term in the Kalman filter, which is determined based on the estimated residual error signal variance and the last prior state misadjustment covariance. The current filter transient quantity refers to a transient term in the Kalman filter, which is determined based on the current error signal.
[0115] For example, step S480 can include determining a current filter equivalent step based on the estimated residual error signal variance and the last prior state misadjustment covariance; determining the current filter steady-state quantity based on the current filter equivalent step and the estimated residual error signal variance, and determining the current filter transient quantity based on the current filter equivalent step and the current error signal.
[0116] Specifically, the last prior state misadjustment covariance R m (n-1, k) is divided by the estimated residual error signal variance R e (n, k) to obtain a current filter equivalent step μ(n, k), that is: The current filter equivalent step μ(n, k) is multiplied by the estimated residual error signal variance R e (n, k) to obtain the current filter steady-state quantity, that is, μ(n, k)R e (n, k). The current filter equivalent step μ(n, k) is multiplied by the square of the current error signal E(n, k) to obtain the current filter transient quantity, that is, μ(n, k)E 2 (n, k).
[0117] S490, determining a current proportion of the estimated residual echo signal variance in the estimated residual error signal variance, and determining a current transient quantity weight and a current steady-state quantity weight based on the last echo path transfer function, the current proportion and the filter length.
[0118] Specifically, the estimated residual echo signal variance is divided by the estimated residual error signal variance to obtain a current proportion γ(n, k), that is: The current proportion γ(n, k) is used to represent the proportion of residual echo in the residual error signal, and its value is between 0 and 1. The current transient quantity weight is determined based on the last echo path transfer function, the current proportion and the filter length. The current transient quantity weight is used to adjust the proportion of the current filter transient quantity in the prior state misadjustment covariance. The current steady-state quantity weight is determined based on the current transient quantity weight. The sum of the current transient quantity weight and the current steady-state quantity weight is equal to 1.
[0119] Exemplarily, the "determining the current transient quantity weight and the current steady quantity weight based on the last echo path transfer function, the current proportion and the filter length" in the step S490 can include: determining a current transient adjustment weight based on the last echo path transfer function and the filter length; determining the current transient quantity weight based on the current transient adjustment weight, the current proportion and the filter length, and determining the current steady quantity weight based on the current transient quantity weight.
[0120] Specifically, based on the filter energy upper limit value H max (the value is pre-set based on actual requirements), the last echo path transfer function is normalized, and the current transient adjustment weight s(n, k) is determined based on the normalization result and the filter length L. For example, the current transient adjustment weight s(n, k) is determined by the formula , wherein the normalization function
[0121] The current transient adjustment weight s(n, k) is used to adjust the proportion of the current filter transient quantity in the current a priori state misadjustment covariance. When a transient change such as a path mutation occurs, a larger current transient adjustment weight can quickly track the system change, and the value interval is [1, L]. The current proportion γ(n, k) is divided by the filter length, and the division result is multiplied by the current transient adjustment weight s(n, k) to obtain the current transient quantity weight. The result of subtracting the current transient quantity weight from 1 is determined as the current steady quantity weight.
[0122] S491, determining the second a priori state misadjustment covariance corresponding to the current frequency in the current time window based on the current transient quantity weight, the current steady quantity weight, the current filter steady quantity and the current filter transient quantity.
[0123] Specifically, the current transient quantity weight is multiplied by the current filter transient quantity, and the current steady quantity weight is multiplied by the current filter steady quantity, and the two multiplication results are added to obtain the second a priori state misadjustment covariance R m2 (n, k). For example, the second a priori state misadjustment covariance R m2 (n, k) is determined by the formula:
[0124] S492, determining the current a priori state misadjustment covariance corresponding to the current frequency in the current time window based on the first a priori state misadjustment covariance and the second a priori state misadjustment covariance.
[0125] Specifically, the first prior state misadjustment covariance and the second prior state misadjustment covariance are compared, and the current prior state misadjustment covariance is determined based on a comparison result. It should be noted that when s(n, k) = 1, the first prior state misadjustment covariance is equal to the second prior state misadjustment covariance.
[0126] Exemplarily, the step S492 can include: determining the maximum value in the first prior state misadjustment covariance and the second prior state misadjustment covariance as the current prior state misadjustment covariance corresponding to the current frequency in the current time window.
[0127] Specifically, the current prior state misadjustment covariance R m (n, k) = max{R m1 (n, k), R m2 (n, k)}. By comparing the first prior state misadjustment covariance and the second prior state misadjustment covariance, the transient characteristics of the prior state misadjustment covariance R m (n, k) can be dynamically adjusted, and the tracking ability of the algorithm can be accelerated when transient changes such as path mutations occur. The greater the prior state misadjustment covariance, the greater the echo suppression, so that the situation that the echo suppression amount is insufficient when transient changes such as path mutations occur can be avoided.
[0128] S493, based on the current prior state misadjustment covariance, the current reference signal, the current near-end signal variance, the current constraint weight, and the current error signal variance, determining a current Kalman gain corresponding to the current frequency in the current time window.
[0129] S494, based on the last echo path transfer function, the current Kalman gain, and the current error signal, determining a current echo path transfer function estimated by the Kalman filter corresponding to the current frequency in the current time window.
[0130] S495, based on the current collected signal, the current reference signal, and the current echo path transfer function, performing echo cancellation to determine a target signal corresponding to the current frequency in the current time window after echo cancellation.
[0131] S496, performing inverse transformation on the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal.
[0132] The technical scheme of the embodiment of the present disclosure determines the second prior state misadjustment covariance corresponding to the current frequency in the current time window based on the current transient quantity weight, the current steady-state quantity weight, the current filter steady-state quantity and the current filter transient quantity, and determines the current prior state misadjustment covariance corresponding to the current frequency in the current time window based on the first prior state misadjustment covariance and the second prior state misadjustment covariance, so that the transient characteristics of the prior state misadjustment covariance can be dynamically adjusted, the algorithm tracking capability can be accelerated when a path mutation or other transient change occurs, and the insufficient echo suppression quantity when the path mutation or other transient change occurs can be avoided, thereby further improving the echo cancellation effect.
[0133] FIG. 5 is a structural schematic diagram of an echo cancellation device provided by an embodiment of the present disclosure. As shown in FIG. 5, the device specifically comprises: an audio signal acquisition module 510, an audio signal transformation module 520, an echo cancellation module 530 and a target signal inverse transformation module 540.
[0134] The audio signal acquisition module 510 is configured to acquire a first audio signal collected by an audio collection device and a second audio signal output by an audio output device, wherein the first audio signal comprises a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal of the second audio signal transmitted through an echo path. The audio signal transformation module 520 is configured to transform the first audio signal to obtain a collection signal corresponding to each frequency in each time window, and transform the second audio signal to obtain a reference signal corresponding to each frequency in each time window. The echo cancellation module 530 is configured to perform echo cancellation on the collection signal and the reference information corresponding to each frequency in each time window based on a Kalman filter to determine a target signal corresponding to each frequency in each time window after echo cancellation, wherein an optimization function in the Kalman filter is determined based on a posterior state misadjustment covariance, an echo path transfer function change quantity and a constraint weight corresponding to the echo path transfer function change quantity, the echo path transfer function change quantity is determined based on an error signal variance and a Kalman gain, and the constraint weight is determined based on a near-end signal variance and a frequency. The target signal inverse transformation module 540 inversely transforms the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal.
[0135] The technical scheme provided by the embodiments of the present disclosure comprises: transforming the first audio signal collected by the audio collection device to obtain a collection signal corresponding to each frequency in each time window, and transforming the second audio signal output by the audio output device to obtain a reference signal corresponding to each frequency in each time window; performing echo cancellation on the collection signal corresponding to each frequency in each time window and the reference information based on a Kalman filter to determine a target signal corresponding to each frequency in each time window after echo cancellation, and inversely transforming the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal. Since the penalty term of the echo path transfer function variation is added to the optimization function of the Kalman filter, and the echo path transfer function variation is determined based on the error signal variance and the Kalman gain, the Kalman gain can be weighted and constrained, the protection effect of the Kalman filter on the near-end speech is improved, and the adaptive constraint of the Kalman gain is realized by determining the constraint weight based on the near-end signal variance and the frequency, so that the echo cancellation effect is ensured on the basis of protecting the near-end speech, and the robustness of the Kalman filter is improved.
[0136] On the basis of the above technical solutions, the echo cancellation module 530 comprises:
[0137] The current error signal determination unit is configured to determine a current error signal corresponding to a current frequency in a current time window based on a current collection signal corresponding to the current frequency in the current time window, a current reference signal, and a last echo path transfer function corresponding to the current frequency in a last time window estimated by the Kalman filter.
[0138] The current near-end signal variance determination unit is configured to determine a current near-end signal variance corresponding to the current frequency in the current time window based on the current error signal corresponding to the current frequency in the current time window, the current reference signal, a current correlation coefficient between the current error signal and the current reference signal, a current error signal variance, and a current reference signal variance.
[0139] The current constraint weight determination unit is configured to determine a current constraint weight corresponding to the current frequency in the current time window based on the current near-end signal variance and the current frequency.
[0140] The current prior state misadjustment covariance determination unit is configured to determine a current prior state misadjustment covariance corresponding to the current frequency in the current time window based on a last filter disturbance variance corresponding to the current frequency in the last time window and a last posterior state misadjustment covariance.
[0141] A current Kalman gain determination unit is configured to determine a current Kalman gain corresponding to a current frequency within a current time window based on the current a priori state misalignment covariance, the current reference signal, the current near-end signal variance, the current constraint weight, and a current error signal variance;
[0142] A current echo path transfer function determination unit is configured to determine a current echo path transfer function estimated by a Kalman filter corresponding to the current frequency within the current time window based on the last echo path transfer function, the current Kalman gain, and the current error signal.
[0143] A target signal determination unit is configured to perform echo cancellation based on the current captured signal, the current reference signal, and the current echo path transfer function to determine a target signal corresponding to the current frequency within the current time window after echo cancellation.
[0144] Based on the above technical solutions, the current near-end signal variance determination unit is specifically configured to:
[0145] The current error signal variance, the current reference signal variance, and a current correlation coefficient between the current error signal and the current reference signal are updated based on the current error signal and the current reference signal corresponding to the current frequency within the current time window; and the current near-end signal variance corresponding to the current frequency within the current time window is determined based on the updated current error signal variance, the current reference signal variance, and the current correlation coefficient.
[0146] Based on the above technical solutions, the current constraint weight determination unit is specifically configured to:
[0147] The current near-end signal variance is normalized based on a filter energy upper limit value to obtain a normalization result; a normalization coefficient is determined by comparing the current frequency with a preset frequency and based on a comparison result; and the current constraint weight corresponding to the current frequency within the current time window is determined based on the normalization result and the normalization coefficient.
[0148] Based on the above technical solutions, the current a priori state misalignment covariance determination unit includes:
[0149] A first a priori state misalignment covariance determination subunit is configured to determine a first a priori state misalignment covariance corresponding to the current frequency within the current time window based on a last filter disturbance variance corresponding to the current frequency within a last time window and a last a posteriori state misalignment covariance.
[0150] The estimated residual signal variance determination sub-unit is configured to determine an estimated residual echo signal variance based on the previous prior state misadjustment covariance corresponding to the current frequency in the previous time window and the current reference signal, and determine an estimated residual signal variance based on the estimated residual echo signal variance and the current near-end signal variance.
[0151] The current filter transient quantity determination sub-unit is configured to determine a current filter steady-state quantity and a current filter transient quantity based on the estimated residual signal variance, the previous prior state misadjustment covariance, and the current error signal.
[0152] The current steady-state quantity weight determination sub-unit is configured to determine a current proportion of the estimated residual echo signal variance in the estimated residual signal variance, and determine a current transient quantity weight and a current steady-state quantity weight based on a previous echo path transfer function, the current proportion, and a filter length.
[0153] The second prior state misadjustment covariance determination sub-unit is configured to determine a second prior state misadjustment covariance corresponding to the current frequency in the current time window based on the current transient quantity weight, the current steady-state quantity weight, the current filter steady-state quantity, and the current filter transient quantity.
[0154] The current prior state misadjustment covariance determination sub-unit is configured to determine a current prior state misadjustment covariance corresponding to the current frequency in the current time window based on the first prior state misadjustment covariance and the second prior state misadjustment covariance.
[0155] In the above technical solutions, the estimated residual signal variance determination sub-unit is specifically configured to determine an estimated residual signal variance based on the estimated residual echo signal variance, the current near-end signal variance, the current constraint weight, and a current error signal variance, and / or
[0156] The current filter transient quantity determination sub-unit is specifically configured to determine a current filter equivalent step based on the estimated residual signal variance and the previous prior state misadjustment covariance, determine a current filter steady-state quantity based on the current filter equivalent step and the estimated residual signal variance, and determine a current filter transient quantity based on the current filter equivalent step and the current error signal, and / or
[0157] The current steady-state quantity weight determination sub-unit is specifically configured to determine a current transient adjustment weight based on a previous echo path transfer function and a filter length, determine a current transient quantity weight based on the current transient adjustment weight, the current proportion, and the filter length, and determine a current steady-state quantity weight based on the current transient quantity weight, and / or
[0158] The current prior state misadjustment covariance determination subunit is specifically configured to: determine the maximum value in the first prior state misadjustment covariance and the second prior state misadjustment covariance as a current prior state misadjustment covariance corresponding to a current frequency in a current time window.
[0159] On the basis of the above technical solutions, the current Kalman gain determination unit is specifically configured to:
[0160] Based on the current prior state misadjustment covariance and the current reference signal, a current residual echo signal variance is determined; based on the current residual echo signal variance, the current near-end signal variance, the current constraint weight, and a current error signal variance, a current residual signal variance is determined; and based on the current prior state misadjustment covariance, the current reference signal, and the current residual signal variance, a current Kalman gain corresponding to a current frequency in a current time window is determined.
[0161] On the basis of the above technical solutions, the echo cancellation module 530 further includes:
[0162] A current filter perturbation variance determination unit is configured to determine a current filter perturbation variance corresponding to a current frequency in a current time window based on the current echo path transfer function, the last echo path transfer function, and a filter length.
[0163] A current posterior state misadjustment covariance determination unit is configured to determine a current posterior state misadjustment covariance corresponding to a current frequency in a current time window based on the current prior state misadjustment covariance, the current reference signal, the current Kalman gain, the current constraint weight, and a current error signal variance.
[0164] On the basis of the above technical solutions, the current posterior state misadjustment covariance determination unit is specifically configured to:
[0165] Based on the current Kalman gain, the current reference signal, and the current prior state misadjustment covariance, a candidate posterior state misadjustment covariance is determined; based on the current Kalman gain and a current error signal variance, a current echo path transfer function change amount corresponding to a current frequency in a current time window is determined, and the current echo path transfer function change amount is multiplied by the current constraint weight to obtain a target echo path transfer function change amount; and based on the candidate posterior state misadjustment covariance and the target echo path transfer function change amount, a current posterior state misadjustment covariance corresponding to a current frequency in a current time window is determined.
[0166] The echo cancellation device provided in the embodiments of the present disclosure can execute the echo cancellation method provided in any of the embodiments of the present disclosure, and has the corresponding function modules and beneficial effects of executing the echo cancellation method.
[0167] It is noted that the units and modules included in the above apparatus are only divided according to functional logic, and are not limited to the above division, as long as the corresponding functions can be implemented; in addition, the specific names of the functional units are only for the convenience of mutual differentiation, and do not serve to limit the protection scope of the embodiments of the present disclosure.
[0168] FIG. 6 is a structural schematic diagram of an electronic device according to an embodiment of the present disclosure. Referring to FIG. 6, a structural schematic diagram of an electronic device (e.g., a terminal device or a server in FIG. 6) 500 suitable for implementing an embodiment of the present disclosure is shown. The terminal device in the embodiments of the present disclosure can include, but is not limited to, a mobile terminal such as a mobile phone, a notebook computer, a digital broadcast receiver, a PDA (Personal Digital Assistant), a PAD (Tablet Personal Computer), a PMP (Portable Multimedia Player), a vehicle terminal (e.g., a car navigation terminal), and the like, and a fixed terminal such as a digital TV, a desktop computer, and the like. The electronic device shown in FIG. 6 is only an example, and should not impose any limitation on the functions and use range of the embodiments of the present disclosure.
[0169] As shown in FIG. 6, the electronic device 500 can include a processing device (e.g., a central processing unit, a graphics processing unit, etc.) 501, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 502 or loaded into a random access memory (RAM) 503 from a storage device 508. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. An input / output (I / O) interface 505 is also connected to the bus 504.
[0170] Generally, the following devices can be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, and the like; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, and the like; a storage device 508 including, for example, a magnetic tape, a hard disk, and the like; and a communication device 509. The communication device 509 can allow the electronic device 500 to communicate with other devices wirelessly or by wire to exchange data. Although FIG. 6 shows the electronic device 500 with various devices, it is understood that all the shown devices are not required to be implemented or provided. More or fewer devices can be alternatively implemented or provided.
[0171] In particular, according to embodiments of the present disclosure, the processes described above with reference to the flowcharts can be implemented as a computer software program. For example, embodiments of the present disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for executing the methods illustrated by the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network by the communication apparatus 509, or installed from the storage apparatus 508, or installed from the ROM 502. When the computer program is executed by the processing apparatus 501, the above-mentioned functions defined in the methods of the embodiments of the present disclosure are executed.
[0172] The names of the messages or information exchanged between the plurality of devices in the embodiments of the present disclosure are only for illustrative purposes, and are not intended to limit the scope of the messages or information.
[0173] The electronic device provided by the embodiments of the present disclosure and the echo cancellation method provided by the above-described embodiments belong to the same inventive concept, and the technical details not described in detail in the present embodiment can be referred to the above-described embodiments, and the present embodiment has the same beneficial effects as the above-described embodiments.
[0174] The embodiments of the present disclosure provide a computer storage medium, which stores a computer program, and the program is executed by a processor to implement the echo cancellation method provided by the above-described embodiments.
[0175] It should be noted that the computer-readable medium described above can be a computer-readable signal medium or a computer-readable storage medium or any combination thereof. The computer-readable storage medium can be, for example, but not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus or device, or any suitable combination of the above. More specific examples of the computer-readable storage medium can include, but are not limited to, an electrical connection having one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present disclosure, the computer-readable storage medium can be any tangible medium that contains or stores a program used by or in connection with an instruction execution system, apparatus or device. In the present disclosure, the computer-readable signal medium can include a data signal propagated in baseband or propagated as a carrier wave in a propagated data signal, in which the computer-readable program code is contained. Such a propagated data signal can take many forms, including but not limited to, an electromagnetic signal, an optical signal, or any suitable combination of the above. The computer-readable signal medium can also be any computer-readable medium that can send, propagate or transfer the program for use by or in connection with the instruction execution system, apparatus or device. The program code contained in the computer-readable medium can be transmitted by any suitable medium, including but not limited to, wire, cable, RF (radio frequency), etc., or any suitable combination of the above.
[0176] In some embodiments, the client, server, or both can communicate using any current known or future developed network protocol, such as HTTP (HyperText Transfer Protocol), and can be interconnected with any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include local area networks ("LAN"), wide area networks ("WAN"), the Internet, and peer-to-peer networks (e.g., ad hoc peer-to-peer networks), as well as any current known or future developed networks.
[0177] The computer-readable medium described above can be included in the electronic device described above; or can exist separately from the electronic device, and not be assembled into the electronic device.
[0178] The computer readable medium described above carries one or more programs, when the one or more programs are executed by the electronic device, cause the electronic device to: acquire a first audio signal collected by an audio collection device and a second audio signal output by an audio output device, wherein the first audio signal includes a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal of the second audio signal transmitted through an echo path; transform the first audio signal to obtain a collection signal corresponding to each frequency in each time window, and transform the second audio signal to obtain a reference signal corresponding to each frequency in each time window; perform echo cancellation on the collection signal and the reference information corresponding to each frequency in each time window based on a Kalman filter to determine a target signal corresponding to each frequency in each time window after echo cancellation, wherein an optimization function in the Kalman filter is determined based on a posterior state misadjustment covariance, an echo path transfer function variation, and a constraint weight corresponding to the echo path transfer function variation, the echo path transfer function variation is determined based on an error signal variance and a Kalman gain, and the constraint weight is determined based on a near-end signal variance and a frequency; and inverse transform the target signal corresponding to each frequency in each time window to obtain a target audio signal after echo cancellation of the first audio signal.
[0179] Computer program code for carrying out operations of the present disclosure can be written in any combination of one or more programming languages, including an object oriented programming language such as Java, Smalltalk, C++ or the like, and conventional procedural programming languages, such as the "C" programming language or similar programming languages. The program code can execute entirely on the user's computer, partly on the user's computer, as a stand-alone software package, partly on the user's computer and partly on a remote computer or entirely on the remote computer or server. In the latter scenario, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or the connection can be made to an external computer (for example, through the Internet using an Internet Service Provider).
[0180] The computer program product of the first aspect can include one or more non-transitory computer-readable media storing instructions that, when executed, cause one or more processors to perform the operations of the first aspect. The computer program product of the first aspect can include a non-transitory computer-readable medium storing code that, when executed, causes a computer to perform operations for the first aspect.
[0181] The units described in the embodiments of the present disclosure can be implemented by software, or by hardware. In some cases, the name of the unit does not constitute a limitation on the unit itself. For example, the first obtaining unit can also be described as a unit for obtaining at least two Internet protocol addresses.
[0182] The functions described in this specification can be implemented in part or in whole by one or more hardware logic components. For example, and without limitation, illustrative types of hardware logic components that can be used include Field-programmable Gate Arrays (FPGAs), Program-specific Integrated Circuits (ASICs), Program-specific Standard Products (ASSPs), System-on-a-chip systems (SOCs), Complex Programmable Logic Devices (CPLDs), etc.
[0183] In the context of the present disclosure, a machine-readable medium can be a tangible medium that contains or stores a program for use by or in connection with an instruction execution system, apparatus, or device. The machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include but is not limited to an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium will include one or more lines of electrical connections, portable computer disks, hard disk drives, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or Flash memory), optical fiber, portable compact disc read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0184] The above description merely illustrates the preferred embodiments of the disclosure and a principle for applying the technologies. It is understood by those skilled in the art that the disclosed scope of the disclosure is not limited to the technical solutions formed by the specific combinations of the technical features described above, and should also cover other technical solutions formed by the combinations of the technical features described above or their equivalent features without departing from the disclosed concept. For example, the technical solutions formed by the mutual replacement of the above-described features and the technical features with similar functions disclosed in the disclosure (but not limited to) can be formed.
[0185] Further, although operations are depicted in a particular, sequential order, this should not be understood as requiring or implying that the operations are performed in the order illustrated or sequentially. In certain circumstances, multitasking and parallel processing can be advantageous. Likewise, although specific implementation details are included for the purpose of providing a thorough disclosure, these should not be construed as limitations on the scope of the disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single embodiment. Conversely, various features that are described in the context of a single embodiment can also be implemented in multiple embodiments separately or in any suitable sub-combination.
[0186] Although the subject matter has been described in language specific to structural features and / or methodological acts, it is to be understood that the subject defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are disclosed as example forms of implementing the claims.
Claims
1. An echo cancellation method, comprising: Acquire a first audio signal collected by an audio collection device and a second audio signal output by an audio output device, wherein the first audio signal includes a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal of the second audio signal transmitted through an echo path; transforming the first audio signal to obtain a collection signal corresponding to each frequency in each time window, and transforming the second audio signal to obtain a reference signal corresponding to each frequency in each time window; performing echo cancellation based on a Kalman filter, an acquired signal corresponding to each frequency in each time window, and reference information, and determining a target signal corresponding to each frequency in each time window after echo cancellation, wherein the optimization function in the Kalman filter is determined based on a posterior state misalignment covariance, a change in an echo path transfer function, and a constraint weight corresponding to the change in the echo path transfer function, the change in the echo path transfer function is determined based on an error signal variance and a Kalman gain, and the constraint weight is determined based on a near-end signal variance and frequency; An inverse transform is performed on the target signal corresponding to each frequency in each time window to obtain a target audio signal after the echo of the first audio signal is cancelled.
2. The echo cancellation method according to claim 1, wherein: The performing echo cancellation based on the Kalman filter, the collected signal corresponding to each frequency in each time window, and the reference information, and determining the target signal corresponding to each frequency in each time window after the echo cancellation includes: Determine a current error signal corresponding to the current frequency in the current time window based on a current acquisition signal and a current reference signal corresponding to the current frequency in the current time window and a previous echo path transfer function corresponding to the current frequency in the previous time window estimated by a Kalman filter; Determine a current near-end signal variance corresponding to the current frequency within the current time window based on a current error signal corresponding to the current frequency within the current time window, a current reference signal, a current correlation coefficient between the current error signal and the current reference signal, a current error signal variance, and a current reference signal variance; Determining a current constraint weight corresponding to the current frequency in the current time window based on the current near-end signal variance and the current frequency; Determine a current a priori state misalignment covariance corresponding to the current frequency in the current time window based on a previous filter perturbation variance and a previous a posteriori state misalignment covariance corresponding to the current frequency in the previous time window; Determining a current Kalman gain corresponding to a current frequency in a current time window based on the current prior state misalignment covariance, the current reference signal, the current proximal signal variance, the current constraint weight, and the current error signal variance; determining a current echo path transfer function corresponding to a current frequency within a current time window estimated by a Kalman filter based on the previous echo path transfer function, the current Kalman gain, and the current error signal; Echo cancellation is performed based on the current acquisition signal, the current reference signal, and the current echo path transfer function, and a target signal corresponding to a current frequency in a current time window after echo cancellation is determined.
3. The echo cancellation method according to claim 2, wherein: The determining, based on a current error signal corresponding to a current frequency within a current time window, a current reference signal, a current correlation coefficient between the current error signal and the current reference signal, a current error signal variance, and a current reference signal variance, of a current near-end signal corresponding to a current frequency within a current time window, includes: Based on the current error signal and the current reference signal corresponding to the current frequency in the current time window, updating the current error signal variance, the current reference signal variance, and the current correlation coefficient between the current error signal and the current reference signal; Based on the updated current error signal variance, the current reference signal variance, and the current correlation coefficient, a current near-end signal variance corresponding to the current frequency in the current time window is determined.
4. The echo cancellation method according to claim 2 or 3, wherein: The determining, based on the current near-end signal variance and the current frequency, a current constraint weight corresponding to the current frequency in the current time window includes: Normalizing the current near-end signal variance based on the filter energy upper limit value to obtain a normalized result; Comparing the current frequency with a preset frequency and determining a normalization coefficient based on the comparison result; Based on the normalization result and the normalization coefficient, a current constraint weight corresponding to the current frequency in the current time window is determined.
5. The echo cancellation method according to any one of claims 2 to 4, wherein: The determining, based on the last filter disturbance variance and the last a posteriori state imbalance covariance corresponding to the current frequency in the last time window, the current a priori state imbalance covariance corresponding to the current frequency in the current time window includes: Determine a first a priori state imbalance covariance corresponding to the current frequency in the current time window based on a previous filter disturbance variance and a previous a posteriori state imbalance covariance corresponding to the current frequency in the previous time window; Determining an estimated residual echo signal variance based on a previous priori state misalignment covariance corresponding to the current frequency in the previous time window and the current reference signal, and determining an estimated residual signal variance based on the estimated residual echo signal variance and the current near-end signal variance; Determining a current filter steady-state quantity and a current filter transient quantity based on the estimated residual signal variance, the previous priori state misalignment covariance, and the current error signal; Determining a current proportion of the estimated residual echo signal variance in the estimated residual signal variance, and determining a current transient quantity weight and a current steady-state quantity weight based on a previous echo path transfer function, the current proportion, and a filter length; Determining a second priori state imbalance covariance corresponding to a current frequency in a current time window based on the current transient quantity weight, the current steady-state quantity weight, the current filter steady-state quantity, and the current filter transient quantity; Based on the first a priori state imbalance covariance and the second a priori state imbalance covariance, a current a priori state imbalance covariance corresponding to a current frequency in a current time window is determined. The echo cancellation method according to claim 5 , wherein: The determining, based on the estimated residual echo signal variance and the current near-end signal variance, an estimated residual signal variance includes: determining an estimated residual signal variance based on the estimated residual echo signal variance, the current near-end signal variance, the current constraint weight and the current error signal variance; and / or, The determining of a current filter steady-state quantity and a current filter transient quantity based on the estimated residual signal variance, the previous priori state imbalance covariance, and the current error signal comprises: Determining a current filter equivalent step size based on the estimated residual signal variance and the previous prior state misalignment covariance; Determining a current filter steady-state quantity based on the current filter equivalent step size and the estimated residual signal variance, and determining a current filter transient quantity based on the current filter equivalent step size and the current error signal; and / or, The determining of the current transient quantity weight and the current steady-state quantity weight based on the previous echo path transfer function, the current proportion and the filter length includes: Determining the current transient adjustment weight based on the previous echo path transfer function and filter length; Determining a current transient quantity weight based on the current transient adjustment weight, the current proportion, and the filter length, and determining a current steady-state quantity weight based on the current transient quantity weight; and / or, The determining, based on the first a priori state imbalance covariance and the second a priori state imbalance covariance, a current a priori state imbalance covariance corresponding to a current frequency in a current time window includes: The maximum value of the first a priori state imbalance covariance and the second a priori state imbalance covariance is determined as the current a priori state imbalance covariance corresponding to the current frequency in the current time window.
7. The echo cancellation method according to any one of claims 2 to 6, wherein: The determining, based on the current prior state misalignment covariance, the current reference signal, the current proximal signal variance, the current constraint weight, and the current error signal variance, a current Kalman gain corresponding to a current frequency in a current time window includes: determining a current residual echo signal variance based on the current prior state misalignment covariance and the current reference signal; Determining a current residual signal variance based on the current residual echo signal variance, the current near-end signal variance, the current constraint weight, and the current error signal variance; A current Kalman gain corresponding to a current frequency in a current time window is determined based on the current prior state misalignment covariance, the current reference signal, and the current residual signal variance.
8. The echo cancellation method according to any one of claims 2 to 7, further comprising: determining a current filter disturbance variance corresponding to a current frequency in a current time window based on the current echo path transfer function, the previous echo path transfer function, and a filter length; A current a posteriori state imbalance covariance corresponding to a current frequency in a current time window is determined based on the current a priori state imbalance covariance, the current reference signal, the current Kalman gain, the current constraint weight, and the current error signal variance.
9. The echo cancellation method according to claim 8, wherein: The determining, based on the current a priori state imbalance covariance, the current reference signal, the current Kalman gain, the current constraint weight, and the current error signal variance, a current a posteriori state imbalance covariance corresponding to the current frequency in the current time window includes: Determining a to-be-selected a posteriori state imbalance covariance based on the current Kalman gain, the current reference signal, and the current a priori state imbalance covariance; determining a current echo path transfer function change corresponding to a current frequency within a current time window based on the current Kalman gain and the current error signal variance, and multiplying the current echo path transfer function change by the current constraint weight to obtain a target echo path transfer function change; Based on the candidate a posteriori state imbalance covariance and the target echo path transfer function change, a current a posteriori state imbalance covariance corresponding to the current frequency in the current time window is determined.
10. An echo cancellation device, comprising: an audio signal acquisition module configured to acquire a first audio signal collected by an audio acquisition device and a second audio signal output by an audio output device, wherein the first audio signal includes a near-end audio signal and a far-end audio signal, and the far-end audio signal is an echo signal of the second audio signal transmitted through an echo path; an audio signal conversion module configured to convert the first audio signal to obtain a sampled signal corresponding to each frequency in each time window, and to convert the second audio signal to obtain a reference signal corresponding to each frequency in each time window; an echo cancellation module configured to perform echo cancellation based on a Kalman filter, an acquired signal corresponding to each frequency in each time window, and reference information, and determine a target signal corresponding to each frequency in each time window after echo cancellation, wherein the optimization function in the Kalman filter is determined based on a posterior state misalignment covariance, a change in an echo path transfer function, and a constraint weight corresponding to the change in the echo path transfer function, the change in the echo path transfer function is determined based on an error signal variance and a Kalman gain, and the constraint weight is determined based on a near-end signal variance and frequency; The target signal inverse transformation module is configured to perform inverse transformation on the target signal corresponding to each frequency in each time window to obtain the target audio signal after the echo of the first audio signal is eliminated.
11. An electronic device comprising: one or more processors; A storage device configured to store one or more programs, wherein When the one or more programs are executed by the one or more processors, the one or more processors are enabled to implement the echo cancellation method according to any one of claims 1 to 9.
12. A storage medium containing computer-executable instructions, wherein: When the computer executable instructions are executed by a computer processor, the computer executable instructions are used to perform the echo cancellation method according to any one of claims 1 to 9.
Citation Information
Patent Citations
Self-adaptive sound echo cancellation method based on frequency domain Kalman filtering
CN108806709A
Variable-step-size echo cancellation method based on rapid convergence characteristic
CN109754813A
Method and device for eliminating echo signal, computing equipment and storage medium
CN113763977A
Echo cancellation method and device, equipment and storage medium
CN116647621A
Echo cancellation method and device, equipment and storage medium
CN118136037A