Noise joint processing method and device, electronic equipment and readable storage medium

By performing linear filtering on far-end and near-end signals and fusing sub-band feature signals, the problem of difficulty in jointly processing echo, noise and reverberation in existing technologies is solved, achieving efficient noise reduction on small embedded platforms.

CN116246646BActive Publication Date: 2026-04-17SHENZHEN EMEET TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
SHENZHEN EMEET TECH CO LTD
Filing Date
2023-02-15
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies struggle to process echoes, noise, and reverberation together, and the computational demands on small embedded platforms are too high to achieve good noise reduction results.

Method used

By acquiring far-end and near-end signals and performing linear filtering, sub-band feature signals and speech signals are extracted, spliced ​​and fused, the target sub-band gain and speech filter coefficients are determined, and nonlinear noise residual signals are filtered out to achieve joint noise processing.

Benefits of technology

It reduces algorithm complexity, ensures the integrity of the sound signal, improves the full-duplex effect of echo cancellation, and achieves good noise suppression and reverberation suppression.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116246646B_ABST
    Figure CN116246646B_ABST
Patent Text Reader

Abstract

This application discloses a noise joint processing method, apparatus, electronic device, and readable storage medium. The noise joint processing method includes: extracting a first sub-band feature signal and a first speech signal from a nonlinear noise residual signal, and extracting a second sub-band feature signal and a second speech signal from a far-end signal; fusing the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and fusing the first speech signal and the second speech signal to obtain a second input feature signal; fusing the first input feature signal and the second input feature signal to obtain a fused feature signal; determining a target sub-band gain and target speech filter coefficients based on the fused feature signal, the first input feature signal, and the second input feature signal; and obtaining a sound signal based on the target speech filter coefficients and the target sub-band gain. This application solves the technical problem of difficulty in jointly processing echo, noise, and reverberation while achieving good noise reduction effect.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of communication technology, and in particular to a noise joint processing method, apparatus, electronic device, and readable storage medium. Background Technology

[0002] In recent years, with the rapid development of the internet, the demand for remote work, online healthcare, and online education has exploded, making products and technologies with good call performance increasingly important. Call-related technologies mainly consist of three key components: echo cancellation, noise suppression, and reverberation suppression. Current mainstream echo cancellation technologies generally use adaptive filters for linear echo processing, utilizing signal similarity to determine residual echoes and further suppress them. However, speakers experience vibration and distortion during playback, making it difficult to express this using general mathematical models. While suppressing residual echoes, it also suppresses useful near-end voice signals, failing to achieve full-duplex performance. Furthermore, noise suppression and reverberation suppression have generally adopted deep learning methods in recent years. However, deep learning suffers from high computational costs and data mismatches between modules, making it difficult to run all three functional modules simultaneously on a small embedded platform. Summary of the Invention

[0003] The main objective of this application is to provide a noise joint processing method, apparatus, electronic device, and readable storage medium, aiming to solve the technical problem in the prior art that it is difficult to jointly process echo, noise, and reverberation and achieve good noise reduction effect.

[0004] To achieve the above objectives, this application provides a joint noise processing method, the joint noise processing method comprising:

[0005] Acquire the far-end signal and the near-end signal, and perform linear filtering on the near-end signal and the far-end signal to obtain a nonlinear noise residual signal;

[0006] Extract the first sub-band feature signal and the first speech signal from the nonlinear noise residual signal, and extract the second sub-band feature signal and the second speech signal from the far-end signal;

[0007] The first sub-band feature signal and the second sub-band feature signal are spliced ​​and fused to obtain the first input feature signal, and the first speech signal and the second speech signal are spliced ​​and fused to obtain the second input feature signal;

[0008] The first input feature signal and the second input feature signal are fused to obtain a fused feature signal;

[0009] Based on the fused feature signal and the first input feature signal, the target subband gain is determined, and based on the fused feature signal and the second input feature signal, the target speech filter coefficients are determined.

[0010] Based on the target speech filter coefficients and the target subband gain, the nonlinear noise signal in the nonlinear noise residual signal is filtered out to obtain the sound signal after noise joint processing.

[0011] To achieve the above objectives, this application also provides a noise joint processing apparatus, the noise joint processing apparatus comprising:

[0012] A linear processing module is used to acquire far-end signals and near-end signals, and to perform linear filtering on the near-end signals and the far-end signals to obtain nonlinear noise residual signals;

[0013] The signal extraction module is used to extract a first sub-band feature signal and a first speech signal from the nonlinear noise residual signal, and to extract a second sub-band feature signal and a second speech signal from the far-end signal.

[0014] The signal splicing module is used to splice and fuse the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and to splice and fuse the first speech signal and the second speech signal to obtain a second input feature signal;

[0015] An input feature determination module is used to fuse the first input feature signal and the second input feature signal to obtain a fused feature signal;

[0016] The coefficient and gain determination module is used to determine the target subband gain based on the fused feature signal and the first input feature signal, and to determine the target speech filter coefficients based on the fused feature signal and the second input feature signal.

[0017] The nonlinear processing module is used to filter out the nonlinear noise signal in the nonlinear noise residual signal according to the target speech filter coefficients and the target subband gain, so as to obtain the sound signal after joint noise processing.

[0018] This application also provides an electronic device, the electronic device comprising: a memory, a processor, and a program of the noise joint processing method stored in the memory and executable on the processor, wherein when the program of the noise joint processing method is executed by the processor, it can implement the steps of the noise joint processing method as described above.

[0019] This application also provides a computer-readable storage medium storing a program for implementing a noise joint processing method, wherein when the program for the noise joint processing method is executed by a processor, it implements the steps of the noise joint processing method as described above.

[0020] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the noise joint processing method described above.

[0021] This application provides a noise joint processing method, apparatus, electronic device, and readable storage medium. The noise joint processing method includes: acquiring a far-end signal and a near-end signal; performing linear filtering on the near-end signal and the far-end signal to obtain a nonlinear noise residual signal; extracting a first sub-band feature signal and a first speech signal from the nonlinear noise residual signal, and extracting a second sub-band feature signal and a second speech signal from the far-end signal; concatenating and fusing the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and concatenating and fusing the first speech signal and the second speech signal to obtain a second input feature signal; fusing the first input feature signal and the second input feature signal to obtain a fused feature signal; determining a target sub-band gain based on the fused feature signal and the first input feature signal, and determining target speech filter coefficients based on the fused feature signal and the second input feature signal; and filtering out nonlinear noise signals in the nonlinear noise residual signal based on the target speech filter coefficients and the target sub-band gain to obtain a noise-jointly processed sound signal. Because it can perform linear filtering on both far-end and near-end signals, it initially performs linear noise reduction on the signals acquired by the microphone, thus obtaining the nonlinear noise residual signal. Furthermore, because it can extract sub-band signals and speech signals separately, and the sub-band signal is a coarse-grained extraction of the far-end signal and the nonlinear noise residual signal, it reduces the dimensionality of the input features of the sub-band signal, thereby reducing the algorithm's complexity. Then, it extracts and processes the speech signal, ensuring that no effective sound signal is lost. Further, by splicing the sub-band signals of the far-end signal and the nonlinear noise residual signal with the speech signal, the sub-band signal and the speech signal are fused. The target sub-band gain and target speech filter coefficients can then be determined using the fused signal, ensuring no signal loss. Finally, the residual noise signal in the nonlinear noise residual signal is filtered out using the filter coefficients to obtain the noise-suppressed sound signal. Through linear and nonlinear filtering of the far-end and near-end signals, noise suppression and reverberation suppression are achieved, and the full-duplex effect of echo cancellation is improved. Additionally, by extracting sub-band features, the complexity of noise processing is reduced. Therefore, this application solves the technical problem in the prior art of being unable to jointly process echo, noise and reverberation and achieve good noise reduction effect. Attached Figure Description

[0022] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0023] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0024] Figure 1 This is a flowchart illustrating the first embodiment of the noise joint processing method of this application;

[0025] Figure 2 This is a flowchart illustrating the second embodiment of the noise joint processing method of this application;

[0026] Figure 3 This is a flowchart illustrating the third embodiment of the noise joint processing method of this application;

[0027] Figure 4 This is a schematic diagram of the structure of an embodiment of the noise joint processing device of this application;

[0028] Figure 5 This is a schematic diagram of a scenario for the noise joint processing method of this application;

[0029] Figure 6 This is a schematic diagram of subband feature signal extraction in the joint noise processing method of this application;

[0030] Figure 7 This is a flowchart illustrating the process of obtaining the target subband gain and target speech filter coefficients in the joint noise processing of this application;

[0031] Figure 8 This is a schematic diagram of the device structure of the hardware operating environment involved in the noise joint processing method in the embodiments of this application.

[0032] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0033] To make the above-mentioned objectives, features, and advantages of this application more apparent and understandable, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0034] Example 1

[0035] Reference Figure 1This application provides a noise joint processing method. In a first embodiment of the noise joint processing method, the noise joint processing method includes:

[0036] Step S10: Acquire the far-end signal and the near-end signal, and perform linear filtering on the near-end signal and the far-end signal to obtain a nonlinear noise residual signal;

[0037] Step S20: Extract the first sub-band feature signal and the first speech signal from the nonlinear noise residual signal, and extract the second sub-band feature signal and the second speech signal from the far-end signal;

[0038] Step S30: The first sub-band feature signal and the second sub-band feature signal are spliced ​​and fused to obtain the first input feature signal, and the first speech signal and the second speech signal are spliced ​​and fused to obtain the second input feature signal;

[0039] Step S40: Fuse the first input feature signal and the second input feature signal to obtain a fused feature signal;

[0040] Step S50: Determine the target subband gain based on the fused feature signal and the first input feature signal, and determine the target speech filter coefficients based on the fused feature signal and the second input feature signal;

[0041] Step S60: Based on the target speech filter coefficients and the target subband gain, filter out the nonlinear noise signal in the nonlinear noise residual signal to obtain the sound signal after noise joint processing.

[0042] In this embodiment, it should be noted that in a call scenario, the main interference signals affecting the call are: echo, noise, and reverberation. The far-end signal is the electrical signal collected by the speaker, and the near-end signal is all signals received by the microphone. Both the far-end and near-end signals include echo, noise, and reverberation. (Refer to...) Figure 5 , Figure 5 This is a schematic diagram of a joint noise processing method. 1 represents a microphone, and 2 represents a speaker. The near-end signal collected by microphone 1 and the far-end signal collected by speaker 2 are fed into deep learning model 3. Deep learning model 3 performs AEC (Acoustic Echo Cancellation), Dereverb (reverberation removal), and NS (Noise Suppression) on the near-end signal. The sound signal from speaker 2 travels through an echo path to microphone 1. During echo cancellation, the far-end signal d needs to be considered. k .

[0043] As an example, the expression for the near-end signal is:

[0044] X k =echo+noise+voice*rir

[0045] X in the expression k The deep learning model extracts the voice signal from the near-end signal, achieving echo cancellation, noise suppression, and reverberation suppression. Here, the near-end signal is represented by: echo (the sound signal from the speaker reaching the microphone via the echo path), noise (additive noise), voice (the clean sound signal after echo cancellation, noise suppression, and reverberation suppression), rir (the room impulse response), and voice*rir (the convolution of the voice signal and the room impulse response rir).

[0046] Additionally, it should be noted that the linear filtering of the near-end signal can be performed using a block-based adaptive filter to obtain a nonlinear noise residual signal. This nonlinear noise residual signal is the signal after linear filtering and includes both nonlinear noise and audio signals. The first sub-band feature signal is a coarse-grained signal extracted from the nonlinear noise residual signal. The first audio signal is a mid-to-low frequency signal extracted from the nonlinear noise residual signal, typically within the frequency range of human voice concentration, generally between 0Hz and 2000Hz. This mid-to-low frequency signal can be adjusted according to the audio signal of interest. Correspondingly, the second sub-band feature is a coarse-grained signal extracted from the far-end signal, and the second audio feature is a coarse-grained signal extracted from the far-end signal. The low-to-mid frequency signal extracted from the sound is generally within the frequency range of human voice concentration, typically between 0Hz and 2000Hz. This range can be adjusted based on the sound signal of interest. The first input feature signal is obtained by splicing and fusing the first sub-band feature and the second sub-band feature. The first and second sub-band features can be fused and encoded to obtain the first input feature signal. The first and second input feature signals can be input into a recurrent neural network. The recurrent neural network processes the first and second input features to obtain a fused feature signal. The fused feature signal is a compressed signal obtained by fusing the far-end signal and the nonlinear noise residual signal. This feature signal can reduce the amount of computation and ensure that no sound signal is lost.

[0047] The target subband gain is a coarse-grained sound signal free of echo, reverberation, and noise, while the target speech filter coefficients are fine-grained sound signals free of echo, reverberation, and noise. Both the target subband gain and the target speech filter coefficients are used to filter nonlinear noise. The target subband gain includes target subband signal echo gain, target subband signal reverberation gain, and target subband signal noise gain. The target speech filter coefficients include target speech signal echo filter coefficients, target speech signal reverberation filter coefficients, and target speech signal noise filter coefficients. The target subband signal reverberation gain includes the gain at the current time point and the gain at the previous time point, and the target speech signal reverberation filter coefficients include the filter coefficients at the current time point and the previous time point, thereby achieving a good reverberation suppression effect. The sound signal is a clean sound signal that has undergone joint noise processing.

[0048] As an example, steps S10 to S60 include: acquiring a far-end signal and a near-end signal; performing linear filtering on the far-end signal and the near-end signal to obtain a nonlinear noise residual signal; decomposing the nonlinear noise residual signal to obtain a first sub-band feature signal and extracting a first speech signal from the nonlinear noise residual signal; decomposing the far-end signal to obtain a second sub-band feature signal and extracting a second speech signal from the far-end signal; concatenating the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal and concatenating the first speech signal and the second speech signal to obtain a second input feature signal; fusing the first input feature signal and the second input feature signal to obtain a fused feature signal; calculating the target sub-band gain based on the fused feature signal and the first input feature signal and calculating the target speech filter coefficients based on the feature signal and the second input feature signal; and filtering out the nonlinear noise residual signal based on the target speech filter coefficients and the target sub-band gain to obtain a noise-suppressed sound signal. For example, linear filtering can use a block adaptive filter to filter out the near-end signal and the far-end signal, which can suppress some linear echoes.

[0049] As an example, the steps for obtaining a noise-free residual signal by performing linear filtering on the near-end signal and the far-end signal are as follows:

[0050] The near-end and far-end signals are divided into blocks to obtain the near-end and far-end signals. Each block is called a frame. The Kth block X of the near-end signal is read. k , and the data block d of the Kth block of the remote signal k In this embodiment, the frame length of the data block can be set to 32ms, and the frame shift can be set to 16ms. The near-end signal X of the Kth block... kThe buffer of the input block adaptive filter has a maximum length of P. The formula for performing a Fast Fourier Transform on the near-end signal is as follows:

[0051]

[0052] Xf in the Fourier transform formula k The result of performing a Fourier transform on the Kth proximal signal, Xf k Characterized as a near-end signal in the frequency domain, FFT (Fast Fourier Transform) is the Fast Fourier Transform, X k-p For the KP-th data block of the near-end signal, X k-p+1 For the (K-P+1)th data block of the near-end signal, X k-1 For the (K-1)th data block of the near-end signal, X k For the Kth data block of the near-end signal, further, the current filter coefficients of the block adaptive filter are W. k For the current filter coefficients W of the block adaptive filter k and Xf k Perform IFFT (Inverse Fast)

[0053] The formula for the Fourier Transform (Inverse Fast Fourier Transform) is as follows:

[0054]

[0055] For the current filter coefficients W k and Xf k The result of the IFFT, y k According to W k and Xf k The error between the filter's far-end signal and the original signal obtained after IFFT is calculated using the following formula:

[0056] e = y k -d k

[0057] e represents the error between the far-end signal and the far-end signal of the filter, and y represents the error between the far-end signal and the far-end signal. k For the signal at the far end of the filter, d k For the far-end signal of the loudspeaker, the formula for calculating the gradient of the block adaptive filter is as follows:

[0058]

[0059] Let be the gradient of the block adaptive filter, where To Perform a Fast Fourier Transform. To Performing an inverse Fast Fourier transform, the formula for updating the filter coefficients is as follows:

[0060]

[0061] W k+1 For the updated filter coefficients, W k Here, μ represents the current filter coefficients, and μ is the preset update step size.

[0062] To Perform a Fast Fourier Transform.

[0063] This application provides a noise joint processing method, which includes: acquiring a far-end signal and a near-end signal; performing linear filtering on the near-end signal and the far-end signal to obtain a nonlinear noise residual signal; extracting a first sub-band feature signal and a first speech signal from the nonlinear noise residual signal, and extracting a second sub-band feature signal and a second speech signal from the far-end signal; concatenating and fusing the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and concatenating and fusing the first speech signal and the second speech signal to obtain a second input feature signal; fusing the first input feature signal and the second input feature signal to obtain a fused feature signal; determining a target sub-band gain based on the fused feature signal and the first input feature signal, and determining target speech filter coefficients based on the fused feature signal and the second input feature signal; and filtering out nonlinear noise signals in the nonlinear noise residual signal based on the target speech filter coefficients and the target sub-band gain to obtain a noise-jointly processed sound signal. Because it can perform linear filtering on both far-end and near-end signals, it initially performs linear noise reduction on the signals acquired by the microphone, thus obtaining the nonlinear noise residual signal. Furthermore, because it can extract sub-band signals and speech signals separately, and the sub-band signal is a coarse-grained extraction of the far-end signal and the nonlinear noise residual signal, it reduces the dimensionality of the input features of the sub-band signal, thereby reducing the algorithm's complexity. Then, it extracts and processes the speech signal, ensuring that no effective sound signal is lost. Further, by splicing the sub-band signals of the far-end signal and the nonlinear noise residual signal with the speech signal, the sub-band signal and the speech signal are fused. The target sub-band gain and target speech filter coefficients can then be determined using the fused signal, ensuring no signal loss. Finally, the residual noise signal in the nonlinear noise residual signal is filtered out using the filter coefficients to obtain the noise-suppressed sound signal. Through linear and nonlinear filtering of the far-end and near-end signals, noise suppression and reverberation suppression are achieved, and the full-duplex effect of echo cancellation is improved. Additionally, by extracting sub-band features, the complexity of noise processing is reduced. Therefore, this application solves the technical problem in the prior art of being unable to jointly process echo, noise and reverberation and achieve good noise reduction effect.

[0064] Example 2

[0065] Furthermore, referring to Figure 2Based on the above embodiments of this application, in another embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, the steps of extracting the first sub-band feature signal and the first speech signal from the nonlinear noise residual signal, and extracting the second sub-band feature signal and the second speech signal from the far-end signal, include:

[0066] Step A10: Based on the preset sampling frequency and the preset number of sub-band features, extract the first sub-band feature signal from the nonlinear noise residual signal and extract the second sub-band feature signal from the far-end signal.

[0067] Step A20: Extract the first speech signal from the nonlinear noise residual signal and extract the second speech signal from the far-end signal according to the preset spectrum.

[0068] In this embodiment, it should be noted that the first sub-band feature signal includes at least one of a preset sampling frequency and a preset number of sub-bands, and the second speech signal includes a preset spectrum. The preset sampling frequency and the preset number of sub-bands are used to extract the sub-band signal, and the preset spectrum is used to extract the speech signal. The extraction of the first and second sub-band feature signals can be performed using Fourier transform. First, a preset sampling frequency can be set, the length of the Fourier transform and the number of signal frequency points can be determined, and then the number of frequency points in each sub-band can be calculated based on the number of sub-band feature signals to determine the starting frequency point of the sub-band. Thus, the sub-band feature signal is determined based on the number of frequency points and the starting frequency point of the sub-band. The preset spectrum can be determined based on the desired human voice frequency band, thereby extracting either the first or second speech signal based on the preset spectrum.

[0069] As an example, steps A10 to A20 include: determining the starting frequency point in the nonlinear noise residual signal and the starting frequency point in the far-end signal based on a preset sampling frequency and a preset number of sub-band features; extracting first sub-band features based on the nonlinear noise residual signal and its starting frequency point; extracting second sub-band features based on the far-end signal and its starting frequency point; and extracting a first speech signal from the nonlinear noise residual signal and a second speech signal from the far-end signal based on the preset spectrum. In this embodiment, by extracting the first sub-band feature signal, the first speech signal, the second sub-band feature signal, and the second speech signal from the far-end noise and the nonlinear noise residual signal respectively, the integrity of the sound is ensured while reducing the computational load of noise processing.

[0070] As an example, refer to Figure 6 , Figure 6This diagram illustrates the extraction of sub-band feature signals in a joint noise processing method. The steps for extracting the first and second sub-band feature signals are as follows:

[0071] Perform a short-time Fourier transform on the near-end signal to obtain a frequency-domain near-end signal converted from the time domain to the frequency domain. Perform a short-time Fourier transform on the far-end signal to obtain a frequency-domain far-end signal converted from the time domain to the frequency domain. Further, refer to... Figure 6 Line 1 contains the amplitude spectrum `mag` of the near-end or far-end signal in the frequency domain. Specifically, when extracting the first sub-band feature signal, the amplitude spectrum of the near-end signal is determined; when extracting the second sub-band feature signal, the amplitude spectrum of the far-end signal is determined. This determines the length of the Fourier transform (`fft_size`) and the number of frequency points in the Fourier transform (`fft_size / 2` + 1). (Refer to...) Figure 6 Line 2, the content of which represents the calculation of the maximum sub-band scale (sub-band scale) based on the preset sampling frequency sr; refer to Figure 6 Line 3 contains information about calculating the subband step size (step) based on the preset number of subband features (nb_bands) and the maximum subband scale (sub_high). Further details can be found by referring to... Figure 6 Lines 5 to 11 contain information representing the calculation of the upper frequency limit f of each sub-band, and based on the upper frequency limit f, the number of frequency points in each sub-band is further calculated as bands_widths (sub-band width); see reference Figure 6 Line 12 contains information representing the calculation of the sub-band starting frequency b_pts based on the number of frequency points in each sub-band; see reference. Figure 6 Rows 13 to 15 contain information representing the construction of a frequency point mask matrix freq2band (freq2 frequency band) based on the sub-band starting frequency point b_pts and the number of sub-band frequency points bands_widths; refer to... Figure 6 Line 17 contains the following: based on the amplitude spectrum `mag` and the frequency mask matrix `freq2band`, the sub-band feature signal `band_feat` is calculated. When the amplitude spectrum `mag` represents the amplitude spectrum of the near-end signal in the frequency domain, the sub-band feature signal `band_feat` can be the first sub-band feature signal; when the amplitude spectrum `mag` represents the amplitude spectrum of the far-end signal in the frequency domain, the sub-band feature signal `band_feat` can be the second sub-band feature signal. For example, where... Figure 6In the algorithm freq2band (frequency conversion band algorithm), when the preset sampling frequency is 16000, the Fourier transform length is 512, and the preset number of sub-band features is 16, the number of frequency points for the Fast Fourier Transform is 257. Therefore, the starting frequency point of each sub-band is:

[0072] [0,2,4,7,11,15,21,37,48,61,79,100,127,161,203]

[0073] The number of frequency points in each sub-band is:

[0074] [2,2,3,4,4,6,7,9,11,13,18,21,27,34,42,54]

[0075] Based on the sub-band start frequency and the number of sub-band frequency points, construct a frequency point mask matrix with a size of [fft_size / 2, nb_bands]. Calculate the sub-band signal feature band_feat based on the frequency point mask matrix.

[0076] The steps of concatenating and fusing the first sub-band feature signal and the second sub-band feature signal to obtain the first input feature signal, and concatenating and fusing the first speech signal and the second speech signal to obtain the second input feature signal, include:

[0077] Step B10: The first sub-band feature signal and the second sub-band feature signal are fused and encoded to obtain the first input feature signal;

[0078] Step B20: The first speech signal and the second speech signal are fused and encoded to obtain the second input feature signal.

[0079] As an example, steps B10 to B20 include: fusing and encoding the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal; and fusing and encoding the first speech signal and the second speech signal to obtain a second input feature signal. Exemplarily, the first sub-band feature signal and the second sub-band feature signal can be fed into a one-dimensional convolution to perform sub-band signal fusion encoding, converting it into a first input feature signal learned by the neural network; the first speech signal and the second speech signal can be fed into a one-dimensional convolution to perform speech signal fusion encoding, converting it into a second input feature signal learned by the neural network. Here, the first input feature signal is a sub-band fusion feature learned by the neural network, and the second input feature signal is a speech signal fusion feature learned by the neural network.

[0080] The step of fusing the first input feature signal and the second input feature signal to obtain a fused feature signal includes:

[0081] Step C10: The first input feature signal is compressed and encoded step by step to obtain the target sub-band coded feature;

[0082] Step C20: The second input feature signal is compressed and encoded step by step to obtain the target speech coding features;

[0083] Step C30: The target subband coding features and the target speech coding features are concatenated and fused to obtain the fused feature signal.

[0084] In this embodiment, it should be noted that the first input feature signal includes at least one of the following: a first intermediate sub-band coding feature, a second intermediate sub-band coding feature, a third intermediate sub-band coding feature, and a target sub-band coding feature. The second input feature includes a target speech coding feature. Since the far-end signal and the nonlinear noise residual signal are prone to signal loss after multiple encodings, the first intermediate sub-band coding feature, the second intermediate sub-band coding feature, the third intermediate sub-band coding feature, and the target sub-band coding feature can be used to decode and recover the first input feature signal, and the target speech coding feature can be used to recover the second input feature signal.

[0085] As an example, steps C10 to C30 include: encoding the first input feature signal to obtain a first intermediate sub-band coded feature; encoding the first intermediate sub-band coded feature to obtain a second intermediate sub-band coded feature; encoding the second intermediate sub-band coded feature to obtain a third intermediate sub-band coded feature; encoding the third intermediate sub-band coded feature to obtain a target sub-band coded feature; encoding the second input feature signal to obtain a first intermediate speech coded feature; encoding the first intermediate speech coded feature to obtain a target speech coded feature; and fusing the target sub-band coded feature and the target speech coded feature to obtain a fused feature signal. For example, the first input feature signal can be convolved in two dimensions to obtain a first intermediate sub-band coding feature; the first intermediate sub-band coding feature can be convolved in two dimensions to obtain a second intermediate sub-band coding feature; the second intermediate sub-band coding feature can be convolved in two dimensions to obtain a third intermediate sub-band coding feature; the third intermediate sub-band coding feature can be convolved in two dimensions to obtain a target sub-band coding feature; the second input feature signal can be convolved in two dimensions to obtain a first intermediate speech coding feature; the first intermediate speech coding feature can be convolved in two dimensions to obtain a target speech coding feature; the target sub-band coding feature and the target speech coding feature can be fused and encoded to obtain a fused feature signal. For example, the target sub-band coding feature and the target speech coding feature can be input into a recurrent neural network, and then the target sub-band coding feature and the target speech coding feature can be fused together to obtain a fused feature signal.

[0086] In this embodiment, a first sub-band feature signal, a first speech signal, a second sub-band feature signal, and a second speech signal are extracted from the far-end noise and the nonlinear noise residual signal, respectively. The first and second sub-band feature signals are then fused to obtain a first input feature, and the first and second speech signals are fused to obtain a second input feature. Furthermore, the first and second input features are progressively compressed and then fused again to obtain a fused feature signal. Since the first, second, first, and second sub-band feature signals, the first speech signal, and the second speech signal can be extracted, and the first sub-band feature signal is a coarse-dimensional extraction of the nonlinear noise residual signal, and the second sub-band feature signal is a coarse-dimensional extraction of the far-end signal, the computational load for noise processing is reduced. Furthermore, the extraction of sub-band signals improves the smoothness of the spectrum processing. Additionally, the dimensions of the extracted sub-band signals can be adjusted according to the current device configuration. Moreover, by extracting speech signals from the far-end signal and the nonlinear noise residual signal, the integrity of the sound is ensured.

[0087] Example 3

[0088] Furthermore, referring to Figure 3Based on the above embodiments of this application, in another embodiment of this application, the same or similar content as the above embodiments can be referred to the above description, and will not be repeated hereafter. Based on this, the step of determining the target sub-band gain according to the fused feature signal and the first input feature signal includes:

[0089] Step D10: Perform a linear transformation on the fused feature signal to obtain the linear transformation feature;

[0090] Step D20: The linear transform feature and the target subband coding feature are fused and decoded step by step to obtain the target subband gain.

[0091] In this embodiment, it should be noted that the first input feature signal includes target subband coding features, the fused feature signal includes linear transform features, and the target subband gain includes at least one of a first intermediate subband signal gain, a second intermediate subband signal gain, a third intermediate subband signal gain, and a fourth intermediate subband signal gain. The first intermediate subband signal gain, the second intermediate subband signal gain, the third intermediate subband signal gain, and the fourth intermediate subband signal gain are intermediate results for generating the target subband gain. The process of determining the target subband gain is the process of recovering the sound signal, that is, performing step-by-step fusion decoding on the target subband coding features.

[0092] As an example, steps D10 to D20 include: performing a linear transformation on the fused feature signal to obtain a linearly transformed feature; performing a linear transformation on the fused feature signal to obtain the linearly transformed feature; fusing and decoding the linearly transformed feature and the target subband coding feature to obtain a first intermediate subband signal gain; fusing and decoding the first intermediate subband signal gain and the target subband coding feature to obtain a second intermediate subband signal gain; fusing and decoding the second intermediate subband signal gain and the third intermediate subband coding feature to obtain a third intermediate subband signal gain; fusing and decoding the third intermediate subband signal gain and the second intermediate subband coding feature to obtain a fourth intermediate subband signal gain; and fusing and decoding the fourth intermediate subband signal gain and the first intermediate subband coding feature to obtain the target subband gain. Exemplarily, the fused feature signal can be input into a linear layer for linear transformation, and the fusing and decoding of the linearly transformed feature and the target subband coding feature can be performed by adding the linearly transformed feature and the target subband coding feature and then performing a deconvolution to obtain the first intermediate subband signal gain.

[0093] The step of determining the target speech filter coefficients based on the fused feature signal and the second input feature signal includes:

[0094] Step E10: Perform a linear transformation on the fused feature signal to obtain the linear transformation feature;

[0095] Step E20: Recover and decode the linear transform features to obtain the coefficients of the first intermediate speech filter;

[0096] Step E30: Perform a linear transformation on the coefficients of the first intermediate speech filter to obtain the coefficients of the second intermediate speech filter;

[0097] Step E40: The second intermediate speech filter coefficients and the target speech coding features are concatenated and fused to obtain the target speech filter coefficients.

[0098] In this embodiment, it should be noted that the second input feature includes target speech coding features, the fused feature signal includes linear transformation features, and the target speech filter coefficients include at least one of a first intermediate speech filter coefficient and a second intermediate speech filter coefficient. The first intermediate speech filter coefficient and the second intermediate speech filter coefficient are intermediate results generated when decoding and recovering the sound signal, and the target speech filter coefficient is a speech signal without interference.

[0099] As an example, steps E10 to E40 include: performing a linear transformation on the fused feature signal to obtain the linearly transformed feature; performing recovery decoding on the linearly transformed feature to obtain the first intermediate speech filter coefficients; performing a linear transformation on the first intermediate speech filter coefficients to obtain the second intermediate speech filter coefficients; and concatenating and fusing the second intermediate speech filter coefficients and the target speech coding feature to obtain the target speech filter coefficients. The fused feature signal can be input into a linear layer for linear transformation.

[0100] The step of filtering out nonlinear noise signals from the nonlinear noise residual signal based on the target speech filter coefficients and the target subband gain to obtain the noise-jointly processed sound signal includes:

[0101] Step F10: Convert the nonlinear noise residual signal into a frequency domain signal to obtain a complex spectrum nonlinear noise residual signal;

[0102] Step F20: Interpolate the target subband gain to obtain the target frequency subband signal gain;

[0103] Step F30: Nonlinearly filter out the complex spectrum nonlinear noise residual signal according to the target frequency subband signal gain to obtain the first sound signal;

[0104] Step F40: The first sound signal is nonlinearly filtered out according to the target speech filter coefficients to obtain the sound signal after noise joint processing.

[0105] In this embodiment, it should be noted that the first audio signal is obtained through nonlinear filtering. Because the target subband gain is a coarse-grained clean subband characteristic signal, that is, the bandwidth of the target subband gain, the filtered interference signal is not filtered out for interference signals at every frequency point. Therefore, the first audio signal is a signal obtained after coarse noise processing. The interference signal includes at least one of echo, noise, and reverberation. The target speech filter coefficients are clean speech signals concentrated in the human voice frequency band. Nonlinear filtering of the first audio signal based on the target speech filter coefficients further filters out the residual echo, noise, and reverberation in the first audio signal to achieve the effects of echo cancellation, noise suppression, and reverberation suppression.

[0106] As an example, steps F10 to F40 include: converting the nonlinear noise residual signal into a frequency domain signal to obtain a complex spectrum nonlinear noise residual signal; interpolating the target subband gain to compensate for the frequency points of the target subband gain to obtain the target frequency point subband signal gain; multiplying the target frequency point subband signal gain with the complex spectrum nonlinear noise residual signal to filter out nonlinear noise and obtain a first sound signal; and multiplying the target speech filter coefficients with the first sound signal to filter out nonlinear noise and obtain a sound signal after noise joint processing. When the target frequency subband signal gain is multiplied by the complex spectrum nonlinear noise residual signal, the remaining signal is the audio signal corresponding to the target subband gain. When the target speech filter coefficients are multiplied by the complex spectrum nonlinear noise residual signal, the remaining signal is the audio signal corresponding to the target speech filter coefficients. The nonlinear noise residual signal can be converted into a frequency domain signal by performing a Fourier transform. Both the first and second audio signals include echo cancellation and noise and reverberation suppression processes. After processing by the target subband gain and the target speech filter coefficients, the nonlinear noise residual signal is obtained as an audio signal after echo cancellation, noise suppression, and reverberation suppression.

[0107] As an example, the formula for multiplying the subband signal gain at the target frequency by the residual complex spectrum nonlinear noise signal is as follows:

[0108] Y G (k1,f1)=E(k2,f2)*G(k3,f3)

[0109] Y G(k1,f1) represents the first audio signal, E(k2,f2) represents the complex spectrum nonlinear noise residual signal, G(k3,f3) represents the target frequency subband signal gain, k1 represents the data block corresponding to the first audio signal, f1 represents the frequency point corresponding to the first audio signal, k2 represents the data block corresponding to the complex spectrum nonlinear noise residual signal, f2 represents the frequency point corresponding to the complex spectrum nonlinear noise residual signal, k3 represents the data block corresponding to the target frequency subband signal gain, and f3 represents the frequency point corresponding to the target frequency subband signal gain. For example, the frame length of the data block k1 corresponding to the first audio signal can be set to 32ms, and the frame shift can be set to 16ms. Similarly, the frame lengths of k2 and k3 can also be set to 32ms, and the frame shift can also be set to 16ms.

[0110] As an example, the formula for multiplying the target speech filter coefficients by the first sound signal is:

[0111]

[0112] Y(k4,f df1 () is the second sound signal. Y represents the target speech filter coefficients. G (k6,f df3 k1 is the second input feature signal, which can be the complex spectrum fusion feature of the speech signal; k2 is the data block corresponding to the second sound signal; k3 is the data block corresponding to the target speech filter coefficients; k4 is the data block corresponding to the second input feature signal; f is the data block corresponding to the second input feature signal. df1 The frequency point corresponding to the second sound signal, f df2 The frequency points corresponding to the target speech filter coefficients, f df3 The frequency point corresponding to the second input feature signal.

[0113] As an example, refer to Figure 7 , Figure 7 This is a flowchart for obtaining the target subband gain and target speech filter coefficients in joint noise processing, with the far-end signal d. k and near-end microphone signal X k All signals enter a block-domain adaptive filter, which can adjust the far-end signal d. k and near-end signal X k Linear filtering is performed to obtain the nonlinear noise residual signal e. kThe nonlinear residual noise signal is subjected to STFT (short-time Fourier transform) to extract the first sub-band feature signal and the first speech signal. The far-end signal is subjected to STFT to extract the second sub-band feature signal and the second speech signal. The STFT is a piecewise FFT transform. The first and second sub-band feature signals are concatenated to obtain the sub-band concatenated signal. The first and second speech signals are concatenated to obtain the speech concatenated signal. The sub-band concatenated signal is fed into a one-dimensional convolution for sub-band fusion coding to obtain the first input feature signal. The speech concatenated signal is fed into a one-dimensional convolution for speech complex spectrum fusion coding to obtain the second input feature signal. The first input feature signal is subjected to a first two-dimensional convolution Sconv2D_1 to obtain the first intermediate sub-band coding feature e1. The intermediate subband coding feature e1 is subjected to a second two-dimensional convolution Sconv2D_2 and downsampled to obtain the second intermediate subband coding e2. The second intermediate subband coding e2 is subjected to a third two-dimensional convolution Sconv2D_3 and downsampled to obtain the second intermediate subband coding e3. The third intermediate subband coding e3 is subjected to a fourth two-dimensional convolution Sconv2D_4 to obtain the target subband coding feature e4. The second input feature signal is subjected to a first two-dimensional convolution Cconv2D_1 to obtain the first intermediate speech coding feature c1. The first intermediate speech coding feature c1 is subjected to a second two-dimensional convolution Cconv2D_2 to obtain the second intermediate speech coding feature c2. The second intermediate speech coding feature c2 is fed into the linear layer Linear to obtain the target speech coding feature c3. The target speech coding feature c3 and the target sub-band coding feature e4 are concatenated and input into a GRU (Gate Recurrent Unit) to obtain a fused feature signal emb. The target sub-band coding feature e4 and the fused feature signal emb are then fed into a linear layer to obtain the first intermediate sub-band signal gain te4+e4. te4+e4 is obtained by deconvolving the first intermediate sub-band signal gain and adding it back. The first intermediate sub-band signal gain te4+e4 is then subjected to a first two-dimensional deconvolution STconv2D_1 to obtain the second intermediate sub-band signal gain te3+e3. The second intermediate sub-band signal gain te3+e3 is then subjected to a second two-dimensional deconvolution STconv2D_2. The third intermediate sub-band signal gain te3+e3 is then subjected to a third two-dimensional deconvolution STconv2D_3 to obtain the fourth intermediate sub-band signal gain te1+e1. The fourth intermediate sub-band signal gain te1+e1 is then subjected to a fourth two-dimensional deconvolution to obtain the target sub-band gain G. band (k,b), where G bandIn (k,b), k represents the Kth data block, and b represents the frequency band of the sub-band feature signal. The target speech coding feature c3 and the fused feature signal emb are fed into the linear layer to obtain the first intermediate speech filter coefficient tc1. The first intermediate speech filter coefficient tc1 is then fed into the GRU to obtain the second intermediate speech filter coefficient tc2. The second intermediate speech filter coefficient tc2 is then fed into the linear layer to obtain the target speech filter coefficient C. df (k,i,f df ), where, in C df (k,i,f df In this context, k represents the k-th data block, i represents the coefficient of the i-th target speech filter, and f df The frequency points corresponding to the target speech filter coefficients.

[0114] In this embodiment, the nonlinear noise signal in the nonlinear noise residual signal is processed by obtaining the target subband gain to obtain a first sound signal that has undergone echo cancellation, noise suppression, and reverberation suppression. This reduces the computational load of noise processing and achieves simultaneous processing of echo, noise, and reverberation. Then, the first sound signal is processed by obtaining the target speech filter coefficients to obtain a sound signal that has undergone echo cancellation, noise suppression, and reverberation suppression. This achieves simultaneous processing of echo, noise, and reverberation. After two nonlinear filtering steps, a good noise reduction effect can be achieved.

[0115] Example 4

[0116] Reference Figure 4 This application also provides a noise joint processing device, the noise joint processing device comprising:

[0117] Linear processing module 10 is used to acquire far-end signals and near-end signals, and perform linear filtering on the near-end signals and far-end signals to obtain nonlinear noise residual signals;

[0118] Signal extraction module 20 is used to extract a first sub-band feature signal and a first speech signal from the nonlinear noise residual signal, and to extract a second sub-band feature signal and a second speech signal from the far-end signal.

[0119] The signal splicing module 30 is used to splice and fuse the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and to splice and fuse the first speech signal and the second speech signal to obtain a second input feature signal;

[0120] The input feature determination module 40 is used to fuse the first input feature signal and the second input feature signal to obtain a fused feature signal;

[0121] The coefficient and gain determination module 50 is used to determine the target subband gain based on the fused feature signal and the first input feature signal, and to determine the target speech filter coefficients based on the fused feature signal and the second input feature signal.

[0122] The nonlinear processing module 60 is used to filter out the nonlinear noise signal in the nonlinear noise residual signal according to the target speech filter coefficients and the target subband gain, so as to obtain the sound signal after joint noise processing.

[0123] Optionally, the signal extraction module 20 is further configured to:

[0124] Based on the preset sampling frequency and the preset number of sub-band features, the first sub-band feature signal is extracted from the nonlinear noise residual signal and the second sub-band feature signal is extracted from the far-end signal;

[0125] Based on the preset spectrum, the first speech signal is extracted from the nonlinear noise residual signal, and the second speech signal is extracted from the far-end signal.

[0126] Optionally, the signal splicing module 30 is further configured to:

[0127] The first sub-band feature signal and the second sub-band feature signal are fused and encoded to obtain the first input feature signal;

[0128] The first speech signal and the second speech signal are fused and encoded to obtain the second input feature signal.

[0129] Optionally, the input feature determination module 40 is further configured to:

[0130] The target sub-band coded features are obtained by progressively compressing and encoding the first input feature signal.

[0131] The target speech coding features are obtained by progressively compressing and encoding the second input feature signal.

[0132] The target subband coding features and the target speech coding features are fused and encoded to obtain the fused feature signal.

[0133] Optionally, the coefficient and gain determination module 50 is further configured to:

[0134] The fused feature signal is linearly transformed to obtain the linearly transformed feature;

[0135] The linear transform features and the target subband coding features are fused and decoded step by step to obtain the target subband gain.

[0136] Optionally, the coefficient and gain determination module 50 is further configured to:

[0137] The fused feature signal is linearly transformed to obtain the linearly transformed feature;

[0138] The linear transform features are recovered and decoded to obtain the coefficients of the first intermediate speech filter;

[0139] The coefficients of the first intermediate speech filter are linearly transformed to obtain the coefficients of the second intermediate speech filter.

[0140] The second intermediate speech filter coefficients and the target speech coding features are concatenated and fused to obtain the target speech filter coefficients.

[0141] Optionally, the nonlinear processing module 60 is further configured to:

[0142] The nonlinear noise residual signal is converted into a frequency domain signal to obtain a complex spectrum nonlinear noise residual signal;

[0143] Interpolate the target subband gain to obtain the target frequency subband signal gain;

[0144] The complex spectrum nonlinear noise residual signal is nonlinearly filtered out based on the subband signal gain at the target frequency point to obtain the first sound signal;

[0145] The first sound signal is nonlinearly filtered out according to the target speech filter coefficients to obtain a sound signal after joint noise processing.

[0146] The noise combined processing apparatus provided in this application employs the noise combined processing method in the above embodiments, solving the technical problem in the prior art that it is difficult to combine echo, noise, and reverberation for combined processing and achieve good noise reduction effect. Compared with the prior art, the beneficial effects of the noise combined processing method provided in this application are the same as those of the noise combined processing method provided in the above embodiments, and other technical features in this noise combined processing apparatus are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0147] Example 5

[0148] This invention provides an electronic device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the platform interface flow control method in Embodiment 1 above.

[0149] The following is for reference. Figure 8 The diagram illustrates a structural schematic of an electronic device suitable for implementing embodiments of the present disclosure. The electronic devices in the embodiments of the present disclosure may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (portable Android devices), PMPs (portable media players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 8 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0150] like Figure 8 As shown, the electronic device may include a processing unit 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in ROM (Read-Only Memory) 1002 or a program loaded from storage device 1003 into RAM (Random Access Memory) 1004. RAM 1004 also stores various programs and data required for the operation of the electronic device. The processing unit 1001, ROM 1002, and RAM 1004 are interconnected via bus 1005. Input / output (I / O) interface 1006 is also connected to the bus.

[0151] Typically, the following systems can be connected to I / O interface 1006: input devices 1007 including, for example, touchscreens, touchpads, keyboards, mice, image sensors, microphones, tachometers, gyroscopes, etc.; output devices 1008 including, for example, LCDs (Liquid Crystal Displays), speakers, vibrators, etc.; storage devices 1003 including, for example, magnetic tapes, hard disks, etc.; and communication devices 1009. The communication devices allow electronic devices to communicate wirelessly or wiredly with other devices to exchange data. Although electronic devices with various systems are shown in the figures, it should be understood that it is not required to implement or possess all the systems shown. More or fewer systems may be implemented alternatively.

[0152] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication device 1009, or installed from storage device 1003, or installed from ROM 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of embodiments of this disclosure.

[0153] The electronic device provided by this invention employs the noise joint processing method in the above embodiments, solving the technical problem in the prior art of difficulty in jointly processing echo, noise, and reverberation while achieving good noise reduction effect. Compared with the prior art, the beneficial effects of the electronic device provided by the embodiments of this invention are the same as those of the noise joint processing method provided in the above embodiments, and other technical features of this electronic device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0154] It should be understood that various parts of this disclosure can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics may be combined in any suitable manner in one or more embodiments or examples.

[0155] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0156] Example 6

[0157] This embodiment provides a computer-readable storage medium having computer-readable program instructions stored thereon, which are used to perform the noise joint processing method in Embodiment 1 above.

[0158] The computer-readable storage medium provided in this application embodiment may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor devices, apparatuses, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections with one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable EPROM (Electrical Programmable Read Only Memory) or flash memory, optical fiber, portable compact disk CD-ROM (compact discread-only memory), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution device, apparatus, or apparatus. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0159] The aforementioned computer-readable storage medium may be included in an electronic device or may exist independently without being assembled into an electronic device.

[0160] The aforementioned computer-readable storage medium carries one or more programs that, when executed by an electronic device, cause the electronic device to: acquire a far-end signal and a near-end signal; perform linear filtering on the near-end signal and the far-end signal to obtain a nonlinear noise residual signal; extract a first sub-band feature signal and a first speech signal from the nonlinear noise residual signal, and extract a second sub-band feature signal and a second speech signal from the far-end signal; concatenate and fuse the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and concatenate and fuse the first speech signal and the second speech signal to obtain a second input feature signal; fuse the first input feature signal and the second input feature signal to obtain a fused feature signal; determine a target sub-band gain based on the fused feature signal and the first input feature signal, and determine target speech filter coefficients based on the fused feature signal and the second input feature signal; and filter out nonlinear noise signals in the nonlinear noise residual signal based on the target speech filter coefficients and the target sub-band gain to obtain a sound signal after noise joint processing.

[0161] Computer program code for performing the operations of this disclosure can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a LAN (local area network) or WAN (wide area network)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0162] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of devices, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based device that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0163] The modules described in the embodiments of this disclosure can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0164] The computer-readable storage medium provided in this application stores computer-readable program instructions for executing the above-described noise joint processing method, solving the technical problem in the prior art of difficulty in jointly processing echo, noise, and reverberation while achieving good noise reduction effect. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as the beneficial effects of the noise joint processing method provided in the above-described embodiments, and will not be repeated here.

[0165] Example 7

[0166] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the steps of the noise joint processing prediction method described above.

[0167] The computer program product provided in this application solves the technical problem in the prior art of jointly processing echo, noise, and reverberation while achieving good noise reduction effect. Compared with the prior art, the beneficial effects of the computer program product provided in the embodiments of this application are the same as the beneficial effects of the noise joint processing method provided in the above embodiments, and will not be repeated here.

[0168] The above are merely preferred embodiments of this application and do not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent scope of this application.

Claims

1. A method of joint processing of noise, characterized in that, The noise joint processing method includes: Acquire the far-end signal and the near-end signal, and perform linear filtering on the near-end signal and the far-end signal to obtain a nonlinear noise residual signal; Extract the first sub-band feature signal and the first speech signal from the nonlinear noise residual signal, and extract the second sub-band feature signal and the second speech signal from the far-end signal; The first sub-band feature signal and the second sub-band feature signal are spliced ​​and fused to obtain the first input feature signal, and the first speech signal and the second speech signal are spliced ​​and fused to obtain the second input feature signal; The first input feature signal and the second input feature signal are fused to obtain a fused feature signal; Based on the fused feature signal and the first input feature signal, the target subband gain is determined, and based on the fused feature signal and the second input feature signal, the target speech filter coefficients are determined. Based on the target speech filter coefficients and the target subband gain, the nonlinear noise signal in the nonlinear noise residual signal is filtered out to obtain the sound signal after joint noise processing; The first input feature signal includes target subband coding features, the fused feature signal includes linear transform features, and the step of determining the target subband gain based on the fused feature signal and the first input feature signal includes: The fused feature signal is linearly transformed to obtain the linearly transformed feature; the linearly transformed feature and the target subband coding feature are fused and decoded step by step to obtain the target subband gain; The second input feature includes target speech coding features, the fused feature signal includes linear transform features, and the step of determining the target speech filter coefficients based on the fused feature signal and the second input feature signal includes: The fused feature signal is linearly transformed to obtain the linearly transformed feature; the linearly transformed feature is recovered and decoded to obtain the first intermediate speech filter coefficients; the first intermediate speech filter coefficients are linearly transformed to obtain the second intermediate speech filter coefficients; the second intermediate speech filter coefficients and the target speech coding feature are concatenated and fused to obtain the target speech filter coefficients.

2. The noise co-processing method of claim 1, wherein, The first sub-band feature signal includes a preset sampling frequency and a preset number of sub-bands, and the second speech signal includes a preset spectrum. The steps of extracting the first sub-band feature signal and the first speech signal from the nonlinear noise residual signal, and extracting the second sub-band feature signal and the second speech signal from the far-end signal, include: Based on the preset sampling frequency and the preset number of sub-bands, the first sub-band feature signal is extracted from the nonlinear noise residual signal and the second sub-band feature signal is extracted from the far-end signal; Based on the preset spectrum, the first speech signal is extracted from the nonlinear noise residual signal, and the second speech signal is extracted from the far-end signal.

3. The noise co-processing method of claim 1, wherein, The steps of splicing and fusing the first sub-band feature signal and the second sub-band feature signal to obtain the first input feature signal, and splicing and fusing the first speech signal and the second speech signal to obtain the second input feature signal, include: The first sub-band feature signal and the second sub-band feature signal are fused and encoded to obtain the first input feature signal; The first speech signal and the second speech signal are fused and encoded to obtain the second input feature signal.

4. The noise co-processing method of claim 1, wherein, The first input feature signal includes target sub-band coding features, and the second input feature signal includes target speech coding features. The step of fusing the first input feature signal and the second input feature signal to obtain the fused feature signal includes: The first input feature signal is compressed and encoded step by step to obtain the target sub-band encoded feature; The second input feature signal is compressed and encoded step by step to obtain the target speech coding features; The target subband coding features and the target speech coding features are fused and encoded to obtain the fused feature signal.

5. The noise co-processing method of claim 1, wherein, The sound signal includes a first sound signal, and the nonlinear noise residual signal includes a complex spectrum nonlinear noise residual signal. The step of filtering out nonlinear noise signals from the nonlinear noise residual signal based on the target speech filter coefficients and the target subband gain to obtain the noise-jointly processed audio signal includes: The nonlinear noise residual signal is converted into a frequency domain signal to obtain a complex spectrum nonlinear noise residual signal; Interpolate the target subband gain to obtain the target frequency subband signal gain; The complex spectrum nonlinear noise residual signal is nonlinearly filtered out based on the subband signal gain at the target frequency point to obtain the first sound signal; The first sound signal is nonlinearly filtered out according to the target speech filter coefficients to obtain a sound signal after joint noise processing.

6. A noise joint processing device, characterized in that, The noise processing device includes: A linear processing module is used to acquire far-end signals and near-end signals, and to perform linear filtering on the near-end signals and the far-end signals to obtain nonlinear noise residual signals; The signal extraction module is used to extract a first sub-band feature signal and a first speech signal from the nonlinear noise residual signal, and to extract a second sub-band feature signal and a second speech signal from the far-end signal. The signal splicing module is used to splice and fuse the first sub-band feature signal and the second sub-band feature signal to obtain a first input feature signal, and to splice and fuse the first speech signal and the second speech signal to obtain a second input feature signal; An input feature determination module is used to fuse the first input feature signal and the second input feature signal to obtain a fused feature signal; The coefficient and gain determination module is used to determine the target subband gain based on the fused feature signal and the first input feature signal, and to determine the target speech filter coefficients based on the fused feature signal and the second input feature signal; wherein, the first input feature signal includes target subband coding features, the fused feature signal includes linear transform features, the second input feature signal includes target speech coding features, and the fused feature signal includes linear transform features; The nonlinear processing module is used to filter out the nonlinear noise signal in the nonlinear noise residual signal according to the target speech filter coefficients and the target subband gain, so as to obtain the sound signal after joint noise processing; The coefficient and gain determination module is also used to perform a linear transformation on the fused feature signal to obtain the linear transformation feature; and to perform step-by-step fusion decoding on the linear transformation feature and the target subband coding feature to obtain the target subband gain. The coefficient and gain determination module is further configured to perform a linear transformation on the fused feature signal to obtain the linear transformation feature; perform recovery decoding on the linear transformation feature to obtain the first intermediate speech filter coefficients; perform a linear transformation on the first intermediate speech filter coefficients to obtain the second intermediate speech filter coefficients; and concatenate and fuse the second intermediate speech filter coefficients with the target speech coding feature to obtain the target speech filter coefficients.

7. An electronic device, characterized in that, The electronic device includes: At least one processor; and a memory communicatively connected to the at least one processor; The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the steps of the noise joint processing method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program for implementing a noise co-processing method, which is executed by a processor to implement the steps of the noise co-processing method as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method of adaptive full duplex full frequency band echo cancellation

    CN101562669A

  • Residual echo cancellation method and system based on deep neural network

    CN114566176A