Voice noise reduction method and device based on non-negative matrix factorization and time-frequency masking, and medium

By combining non-negative matrix decomposition and time-frequency masking, speech noise reduction is achieved based on the MDCT spectrum coefficient, which solves the general performance problem in the prior art at low signal-to-noise ratio, and improves the noise reduction effect and sound quality.

CN120220711APending Publication Date: 2025-06-27CHONGQING BAIRUI INTERNET ELECTRONICS TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510415734.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-03
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

In the existing Bluetooth audio technology, the speech noise reduction method has a relatively average performance at a lower signal-to-noise ratio, and it is difficult to accurately estimate the base matrix and activation matrix, resulting in more residual noise.

Method used

Combining the speech noise reduction method of non-negative matrix decomposition and time-frequency masking, based on the MDCT spectrum coefficient, the NMF amplitude spectrum gain is calculated through non-negative matrix decomposition, and the NN amplitude spectrum gain is calculated through the pre-trained deep neural network, and the noise reduction gain is output after fusion.

Benefits of technology

The sound quality reduction caused by using noisy voice phase reconstruction is avoided, and the advantages of non-negative matrix decomposition and time-frequency masking are fully utilized, which not only avoids excessive residual noise, but also improves the quality of high-frequency signals and ensures sound quality.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120220711A_ABST
    Figure CN120220711A_ABST
Patent Text Reader

Abstract

The invention discloses a voice noise reduction method and device in combination with non-negative matrix factorization and time-frequency masking and a storage medium, and belongs to the technical field of Bluetooth audios, and the method comprises the steps: inputting monaural voice PCM data, and executing low-delay improved discrete cosine transform to obtain an MDCT spectral coefficient; calculating sub-band energy according to the MDCT spectrum coefficient, and constructing a noisy voice amplitude spectrum observation matrix; performing non-negative matrix factorization on the noisy voice amplitude spectrum observation matrix, and calculating an NMF amplitude spectrum gain; performing feature extraction according to the sub-band energy, and calculating an NN amplitude spectrum gain through a pre-trained deep neural network; fusing the NMF amplitude spectrum gain and the NN amplitude spectrum gain, and outputting a noise reduction gain; obtaining a noise reduction spectral coefficient according to the noise reduction gain and the MDCT spectral coefficient; and according to the noise reduction spectrum coefficient, continuing to execute the coding process, and outputting a noise reduction voice code stream. According to the invention, based on the MDCT spectral coefficient, voice noise reduction is realized by combining non-negative matrix factorization and time-frequency masking, and the tone quality is ensured.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of Bluetooth audio technology, and particularly relates to a voice noise reduction method, device and storage medium combining non-negative matrix factorization and time-frequency masking. Background Art

[0002] The current mainstream Bluetooth audio encoders are as follows: SBC: Mandatory requirement of the A2DP protocol, most widely used, and must be supported by all Bluetooth audio devices, but the sound quality is average; AAC-LC: Good sound quality and relatively widely used, supported by many mainstream mobile phones. However, compared with SBC, it has a larger memory footprint and higher computational complexity. Many Bluetooth devices are based on embedded platforms with limited battery capacity, poor processor computing power and limited memory. Moreover, its patent fee is relatively high; aptX series: Good sound quality, but high bit rate. aptX requires a bit rate of 384 kbps, while the bit rate of aptX-HD is 576 kbps, and it is a technology exclusive to Qualcomm, which is relatively closed; LDAC: Good sound quality, but also high bit rate, which are 330 kbps, 660 kbps and 990 kbps respectively. Due to the extremely complex wireless environment where Bluetooth devices are located, it is somewhat difficult to stably support such high bit rates, and it is a technology exclusive to Sony, which is also very closed; LHDC: Good sound quality, but also high bit rate, typically including 400 kbps, 600 kbps and 900 kbps. Such high bit rates pose high requirements for the Bluetooth baseband / RF design. For the above reasons, the Bluetooth Special Interest Group (Bluetooth Sig) jointly launched LC3 with many manufacturers, mainly for low-power Bluetooth, and can also be used for classic Bluetooth. It has the advantages of low latency, high sound quality and coding gain, and no patent fee in the Bluetooth field, and has received wide attention from manufacturers.

[0003] In many Bluetooth applications, such as Bluetooth calls, Bluetooth microphones and recordings, noise reduction is required.

[0004] Nonnegative Matrix Factorization (NMF) is abbreviated as NMF, which makes all matrix components after decomposition non-negative values and simultaneously realizes non-linear dimensionality reduction. NMF has gradually become one of the most popular multi-dimensional data processing tools in research fields such as signal processing, biomedical engineering, pattern recognition, computer vision and image engineering.

[0005] Nonnegative matrix factorization has certain applications in voice noise reduction, but its performance is average at low signal-to-noise ratios, it is difficult to accurately estimate the basis matrix and activation matrix, and there is more residual noise after noise reduction.

[0006] Speech enhancement methods based on time-frequency masking are widely used. Typical methods include the Ideal Ratio Mask (IRM), Ideal Binary Mask (IBM), and Ideal Ratio Mask in the complex domain, etc.

[0007] Applying deep learning to the time-frequency masking algorithm has achieved good results. However, the neural network required by the time-frequency masking algorithm based on frequency bin gain is relatively complex. To simplify the operation, the time-frequency masking based on sub-band gain can greatly simplify the complexity of the neural network, thereby saving the amount of computation. For example, if the sampling rate of speech is 48 kHz, each frame with a length of 10 ms corresponds to 480 sampling points. After the low-latency modified discrete cosine transform, 480 spectral coefficients are obtained, and the neural network needs to output 480 gains. If calculated in sub-bands, taking 64 sub-bands as an example, the neural network only needs to output 64 sub-band gains, which greatly reduces the complexity. To conform to the human auditory theory, this sub-band division is usually finer in the low frequency and coarser in the high frequency, which directly results in better bass sound quality of the denoised speech and certain distortion in the high-frequency part. Summary of the Invention

[0008] Aiming at the above technical problems existing in the prior art, this application provides a speech denoising method, device, and storage medium that combines non-negative matrix factorization and time-frequency masking. Based on the MDCT spectral coefficients, speech denoising is achieved by combining non-negative matrix factorization and time-frequency masking, avoiding the reduction in sound quality caused by the need to use the phase of the noisy speech for reconstruction in the prior art, and also making full use of the respective advantages of non-negative matrix factorization and time-frequency masking, avoiding excessive residual noise and improving the quality of high-frequency signals.

[0009] To achieve the above object, the first technical solution adopted in this application is: to provide a speech denoising method that combines non-negative matrix factorization and time-frequency masking, including: inputting mono-channel speech PCM data and performing a low-latency modified discrete cosine transform to obtain MDCT spectral coefficients; calculating sub-band energy based on the MDCT spectral coefficients and constructing a noisy speech amplitude spectrum observation matrix; performing non-negative matrix factorization on the noisy speech amplitude spectrum observation matrix to calculate the NMF amplitude spectrum gain; performing feature extraction based on the sub-band energy and calculating the NN amplitude spectrum gain through a pre-trained deep neural network; fusing the NMF amplitude spectrum gain and the NN amplitude spectrum gain to output a denoising gain; obtaining denoised spectral coefficients based on the denoising gain and the MDCT spectral coefficients; and continuing to perform the encoding process based on the denoised spectral coefficients to output a denoised speech bitstream.

[0010] The second technical solution adopted in this application is: to provide a speech noise reduction device that combines non-negative matrix factorization and time-frequency masking, including: a module for inputting monophonic speech PCM data and performing low-latency improved discrete cosine transform to obtain MDCT spectral coefficients; a module for calculating sub-band energy based on the MDCT spectral coefficients and constructing an observation matrix of the noisy speech amplitude spectrum; a module for performing non-negative matrix factorization on the observation matrix of the noisy speech amplitude spectrum to calculate the NMF amplitude spectrum gain; a module for performing feature extraction based on the sub-band energy and calculating the NN amplitude spectrum gain through a pre-trained deep neural network; a module for fusing the NMF amplitude spectrum gain and the NN amplitude spectrum gain to output a noise reduction gain; a module for obtaining noise-reduced spectral coefficients based on the noise reduction gain and the MDCT spectral coefficients; and a module for continuing the encoding process based on the noise-reduced spectral coefficients and outputting a noise-reduced speech bitstream.

[0011] The third technical solution adopted in this application is: to provide a computer-readable storage medium that stores computer instructions, where the computer instructions are operated to execute the speech noise reduction method that combines non-negative matrix factorization and time-frequency masking in Solution 1.

[0012] The beneficial effects that can be achieved by the technical solutions of this application are: the technical solutions of this application can be applied to both classic Bluetooth (BR, EDR) and low-power Bluetooth (BLE). Based on the MDCT spectral coefficients, speech noise reduction is achieved by combining non-negative matrix factorization and time-frequency masking. Among them, based on the MDCT spectral coefficients, it avoids the reduction in sound quality caused by the need to use the phase of the noisy speech for reconstruction in the prior art; it fully utilizes the respective advantages of non-negative matrix factorization and time-frequency masking, both avoiding excessive residual noise and improving the quality of high-frequency signals, thus ensuring the sound quality. Description of the Drawings

[0013] Figure 1 is a schematic flowchart of a specific implementation manner of the speech noise reduction method that combines non-negative matrix factorization and time-frequency masking in this application;

[0014] Figure 2 is a schematic diagram of a specific example of the speech noise reduction that combines non-negative matrix factorization and time-frequency masking in this application;

[0015] Figure 3 is a schematic diagram of a specific implementation manner of the speech noise reduction device that combines non-negative matrix factorization and time-frequency masking in this application. Detailed Implementation Manner

[0016] Next, specific embodiments will be used to elaborate in detail on the technical solution of the present application and how the technical solution of the present application solves the above technical problems. The specific embodiments described below can be combined with each other to form new embodiments. For the same or similar ideas or processes described in one embodiment, they may not be repeated in some other embodiments. Next, the embodiments of the present application will be described in conjunction with the accompanying drawings.

[0017] Next, the specific implementation manner of the present application will be elaborated in detail by taking a sampling rate of 48 kHz as an example. The principles of other sampling rates are similar to those of the 48 kHz sampling rate and will not be repeated.

[0018] Figure 1 It is a schematic flowchart of a specific implementation manner of the speech noise reduction method of the present application that combines non - negative matrix factorization and time - frequency masking.

[0019] In Figure 1 In a specific implementation manner shown, the speech noise reduction method of the present application that combines non - negative matrix factorization and time - frequency masking includes process S101: inputting mono - channel speech PCM data and performing a low - latency improved discrete cosine transform to obtain MDCT spectral coefficients.

[0020] In a specific embodiment of the present application, inputting mono - channel speech PCM data and performing a low - latency improved discrete cosine transform to obtain MDCT spectral coefficients includes: performing frame division on the mono - channel speech PCM data and windowing the obtained audio frames; and performing a low - latency improved discrete cosine transform on the windowed audio frames to obtain MDCT spectral coefficients.

[0021] In this specific embodiment, at the Bluetooth transmitter end, mono - channel speech PCM data is input, then frame division and windowing operations are performed on the mono - channel noisy speech PCM data, and then a low - latency improved discrete cosine transform is performed on the windowed audio frames to obtain MDCT spectral coefficients.

[0022] Specifically, input one frame of speech signal and output one frame of spectral coefficients:

[0023] t(n) = x s (Z - N F + n), for n = 0…2·N F -1 - Z

[0024] t(2N F - Z + n) = 0, for n = 0…Z - 1

[0025]

[0026] where, x s (n) is the input audio signal, It is the analysis window in LC3, and X(k) is the MDCT spectral coefficient.

[0027] In a specific example of the present application, taking a sampling rate of 48 kHz and a frame length of 10 ms as an example, the MDCT spectral coefficients are: X(k), k = 0, …, N F -1, where N F = 480.

[0028] In Figure 1 In a specific embodiment shown, the speech noise reduction method combining non - negative matrix factorization and time - frequency masking in the present application includes process S102. According to the MDCT spectral coefficients, calculate the sub - band energy and construct the noisy speech amplitude spectrum observation matrix.

[0029] In a specific embodiment of the present application, calculating the sub - band energy and constructing the noisy speech amplitude spectrum observation matrix according to the MDCT spectral coefficients includes: dividing the MDCT spectral coefficients into sub - bands to obtain the frequency bin index table corresponding to each sub - band; obtaining the sub - band energy according to the MDCT spectral coefficients and the frequency bin index table; and calculating the amplitude spectrum according to the MDCT spectral coefficients and constructing the noisy speech amplitude spectrum observation matrix according to the amplitude spectrum.

[0030] Specifically, first divide the MDCT spectral coefficients into sub - bands. Divide the above - mentioned MDCT spectral coefficients into 64 sub - bands, and the sub - band numbers are denoted as b = 0, 1, 2, …, 63. This sub - band division conforms to the human auditory characteristics, that is, fine in the low - frequency range and rough in the high - frequency range. Taking the division of the MDCT spectral coefficients into 64 sub - bands as an example in the present application, in practical applications, the sub - band size can be flexibly configured according to requirements to balance the noise reduction performance and the computational complexity requirements. Then obtain the frequency bin index table corresponding to each sub - band as follows: Index[] = [01,2,3,4,5,6,7,8,9,10,11,12,13,14,15,16,17,18,20,22,24,26,28,30,32,34,36,39,42,45,48,51,55,59,63,67,71,76,81,86,92,98,105,112,119,127,135,144,154,164,175,186,198,211,225,240,256,273,291,310,330,352,375,400]

[0031] Then, according to the MDCT spectral coefficients and the frequency bin index table, calculate the sub - band energy:

[0032]

[0033] At the same time, also calculate the amplitude spectrum |X(k)| according to the MDCT spectral coefficients, and then construct the noisy speech amplitude spectrum observation matrix Vnoisy 。

[0034] In Figure 1 a specific embodiment shown, the speech noise reduction method combining non - negative matrix factorization and time - frequency masking of the present application includes process S103, performing non - negative matrix factorization on the noisy speech magnitude spectrum observation matrix and calculating the NMF magnitude spectrum gain.

[0035] In a specific embodiment of the present application, performing non - negative matrix factorization on the noisy speech magnitude spectrum observation matrix and calculating the NMF magnitude spectrum gain includes: respectively fixing the speech basis matrix and the noise basis matrix to perform non - negative matrix factorization on the noisy speech magnitude spectrum observation matrix to obtain the speech activation matrix and the noise activation matrix; multiplying the speech basis matrix by the speech activation matrix to obtain the speech frequency bin energy value; multiplying the noise basis matrix by the noise activation matrix to obtain the noise frequency bin energy value; and obtaining the NMF magnitude spectrum gain according to the speech frequency bin energy value and the noise frequency bin energy value.

[0036] In this specific embodiment, the process of calculating the NMF magnitude spectrum gain based on the non - negative matrix factorization algorithm is as follows:

[0037] First, fix the speech basis matrix W speech , and perform non - negative matrix factorization on the noisy speech magnitude spectrum observation matrix to obtain the speech activation matrix H s ′ peech 。

[0038] Then, fix the noise basis matrix W noise , and perform non - negative matrix factorization on the noisy speech magnitude spectrum observation matrix to obtain the noise activation matrix H ′ noise 。

[0039] Multiply the speech basis matrix by the speech activation matrix to calculate the speech frequency bin energy value:

[0040] V s ′ peech (b)=W speech H s ′ peech

[0041] Multiply the noise basis matrix by the noise activation matrix to calculate the noise frequency bin energy value:

[0042] N n ′ oise (b)=W noise H ′ noise

[0043] Then, according to the speech frequency bin energy value and the noise frequency bin energy value, calculate the NMF magnitude spectrum gain, also known as the frequency bin noise reduction gain:

[0044]

[0045] In a specific embodiment of the present application, the training process of the speech basis matrix and the noise basis matrix includes: selecting clean speech and noise as the training set; respectively performing frame division, windowing, and low-latency modified discrete cosine transform on the clean speech and the noise to obtain clean speech spectrum coefficients and noise spectrum coefficients; respectively calculating the corresponding sub-band energies according to the clean speech spectrum coefficients and the noise spectrum coefficients, and correspondingly constructing a clean speech sub-band energy observation matrix and a noise sub-band energy observation matrix; and respectively performing non-negative matrix factorization on the clean speech sub-band energy observation matrix and the noise sub-band energy observation matrix to obtain the speech basis matrix and the noise basis matrix.

[0046] The following briefly describes the training method of the speech basis matrix W speech and the noise basis matrix W noise as follows:

[0047] (1) Select clean speech and noise as the training set (hereinafter, the training process of clean speech is taken as an example, and the training process of noise is similar and will not be elaborated);

[0048] (2) Perform frame division, windowing, and low-latency modified discrete cosine transform on the clean speech to obtain spectrum coefficients;

[0049] (3) Calculate the sub-band energy and construct the observation matrix V;

[0050] (4) Perform non-negative matrix factorization on the observation matrix V to obtain the basis matrix W (i.e., the speech basis matrix W speech ) and the activation matrix H.

[0051] For the non-negative matrix V, non-negative matrix factorization is to find an optimal solution such that the product of the decomposed W and H is an approximation of V. The steps of non-negative matrix factorization are briefly described as follows:

[0052] (1) Construct the observation matrix V KL , which has K rows and L columns, where each column is the magnitude spectrum of one frame, and there are L columns in total.

[0053] (2) Initialize the basis matrix W KM .

[0054] (3) Initialize the activation matrix H ML , in a typical scenario, the above parameters can take values of K = 480, M = 240, and L = 3 - 5.

[0055] (4) Iterative solution based on the KL divergence criterion is briefly described as follows:

[0056]

[0057] where denotes element-wise multiplication, T denotes the transpose operation of a matrix, and I denotes the identity matrix.

[0058] During the iteration process, calculate the KL divergence between V and WH, and stop when it is less than a preset threshold. To avoid excessive calculation time caused by abnormal situations, the number of iterations can also be preset, and stop when the preset number of iterations is reached, so as to obtain the basis matrix and the activation matrix.

[0059] In Figure 1 a specific embodiment shown, the speech noise reduction method combining non-negative matrix factorization and time-frequency masking of the present application includes process S104, performing feature extraction according to the subband energy, and calculating the NN magnitude spectrum gain through a pre-trained deep neural network.

[0060] In this specific embodiment, calculating the magnitude spectrum gain based on time-frequency masking according to the subband energy can ensure that the reconstructed speech characteristics conform to the human ear auditory characteristics.

[0061] In a specific embodiment of the present application, performing feature extraction according to the subband energy and calculating the NN magnitude spectrum gain through a pre-trained deep neural network includes: performing feature extraction according to the subband energy to obtain Bark features; inputting the Bark features into a pre-trained deep neural network to output the subband noise reduction gain; interpolating the subband noise reduction gain to obtain the NN magnitude spectrum gain.

[0062] In this specific embodiment, first perform feature extraction according to the subband energy to obtain Bark features, and then input the Bark features into a pre-trained deep neural network to output the subband noise reduction gain and then interpolate the subband noise reduction gain to obtain the NN magnitude spectrum gain that is, the gain of all frequency bins. Among them, the Bark features include BFCC (Bark Frequency Cepstral Coefficients, which are feature data extracted based on the Bark filter bank), a predetermined number of time differences of BFCC in front, and a predetermined number of second-order time differences of BFCC in front.

[0063] In a specific embodiment of the present application, the training process of the pre-trained deep neural network includes: obtaining a predetermined number of clean voices and noises, and mixing the clean voices and noises to obtain noisy voices; respectively performing low-latency improved discrete cosine transform on the clean voices and the noisy voices, and outputting clean spectral coefficients and noisy spectral coefficients; respectively calculating the clean voice subband energy, the noisy voice subband energy, and the ideal subband gain according to the clean spectral coefficients and the noisy spectral coefficients; obtaining the Bark feature of the noisy voice according to the noisy voice subband energy; inputting the Bark feature of the noisy voice into the deep neural network, and training the deep neural network with the ideal subband gain as the target. When the loss level reaches the expectation, freeze the deep neural network to obtain the pre-trained deep neural network.

[0064] Specifically, first, obtain a certain number of clean voices and noises, and then mix the clean voices with the noises to obtain noisy voices.

[0065] Then perform feature extraction on the clean voices and the noisy voices. Specifically, first perform low-latency improved discrete cosine transform on the clean voices and the noisy voices respectively to output their corresponding spectral coefficients (the specific calculation method is the same as above, which is omitted here and will not be elaborated further), and then calculate the subband energy of the clean voices and the noisy voices respectively (the specific calculation method is the same as above, which is omitted here and will not be elaborated further). Among them, the clean voice is: x clean (k), the clean voice spectral coefficient is: X clean (k), the clean voice subband energy is: Energy subband,clean (b); the noisy voice is: x noisy (k), the noisy voice spectral coefficient is: X noisy (k), the noisy voice subband energy is: Energy subband,noisy (b).

[0066] After calculating the clean voice subband energy and the noisy voice subband energy, calculate the ideal subband gain:

[0067]

[0068] Then perform transformation according to the noisy voice subband energy to obtain the Bark feature of the noisy voice. Specifically:

[0069] First perform logarithmic transformation on the noisy voice subband energy:

[0070]

[0071] Then perform DCT transformation (Discrete Cosine Transform) to obtain BFCC:

[0072]

[0073] There are 84 Bark features of noisy speech, including 64 BFCCs, the first 10 temporal differences of BFCCs, and the first 10 second-order temporal differences of BFCCs. Specifically:

[0074] The first 10 temporal differences of BFCCs are:

[0075] BFCC diff (k) = BFCC curr (k) - BFCC lastlast (k), k = 0, …, 9

[0076] The first 10 second-order temporal differences of BFCCs are:

[0077] BFCC diff2 (k) = BFCC curr (k) - 2 * BFCC last (k) + BFCC lastlast (k), k = 0, …, 9

[0078] Input the above features into a deep neural network and train the deep neural network with the ideal subband gain as the target. When the loss level reaches the expectation, freeze the deep neural network to obtain a pre-trained neural network.

[0079] Among them, the choice of the deep neural network is not limited in this application. Considering the forward and backward correlation characteristics of speech frames, a recurrent neural network (RNN) is preferably selected.

[0080] When performing backpropagation, the loss function used is defined as:

[0081]

[0082] In a specific embodiment of this application, interpolate the subband noise reduction gain to obtain the NN magnitude spectrum gain, including: performing a predetermined operation on the subband noise reduction gain and the waveform symmetric about the subband center frequency bin to obtain the NN magnitude spectrum gain.

[0083] Specifically,

[0084]

[0085] Among them, Weight b (k) is the waveform symmetric about the subband center frequency bin, which satisfies the following conditions;

[0086]

[0087] In Figure 1In a specific embodiment shown, the speech noise reduction method combining non-negative matrix factorization and time-frequency masking of the present application includes process S105, which fuses the NMF magnitude spectrum gain and the NN magnitude spectrum gain and outputs a noise reduction gain.

[0088] In this specific embodiment, the NMF magnitude spectrum gain and the NN magnitude spectrum gain are fused to obtain the NR magnitude spectrum gain, that is, the final noise reduction gain:

[0089]

[0090] In Figure 1 In a specific embodiment shown, the speech noise reduction method combining non-negative matrix factorization and time-frequency masking of the present application includes process S106, which obtains the noise reduction spectrum coefficients according to the noise reduction gain and the MDCT spectrum coefficients.

[0091] In this specific embodiment, the noise reduction gain is applied to the MDCT spectrum coefficients, and the new MDCT spectrum coefficients, that is, the noise reduction spectrum coefficients, are output:

[0092]

[0093] In Figure 1 In a specific embodiment shown, the speech noise reduction method combining non-negative matrix factorization and time-frequency masking of the present application includes process S107, which continues to execute the encoding process according to the noise reduction spectrum coefficients and outputs a noise reduction speech code stream.

[0094] In a specific embodiment of the present application, continuing to execute the encoding process according to the noise reduction spectrum coefficients and outputting a noise reduction speech code stream includes: continuing to complete modules such as transform domain noise shaping, time domain noise shaping, quantization, noise level estimation, arithmetic coding, and residual coding according to the noise reduction spectrum coefficients, and outputting a noise reduction speech code stream.

[0095] Figure 2 It is a schematic diagram of a specific example of the speech noise reduction combining non-negative matrix factorization and time-frequency masking of the present application.

[0096] As Figure 2In a specific example shown below, the speech noise reduction process of this application that combines non - negative matrix factorization and time - frequency masking is as follows: First, input audio data (PCM), and perform a low - latency improved discrete cosine transform to obtain MDCT spectral coefficients. Then, based on the MDCT spectral coefficients, construct an amplitude - spectrum observation matrix and perform feature extraction respectively, and perform non - negative matrix factorization on the amplitude - spectrum observation matrix to obtain activation matrices of speech and noise. Then, calculate the frequency - bin energies of speech and noise, so as to obtain the NMF amplitude - spectrum gain. At the same time, the extracted features are also input into a pre - trained deep neural network to obtain the NN amplitude - spectrum gain. Then, fuse the NMF amplitude - spectrum gain and the NN amplitude - spectrum gain to obtain the NR amplitude - spectrum gain, and apply the gain to obtain the noise - reduced spectral coefficients, and continue to perform the encoding process on the noise - reduced spectral coefficients to output the noise - reduced bitstream. The advantages of this application in combining non - negative matrix factorization and time - frequency masking for speech noise reduction are as follows:

[0097] (1) As a supervised algorithm, non - negative matrix factorization obtains prior information through offline training of noise and clean speech, ensuring the speech quality in the real - time enhancement stage;

[0098] (2) Time - frequency masking based on sub - band gain ensures that the reconstructed speech characteristics conform to the human auditory characteristics;

[0099] (3) By combining non - negative matrix factorization and time - frequency masking, the advantages of both are fully utilized, and the disadvantages of both are avoided. That is, both excessive residual noise is avoided and the quality of high - frequency signals is improved.

[0100] In the speech noise reduction method of this application that combines non - negative matrix factorization and time - frequency masking, based on the MDCT spectral coefficients, speech noise reduction is achieved by combining non - negative matrix factorization and time - frequency masking. Among them, based on the MDCT spectral coefficients, it avoids the reduction in sound quality caused by using the phase of noisy speech for reconstruction in the prior art; fully utilizes the respective advantages of non - negative matrix factorization and time - frequency masking, both avoids excessive residual noise and improves the quality of high - frequency signals, ensuring the sound quality.

[0101] Figure 3 It is a schematic diagram of a specific implementation manner of the speech noise reduction device of this application that combines non - negative matrix factorization and time - frequency masking.

[0102] In Figure 3In a specific embodiment shown, the speech noise reduction device combining non-negative matrix factorization and time-frequency masking of the present application includes: a module 301 for inputting mono-channel speech PCM data and performing a low-latency improved discrete cosine transform to obtain MDCT spectral coefficients; a module 302 for calculating sub-band energies based on the MDCT spectral coefficients and constructing a noisy speech amplitude spectrum observation matrix; a module 303 for performing non-negative matrix factorization on the noisy speech amplitude spectrum observation matrix to calculate the NMF amplitude spectrum gain; a module 304 for performing feature extraction based on the sub-band energies and calculating the NN amplitude spectrum gain through a pre-trained deep neural network; a module 305 for fusing the NMF amplitude spectrum gain and the NN amplitude spectrum gain and outputting a noise reduction gain; a module 306 for obtaining noise-reduced spectral coefficients based on the noise reduction gain and the MDCT spectral coefficients; and a module 307 for continuing to perform an encoding process based on the noise-reduced spectral coefficients and outputting a noise-reduced speech bitstream.

[0103] In a specific embodiment of the present application, inputting mono-channel speech PCM data and performing a low-latency improved discrete cosine transform to obtain MDCT spectral coefficients includes: performing frame division on the mono-channel speech PCM data and windowing the obtained audio frames; and performing a low-latency improved discrete cosine transform on the windowed audio frames to obtain MDCT spectral coefficients.

[0104] In a specific embodiment of the present application, calculating sub-band energies based on the MDCT spectral coefficients and constructing a noisy speech amplitude spectrum observation matrix includes: dividing the MDCT spectral coefficients into sub-bands to obtain a frequency bin index table corresponding to each sub-band; obtaining sub-band energies based on the MDCT spectral coefficients and the frequency bin index table; and calculating an amplitude spectrum based on the MDCT spectral coefficients and constructing a noisy speech amplitude spectrum observation matrix based on the amplitude spectrum.

[0105] In a specific embodiment of the present application, performing non-negative matrix factorization on the noisy speech amplitude spectrum observation matrix to calculate the NMF amplitude spectrum gain includes: respectively fixing the speech basis matrix and the noise basis matrix and performing non-negative matrix factorization on the noisy speech amplitude spectrum observation matrix to obtain a speech activation matrix and a noise activation matrix; multiplying the speech basis matrix by the speech activation matrix to obtain speech frequency bin energy values; multiplying the noise basis matrix by the noise activation matrix to obtain noise frequency bin energy values; and obtaining the NMF amplitude spectrum gain based on the speech frequency bin energy values and the noise frequency bin energy values.

[0106] In a specific embodiment of the present application, the training process of the voice basis matrix and the noise basis matrix includes: selecting clean speech and noise as the training set; respectively performing frame division, windowing, and low-latency improved discrete cosine transform on the clean speech and the noise to obtain clean speech spectral coefficients and noise spectral coefficients; respectively calculating the corresponding sub-band energies according to the clean speech spectral coefficients and the noise spectral coefficients, and correspondingly constructing a clean speech sub-band energy observation matrix and a noise sub-band energy observation matrix; and respectively performing non-negative matrix factorization on the clean speech sub-band energy observation matrix and the noise sub-band energy observation matrix to obtain the voice basis matrix and the noise basis matrix.

[0107] In a specific embodiment of the present application, according to the sub-band energy, feature extraction is performed, and the NN magnitude spectrum gain is calculated through a pre-trained deep neural network, including: according to the sub-band energy, feature extraction is performed to obtain Bark features; the Bark features are input into the pre-trained deep neural network, and the sub-band noise reduction gain is output; the sub-band noise reduction gain is interpolated to obtain the NN magnitude spectrum gain.

[0108] In a specific embodiment of the present application, the training process of the pre-trained deep neural network includes: obtaining a predetermined number of clean speech and noise, and mixing the clean speech and the noise to obtain noisy speech; respectively performing low-latency improved discrete cosine transform on the clean speech and the noisy speech, and outputting clean spectral coefficients and noisy spectral coefficients; respectively calculating the clean speech sub-band energy, the noisy speech sub-band energy, and the ideal sub-band gain according to the clean spectral coefficients and the noisy spectral coefficients; obtaining the Bark features of the noisy speech according to the noisy speech sub-band energy; inputting the Bark features of the noisy speech into the deep neural network, and training the deep neural network with the ideal sub-band gain as the target. When the loss level reaches the expectation, the deep neural network is frozen to obtain the pre-trained deep neural network.

[0109] In a specific embodiment of the present application, interpolating the sub-band noise reduction gain to obtain the NN magnitude spectrum gain includes: performing a predetermined operation on the sub-band noise reduction gain and the waveform symmetric to the sub-band center frequency bin to obtain the NN magnitude spectrum gain.

[0110] The speech noise reduction device combining non-negative matrix factorization and time-frequency masking provided by the present application can be used to perform the speech noise reduction method combining non-negative matrix factorization and time-frequency masking described in any of the above embodiments. The implementation principle and technical effect are similar and will not be elaborated here.

[0111] In the voice noise reduction device combining non - negative matrix factorization and time - frequency masking of the present application, based on the MDCT spectral coefficients, voice noise reduction is achieved by combining non - negative matrix factorization and time - frequency masking. Among them, based on the MDCT spectral coefficients, it avoids the reduction of sound quality caused by the need to use the phase of the noisy speech for reconstruction in the prior art; it fully utilizes the respective advantages of non - negative matrix factorization and time - frequency masking, avoiding excessive residual noise and improving the quality of high - frequency signals, thus ensuring the sound quality.

[0112] In a specific embodiment of the present application, a computer - readable storage medium stores computer instructions, and the computer instructions are operated to execute the voice noise reduction method combining non - negative matrix factorization and time - frequency masking described in any embodiment. Among them, the storage medium can be directly in hardware, in a software module executed by a processor, or in a combination of both.

[0113] The software module can reside in a RAM memory, flash memory, ROM memory, EPROM memory, EEPROM memory, register, hard disk, removable disk, CD - ROM, or any other form of storage medium known in the art. The exemplary storage medium is coupled to the processor such that the processor can read information from and write information to the storage medium.

[0114] The processor can be a Central Processing Unit (CPU for short), or other general - purpose processors, Digital Signal Processors (DSP for short), Application Specific Integrated Circuits (ASIC for short), Field Programmable Gate Arrays (FPGA for short), or other programmable logic devices, discrete gate or transistor logic, discrete hardware components, or any combination thereof. The general - purpose processor can be a microprocessor, but in an alternative, the processor can be any conventional processor, controller, microcontroller, or state machine. The processor can also be implemented as a combination of computing devices, such as a combination of a DSP and a microprocessor, multiple microprocessors, one or more microprocessors combined with a DSP core, or any other such configuration. In an alternative, the storage medium can be integrated with the processor. The processor and the storage medium can reside in an ASIC. The ASIC can reside in a user terminal. In an alternative, the processor and the storage medium can reside in the user terminal as discrete components.

[0115] In a specific embodiment of the present application, a computer device includes a processor and a memory, and the memory stores computer instructions. Among them, the processor operates the computer instructions to execute the speech noise reduction method combining non-negative matrix factorization and time-frequency masking described in any embodiment.

[0116] In the embodiments provided by the present application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces. The indirect couplings or communication connections of the devices or units can be in electrical, mechanical or other forms.

[0117] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0118] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structural transformation made by using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, is equally included in the patent protection scope of the present application.

Claims

1. A speech denoising method combining non-negative matrix decomposition and time-frequency masking, characterized in that: include: Input monophonic speech PCM data and perform low-delay modified discrete cosine transform to obtain MDCT spectrum coefficients; According to the MDCT spectrum coefficients, subband energy is calculated, and a noisy speech amplitude spectrum observation matrix is ​​constructed; Performing non-negative matrix decomposition on the noisy speech amplitude spectrum observation matrix to calculate the NMF amplitude spectrum gain; Based on the sub-band energy, feature extraction is performed and NN amplitude spectrum gain is calculated through a pre-trained deep neural network; The NMF amplitude spectrum gain and the NN amplitude spectrum gain are fused to output a noise reduction gain; Obtaining a noise reduction spectrum coefficient according to the noise reduction gain and the MDCT spectrum coefficient; as well as The encoding process is continued according to the noise reduction spectrum coefficients to output a noise reduction speech code stream.

2. The speech denoising method combining non-negative matrix decomposition and time-frequency masking according to claim 1, characterized in that: The input monophonic speech PCM data and the low-delay improved discrete cosine transform are performed to obtain MDCT spectrum coefficients, including: Performing framing on the monophonic voice PCM data and performing windowing on the obtained audio frames; and A low-delay modified discrete cosine transform is performed on the windowed audio frame to obtain the MDCT spectrum coefficients.

3. The speech denoising method combining non-negative matrix decomposition and time-frequency masking according to claim 1, characterized in that: The subband energy is calculated according to the MDCT spectrum coefficients, and a noisy speech amplitude spectrum observation matrix is ​​constructed, including: Dividing the MDCT spectrum coefficients into sub-bands to obtain a frequency bin index table corresponding to each of the sub-bands; Obtaining the subband energy according to the MDCT spectrum coefficients and the frequency bin index table; and The amplitude spectrum is calculated according to the MDCT spectrum coefficients, and the noisy speech amplitude spectrum observation matrix is ​​constructed according to the amplitude spectrum.

4. The speech denoising method combining non-negative matrix decomposition and time-frequency masking according to claim 1, characterized in that: The performing non-negative matrix decomposition on the noisy speech amplitude spectrum observation matrix to calculate the NMF amplitude spectrum gain includes: The speech basic matrix and the noise basic matrix are fixed respectively to perform non-negative matrix decomposition on the noisy speech amplitude spectrum observation matrix to obtain a speech activation matrix and a noise activation matrix; Multiplying the speech basic matrix by the speech activation matrix to obtain a speech bin energy value; Multiplying the noise basic matrix by the noise activation matrix to obtain noise frequency bin energy values; The NMF amplitude spectrum gain is obtained according to the energy value of the speech frequency bin and the energy value of the noise frequency bin.

5. The speech denoising method combining non-negative matrix decomposition and time-frequency masking according to claim 1, characterized in that: The method of performing feature extraction according to the sub-band energy and calculating the NN amplitude spectrum gain through a pre-trained deep neural network includes: Perform feature extraction according to the sub-band energy to obtain Bark features; Inputting the Bark feature into the pre-trained deep neural network, and outputting a sub-band noise reduction gain; The sub-band noise reduction gain is interpolated to obtain the NN amplitude spectrum gain.

6. The speech denoising method combining non-negative matrix decomposition and time-frequency masking according to claim 1, characterized in that: The training process of the speech basic matrix and the noise basic matrix includes: Select pure speech and noise as training sets; Performing framing, windowing and low-delay improved discrete cosine transform on the clean speech and the noise respectively to obtain clean speech spectrum coefficients and noise spectrum coefficients; Calculating corresponding subband energies according to the clean speech spectrum coefficients and the noise spectrum coefficients respectively, and constructing a clean speech subband energy observation matrix and a noise subband energy observation matrix accordingly; and Non-negative matrix decomposition is performed on the clean speech subband energy observation matrix and the noise subband energy observation matrix respectively to obtain the speech basic matrix and the noise basic matrix.

7. The speech denoising method combining non-negative matrix decomposition and time-frequency masking according to claim 1, characterized in that: The training process of the pre-trained deep neural network includes: Acquire a predetermined amount of clean speech and noise, and mix the clean speech and the noise to obtain noisy speech; Performing low-delay improved discrete cosine transform on the clean speech and the noisy speech respectively, and outputting clean spectral coefficients and noisy spectral coefficients; Calculating clean speech subband energy, noisy speech subband energy and ideal subband gain according to the clean spectrum coefficients and the noisy spectrum coefficients respectively; Obtaining a Bark feature of the noisy speech according to the noisy speech subband energy; The Bark feature of the noisy speech is input into a deep neural network, and the deep neural network is trained with the ideal subband gain as a target. When the loss level reaches an expected level, the deep neural network is frozen to obtain the pre-trained deep neural network.

8. The method for speech denoising combining non-negative matrix decomposition and time-frequency masking according to claim 5, characterized in that: The interpolating the sub-band noise reduction gain to obtain the NN amplitude spectrum gain includes: The sub-band noise reduction gain is subjected to a predetermined operation with a waveform symmetrical with respect to a sub-band center frequency bin to obtain the NN amplitude spectrum gain.

9. A speech noise reduction device combining non-negative matrix decomposition and time-frequency masking, characterized in that: include: A module for inputting monophonic speech PCM data and performing low-delay modified discrete cosine transform to obtain MDCT spectrum coefficients; A module for calculating subband energy according to the MDCT spectrum coefficients and constructing a noisy speech amplitude spectrum observation matrix; A module for performing non-negative matrix decomposition on the noisy speech amplitude spectrum observation matrix to calculate the NMF amplitude spectrum gain; A module for performing feature extraction based on the sub-band energy and calculating the NN amplitude spectrum gain through a pre-trained deep neural network; A module for fusing the NMF amplitude spectrum gain and the NN amplitude spectrum gain to output a noise reduction gain; A module for obtaining a noise reduction spectrum coefficient according to the noise reduction gain and the MDCT spectrum coefficient; A module for continuing the encoding process according to the noise reduction spectrum coefficients and outputting a noise reduction speech code stream.

10. A computer-readable storage medium storing computer instructions, wherein the computer instructions are operated to execute the speech denoising method combining non-negative matrix decomposition and time-frequency masking as described in any one of claims 1-8.