Audio noise reduction filtering method, noise reduction filtering device, electronic device and storage medium

By combining preset neural networks and traditional signal processing, the problem of speech quality loss under low signal-to-noise ratio is solved, and effective noise reduction and voice quality improvement under low signal-to-noise ratio conditions are achieved.

CN114495960BActive Publication Date: 2025-08-08ZHEJIANG DAHUA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111605349.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-25
Publication Date
2025-08-08
Estimated Expiration
2041-12-25

AI Technical Summary

Technical Problem

When the signal-to-noise ratio of existing audio noise reduction technology decreases seriously, voice quality loss occurs and word stuttering occurs.

Method used

The preset neural network is used to combine the traditional signal processing method, and the filtering weight coefficient is calculated by obtaining the characteristic parameters of the audio input signal, and the preset neural network is trained using the generation value to obtain the filtered audio signal.

Benefits of technology

Under low signal-to-noise ratio conditions, it effectively reduces noise, improves voice quality, and steadily ensures voice processing effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114495960B_ABST
    Figure CN114495960B_ABST
Patent Text Reader

Abstract

The present application discloses an audio noise reduction filtering method, a noise reduction filter device, an electronic device, and a computer storage medium, relating to the field of audio signal processing technology. The method comprises: obtaining characteristic parameters of an audio input signal using a preset neural network; calculating a filter weight coefficient based on the characteristic parameters; processing the audio input signal based on the filter weight coefficient to obtain a filtered audio signal; calculating a cost value based on the filtered audio signal and a real signal; and training the preset neural network using the cost value. Through the above-mentioned methods, the audio noise reduction filtering method of the present application can effectively reduce noise in an audio system and improve voice quality.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of audio signal processing, and in particular to an audio noise reduction filtering method, a noise reduction filtering device, an electronic device, and a computer storage medium. Background Art

[0002] In real life, when people use mobile devices such as mobile phones to turn on hands-free calls or video conferencing terminals for video conferencing, there are various noises in the on-site environment. In addition to collecting the target signal, the microphone also collects environmental noise. Therefore, filtering technology is needed to suppress the noise. However, when the signal-to-noise ratio of current noise reduction technology becomes low, the voice will be severely lost, and the processed voice will be interrupted. Summary of the Invention

[0003] The main technical problem solved by this application is to provide an audio noise reduction filtering method, a noise reduction filtering device, an electronic device and a computer storage medium to reduce noise and improve the quality of speech.

[0004] To solve the above technical problems, the present application adopts a technical solution: providing an audio noise reduction filtering method. The method includes:

[0005] A preset neural network is used to obtain characteristic parameters of an audio input signal; a filter weight coefficient is calculated based on the characteristic parameters; the audio input signal is processed based on the filter weight coefficient to obtain a filtered audio signal; a cost value is calculated based on the filtered audio signal and a real signal; and the preset neural network is trained using the cost value.

[0006] To solve the above technical problems, another technical solution adopted by this application is to provide a noise reduction filter device. The noise reduction filter device includes:

[0007] A preset neural network module is used to obtain characteristic parameters of the audio input signal; a calculation module is connected to the preset neural network module and is used to calculate the filter weight coefficient based on the characteristic parameters; a filter module is connected to the calculation module and is used to process the audio input signal based on the filter weight coefficient to obtain a filtered audio signal; the calculation module is further used to calculate the cost value of the filtered audio signal and the real signal, and send the cost value to the preset neural network module for training using the cost value.

[0008] In order to solve the above technical problems, another technical solution adopted in this application is: to provide an electronic device, which includes a processor and a memory connected to the processor, the memory storing program data, and the processor executing the program data stored in the memory to perform: obtaining characteristic parameters of an audio input signal using a preset neural network; calculating a filtering weight coefficient based on the characteristic parameters; processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; calculating a cost value based on the filtered audio signal and a real signal; and training the preset neural network using the cost value.

[0009] To solve the above technical problems, another technical solution adopted in this application is: providing a computer storage medium, which stores program instructions internally, and the program instructions are executed to achieve: using a preset neural network to obtain characteristic parameters of an audio input signal; calculating a filtering weight coefficient based on the characteristic parameters; processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; calculating a cost value based on the filtered audio signal and a real signal; and using the cost value to train the preset neural network.

[0010] The beneficial effects of the present application are as follows: Different from the existing technology, the audio noise reduction filtering method of the present application uses a combination of a preset neural network and traditional signal processing to process audio sounds. The advantage of the preset neural network is that it can better solve the characteristic parameters that are difficult to estimate in traditional signal processing. The preset neural network is trained to obtain a final preset neural network model, thereby obtaining the filter weight coefficient to perform traditional signal processing on the audio sound, which can robustly ensure voice quality. The present application combines the preset neural network with traditional signal processing so that the two complement each other. When the signal-to-noise ratio becomes low, it can also reduce the noise of the audio input signal, thereby improving the voice quality of the final audio output. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 This is a structural diagram of an embodiment of the noise reduction filter device of the present application;

[0012] Figure 2 This is a flow chart of an embodiment of the audio noise reduction filtering method of the present application;

[0013] Figure 3 yes Figure 2 A specific flow chart of step S101;

[0014] Figure 4 yes Figure 2 A specific flow chart of step S101;

[0015] Figure 5 yes Figure 2 A specific flow chart of step S102;

[0016] Figure 6 yes Figure 5 A specific flow chart of step S401;

[0017] Figure 7 yes Figure 5 A specific flow chart of step S402;

[0018] Figure 8 This is a flow chart of another embodiment of the audio noise reduction filtering method of the present application;

[0019] Figure 9 This is a schematic diagram of an implementation scheme of an embodiment of the audio noise reduction filtering method of the present application;

[0020] Figure 10 This is a structural diagram of an embodiment of an electronic device of the present application;

[0021] Figure 11 It is a structural diagram of an embodiment of the computer storage medium of the present application. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] This application first proposes a noise reduction filter device 100, such as Figure 1 As shown, Figure 1 : is a schematic structural diagram of an embodiment of a noise reduction filter device of the present application. The noise reduction filter device 100 of this embodiment includes:

[0024] The preset neural network module 110 is used to obtain characteristic parameters of the audio input signal; the calculation module 120 is connected to the preset neural network module 110 and is used to calculate the filter weight coefficient based on the characteristic parameters; the filter module 130 is connected to the calculation module 120 and is used to process the audio input signal based on the filter weight coefficient to obtain a filtered audio signal; the calculation module 120 is further used to calculate the cost value of the filtered audio signal and the real signal, and send the cost value to the preset neural network module 110 for training using the cost value.

[0025] The preset neural network module 110 is used to train the characteristic parameters of the audio input signal. The characteristic parameters of the audio input signal to be processed are input into the preset neural network module 110 for training, and three processed characteristic parameters can be obtained, namely the noise covariance matrix, the received signal covariance matrix and the prior signal-to-noise ratio.

[0026] Optionally, the preset neural network module 110 can adopt various common neural networks, such as recurrent neural network (RNN), convolutional neural network (CNN) and convolutional recurrent neural network (CRNN).

[0027] Taking CNN as an example, the default neural network module can be viewed as an end-to-end black box, with a hidden layer in the middle. The hidden layer can include convolutional layers and pooling layers, with one end being the input layer and the other being the output layer. When the input layer inputs the feature parameters of the audio input signal to be processed, normalization is used in the input layer for ease of calculation. In the convolutional layer in the hidden layer, feature extraction is performed on the feature parameters to enhance the characteristics of the original signal and reduce noise. The pooling layer in the hidden layer preserves useful information while minimizing the amount of data. Finally, the output layer outputs the three processed feature parameters after training. The parameters in the hidden layer are updated with each training session.

[0028] The calculation module 120 is connected to the filter module 130 at one end, and uses the noise covariance matrix, the received signal covariance matrix and the prior signal-to-noise ratio to calculate the inter-frame correlation coefficient and the filtering weight coefficient. The other end is connected to the preset neural network module 110, and obtains the noise covariance matrix, the received signal covariance matrix and the prior signal-to-noise ratio from the preset neural network module 110, and inputs the audio signal filtered by the filter module 130 and the real signal into the cost function to calculate the cost value, and outputs it to the preset neural network module 110, so that the preset neural network module 110 continues to train until the final preset neural network module 110 is obtained.

[0029] Filter module 130 is connected to the preset neural network module 110. It trains the preset neural network module 110 to obtain three processed feature parameters. It then calculates the filter weight parameters based on a formula and inputs them into filter module 130. The filter module 130 then processes the audio input signal to obtain a filtered audio signal. Filter module 130 can be any audio filter and is not limited here.

[0030] The preset neural network module 110 and the filter module 130 are used to process the audio input signal in combination, which solves the problem that some filter weight coefficients in traditional signal processing methods are difficult to estimate, effectively balances the noise reduction effect and voice quality, and improves the final effect of voice processing.

[0031] This application further proposes an audio noise reduction filtering method, such as Figure 2 As shown, Figure 2 This is a flow chart of an embodiment of an audio noise reduction filtering method according to an embodiment of the present application. The method can be used in the above-mentioned noise reduction filtering device 100, and specifically includes steps S101 to S105:

[0032] Step S101: Acquire characteristic parameters of an audio input signal using a preset neural network.

[0033] Obtain sample audio, obtain feature parameters of the audio input signal to be processed, input them into a preset neural network for training, and obtain feature parameters of the processed audio input signal from the preset neural network. The obtained processed feature parameters include: noise covariance matrix, received signal covariance matrix and prior signal-to-noise ratio.

[0034] Optionally, this embodiment can be implemented by Figure 3 The method shown implements step S101, and the specific implementation steps include steps S201 to S202:

[0035] Step S201: Based on the audio input signal, obtain the real part and the imaginary part of the audio input signal.

[0036] Model Y receiving signal from microphone k.l =X k,l +N k,l For example, where X k,l represents the target signal, N k,l represents the noise signal, Y k.l Represents the audio input signal of the microphone, k represents the frequency point, and l represents the time frame. Since the operation is the same for each frequency point, we omit the symbols of the frequency points in the following text.

[0037] For noise reduction algorithms, all noise reduction methods can be regarded as calculating a weight vector for the microphone audio input signal, and recovering the target signal through the weight vector, that is:

[0038]

[0039] w l is the filter weight coefficient, is the filtered audio signal.

[0040] In the multi-frame algorithm, formula (1) can be changed to:

[0041]

[0042] in:

[0043] y l =[Y l ,Y l-1 ,…,Y l-N+1 ] T (3)

[0044] in* T represents the transpose of the matrix, * H represents the conjugate transpose of the matrix, w l The representation method is the same as above; N is usually 4, which means that 4 historical frames are taken.

[0045] Now assume that the target signal and the noise signal in the signal received by the microphone are unrelated, then:

[0046] Φ y,l =Φ x,l +Φ n.l (4)

[0047] where Φ x,l represents the target signal covariance matrix, Φ n.l represents the noise covariance matrix, Φ y,l represents the received signal covariance matrix.

[0048] For the noise covariance matrix Φ n.l and the received signal covariance matrix Φ y,l , based on the audio input signal Y l , use formula (5) to obtain the real and imaginary parts of the audio input signal value.

[0049] y c,l =[Real(Y l ),Imag(Y l )] T (5)

[0050] Among them, Real(Y l ) represents the audio input signal Y l Take the real part, Imag(Y l ) represents the audio input signal Y l Take the imaginary part, y c,l Indicates the audio input signal Y l The matrix of real and imaginary parts, * T Represents the transpose of a matrix.

[0051] Step S202: Obtain a noise covariance matrix and a received signal covariance matrix based on the real part and the imaginary part.

[0052] Model Y receiving signal from microphone k.l =X k,l +N k,l For example, based on the above audio input signal Y l The real part and the imaginary part can be mapped using a preset neural network and then arranged according to the Hermitian matrix to obtain the noise covariance matrix and the received signal covariance matrix. The estimated values of the noise covariance matrix and the received signal covariance matrix can be obtained using formulas (6) and (7):

[0053]

[0054] Among them, Hermitian{·} means that the values in the brackets are arranged according to the format of the Hermitian matrix. Expressed as the estimated value of the covariance matrix of the received signal, represents the estimated value of the noise covariance matrix, y c,l Indicates the audio input signal Y l The matrix of real and imaginary parts, Indicates different mapping methods of the preset neural network. The preset neural network can adopt various common neural networks, such as RNN, CNN, CRNN, etc.

[0055] Optionally, this embodiment can be implemented by Figure 4 The method shown implements step S101, and the specific implementation steps include steps S301 to S302:

[0056] Step S301: Obtain the absolute value of an audio input signal, and calculate the base 10 logarithm of the absolute value.

[0057] Model Y receiving signal from microphone k.l =X k,l +N k,l For example, based on the audio input signal Y l , and get the logarithm log 10 |Y l |value.

[0058] Step S302: obtaining a priori signal-to-noise ratio based on the logarithm.

[0059] Based on the above logarithms, the mapping method of the preset neural network is used to obtain the priori signal-to-noise ratio. The estimated value of the priori signal-to-noise ratio can be obtained using formula (8):

[0060]

[0061] in, represents the estimate of the prior signal-to-noise ratio, Indicates different mapping methods of the neural network used. The preset neural network can adopt various common neural networks, such as RNN, CNN, CRNN, etc.

[0062] Step S102: Calculate the filtering weight coefficient based on the characteristic parameters.

[0063] The preset neural network module calculates the filter weight coefficient through a formula based on the characteristic parameters of the audio input signal. The characteristic parameters of the audio input signal include: noise covariance matrix, received signal covariance matrix and prior signal-to-noise ratio.

[0064] Optionally, this embodiment can be implemented by Figure 5 The method shown implements step S102, and the specific implementation steps include steps S401 to S402:

[0065] Step S401: Calculate the inter-frame correlation coefficient based on the noise covariance matrix, the received signal covariance matrix and the priori signal-to-noise ratio.

[0066] Model Y receiving signal from microphone k.l =X k,l +N k,l As an example, it is assumed that the multi-frame target signal can be decomposed as follows:

[0067] x l =γ x,l X l +x′ l (9)

[0068] Among them, γ x,l X l Indicates the correlation components between signals in multiple frames, x′ l represents the non-correlated components in the multi-frame signal, γ x,l represents the inter-frame correlation coefficient.

[0069] For speech signals, the quality of the speech signals can be guaranteed when the correlation components between the speech signals are guaranteed.

[0070] The calculation module calculates the inter-frame correlation coefficient based on the noise covariance matrix, the received signal covariance matrix and the estimated value of the prior signal-to-noise ratio obtained by the preset neural network as the true value.

[0071] Optionally, this embodiment can be implemented by Figure 6 The method shown implements step S401, and the specific implementation steps include steps S501 to S506:

[0072] Step S501: obtaining a sum of a priori signal-to-noise ratio and a reciprocal of the priori signal-to-noise ratio, and obtaining a first product of the sum, a received signal covariance matrix, and a preset matrix.

[0073] Step S502: Obtain the transpose of the preset matrix, the second product of the received signal covariance matrix and the preset matrix.

[0074] Step S503: Obtain the third product of the inverse of the priori signal-to-noise ratio, the noise covariance matrix, and the preset matrix.

[0075] Step S504: Obtain the transpose of the preset matrix, the noise covariance matrix, and the fourth product of the preset matrix.

[0076] Step S505: Obtain a first quotient of the first product and the second product, and obtain a second quotient of the third product and the fourth product.

[0077] Step S506: Obtain the difference between the first quotient and the second quotient to obtain the inter-frame correlation coefficient, wherein the preset matrix e=[1,0,…,0] T .

[0078] Model Y receiving signal from microphone k.l =X k,l +N k,l For example, steps S501 to S506 can be implemented using formula (10):

[0079]

[0080] Among them, γ x,l Expressed as the inter-frame correlation coefficient, ξ l represents the prior signal-to-noise ratio, γ x,l represents the correlation coefficient between the frame signals, Φ y,l represents the received signal covariance matrix, Φ x,l represents the target received signal covariance matrix, Φ n.l Represents the noise covariance matrix, the preset matrix e=[1,0,…,0] T , e T Represents the transpose of the preset matrix e.

[0081] Step S402: Calculate filtering weight coefficients based on the inter-frame correlation coefficient and the noise covariance matrix.

[0082] Model Y receiving signal from microphone k.l =X k,l +N k,l For example, the calculation module calculates the filtering weight coefficient based on the inter-frame correlation coefficient obtained above and the estimated value of the noise covariance matrix obtained by the preset neural network as the true value.

[0083] Optionally, this embodiment can be implemented by Figure 7 The method shown implements step S402, and the specific implementation steps include steps S601 to S603:

[0084] Step S601: Obtain the fifth product of the inverse matrix of the noise covariance matrix and the inter-frame correlation coefficient.

[0085] Step S602: Obtain the conjugate transpose of the inter-frame correlation coefficient, the inverse matrix of the noise covariance matrix, and the sixth product of the inter-frame correlation coefficient.

[0086] Step S603: Obtain a third quotient of the fifth product and the sixth product to obtain a filtering weight coefficient.

[0087] Model Y receiving signal from microphone k.l =X k,l +N k,l For example, according to the definition of minimum variance distortion-free response:

[0088]

[0089] Steps S601 to S603 can be implemented using formula (12):

[0090]

[0091] Among them, Φ n.l represents the noise covariance, Represents Φ n.l The inverse of a matrix, Represents the estimated value of the filter weight coefficient, γ x,l represents the correlation coefficient between frame signals, Represents the conjugate transpose of the correlation coefficient between frame signals.

[0092] Step S103: Process the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal.

[0093] The filter module processes the audio input signal based on the filter weight coefficient calculated by the calculation module to obtain a filtered audio signal.

[0094] Step S104: Calculate a cost value based on the filtered audio signal and the real signal.

[0095] The calculation module calculates a cost value based on the filtered audio signal and the real signal.

[0096] Step S105: using the cost value to train the preset neural network.

[0097] The preset neural network module trains the preset neural network using the cost value until the preset neural network module converges or reaches a preset training number of times, and the trained preset neural network module is used to process subsequent audio input signals.

[0098] This application further proposes an audio noise reduction filtering method, such as Figure 8 As shown, Figure 8 This is a flow chart of another embodiment of the audio noise reduction filtering method of the present application, wherein the specific implementation steps include steps S701 to S706:

[0099] Step S701: Acquire characteristic parameters of an audio input signal using a preset neural network.

[0100] Step S701 is the same as step S101 and will not be described again.

[0101] Step S702: Calculate the filtering weight coefficient based on the characteristic parameters.

[0102] Step S702 is the same as step S102 and will not be described again.

[0103] Step S703: Process the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal.

[0104] Step S703 is the same as step S103 and will not be described again.

[0105] Step S704: Constructing a cost function for the preset neural network.

[0106] The cost function adopts formula (13):

[0107]

[0108] in, represents the real signal, Represents the filtered audio signal.

[0109] For neural networks, a training target is required when conducting network training. There will be three training targets for the three characteristic parameters mentioned above, which will be unfriendly to network training. Therefore, in this scheme, the three characteristic parameters can be calculated using formula (12) to obtain a weight coefficient, and the weight coefficient is multiplied by the unprocessed signal to obtain the processed signal, and the real signal is used as the training target.

[0110] Step S705: Calculate the cost values of the filtered audio signal and the real signal using the cost function.

[0111] The filtered audio signal and the true signal are input into the cost function to obtain the cost value.

[0112] Step S706: Using the cost value to train the preset neural network.

[0113] Step S706 is the same as step S105 and will not be described again.

[0114] Optionally, the audio noise reduction filtering method of this embodiment further includes step S707:

[0115] Step S707: In response to the preset neural network converging, the audio input signal is processed using the corresponding filter weight coefficient to obtain a target signal.

[0116] When the preset neural network converges or reaches the preset number of training times, the cost value is the current minimum value, and the trained preset neural network model is obtained. The audio input signal is processed by the filtering weight coefficient obtained by the preset neural network model at this time to obtain the target signal.

[0117] In an application scenario, such as Figure 9 As shown, Figure 9 This is a schematic diagram of an embodiment of the audio noise reduction filtering method of the present application. The dotted line portion in the figure represents the process flow required for the preset neural network training, which is not required in the actual inference process.

[0118] like Figure 9 As shown, the audio input signal is sent to the preset neural network module 110 for training to obtain three different processed feature parameters, namely the noise covariance matrix, the input signal covariance matrix, and the prior signal-to-noise ratio; the processed feature parameters are calculated according to the formula by the calculation module 120 to obtain the inter-frame correlation coefficient; the filtering weight coefficient is calculated according to the inter-frame correlation coefficient and the prior signal-to-noise ratio; based on the filtering weight coefficient, the signal is filtered and output using the filter module 130, and the filtered signal and the real signal are sent to the cost function in the calculation module 120 to calculate the cost value, and then sent back to the preset neural network module 110.

[0119] Optionally, the present application further proposes an electronic device 200. Figure 10 As shown, Figure 10 2 is a schematic structural diagram of an embodiment of an electronic device 200 of the present application. The electronic device 200 includes a processor 201 and a memory 202 connected to the processor 201 .

[0120] Processor 201 may also be referred to as a CPU (Central Processing Unit). Processor 201 may be an integrated circuit chip having signal processing capabilities. Processor 201 may also be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, or a discrete hardware component. A general-purpose processor may be a microprocessor or any conventional processor.

[0121] The memory 202 is used to store program data required for the processor 201 to run.

[0122] The processor 201 is used to execute the program data stored in the memory 202 to achieve: using a preset neural network to obtain characteristic parameters of the audio input signal; calculating the filtering weight coefficient based on the characteristic parameters; processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; calculating a cost value based on the filtered audio signal and a real signal; and using the cost value to train the preset neural network.

[0123] Optionally, the present application further proposes a computer storage medium 300. Figure 11 As shown, Figure 11 This is a structural diagram of an embodiment of the computer storage medium 300 of the present application.

[0124] The computer storage medium 300 of the embodiment of the present application internally stores program instructions 310, which are executed to implement: obtaining characteristic parameters of an audio input signal using a preset neural network; calculating a filtering weight coefficient based on the characteristic parameters; processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; calculating a cost value based on the filtered audio signal and a real signal; and training the preset neural network using the cost value.

[0125] The program instructions 310 may be stored in the aforementioned storage medium in the form of a program file as a software product, so that a computer device (which may be a personal computer, server, or network device, etc.) or a processor executes all or part of the steps of the various embodiments of the present application. The aforementioned storage medium includes: a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk, etc., which can store program code, or a terminal device such as a computer, server, mobile phone, or tablet.

[0126] Different from the existing technology, the audio noise reduction filtering method of the present application uses a combination of a preset neural network and traditional signal processing to process audio sounds. The preset neural network can better solve the characteristic parameters that are difficult to estimate in traditional signal processing. The preset neural network is trained to obtain the final preset neural network model, thereby obtaining the filter weight coefficient to perform traditional signal processing on the audio sound, which can robustly ensure the voice quality. The present application combines the preset neural network with traditional signal processing so that the two complement each other. When the signal-to-noise ratio becomes low, it can also reduce the noise of the audio input signal and improve the voice quality of the final audio output.

[0127] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.

Claims

1. An audio noise reduction filtering method, characterized in that: include: Using a preset neural network to obtain characteristic parameters of the audio input signal; Calculating a filtering weight coefficient based on the characteristic parameters; Processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; Calculating a cost value based on the filtered audio signal and the real signal; Training the preset neural network using the cost value; The characteristic parameters include: a noise covariance matrix, a received signal covariance matrix, and a priori signal-to-noise ratio, and the calculation of the filter weight coefficient based on the characteristic parameters includes: An inter-frame correlation coefficient is calculated based on the noise covariance matrix, the received signal covariance matrix and the priori signal-to-noise ratio; and a filtering weight coefficient is calculated based on the inter-frame correlation coefficient and the noise covariance matrix.

2. The audio noise reduction filtering method according to claim 1, wherein: The calculating the inter-frame correlation coefficient based on the noise covariance matrix, the received signal covariance matrix and the priori signal-to-noise ratio includes: Obtaining a sum of the priori signal-to-noise ratio and the inverse of the priori signal-to-noise ratio, and obtaining a first product of the sum, the received signal covariance matrix, and a preset matrix; Obtaining a transpose of the preset matrix, the received signal covariance matrix, and a second product of the preset matrix; Obtaining a third product of the inverse of the priori signal-to-noise ratio, the noise covariance matrix, and the preset matrix; Obtaining a fourth product of the transpose of the preset matrix, the noise covariance matrix, and the preset matrix; Obtaining a first quotient of the first product and the second product, and obtaining a second quotient of the third product and the fourth product; Obtaining a difference between the first quotient value and the second quotient value to obtain the inter-frame correlation coefficient; Wherein, the preset matrix e=[1,0,…,0] T .

3. The audio noise reduction filtering method according to claim 2, wherein: The calculating of the filtering weight coefficient based on the inter-frame correlation coefficient and the noise covariance matrix includes: Obtaining a fifth product of an inverse matrix of the noise covariance matrix and the inter-frame correlation coefficient; Obtaining a sixth product of the conjugate transpose of the inter-frame correlation coefficient, the inverse matrix of the noise covariance matrix, and the inter-frame correlation coefficient; A third quotient of the fifth product and the sixth product is obtained to obtain the filtering weight coefficient.

4. The audio noise reduction filtering method according to claim 1, wherein: Further including: Constructing a cost function for the preset neural network; The calculating the cost value based on the filtered audio signal and the real signal includes: Calculating a cost value of the cost function based on the filtered audio signal and the true signal; Among them, the cost function adopts the formula: in, represents the real signal, Represents the filtered audio signal.

5. The audio noise reduction filtering method according to claim 1, wherein: Further including: In response to the preset neural network converging, the audio input signal is processed using the corresponding filtering weight coefficient to obtain a target signal.

6. The audio noise reduction filtering method according to claim 1, wherein: The method of obtaining characteristic parameters of the audio input signal by using a preset neural network includes: Based on the audio input signal, obtaining a real part and an imaginary part of the audio input signal; The noise covariance matrix and the received signal covariance matrix are obtained based on the real part and the imaginary part.

7. The audio noise reduction filtering method according to claim 1, wherein: The method of obtaining characteristic parameters of the audio input signal by using a preset neural network includes: Obtaining an absolute value of the audio input signal, and calculating a base-10 logarithm of the absolute value; The priori signal-to-noise ratio is obtained based on the logarithm.

8. A noise reduction filter device, characterized in that: include: A neural network module is preset to obtain characteristic parameters of an audio input signal; A calculation module, connected to the preset neural network module, for calculating a filter weight coefficient based on the characteristic parameters, wherein the characteristic parameters include: a noise covariance matrix, a received signal covariance matrix, and a priori signal-to-noise ratio, and calculating the filter weight coefficient based on the characteristic parameters includes: calculating an inter-frame correlation coefficient based on the noise covariance matrix, the received signal covariance matrix, and the priori signal-to-noise ratio; and calculating the filter weight coefficient based on the inter-frame correlation coefficient and the noise covariance matrix. a filter module, connected to the calculation module, configured to process the audio input signal based on the filter weight coefficient to obtain a filtered audio signal; The calculation module is further used to calculate the cost values of the filtered audio signal and the real signal, send the cost values to the preset neural network module, and use the cost values for training.

9. An electronic device, characterized in that: The electronic device includes a processor and a memory connected to the processor, wherein the memory stores program data, and the processor executes the program data stored in the memory to implement: Using a preset neural network to obtain characteristic parameters of the audio input signal; Calculating a filter weight coefficient based on the characteristic parameters, wherein the characteristic parameters include: a noise covariance matrix, a received signal covariance matrix, and a priori signal-to-noise ratio, and calculating the filter weight coefficient based on the characteristic parameters includes: calculating an inter-frame correlation coefficient based on the noise covariance matrix, the received signal covariance matrix, and the priori signal-to-noise ratio; and calculating the filter weight coefficient based on the inter-frame correlation coefficient and the noise covariance matrix; Processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; Calculating a cost value based on the filtered audio signal and the real signal; The preset neural network is trained using the cost value.

10. A computer storage medium, characterized in that Program instructions are stored therein, and the program instructions are executed to implement: Using a preset neural network to obtain characteristic parameters of the audio input signal; Calculating a filter weight coefficient based on the characteristic parameters, wherein the characteristic parameters include: a noise covariance matrix, a received signal covariance matrix, and a priori signal-to-noise ratio, and calculating the filter weight coefficient based on the characteristic parameters includes: calculating an inter-frame correlation coefficient based on the noise covariance matrix, the received signal covariance matrix, and the priori signal-to-noise ratio; and calculating the filter weight coefficient based on the inter-frame correlation coefficient and the noise covariance matrix; Processing the audio input signal based on the filtering weight coefficient to obtain a filtered audio signal; Calculating a cost value based on the filtered audio signal and the real signal; The preset neural network is trained using the cost value.

Citation Information

Patent Citations

  • Audio identification method and apparatus, echo cancellation method and apparatus, and device

    CN108429994A

  • Prior signal-to-noise ratio computation method, electronic equipment and storage medium

    CN110634500A