Audio processing method and device and storage medium

By inputting the covariance matrix of audio data into the neural network model to obtain beamformer parameters, the problem of numerical instability in multi-channel speech dereverberation is solved, and frame-level dereverberation processing and effect optimization are achieved.

CN120151740APending Publication Date: 2025-06-13BEIJING XIAOMI MOBILE SOFTWARE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311707743.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-12
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The prior art has numerical instability in the multi-channel voice dereverberation process, resulting in unsatisfactory dereverberation effect in some cases.

Method used

By determining the covariance matrix of the audio data to be processed, it is input to the neural network model to obtain the beamformer parameters, which are used for frame-level dereverberation processing.

Benefits of technology

The frame-level dereverberation processing is realized, and the audio dereverberation effect is optimized, avoiding the problem of numerical instability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120151740A_ABST
    Figure CN120151740A_ABST
Patent Text Reader

Abstract

The invention relates to an audio processing method and device and a storage medium. The audio processing method comprises the following steps: determining a first covariance matrix of to-be-processed audio data amplitude; inputting the first covariance matrix into a first neural network model to obtain a beam former parameter, the beam former parameter being used for dereverberation processing of audio data; and de-reverberation processing is carried out on the audio data to be processed based on the beam former parameters to obtain audio data after de-reverberation processing. According to the invention, frame-level de-reverberation processing is realized, so that the audio de-reverberation effect is optimized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of audio processing, and in particular, to an audio processing method, apparatus, and storage medium. Background Art

[0002] Currently, for multi-channel speech dereverberation, beamforming methods are often used, which can effectively reduce non-linear distortion. However, the processing results exhibit numerical instability. That is, the dereverberation effect is not ideal in some cases. Summary of the Invention

[0003] To overcome the problems existing in the related art, the present disclosure provides an audio processing method, apparatus, and storage medium.

[0004] According to a first aspect of an embodiment of the present disclosure, an audio processing method is provided, including: determining a first covariance matrix of the amplitude of audio data to be processed; inputting the first covariance matrix into a first neural network model to obtain beamformer parameters for performing dereverberation processing on the audio data; and performing dereverberation processing on the audio data to be processed based on the beamformer parameters to obtain dereverberated audio data.

[0005] In an implementation, the first neural network model is a neural network model with a recursive iteration characteristic, and the input of the first neural network model includes a first covariance matrix of the amplitude of the Nth frame of speech data to be processed and an output result of the first neural network model corresponding to the (N - 1)th frame of speech data to be processed, where N is a positive integer greater than 1.

[0006] In an implementation, the inputting the first covariance matrix into the first neural network model to obtain beamformer parameters includes: performing layer normalization processing on the first covariance matrix to obtain a second covariance matrix of the amplitude of the audio data to be processed; and causing the second covariance matrix to sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

[0007] In one embodiment, the audio data to be processed includes speech data and noise data, the second covariance matrix includes the second covariance matrix of the speech data amplitude and the second covariance matrix of the noise data amplitude, and making the second covariance matrix sequentially go through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters includes: splicing the second covariance matrix of the speech data amplitude and the second covariance matrix of the noise data amplitude to obtain a third covariance matrix of the audio data amplitude to be processed; making the third covariance matrix sequentially go through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

[0008] In one embodiment, the first neural network model is trained in the following manner: obtaining first training data, where the first training data is speech of multiple fixed lengths; performing reverberation processing on the first training data to obtain second training data, where the second training data is reverberant speech of multiple fixed lengths, and the reverberant speech includes speech and noise; inputting the second training data into an initial neural network model to obtain an output result; using the minimum mean square error between the output result and the first training data to train the initial neural network model to obtain the first neural network model.

[0009] In one embodiment, determining the first covariance matrix of the audio data amplitude to be processed includes: extracting features of the audio data to be processed, where the features include the amplitude spectrum of the audio data to be processed; inputting the features into a second neural network model to obtain an ideal ratio mask of the audio data to be processed, where the second neural network model includes a convolutional layer, a deconvolutional layer, and a gated recurrent unit; where the audio data to be processed includes speech data and noise data, the ideal ratio mask includes an ideal ratio mask of the speech data and an ideal ratio mask of the noise data, the ideal ratio mask of the speech data represents the proportion of the speech data in the audio data to be processed, and the ideal ratio mask of the noise data represents the proportion of the noise data in the audio data to be processed; obtaining the first covariance matrix based on the ideal ratio mask and the audio data to be processed.

[0010] In one embodiment, the audio data to be processed includes audio data of M microphone channels, M is a positive integer, the amplitude spectrum of the audio data to be processed includes M amplitude spectra, and inputting the features into the second neural network model includes: adding each of the M amplitude spectra to a reference amplitude spectrum to obtain M added amplitude spectra, and subtracting each of the M amplitude spectra from the reference amplitude spectrum to obtain M subtracted amplitude spectra; splicing the M amplitude spectra, the M added amplitude spectra, and the M subtracted amplitude spectra; inputting the spliced amplitude spectra into the second neural network model.

[0011] According to a second aspect of the embodiments of the present disclosure, there is provided an audio processing apparatus, including: a determination unit configured to determine a first covariance matrix of the amplitude of audio data to be processed; a processing unit configured to input the first covariance matrix into a first neural network model to obtain beamformer parameters for performing dereverberation processing on the audio data; and perform dereverberation processing on the audio data to be processed based on the beamformer parameters to obtain dereverberated audio data.

[0012] In one embodiment, the first neural network model is a neural network model with recursive iteration characteristics, and the input of the first neural network model includes the first covariance matrix of the amplitude of the Nth frame of speech data to be processed and the output result of the first neural network model corresponding to the (N - 1)th frame of speech data to be processed, where N is a positive integer greater than 1.

[0013] In one embodiment, the processing unit inputs the first covariance matrix into the first neural network model in the following manner to obtain beamformer parameters: perform layer normalization processing on the first covariance matrix to obtain a second covariance matrix of the amplitude of the audio data to be processed; and cause the second covariance matrix to sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

[0014] In one embodiment, the audio data to be processed includes speech data and noise data, the second covariance matrix includes a second covariance matrix of the amplitude of the speech data and a second covariance matrix of the amplitude of the noise data, and the processing unit causes the second covariance matrix to sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer in the following manner to obtain the beamformer parameters: splice the second covariance matrix of the amplitude of the speech data and the second covariance matrix of the amplitude of the noise data to obtain a third covariance matrix of the amplitude of the audio data to be processed; and cause the third covariance matrix to sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

[0015] In one embodiment, the processing unit trains a first neural network model in the following manner: obtaining first training data, where the first training data is a plurality of fixed-length voices; performing reverberation processing on the first training data to obtain second training data, where the second training data is a plurality of fixed-length reverberant voices, and the reverberant voices include voices and noises; inputting the second training data into an initial neural network model to obtain an output result; and training the initial neural network model by using the minimum mean square error between the output result and the first training data to obtain the first neural network model.

[0016] In one embodiment, the determining unit determines a first covariance matrix of the amplitude of the audio data to be processed in the following manner: extracting features of the audio data to be processed, where the features include the amplitude spectrum of the audio data to be processed; inputting the features into a second neural network model to obtain an ideal ratio mask of the audio data to be processed, where the second neural network model includes a convolutional layer, a deconvolutional layer, and a gated recurrent unit; where the audio data to be processed includes speech data and noise data, the ideal ratio mask includes an ideal ratio mask of the speech data and an ideal ratio mask of the noise data, the ideal ratio mask of the speech data represents the proportion of the speech data in the audio data to be processed, and the ideal ratio mask of the noise data represents the proportion of the noise data in the audio data to be processed; and obtaining the first covariance matrix based on the ideal ratio mask and the audio data to be processed.

[0017] In one embodiment, the audio data to be processed includes audio data of M microphone channels, M being a positive integer, the amplitude spectrum of the audio data to be processed includes M amplitude spectra, and the determining unit inputs the features into the second neural network model in the following manner: adding each of the M amplitude spectra to a reference amplitude spectrum to obtain M added amplitude spectra, and subtracting each of the M amplitude spectra from the reference amplitude spectrum to obtain M subtracted amplitude spectra; splicing the M amplitude spectra, the M added amplitude spectra, and the M subtracted amplitude spectra; and inputting the spliced amplitude spectra into the second neural network model.

[0018] According to a third aspect of the embodiments of the present disclosure, there is provided an electronic device, including: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the audio processing method in the first aspect and any one of the embodiments of the first aspect.

[0019] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions, and when the instructions are executed by a processor, the audio processing method in the first aspect or any one of the embodiments of the first aspect is executed.

[0020] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: By determining the covariance matrix of the audio data to be processed, inputting the covariance matrix into a neural network model to obtain beamformer parameters, and using these parameters to perform dereverberation processing on the audio data to be processed, frame-level dereverberation processing is achieved to optimize the effect of audio dereverberation.

[0021] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure and, together with the specification, are used to explain the principles of the present disclosure.

[0023] Figure 1 is a flowchart of an audio processing method shown according to an exemplary embodiment.

[0024] Figure 2 is a flowchart of a method for generating beamformer parameters shown according to an exemplary embodiment.

[0025] Figure 3 is a flowchart of a method for generating beamformer parameters shown according to an exemplary embodiment.

[0026] Figure 4 is a flowchart of a method for training a first neural network model shown according to an exemplary embodiment.

[0027] Figure 5 is a flowchart of a method for determining a first covariance matrix shown according to an exemplary embodiment.

[0028] Figure 6 is a schematic diagram of a second neural network model shown according to an exemplary embodiment.

[0029] Figure 7 is a flowchart of a method for determining an ideal ratio mask shown according to an exemplary embodiment.

[0030] Figure 8 is a flowchart of an audio processing method shown according to an exemplary embodiment.

[0031] Figure 9 is a flowchart of an audio processing method shown according to an exemplary embodiment.

[0032] Figure 10 is a block diagram of an audio processing apparatus shown according to an exemplary embodiment.

[0033] Figure 11It is a block diagram of an audio processing device shown according to an exemplary embodiment. Detailed implementation mode

[0034] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure.

[0035] Currently, for multi-channel speech dereverberation, beamforming methods are often used, which can effectively reduce non-linear distortion. However, there is a phenomenon of numerical instability in the processing results. That is, the dereverberation effect is not ideal in some cases.

[0036] In one implementation manner, the speech dereverberation problem can be divided into single-channel speech dereverberation and multi-channel speech dereverberation according to the number of microphones used. Multi-channel speech dereverberation can adopt beamforming methods, which can reduce non-linear distortion and is beneficial to improving the recognition accuracy of the speech recognition backend. For example, a feedforward network (FFN) and a bi-directional long short-term memory network (BLSTM) are used for mask estimation. Among them, the mask can also be called a mask, which is equivalent to covering a mask on the original input data to shield or select some specific elements. The mask in this implementation manner includes the speech part mask M estimated using FNN and BLSTM S , and the noise part mask M calculated by 1 - M S . According to M N and M S , the covariance matrices of speech and noise are calculated respectively. As shown in Formula 1: N

[0037]

[0038] Among them, Φ VV represents the covariance matrix, V belongs to (S, N), that is, V can be S or N. When V is S, Φ SS represents the covariance matrix of speech, and when V is N, Φ NN represents the covariance matrix of noise. ∑ represents the summation formula, represents starting from t = 1 until t = T, finding the sum of M v (t)Y(t)Y H . That is, finding the sum of M v (1)Y(1)Y H (1)+M​v (2)Y(2)Y H( 2)+···M v (T)Y(T)Y H (T)。M v 2 (t) represents M v (t) squared. Where v belongs to (S, N), when v is S, M S represents the speech part mask, when V is N, M N represents the noise part mask. H represents the conjugate transpose symbol, that is, finding the conjugate matrix of the matrix and then inverting the conjugate matrix. Y(t) represents the amplitude spectrum of the multi-channel audio. The amplitude of the audio signal changes with frequency, and the corresponding spectrogram is called the amplitude spectrum. Y H (t) means first finding the conjugate matrix of Y(t) and then inverting the conjugate matrix.

[0039] After obtaining the covariance matrices of speech and noise, the parameters of the beamformer (minimum variance distortionless response, MVDR) can be calculated through Formula 2. Formula 2 is as follows:

[0040]

[0041] Where, W MVDR represents the parameters of the wave velocity former. d represents the spatial filtering coefficient vector of the desired signal. For example, the desired signal is the speech signal. H is the conjugate device symbol, d H represents finding the conjugate vector of d and then inverting it. Φ NN -1 represents inverting Φ NN Inverting.

[0042] After obtaining the parameters of the beamformer, the enhanced speech signal can be calculated through Formula 3, that is, the reverberation of the multi-channel speech is removed. Formula 3 is as follows:

[0043]

[0044] Where, represents the enhanced speech information signal. Where, t represents the time domain corresponding to the speech signal, and f represents the frequency domain corresponding to the speech signal. Y(t,f) represents the amplitude spectrum obtained by performing a short-time Fourier transform (STFT) on Y(t).

[0045] However, in this embodiment, since the inverse of some matrices cannot be calculated, numerical instability may occur. On the other hand, since the beamformer parameters can only be obtained at the segment level or sentence level according to the formula calculation, the noise at the frame level cannot be eliminated, and thus the reverberation cancellation effect is not good.

[0046] Therefore, the present disclosure provides an audio processing method. By determining the covariance matrix of the audio data to be processed and inputting the covariance matrix into a neural network model, beamformer parameters at the frame level can be obtained, and the audio data to be processed is de-reverberated using the parameters to optimize the audio de-reverberation effect.

[0047] Among them, the audio processing method of the present disclosure can be applied to a terminal. A terminal can also be referred to as a terminal device, a mobile station (MS), a mobile terminal (MT), etc. It is a device that provides voice and / or data connectivity to users. For example, a terminal can be a handheld device with wireless connection capabilities, a vehicle-mounted device, etc. Currently, some examples of terminals are: a mobile phone, a pocket personal computer (PPC), a palm computer, a personal digital assistant (PDA), a laptop computer, a tablet computer, a wearable device, or a vehicle-mounted device, etc. In addition, when it is a vehicle-to-everything (V2X) communication system, the terminal device can also be a vehicle-mounted device. It should be understood that the present disclosure embodiments do not limit the specific technologies and specific device forms adopted by the terminal.

[0048] Figure 1 is a flowchart of an audio processing method shown according to an exemplary embodiment. As Figure 1 shown, the audio processing method includes the following steps.

[0049] In step S11, determine the first covariance matrix of the amplitude of the audio data to be processed.

[0050] In one embodiment, the audio data to be processed is multi-channel audio data. The multi-channel audio data represents that multiple microphones simultaneously collect audio data, and each microphone corresponds to one channel. A first covariance matrix of the amplitude of the audio data to be processed can be determined. The first covariance matrix is calculated based on the amplitude spectrum of the audio data and a mask. Among them, the audio data can also be referred to as an audio signal, and the amplitude spectrum of the audio data can be understood as the relationship between the amplitude and frequency of the audio signal. The audio data includes a speech part and a noise part. The speech part in the audio data is called speech data, and the noise part in the audio data is called noise data. It can be understood that in this embodiment, the speech data and the noise data do not need to be separated, and after processing the audio data, a mask for the speech part and a mask for the noise part can be obtained. According to the amplitude spectrum of the audio data and the mask of the speech part, a first covariance matrix of the amplitude of the speech data can be obtained; according to the amplitude spectrum of the audio data and the mask of the noise part, a first covariance matrix of the amplitude of the noise data can be obtained. The audio data corresponds to a first covariance matrix, which can be used to generate beamformer parameters subsequently.

[0051] In step S12, the first covariance matrix is input into the first neural network model to obtain beamformer parameters.

[0052] In one embodiment, the first covariance matrix can be input into the first neural network model to obtain beamformer parameters. The input of the first neural network model is the covariance matrix, and the output is the beamformer parameters. The beamformer parameters obtained through the first neural network model do not need to perform an inverse operation, ensuring the stability of the obtained beamformer parameters. The beamformer parameters are used to perform dereverberation processing on the audio data. Among them, reverberation can be understood as an acoustic characteristic. For example, when a sound source generates sound waves in a room or an auditorium, the observer not only hears the sound waves directly propagated from the sound source, but also hears the reflected waves from the walls, the floor, and the ceiling. The reflected waves can be called reverberation. Dereverberation processing refers to the process of removing reverberation from the audio data.

[0053] In step S13, based on the beamformer parameters, dereverberation processing is performed on the audio data to be processed to obtain the dereverberated audio data.

[0054] In one implementation, reverberation processing can be performed on the audio data to be processed based on beamformer parameters to obtain the audio data after reverberation processing. For example, multiplying the beamformer parameters by the audio data gives the audio data after reverberation processing. Here, the audio data can be the magnitude spectrum of the audio data, that is, multiplying the beamformer parameters by the magnitude spectrum of the audio data gives the magnitude spectrum of the audio data after reverberation processing. Performing an inverse short time Fourier transform (ISTFT) on the magnitude spectrum of the audio data after reverberation processing thus gives the audio data after reverberation processing.

[0055] The present disclosure can obtain frame-level beamformer parameters by determining a first covariance matrix of the magnitude of the audio data to be processed and inputting the first covariance matrix into a first neural network model, and use the parameters to perform reverberation processing on the audio data to be processed to optimize the effect of audio reverberation removal.

[0056] In some implementations, in the audio processing method provided by the present disclosure: the first neural network model is a neural network model with a recursive iteration characteristic, and the input of the first neural network model includes the first covariance matrix of the magnitude of the Nth frame of the speech data to be processed and the output result of the (N - 1)th frame of the speech data to be processed, where N is a positive integer greater than 1.

[0057] In one implementation, the first neural network model has a recursive iteration characteristic. It can be understood that the first neural network model is a recursive iteration process in time. The input of the first neural network model includes the first covariance matrix of the magnitude of the Nth frame of the speech data to be processed and the output result of the first neural network model corresponding to the (N - 1)th frame of the speech data to be processed. Here, N is a positive integer greater than 1. For example, for the first frame of the audio data to be processed, the corresponding first covariance matrix can be input into the first neural network model to obtain an output result. For the second frame of the audio data to be processed, the corresponding first covariance matrix and the output result corresponding to the first frame of the audio data to be processed can be input into the first neural network model. The first neural network model can be a recurrent neural network (RNN), and of course, it can also be other neural network models with a recursive iteration characteristic, which is not limited in the present disclosure.

[0058] By using the first neural network model with a recursive iteration characteristic, the present disclosure can achieve frame-level noise removal, that is, frame-level reverberation processing. Compared with the beamformer parameters calculated according to the formula, which can only perform segment-level and sentence-level reverberation processing, the above embodiments of the present disclosure can remove residual noise and have a better reverberation removal effect.

[0059] In some embodimentsFigure 2 is a flowchart of a method for generating beamformer parameters shown according to an exemplary embodiment. As Figure 2 shown, the audio processing method in the present disclosure generates beamformer parameters by the following steps.

[0060] In step S21, layer normalization is performed on the first covariance matrix to obtain a second covariance matrix of the magnitude of the audio data to be processed.

[0061] In one implementation, layer normalization can be performed on the first covariance matrix. Layer normalization refers to normalizing the inputs of all neurons in batches. For example, making the data follow a normal distribution with a mean of 0 and a variance of 1. To obtain the output result more quickly. For example, as shown in Formulas 4 and 5:

[0062]

[0063]

[0064] Among them, in Formula 4, represents the second covariance matrix of the speech data, and LayerNorm represents the layer normalization symbol. S(t,f)S H (t,f) represents the first covariance matrix of the magnitude of the speech data. represents performing layer normalization on the first covariance matrix of the magnitude of the speech data. Among them, S(t,f) represents the speech data estimation, and the speech data estimation refers to the speech data extracted from the audio data to be processed using the ideal ratio mask (IRM) of the speech data. Among them, IRM can also be called the ideal ratio mask, or ideal ratio masking. IRM is a masking strategy that defines the ideal energy ratio of each component. For example, for the IRM of the speech data, it can represent the ratio of the speech data in the audio data to be processed. The audio data to be processed includes noise data and speech data. Using the IRM of the speech data, the speech data can be extracted from the audio data with straw ropes. That is, the clean speech data without noise is extracted. The IRM of the noise data can represent the ratio of the noise data in the audio data to be processed. The speech data estimation is obtained by multiplying the IRM of the speech data and the audio data to be processed. S H (t,f) represents taking the conjugate transpose of S(t,f), and the function of the conjugate transpose can refer to the above embodiment, which will not be elaborated herein. In Formula 5, represents the second covariance matrix of the magnitude of the noise data, and LayerNorm represents the layer normalization symbol. N(t,f)N H (t,f) represents the first covariance matrix of the magnitude of the noise data. LayerNorm(N(t,f)N H(t, f) represents performing layer normalization on the first covariance matrix of the noise data amplitude. Among them, N(t, f) represents the noise data estimation, and the noise data estimation refers to the noise data extracted from the audio data to be processed using the IRM of the noise data. That is, the noise data estimation is obtained by multiplying the IRM of the noise data and the audio data to be processed. N H (t, f) represents taking the conjugate transpose of N(t, f). The function of the conjugate transpose can be referred to in the above embodiments, and the present disclosure will not elaborate here.

[0065] In step S22, the second covariance matrix is successively passed through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

[0066] In one implementation, the second covariance matrix can be successively passed through an RNN layer, a non-linear deep neural network (DNN) layer, and a linear DNN layer. For example, referring to formula 6 below:

[0067]

[0068] Among them, W RNN-BF (t, f) represents the beamformer parameters obtained through the first neural network model. RNN-DNN means that the second covariance matrix first passes through the RNN layer, then through the non-linear DNN layer, and then through the linear DNN layer. represents the second covariance matrix of the speech data amplitude. represents the second covariance matrix of the noise data amplitude.

[0069] In one implementation, after obtaining the beamformer parameters, the beamformer parameters can be multiplied by the audio data to be processed to obtain the audio data after dereverberation processing. For example, formula 7:

[0070] S finnal (t, f) = W RNN-BF (t, f)Y(t, f)

[0071] Formula 7

[0072] Among them, S finnal (t, f) represents the audio data after dereverberation processing. represents the beamformer parameters obtained through the first neural network model. Y(t, f) represents the amplitude spectrum of the audio data to be processed. The amplitude spectrum of the audio data after dereverberation processing can be obtained by multiplying the beamformer parameters and the amplitude spectrum of the audio data to be processed. Perform ISTFT on the amplitude spectrum of the audio data after dereverberation processing to obtain the audio data after dereverberation processing.

[0073] In the present disclosure, the second covariance matrix is obtained by performing layer normalization on the first covariance matrix, and the second covariance matrix is input into the first neural network model, which can facilitate the calculation of the first neural model, improve the processing speed of the first neural network model, and improve the efficiency of reverberation removal. Making the second covariance matrix pass through the RNN layer, the linear DNN layer, and the non-linear DNN layer in sequence can make the calculation of the beamformer parameters more accurate and the reverberation removal effect better.

[0074] In some embodiments, Figure 3 is a flowchart of a method for generating beamformer parameters shown according to an exemplary embodiment. As Figure 3 shown, the audio processing method in the present disclosure uses the following steps to generate beamformer parameters.

[0075] In step S31, the second covariance matrix of the speech data amplitude and the second covariance matrix of the noise data amplitude are concatenated to obtain a third covariance matrix of the audio data amplitude to be processed.

[0076] In one embodiment, the audio data includes speech data and noise data. The second covariance matrix includes the second covariance matrix of the speech data amplitude and the second covariance data of the noise data amplitude. The second covariance data of the speech data and the second covariance matrix of the noise can be concatenated to obtain a third covariance matrix of the audio data amplitude to be processed. Among them, the concatenation includes horizontal concatenation and vertical concatenation. For example, assume matrix A is Matrix B is Matrix A and matrix B are horizontally concatenated into matrix C1, then matrix C1 is If matrix A and matrix B are vertically concatenated into matrix C2, then matrix C2 is

[0077] In step S32, the third covariance matrix is sequentially passed through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain beamformer parameters.

[0078] In one embodiment, the third covariance matrix can be sequentially passed through the RNN layer, the non-linear DNN layer, and the linear DNN layer to obtain beamformer parameters.

[0079] The implementation of step S32 in this embodiment can refer to the implementation manner of step S22, and the present disclosure will not repeat it here.

[0080] In the present disclosure, by concatenating the second covariance matrix to obtain a third covariance matrix and inputting the third covariance matrix into the first neural network model, the computational complexity of the first neural network model can be reduced, power consumption can be saved, and the efficiency of reverberation removal can be improved.

[0081] In some embodiments, Figure 4 is a flowchart of a first neural network model training method shown according to an exemplary embodiment, as Figure 4 shown, the audio processing method in the present disclosure trains the first neural network model using the following steps.

[0082] In step S41, first training data is obtained.

[0083] In one embodiment, the first training data is speech data. That is, the first training data is speech of multiple fixed degrees. For example, speech data can be obtained from a Chinese speech dataset as the first training data. The Chinese speech dataset can be, for example, the open-source Aishell1.

[0084] In step S42, reverberation processing is performed on the first training data to obtain second training data, and the second training data is reverberant speech of multiple fixed lengths, and the reverberant speech includes speech and noise.

[0085] In one embodiment, reverberation processing can be performed on the first training data. Reverberation processing is the inverse operation of dereverberation processing. It can be understood that reverberation processing means adding noise to speech data. For example, the first training data can be input into a computer programming language library, which is used to generate reverberation, so as to obtain the second training data. The second training data is audio data. That is, the second training data is reverberant speech containing noise corresponding to the first training data. Reverberant speech means speech containing noise. That is, the reverberant speech contains speech and noise. Among them, the computer programming language library can be, for example, the pyroomacoustics open-source python library. Also, for example, noise can be artificially added to the first training data to obtain the second training data.

[0086] In step S43, the second training data is input into the initial neural network model to obtain an output result.

[0087] In one embodiment, the second training data can be input into the initial neural network model to obtain an output result. The output result can be understood as audio data after dereverberation processing of the second training data. Among them, the initial neural network model can be understood as an untrained neural network model.

[0088] In step S44, the initial neural network model is trained using the minimum mean square error between the output result and the first training data to obtain the first neural network model.

[0089] In one embodiment, the parameters of the initial neural network model can be adjusted by using the minimum mean square error between the output result and the first training data, so as to train the first neural network model. For example, the first training data X corresponds to the second training data Y, and Y represents the audio data after reverberation processing of X. Input X into the initial neural network model to obtain the audio data Z. Use the minimum mean square error between X and Z to adjust the parameters of the initial neural network model so that Z and X are as similar as possible. It can be understood that the more similar Z is to X, the better the effect of the first neural network model.

[0090] The present disclosure obtains the first training data, performs reverberation processing on the first training data to obtain the second training data, and uses the first training data, the second training data, and the minimum mean square error to train the initial neural network model, so that de-reverberation processing can be performed according to the first neural network model, which can ensure the numerical stability of the de-reverberation processing and improve the user experience. The effect of the de-reverberation processing is better, and the effect of speech recognition on the audio data obtained after the de-reverberation processing is also better.

[0091] In some embodiments, Figure 5 is a flowchart of a method for determining a first covariance matrix shown according to an exemplary embodiment. As Figure 5 shown, the audio processing method in the present disclosure determines the first covariance matrix by using the following steps.

[0092] In step S51, features of the audio data to be processed are extracted.

[0093] In one embodiment, features of the audio data to be processed can be extracted. The features include the magnitude spectrum of the audio data to be processed. For example, perform STFT on the audio data to be processed to obtain the magnitude spectrum of the audio data to be processed. Among them, the audio data to be processed includes multi-channel audio data to be processed. The magnitude spectrum of the audio data to be processed includes the magnitude spectrum of the audio data to be processed for each channel.

[0094] In one embodiment, the audio data to be processed can be the audio data of M microphone channels. Wherein, M is a positive integer. For example, if the audio data to be processed corresponds to two channels, that is, M is 2. The features can include the magnitude spectrum of the audio data to be processed for each channel, the result of adding the two magnitude spectra, and the result of subtracting the two magnitude spectra. For example, the magnitude spectra of the two channels are |Y 1 | and |Y 2 |, then the features are [|Y 1 |; |Y 2 |; |Y 1 + Y 2 |; |Y 1 - Y 2 |].

[0095] In one implementation, if M is greater than 2, a reference magnitude spectrum can be determined. The reference magnitude spectrum can be set according to the actual situation, or any one of the M magnitude spectra corresponding to the audio data to be processed in the M channels can be used as the reference magnitude spectrum. Each of the M magnitude spectra is added to the reference magnitude spectrum to obtain M added magnitude spectra. Each of the M magnitude spectra is subtracted from the reference magnitude spectrum to obtain M subtracted magnitude spectra. For example, the M magnitude spectra are |Y 1 |, |Y 2 | ··· |Y M |, and the reference magnitude spectrum is |Y R |. The M added magnitude spectra include: |Y R + Y 1 |, |Y R + Y 2 | ··· |Y R + Y M-1 |, |YR + YM|. The M subtracted magnitude spectra include: |Y R - Y 1 |, |YR - Y2| ··· |Y R - Y M-1 |, |Y R - Y M |]. After splicing the reference magnitude spectrum, the M magnitude spectra, the M added magnitude spectra, and the M subtracted magnitude spectra, the following spliced magnitude spectrum is obtained:

[0096] [|Y R |; |Y 1 |; |Y 2 |; ··· |Y M-1 |; |Y M |; |Y R + Y 1 |; |Y R - Y 1 |]; |Y R + Y 2 |; |Y R - Y 2 | ··· |Y R + Y M-1 |; |Y R - Y M-1 |; |Y R + Y M |; |Y R - Y M |].

[0097] In step S52, the features are input into the second neural network model to obtain the ideal ratio mask of the audio data to be processed.

[0098] In one implementation, features can be input into a second neural network model to obtain the ideal ratio mask (IRM) of the speech data to be processed. Among them, the IRM of the speech data to be processed includes the IRM of the speech data and the IRM of the noise data. Features can be input into the second neural network model to obtain the IRM of the speech data. According to 1 - IRM of the speech data = IRM of the noise data, the IRM of the speech data is obtained. The second neural network model includes a convolutional layer, a deconvolutional layer, and a gated recurrent unit (GRU). The second neural network model can be, for example, a Convolutional Recurrent Network (CRN). Figure 6 It is a schematic diagram of the second neural network model shown according to an exemplary embodiment. As Figure 6 shown, the second neural network model can include an encoder composed of 6 convolutional layers, a decoder composed of 6 deconvolutional layers, and two layers of GRU. The number of convolutional layers and GRUs can be set according to the actual situation. Among them, the convolutional layer and the corresponding deconvolutional layer are connected by a skip connection. In this implementation, windowing processing can be performed on the audio data to be processed. Among them, windowing processing refers to truncating the signal using different truncation functions. The truncation function is called a window function, simply referred to as a window. For example, the window function can be a Hanning window function. The window length and window shift of the Hanning window function can be set according to the actual situation. The window length can be understood as the window size. The window shift is used to represent the smoothness of the window shape. For example, the window length can be set to 512 sampling points, and the window shift is 256 sampling points.

[0099] In some embodiments, the spliced magnitude spectrum can be input into the second neural network model to obtain the ideal ratio mask of the audio data to be processed.

[0100] In step S53, a first covariance matrix is obtained based on the ideal ratio mask and the audio data to be processed.

[0101] In one implementation, the IRM of the speech data can be multiplied by the speech data to be processed to obtain the speech data estimate S(t,f). Perform a conjugate transpose operation on S(t,f) to obtain S H (t,f). Multiply S(t,f) and S H (t,f) to obtain the first covariance matrix S(t,f)S H (t,f) of the speech data amplitude. The IRM of the noise data can be multiplied by the speech data to be processed to obtain the noise data estimate N(t,f). Perform a conjugate transpose operation on N(t,f) to obtain N H (t,f). Multiply N(t,f) and N H(t, f) is multiplied to obtain the first covariance matrix N(t, f)N of the noise data amplitude H (t, f).

[0102] The present disclosure extracts features so that the audio data to be processed can be processed to obtain the IRM. Operations such as addition and subtraction are performed on the amplitude spectra of each channel, which can increase the amount of features input to the second neural network model and enable the output result of the second neural network model to be more reliable. The features are input into the second neural network model to obtain the IRM, and the first covariance matrix can be determined through calculation, thereby realizing the reverberation removal processing of the audio data to be processed and improving the efficiency of the reverberation removal processing.

[0103] In some embodiments, Figure 7 is a flowchart of an ideal ratio mask determination method shown according to an exemplary embodiment, as Figure 7 shown, the audio processing method includes the following steps.

[0104] In step S61, M amplitude spectra are respectively added to the reference amplitude spectrum to obtain M added amplitude spectra, and M amplitude spectra are respectively subtracted from the reference amplitude spectrum to obtain M subtracted amplitude spectra.

[0105] In step S62, the M amplitude spectra, the M added amplitude spectra, and the M subtracted amplitude spectra are concatenated.

[0106] In step S63, the concatenated amplitude spectra are input into the second neural network model to obtain the ideal ratio mask of the audio data to be processed.

[0107] The implementation manners of steps S61 to S63 can refer to the embodiments of steps S51 to S52 above, and the present disclosure will not elaborate here.

[0108] In some embodiments, Figure 8 is a flowchart of an audio processing method shown according to an exemplary embodiment, as Figure 8 shown, the audio processing method includes the following steps.

[0109] In step S71, feature extraction is performed on the multi-channel audio data.

[0110] In one embodiment, the multi-channel audio data can be the audio data to be processed.

[0111] In step S72, the features are input into the reverberation removal network model to obtain the IRM of the multi-channel audio data.

[0112] In one embodiment, the reverberation removal network model can be the second neural network model.

[0113] In step S73, a third covariance matrix is obtained based on the IRM and the multi-channel audio data.

[0114] In step S74, the third covariance matrix is input into the fully neural network model to obtain the beamformer parameters.

[0115] In one implementation, the fully neural network model can be the first neural network model.

[0116] In step S75, based on the beamformer parameters and the magnitude spectrum of the multi-channel audio data, the dereverberated audio data is obtained.

[0117] In one implementation, the magnitude spectrum of the multi-channel audio data can be a complex magnitude spectrum. The dereverberated audio can be the audio data after dereverberation processing.

[0118] Steps S71 to S75 in this implementation can refer to the implementation manners in the above embodiments, and the present disclosure will not be elaborated herein.

[0119] The present disclosure determines the covariance matrix of the multi-channel audio data and inputs the covariance matrix into the fully neural network model, and can obtain the frame-level beamformer parameters. Using these parameters to perform dereverberation processing on the multi-channel audio data can optimize the audio dereverberation effect.

[0120] In some implementations, Figure 9 is a flowchart of an audio processing method shown according to an exemplary embodiment. As Figure 9 shown, the audio processing method includes the following steps.

[0121] In step S81, feature extraction is performed on the multi-channel audio data.

[0122] In one implementation, the multi-channel audio data can be the audio data to be processed.

[0123] In step S82, the features are input into the CRN to obtain the IRM of the multi-channel audio data.

[0124] In step S83, a first covariance matrix is obtained based on the IRM and the complex magnitude spectrum of the multi-channel audio data.

[0125] In step S84, layer normalization processing is performed on the first covariance matrix to obtain a second covariance matrix. The second covariance matrix includes the second covariance matrix of the speech data and the second covariance matrix of the noise data. The second covariance matrix of the speech data and the second covariance matrix of the noise data are concatenated to obtain a third covariance matrix.

[0126] In step S85, the third covariance matrix is input into the GRU to obtain the GRU output result.

[0127] In step S86, the GRU output result is input into the non-linear DNN to obtain the non-linear DNN output result.

[0128] In step S87, the non-linear DNN output result is input into the linear DNN to obtain the beamformer parameter W.

[0129] In step S88, the beamformer parameter is multiplied by the multi-channel audio data, and the inverse Fourier transform is performed to obtain the dereverberated audio data.

[0130] In one embodiment, the dereverberated audio can be the audio data after dereverberation processing.

[0131] Steps S81 to S88 in this embodiment can refer to the implementation manners in the above embodiments, and the present disclosure will not elaborate herein.

[0132] The present disclosure can obtain the frame-level beamformer parameter by determining the covariance matrix of the multi-channel audio data and inputting the covariance matrix into the fully neural network model, and use this parameter to perform dereverberation processing on the multi-channel audio data to optimize the audio dereverberation effect.

[0133] Based on the same concept, the embodiment of the present disclosure also provides an audio processing device.

[0134] It can be understood that, in order to implement the above functions, the audio processing device provided by the embodiment of the present disclosure includes the corresponding hardware structure and / or software module for executing each function. Combining the units and algorithm steps of the examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.

[0135] It should be noted that those skilled in the art can understand that the various embodiments / implementations involved in the embodiments of the present disclosure above can be used in combination with the foregoing embodiments, or can be used independently. Whether used independently or in combination with the foregoing embodiments, their implementation principles are similar. In the embodiments of the present disclosure, some embodiments are described in the implementation manner of being used together. Of course, those skilled in the art can understand that such an illustrative description is not a limitation on the embodiments of the present disclosure.

[0136] Figure 10 It is a block diagram of an audio processing device 100 shown according to an exemplary embodiment. Refer to Figure 10, the device 100 includes a determination unit 101 and a processing unit 102.

[0137] Among them, the determination unit 101 is used to determine the first covariance matrix of the amplitude of the audio data to be processed. The processing unit 102 is used to input the first covariance matrix into the first neural network model to obtain beamformer parameters, and the beamformer parameters are used to perform dereverberation processing on the audio data. The audio data to be processed is subjected to dereverberation processing based on the beamformer parameters to obtain the dereverberated audio data.

[0138] In one embodiment, the first neural network model is a neural network model with recursive iteration characteristics, and the input of the first neural network model includes the first covariance matrix of the amplitude of the Nth frame of the speech data to be processed and the output result of the (N - 1)th frame of the speech data to be processed, where N is a positive integer greater than 1.

[0139] In one embodiment, the processing unit 102 inputs the first covariance matrix into the first neural network model in the following manner to obtain beamformer parameters: perform layer normalization processing on the first covariance matrix to obtain a second covariance matrix. The second covariance matrix is sequentially passed through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain beamformer parameters.

[0140] In one embodiment, the audio data to be processed includes speech data and noise data, and the second covariance matrix includes the second covariance matrix of the amplitude of the speech data and the second covariance matrix of the amplitude of the noise data. The processing unit 102 makes the second covariance matrix sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer in the following manner to obtain beamformer parameters: splice the second covariance matrix of the amplitude of the speech data and the second covariance matrix of the amplitude of the noise data to obtain a third covariance matrix of the amplitude of the audio data to be processed. The third covariance matrix is sequentially passed through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain beamformer parameters.

[0141] In one embodiment, the processing unit 102 trains the first neural network model in the following manner: obtain first training data, where the first training data is multiple fixed-length speeches. Perform reverberation processing on the first training data to obtain second training data, where the second training data is multiple fixed-length reverberant speeches, and the reverberant speeches include speech and noise. Input the second training data into the initial neural network model to obtain an output result. Use the minimum mean square error between the output result and the first training data to train the initial neural network model to obtain the first neural network model.

[0142] In one embodiment, the determining unit 101 determines the first covariance matrix of the amplitude of the audio data to be processed in the following manner: extracting the features of the audio data to be processed, where the features include the amplitude spectrum of the audio data to be processed. Inputting the features into a second neural network model to obtain the ideal ratio mask of the audio data to be processed, where the second neural network model includes a convolutional layer, a deconvolutional layer, and a gated recurrent unit. Among them, the audio data to be processed includes speech data and noise data, and the ideal ratio mask includes the ideal ratio mask of the speech data and the ideal ratio mask of the noise data. The ideal ratio mask of the speech data represents the proportion of the speech data in the audio data to be processed, and the ideal ratio mask of the noise data represents the proportion of the noise data in the audio data to be processed. Obtaining the first covariance matrix based on the ideal ratio mask and the audio data to be processed.

[0143] In one embodiment, the audio data to be processed includes audio data of M microphone channels, where M is a positive integer, and the amplitude spectrum of the audio data to be processed includes M amplitude spectra. The determining unit inputs the features into the second neural network model in the following manner: adding the M amplitude spectra to a reference amplitude spectrum respectively to obtain M added amplitude spectra, and subtracting the M amplitude spectra from the reference amplitude spectrum respectively to obtain M subtracted amplitude spectra; splicing the M amplitude spectra, the M added amplitude spectra, and the M subtracted amplitude spectra; inputting the spliced amplitude spectra into the second neural network model.

[0144] Figure 11 It is a block diagram of an audio processing device 200 shown according to an exemplary embodiment.

[0145] As Figure 11 shown, the device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.

[0146] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone call, data communication, camera operation, and recording operation. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above method. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.

[0147] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any applications or methods operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk, or optical disk.

[0148] The power component 206 provides power for various components of the device 200. The power component 206 may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 200.

[0149] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen can be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of touch or swipe actions but also detect the duration and pressure associated with the touch or swipe operation. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera can be a fixed optical lens system or have focal length and optical zoom capabilities.

[0150] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.

[0151] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, which can be a keyboard, a click wheel, buttons, etc. These buttons can include but are not limited to: a home button, a volume button, a power button, and a lock button.

[0152] The sensor assembly 214 includes one or more sensors for providing a status assessment of various aspects of the device 200. For example, the sensor assembly 214 can detect the on / off state of the device 200, the relative positioning of components, such as components for the display and keypad of the device 200. The sensor assembly 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and a change in the temperature of the device 200. The sensor assembly 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor assembly 214 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0153] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on communication standards, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.

[0154] In an exemplary embodiment, the device 200 can be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.

[0155] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 204 including instructions, and the above instructions can be executed by a processor 220 of the device 200 to complete the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, and an optical data storage device, etc.

[0156] The present disclosure can determine the covariance matrix of the audio data to be processed, input the covariance matrix into a neural network model, obtain the beamformer parameters at the frame level, and use the parameters to perform dereverberation processing on the audio data to be processed, so as to optimize the audio dereverberation effect.

[0157] It is understood that in this disclosure, "a plurality of" means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. The singular forms of "a", "an", and "the" are also intended to include the plural forms, unless the context clearly indicates otherwise.

[0158] It can be further understood that terms such as "first", "second", etc. are used to describe various information, but this information should not be limited to these terms. These terms are only used to distinguish information of the same type from each other, and do not represent a specific order or importance. In fact, expressions such as "first", "second", etc. can be used interchangeably. For example, without departing from the scope of this disclosure, the first information can also be referred to as the second information, and similarly, the second information can also be referred to as the first information.

[0159] It can be further understood that although the operations are described in a specific order in the drawings in the embodiments of this disclosure, it should not be understood as requiring the operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In a specific environment, multitasking and parallel processing may be beneficial.

[0160] Those skilled in the art will readily conceive of other embodiments of this disclosure after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this disclosure, which follow the general principles of this disclosure and include known common knowledge or conventional technical means in the technical field not disclosed in this disclosure.

[0161] Regarding the device in the above embodiments, the specific manner in which each module performs operations has been described in detail in the embodiments related to the method, and will not be elaborated here.

Claims

1. An audio processing method, characterized in that, it includes: Determine the first covariance matrix of the amplitude of the audio data to be processed; Input the first covariance matrix into the first neural network model to obtain beamformer parameters, and the beamformer parameters are used to perform dereverberation processing on the audio data; Perform dereverberation processing on the audio data to be processed based on the beamformer parameters to obtain the audio data after dereverberation processing.

2. The method according to claim 1, characterized in that, The first neural network model is a neural network model with recursive iteration characteristics. The input of the first neural network model is the first covariance matrix of the amplitude of the Nth frame of speech data to be processed and the output result of the first neural network model corresponding to the (N - 1)th frame of speech data to be processed, where N is a positive integer greater than 1.

3. The method according to claim 1, characterized in that, The step of inputting the first covariance matrix into the first neural network model to obtain beamformer parameters includes: Perform layer normalization processing on the first covariance matrix to obtain the second covariance matrix of the amplitude of the audio data to be processed; Let the second covariance matrix sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

4. The method according to claim 3, characterized in that, The audio data to be processed includes speech data and noise data. The second covariance matrix includes the second covariance matrix of the amplitude of the speech data and the second covariance matrix of the amplitude of the noise data. The step of letting the second covariance matrix sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters includes: Concatenate the second covariance matrix of the amplitude of the speech data and the second covariance matrix of the amplitude of the noise data to obtain the third covariance matrix of the amplitude of the audio data to be processed; Let the third covariance matrix sequentially pass through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

5. The method according to any one of claims 1-4, characterized in that, The first neural network model is trained in the following manner: Obtain first training data, and the first training data is multiple voices with a fixed length; Perform reverberation processing on the first training data to obtain second training data, and the second training data is multiple reverberant voices with a fixed length, and the reverberant voices include speech and noise; Input the second training data into the initial neural network model to obtain an output result; Use the minimum mean square error between the output result and the first training data to train the initial neural network model to obtain the first neural network model.

6. The method according to claim 1, characterized in that, The step of determining the first covariance matrix of the amplitude of the audio data to be processed includes: Extract the features of the audio data to be processed, and the features include the amplitude spectrum of the audio data to be processed. Input the feature into a second neural network model to obtain an ideal ratio mask for the audio data to be processed, where the second neural network model includes a convolutional layer, a deconvolutional layer, and a gated recurrent unit; Among them, the audio data to be processed includes speech data and noise data, the ideal ratio mask includes an ideal ratio mask for speech data and an ideal ratio mask for noise data, the ideal ratio mask for speech data represents the proportion of speech data in the audio data to be processed, and the ideal ratio mask for noise data represents the proportion of noise data in the audio data to be processed; Obtain the first covariance matrix based on the ideal ratio mask and the audio data to be processed.

7. The method according to claim 6, wherein, the audio data to be processed includes audio data of M microphone channels, M is a positive integer, the amplitude spectrum of the audio data to be processed includes M amplitude spectra, and the inputting the feature into the second neural network model includes: Adding the M amplitude spectra to a reference amplitude spectrum respectively to obtain M added amplitude spectra, and subtracting the M amplitude spectra from the reference amplitude spectrum respectively to obtain M subtracted amplitude spectra; Concatenating the M amplitude spectra, the M added amplitude spectra, and the M subtracted amplitude spectra; Inputting the concatenated amplitude spectra into the second neural network model.

8. An audio processing device, wherein, it includes: a determination unit configured to determine a first covariance matrix of the amplitude of the audio data to be processed; a processing unit configured to input the first covariance matrix into a first neural network model to obtain beamformer parameters for performing dereverberation processing on audio data; and perform dereverberation processing on the audio data to be processed based on the beamformer parameters to obtain dereverberated audio data.

9. The device according to claim 8, wherein, the first neural network model is a neural network model with a recursive iteration characteristic, the input of the first neural network model includes the first covariance matrix of the amplitude of the Nth frame of the audio data to be processed and the output result of the first neural network model corresponding to the (N - 1)th frame of the audio data to be processed, and N is a positive integer greater than 1.

10. The device according to claim 8, wherein, the processing unit inputs the first covariance matrix into the first neural network model in the following manner to obtain beamformer parameters: Performing layer normalization processing on the first covariance matrix to obtain a second covariance matrix of the amplitude of the audio data to be processed; Making the second covariance matrix sequentially go through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

11. The device according to claim 10, wherein, The to-be-processed audio data includes speech data and noise data. The second covariance matrix includes the second covariance matrix of the speech data amplitude and the second covariance matrix of the noise data amplitude. The processing unit makes the second covariance matrix sequentially go through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer in the following manner to obtain the beamformer parameters: Concatenate the second covariance matrix of the speech data amplitude and the second covariance matrix of the noise data amplitude to obtain a third covariance matrix of the to-be-processed audio data amplitude; Make the third covariance matrix sequentially go through a recurrent neural network layer, a non-linear deep neural network layer, and a linear deep neural network layer to obtain the beamformer parameters.

12. The apparatus according to any one of claims 8-11, wherein, The processing unit trains to obtain a first neural network model in the following manner: Obtain first training data, which is multiple fixed-length speeches; Perform reverberation processing on the first training data to obtain second training data, which is multiple fixed-length reverberant speeches, and the reverberant speeches include speech and noise; Input the second training data into an initial neural network model to obtain an output result; Use the minimum mean square error between the output result and the first training data to train the initial neural network model to obtain a first neural network model.

13. The apparatus according to claim 8, wherein, The determining unit determines the first covariance matrix of the to-be-processed audio data amplitude in the following manner: Extract the features of the to-be-processed audio data, and the features include the amplitude spectrum of the to-be-processed audio data; Input the features into a second neural network model to obtain an ideal ratio mask of the to-be-processed audio data, and the second neural network model includes a convolutional layer, a deconvolutional layer, and a gated recurrent unit; Among them, the to-be-processed audio data includes speech data and noise data, the ideal ratio mask includes an ideal ratio mask of speech data and an ideal ratio mask of noise data, the ideal ratio mask of speech data represents the proportion of speech data in the to-be-processed audio data, and the ideal ratio mask of noise data represents the proportion of noise data in the to-be-processed audio data; Based on the ideal ratio mask and the to-be-processed audio data, obtain the first covariance matrix.

14. The apparatus according to claim 13, wherein, The to-be-processed audio data includes audio data of M microphone channels, M is a positive integer, the amplitude spectrum of the to-be-processed audio data includes M amplitude spectra, and the determining unit inputs the features into the second neural network model in the following manner: Respectively add M amplitude spectra to a reference amplitude spectrum to obtain M added amplitude spectra, and respectively subtract M amplitude spectra from the reference amplitude spectrum to obtain M subtracted amplitude spectra; Concatenate M amplitude spectra, M added amplitude spectra, and M subtracted amplitude spectra; Input the concatenated amplitude spectra into the second neural network model.

15. An electronic device, It is characterized in that including: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the method according to any one of claims 1-7.

16. A storage medium, It is characterized in that the storage medium stores instructions which, when executed by a processor, execute the method according to any one of claims 1-7.