Audio enhancement method and apparatus, electronic device, and readable storage medium
By utilizing the complex spectral characteristics of audio data through an audio enhancement method that employs encoders, frequency domain, and time domain models, the robustness and real-time deployment challenges of existing noise reduction schemes are addressed, resulting in high-quality audio signal enhancement.
Patent Information
- Application Number
- CN202210572281.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-23
- Publication Date
- 2026-01-16
- Estimated Expiration
- 2042-05-23
AI Technical Summary
Existing noise reduction solutions are not robust enough to environmental changes and are difficult to deploy in real time, making it impossible to effectively eliminate the impact of environmental noise on the target sound.
An audio enhancement method is adopted, which obtains the target complex spectral features of audio data, and uses an encoder, frequency domain model and time domain model to extract and fuse features to generate complex masking values to reduce noise and maintain the causality and real-time performance of the system.
It achieves good noise reduction in complex environments, while reducing the difficulty of real-time deployment and improving the quality of audio signals.
Smart Images

Figure CN114974292B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the field of artificial intelligence, and particularly relates to an audio enhancement method and device, electronic equipment and a readable storage medium. BACKGROUND
[0002] In the process of recording audio and audio communication through at least two microphones, the microphone will record environmental noise of non-target sound, and the existence of environmental noise will affect the recording effect of the target sound. Therefore, it is necessary to use a certain noise reduction algorithm to eliminate the environmental noise of the recorded noisy sound, so as to obtain high-quality target sound.
[0003] In the prior art, one scheme estimates the noise spectrum from the noisy speech through an unsupervised speech enhancement method, so as to perform noise suppression. Another scheme trains and mines effective information from a large amount of noisy-clean speech data pairs through a supervised algorithm, obtains a noise reduction network model, and performs noise elimination on the noisy sound through the noise reduction network model.
[0004] However, the algorithm of the above-mentioned first scheme needs to make assumptions about environmental noise according to experience, and there are experience parameters, so the robustness to environmental changes is not good enough, and it is difficult to obtain good noise reduction effect in a complex environment. The second scheme only focuses on the training target related to the amplitude spectrum, and ignores the influence of the phase, and the design of the algorithm is usually non-causal, which leads to great difficulty in real-time deployment. SUMMARY
[0005] The purpose of the embodiments of the present application is to provide an audio enhancement method, device, electronic equipment and readable storage medium, which can solve the problems of poor effect and great difficulty in implementation and deployment of the noise reduction scheme in the prior art.
[0006] In a first aspect, the embodiments of the present application provide an audio enhancement method, which comprises:
[0007] obtaining a target complex spectrum feature of initial audio data; the initial audio data comprises audio data of at least two audio channels;
[0008] inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature;
[0009] inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature;
[0010] inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature;
[0011] fusing the second feature and the third feature and inputting them into a decoder of the target model to obtain a complex masking value;
[0012] generate enhanced audio data according to the complex mask value.
[0013] In a second aspect, an embodiment of the present application provides a model training method, which comprises:
[0014] obtaining a target complex spectrum feature of sample audio data; the sample audio data comprises audio data of at least two audio channels;
[0015] inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature;
[0016] inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature;
[0017] inputting the second feature into a time domain model to perform time domain feature extraction on the second feature to obtain a third feature;
[0018] fusing the second feature and the third feature and inputting them into a decoder of the target model to obtain a complex mask value; the complex mask value is used to generate enhanced audio data;
[0019] optimizing the target model according to the complex mask value and a preset loss function to obtain a trained target model.
[0020] In a third aspect, an embodiment of the present application provides an audio enhancement device, which comprises:
[0021] an obtaining module, configured to obtain a target complex spectrum feature of initial audio data; the initial audio data comprises audio data of at least two audio channels;
[0022] a first feature module, configured to input the target complex spectrum feature into an encoder of a target model to obtain a first feature;
[0023] a second feature module, configured to input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature;
[0024] a third feature module, configured to input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature to obtain a third feature;
[0025] a mask value module, configured to fuse the second feature and the third feature and input them into a decoder of the target model to obtain a complex mask value;
[0026] an enhancement module, configured to generate enhanced audio data according to the complex mask value.
[0027] In a fourth aspect, an embodiment of the present application provides a model training apparatus, the apparatus comprising:
[0028] an acquisition module configured to acquire a target complex spectrum feature of sample audio data; the sample audio data comprising audio data of at least two audio channels;
[0029] a first feature module configured to input the target complex spectrum feature into an encoder of a target model to obtain a first feature;
[0030] a second feature module configured to input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature;
[0031] a third feature module configured to input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature to obtain a third feature;
[0032] a masking value module configured to fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex masking value; the complex masking value being used to generate enhanced audio data;
[0033] an optimization module configured to optimize the target model according to the complex masking value and a preset loss function to obtain a trained target model.
[0034] In a fifth aspect, an embodiment of the present application provides an electronic device, the electronic device comprising a processor and a memory, the memory storing programs or instructions executable on the processor, the programs or instructions being executed by the processor to implement the above-mentioned audio enhancement method or the steps of the above-mentioned model training method.
[0035] In a sixth aspect, an embodiment of the present application provides a readable storage medium, the readable storage medium storing programs or instructions, the programs or instructions being executed by a processor to implement the above-mentioned audio enhancement method or the steps of the above-mentioned model training method.
[0036] In a seventh aspect, an embodiment of the present application provides a chip, the chip comprising a processor and a communication interface, the communication interface being coupled to the processor, the processor being configured to run programs or instructions to implement the method according to the first aspect or the second aspect.
[0037] In an eighth aspect, an embodiment of the present application provides a computer program product, the program product being stored in a storage medium, the program product being executed by at least one processor to implement the method according to the first aspect or the second aspect.
[0038] In the embodiment of the present application, an audio enhancement method is provided, comprising: obtaining a target complex spectrum feature of initial audio data; the initial audio data comprises audio data of at least two audio channels; inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; fusing the second feature and the third feature and inputting the fused second feature and third feature into a decoder of the target model to obtain a complex masking value; and generating enhanced audio data according to the complex masking value. In the present application, multi-channel audio data is inputted into a target model, and the multi-channel audio data can be converted into multi-channel complex spectrum features, so that the structural features of the speech time domain and frequency domain can be processed separately and then integrated, the causality of the whole system is maintained while the phase influence of the audio data is considered, so that the enhanced audio data obtained finally not only has good enhancement effect, but also has low real-time deployment difficulty. BRIEF DESCRIPTION OF DRAWINGS
[0039] Figure 1 is a step flowchart of an audio enhancement method provided by an embodiment of the present application;
[0040] Figure 2 is a target model structure schematic diagram provided by an embodiment of the present application;
[0041] Figure 3 is a step flowchart of another audio enhancement method provided by an embodiment of the present application;
[0042] Figure 4 is another target model structure schematic diagram provided by an embodiment of the present application;
[0043] Figure 5 is a frequency domain-time domain model structure schematic diagram provided by an embodiment of the present application;
[0044] Figure 6 is an expansion convolution network model structure schematic diagram provided by an embodiment of the present application;
[0045] Figure 7 is a step flowchart of a model training method provided by an embodiment of the present application;
[0046] Figure 8 is an audio enhancement flowchart provided by an embodiment of the present application;
[0047] Figure 9 is a block diagram of an audio enhancement device provided by an embodiment of the present application;
[0048] Figure 10 is a block diagram of a model training device provided by an embodiment of the present application;
[0049] Figure 11 is an electronic device provided by an embodiment of the present application;
[0050] Figure 12 is a hardware structure schematic diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION
[0051] The technical solutions in the embodiments of the present application will be clearly described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some of the embodiments of the present application, but not all of them. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art belong to the scope of protection of the present application.
[0052] The terms "first", "second", and the like in the specification and claims of the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence. It should be understood that the data used in this way can be interchanged under appropriate circumstances, so that the embodiments of the present application can be implemented in an order other than those illustrated or described herein, and the objects distinguished by "first", "second", etc. are usually a category and do not limit the number of objects, for example, the first object can be one or more. In addition, "and / or" in the specification and claims indicates at least one of the connected objects, and the character " / ", generally indicates that the front and rear associated objects are in an "or" relationship.
[0053] The audio enhancement method provided by the embodiments of the present application will be described in detail below with reference to the drawings, specific embodiments and application scenarios.
[0054] Referring to Figure 1 , Figure 1 A step flowchart of an audio enhancement method provided by an embodiment of the present application is shown, as shown in Figure 1 , specifically comprising the following steps:
[0055] Step 101, obtaining a target complex frequency spectrum feature of initial audio data; the initial audio data includes audio data of at least two audio channels.
[0056] The initial audio data includes audio data of at least two audio channels, wherein the audio data of at least two audio channels can be obtained by synchronously recording the sound in the environment by at least two microphones at different positions. Wherein, the audio data of each audio channel can be composed of a plurality of continuous time audio frames.
[0057] In the embodiment of the present application, the complex spectrum feature extraction can be performed on each obtained initial audio data to obtain the target complex spectrum feature corresponding to each initial audio data. Since the complex spectrum feature is composed of real part features and imaginary part features, the target complex spectrum feature corresponding to each initial audio data contains one channel of imaginary part features and one channel of real part features.
[0058] In step 102, the target complex spectrum feature is input into the encoder of the target model to obtain a first feature.
[0059] Since the target complex spectrum feature of one channel is generated according to each initial audio data, the target complex spectrum features of at least two channels can be obtained according to the initial audio data of at least two channels.
[0060] After obtaining the target complex spectrum features of at least two channels, the target complex spectrum features can be input into the encoder of the pre-trained target model to obtain a first feature.
[0061] In the embodiment of the present application, the target model can include an encoder, a frequency domain model, a time domain model and a decoder, so the target model also includes an encoder, a frequency domain model, a time domain model and a decoder.
[0062] Reference Figure 2 , Figure 2 A target model structure diagram provided by the embodiment of the present application is shown, as shown in Figure 2 The target model can include an encoder 210, a frequency domain-time domain model 300 and a decoder 220, wherein the frequency domain-time domain model 300 includes a frequency domain model 310 and a time domain model 320.
[0063] The target complex spectrum features of at least two channels can be input into the encoder at the same time, the feature extraction of the target complex spectrum features is performed by the encoder, and the first feature extracted by the encoder is output. The feature dimension of the first feature can be different from the feature dimension of the target complex spectrum feature input into the encoder, so as to realize the dimension compression of the target complex spectrum feature. The channel number of the first feature can also be different from the target complex spectrum feature input into the encoder.
[0064] For example, the channel number of the two target complex spectrum features input into the encoder is 4 channels (which contains 2 channels of real part features and 2 channels of imaginary part features), the feature dimension is 201, and the frame number is 100 frames. The channel number of the first feature output by the encoder can be 128, the feature dimension can be 50, and the frame number can be 100 frames. The technical personnel of the present application can adjust the structure and parameters of the encoder to change the expansion or compression ratio of the channel number and the feature dimension of the encoder.
[0065] Step 103, inputting the first feature into a frequency domain model of the target model, performing frequency domain feature extraction on the first feature to obtain a second feature.
[0066] The audio data is composed of audio frames corresponding to a plurality of time points in succession, and each audio frame of each time point records signal amounts corresponding to a plurality of frequency points in a given frequency band, and each audio frame contains frequency domain information.
[0067] In the embodiment of the present application, the first feature output by the encoder can be input into the frequency domain model to analyze the differences between the signal amounts of each frequency point in the data frame of each time point in the first feature, and the feature extraction is performed in the frequency domain range to obtain the second feature.
[0068] Step 104, inputting the second feature into a time domain model to perform time domain feature extraction on the second feature to obtain a third feature.
[0069] The audio signal also contains time domain features, and the differences between the audio frames of adjacent time points in the audio signal can reflect the features of the audio signal in the time domain range. Each data frame in the second feature contains the features of the audio signal in the frequency domain range, and therefore, the differences between the consecutive data frames in the second feature can also reflect the features of the audio signal in the time domain range.
[0070] In the embodiment of the present application, the second feature data can be input into the time domain model, and the time domain model can analyze the differences between the plurality of data frames in the input second feature and extract the third feature of the audio signal in the time domain from the second feature.
[0071] Step 105, fusing the second feature and the third feature and inputting them into a decoder of the target model to obtain a plurality of mask values.
[0072] After obtaining the second feature and the third feature, the second feature and the third feature can be fused and input into the decoder. It should be noted that, in one embodiment, the second feature and the third feature can be fused to obtain a fusion result, and then the fusion result is input into the decoder; in another embodiment, the second feature and the third feature can be jointly input into the decoder.
[0073] Specifically, the above-mentioned fusion manner can include but is not limited to head-to-tail splicing; it can also include phase addition or multiplication, for example, corresponding data frames in the first feature and the second feature can be added or multiplied to obtain a fusion data frame; it can also include weighted addition or multiplication, for example, corresponding data frames in the first feature and the second feature can be added or multiplied with different weights. The skilled in the art can select the fusion manner according to actual needs, and the embodiment of the present application does not make specific limitation thereto.
[0074] In the embodiment of the present application, the decoder can correspond to the encoder, and is configured to restore the channel number of the fusion result of the second feature and the third feature to the same channel number as the target complex spectrum feature, restore the feature dimension of the fusion result of the second feature and the third feature to the same feature dimension as the target complex spectrum feature, and obtain the complex mask value, so as to mask the target complex spectrum feature by the complex mask value to generate the enhanced audio data.
[0075] In step 106, the enhanced audio data is generated according to the complex mask value.
[0076] After the complex mask value is obtained, the initial audio data can be processed by the complex mask value to mask the noise signal strength in the initial audio data, so that the noise signal in the initial audio data is reduced, and the purpose of enhancing the non-noise signal in the initial audio data is achieved.
[0077] It should be noted that, since the initial audio data includes audio data of at least two audio channels, in order to reduce the resource overhead of the system, the complex mask value can be used to process the audio data of one audio channel in the audio data of at least two audio channels, to obtain the enhanced audio data of one audio channel.
[0078] Since the complex mask value includes the real part mask value and the imaginary part mask value, the complex mask value can be fused with the complex spectrum feature of the audio signal, the real part feature and the imaginary part feature in the target complex spectrum feature are masked, the noise or other sound (such as environmental sound) in the audio signal is shielded, and the effect of enhancing other sound (such as human voice) other than the shielded sound is achieved.
[0079] In summary, the audio enhancement method provided in the embodiment of the present application includes: obtaining a target complex spectrum feature of initial audio data; the initial audio data includes audio data of at least two audio channels; inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; fusing the second feature and the third feature and inputting the fusion result into a decoder of the target model to obtain a complex mask value; and generating enhanced audio data according to the complex mask value. The target model can be used to convert the multi-channel audio data into multi-channel complex spectrum features, so that the structure features of the speech time domain and frequency domain are processed separately and then integrated, the phase influence of the audio data is considered, and the causality of the entire system is maintained, so that the enhanced audio data obtained finally not only has a good enhancement effect, but also has a low real-time deployment difficulty.
[0080] Referring toFigure 3 , Figure 3 A flow chart of steps of another audio enhancement method provided by an embodiment of the present application is shown in FIG. 3, which specifically includes the following steps: Figure 3
[0081] Step 201, obtaining target complex spectrum features of initial audio data; the initial audio data includes audio data of at least two audio channels.
[0082] Optionally, step 201 can further include:
[0083] Sub-step 2011, obtaining audio data of at least two audio channels.
[0084] The audio data of the at least two audio channels can be obtained by at least two microphones. Since the at least two microphones are placed at different positions, the distances between the at least two microphones and the sound source are different, and thus there is a certain phase difference between the sound spectrums of the at least two audio channels.
[0085] For example, two microphones A and B obtain speech sounds. Since the distances between the two microphones and the human body are different, the time for the human voice to be conducted to the different microphones is different. If the human voice needs 0.05 seconds to be conducted to microphone A and 0.04 seconds to be conducted to microphone B, the human voice waveform recorded by the two microphones will have a 0.01-second offset in the same time domain, resulting in a phase difference between the audio data of the two audio channels.
[0086] Sub-step 2012, performing Fourier transform on the audio data of each audio channel respectively to obtain initial complex spectrum features of each channel.
[0087] The audio data is composed of a plurality of continuous audio frames. For each audio frame in the audio data, Fourier transform can be performed on the audio frame of each audio channel to obtain the initial complex spectrum features of each audio channel, wherein each initial complex spectrum channel is composed of a real part feature of a channel and an imaginary part feature of the channel. That is, the audio frames of the audio data and the data frames of the complex spectrum features correspond to each other one by one, and each data frame of the complex spectrum features contains a real part feature frame and an imaginary part feature frame.
[0088] For example, there are two audio channels each containing 10 audio frames. Fourier transform can be performed on the audio data of each audio channel to obtain the initial complex spectrum features of the two channels, and each initial complex spectrum feature of each channel contains 10 data frames.
[0089] Further, short-time Fourier transform can be performed on the audio data of each audio channel respectively to obtain the initial complex spectrum features of each channel; the short-time Fourier transform is a transformation of the audio data by a preset window function.
[0090] Short-term Fourier transform (STFT) is a mathematical transform related to Fourier transform, which is used to determine the frequency and phase of the local region sinusoidal wave of the time-varying signal. In fact, a time window function is added before Fourier transform. The short-term Fourier transform uses a fixed window function, and the shape of the window function does not change after it is determined. The resolution of the short-term Fourier transform is also determined. If you want to change the resolution, you need to select a new window function.
[0091] Since the embodiment of the present application can be used for real-time enhancement of audio signals, a short-term Fourier transform with faster speed can be used to improve the speed of the output enhanced audio signal and reduce the output delay.
[0092] In sub-step 2013, the real part features of each complex spectral feature are spliced respectively, and the imaginary part features of each complex feature are spliced to obtain the target complex spectral feature of the multiple channels.
[0093] After performing Fourier transform on the audio data of each audio channel, the complex spectral feature of one channel can be obtained, and the real part feature of one channel and the imaginary part feature of one channel can be contained in the complex spectral feature of each channel. Therefore, the real part features in the complex spectral features of at least two channels can be spliced to obtain the real part features of at least two channels, and the imaginary part features in the complex spectral features of at least two channels can be spliced to obtain the imaginary part features of at least two channels, so as to obtain the target complex spectral feature of at least two channels.
[0094] In the embodiment of the present application, Fourier transform can be performed on each channel of audio data to obtain the double-channel complex spectral feature corresponding to each channel of audio data, which contains real part and imaginary part. Then, the complex spectral features corresponding to multiple channels of audio data are spliced to obtain the target complex spectral feature of multiple channels, so that more feature information can be obtained in the subsequent model during the audio enhancement process, which is helpful to improve the effect of audio enhancement.
[0095] In step 202, the target complex spectral feature is input into the encoder of the target model to obtain a first feature.
[0096] The encoder comprises N input convolutional layers, and the decoder comprises N output convolutional layers, each corresponding to one of the N input convolutional layers. The structures and parameters of the N input convolutional layers can be the same or different. Similarly, the structures and parameters of the N output convolutional layers can be the same or different. There is a one-to-one correspondence between the N input and N output convolutional layers. Specifically, the nth input convolutional layer in the encoder corresponds to the (N+1)th (N-1)th input convolutional layer in the decoder, where n is an integer in the interval [0, N]. That is, the first input convolutional layer in the encoder corresponds to the last input convolutional layer in the decoder, the last input convolutional layer in the encoder corresponds to the first input convolutional layer in the decoder, and so on.
[0097] Reference Figure 4 , Figure 4 This application provides another target model structure schematic diagram, as shown in the embodiment of the present application. Figure 4 As shown, the encoder 210 contains five input convolutional layers 211 to 215, and the decoder 220 contains five output convolutional layers 221 to 225. Input convolutional layer 211 corresponds to output convolutional layer 225, input convolutional layer 212 corresponds to output convolutional layer 224, input convolutional layer 213 corresponds to output convolutional layer 223, input convolutional layer 214 corresponds to output convolutional layer 222, and input convolutional layer 215 corresponds to output convolutional layer 221. It should be noted that the parameters (kernel size, convolutional dimension, number of output channels, stride, activation function, etc.) in each input and output convolutional layer can be flexibly adjusted by technicians according to actual needs, and the parameters in each input and output convolutional layer are not limited to... Figure 4 The exemplary parameters are shown below.
[0098] Optionally, step 202 may also include:
[0099] Sub-step 2021: For each output convolutional layer, combine the output features of the previous layer with the output features of the input convolutional layer corresponding to the output convolutional layer to obtain combined features.
[0100] In this embodiment of the application, the output of each input convolutional layer in the encoder is fed to the next input convolutional layer and the corresponding output convolutional layer, such as... Figure 4As shown, the input of the input convolutional layer 211 is the target complex spectrum feature, the input convolutional layer 211 outputs the result to the next layer input convolutional layer 212 and outputs the result to the corresponding output convolutional layer 225, so that the output convolutional layer 225 decodes the output feature of the output convolutional layer 224 according to the output result of the input convolutional layer 212. The input of the input convolutional layer 215 is the output of the input convolutional layer 214, and the input convolutional layer 215 outputs the result to the frequency-time domain model 300 and the corresponding output convolutional layer 221, so that the output convolutional layer 221 decodes the output feature of the frequency-time domain model 300 according to the output result of the input convolutional layer 215.
[0101] Correspondingly, the input of each layer output convolutional layer in the decoder includes the output feature of the previous layer network module and the output feature of the input convolutional layer corresponding to the output convolutional layer. In the embodiment of the present application, the output feature of the above-mentioned previous layer network module and the output feature of the above-mentioned input convolutional layer corresponding to the output convolutional layer can be input into the output convolutional layer respectively, or the output feature of the above-mentioned previous layer network module and the output feature of the above-mentioned input convolutional layer corresponding to the output convolutional layer can be combined to obtain a combined feature, and then the combined feature is input into the output convolutional layer.
[0102] In substep 2022, the combined feature is input into the output convolutional layer.
[0103] In the embodiment of the present application, the encoder and the decoder can be composed of corresponding multiple pairs of input convolutional layers and output convolutional layers. Through multiple layers of input convolutional layers, higher feature compression rate and higher compression efficiency can be achieved, which helps to improve the effect of audio enhancement and reduce processing delay.
[0104] In step 203, the frequency points of the feature of each audio frame in the target complex spectrum feature are calculated forwardly and circularly in the frequency dimension to obtain a second feature with frequency correlation.
[0105] In the embodiment of the present application, the frequency domain model can include at least one layer of long short-term memory network model (LSTM, Long Short-Term Memory), at least one layer of linear fully connected layer (Linear), and at least one layer of normalization layer (layer norm).
[0106] The long short-term memory network model is a specific form of recurrent neural network (RNN), and the recurrent neural network is a general term for a series of neural networks capable of processing sequence data. Generally, the recurrent neural network encounters great difficulties in processing long-term dependencies (nodes far apart in time series), because calculating the relationship between nodes far apart in time series will cause the problem of gradient disappearance or gradient explosion. In order to solve this problem, the long short-term memory network model can be used. The long short-term memory network model increases the input gate, the forget gate and the output gate, so that the weight of the self-loop is variable. In this way, under the condition that the model parameters are fixed, the integration scale at different times can be dynamically changed, thereby avoiding the problem of gradient disappearance or gradient explosion. The long short-term memory network model can perform forward calculation through the self-loop unit with the gate structure, and set the recursive calculation to be performed along the frequency dimension direction of the input feature, so that the data of each frequency point in the target complex spectrum feature can be recursively calculated, so that the frequency domain model can obtain the second feature with frequency correlation.
[0107] The input gate is used to store the data frame in the first feature input at each time into the Memory Cell. The switch of the input gate determines whether information will be input into the Memory Cell at this time. The forget gate is used to determine whether to forget the data frame in the Memory Cell at each time. The output gate is used to output the data frame in the Memory Cell at each time.
[0108] Further, the at least one long short-term memory network model can be replaced by at least one bidirectional long short-term memory network model (Bi LSTM). Through the bidirectional long short-term memory network model, forward calculation and backward calculation can be performed on the target complex spectrum feature, so as to realize information integration in two sequence directions, so as to fully consider the frequency domain information of the context.
[0109] Referring to Figure 5 , Figure 5 A frequency-time domain model structure diagram provided by an embodiment of the present application is shown. As shown in Figure 5As shown, the frequency-time domain model 300 can include a frequency domain model 310 and a time domain model 320. The frequency domain model 310 can include a bidirectional long short-term memory network model 311, a linear fully connected layer 312, and a normalization layer 313 connected in sequence. The first feature 330 of the target complex spectral feature output by the encoder can be input into the bidirectional long short-term memory network model 311 of the frequency domain model 310, and processed sequentially by the bidirectional long short-term memory network model 311, the linear fully connected layer 312, and the normalization layer 313 to obtain the second feature with frequency correlation output by the frequency domain model 310.
[0110] In this embodiment, a frequency domain model with a long short-term memory network model is used to extract frequency domain features from audio data. This not only avoids the problems of gradient vanishing or gradient dilation, but also enables information integration in two sequence directions. This fully considers the frequency domain information of the context, which helps to improve the speech enhancement effect.
[0111] Step 204: Input the second feature into the time domain model of the target model, and perform time domain feature extraction on the second feature to obtain the third feature.
[0112] In this embodiment, the temporal model may include a temporal recurrent network model and a dilated convolutional network model. The temporal recurrent network model is used to calculate the inter-frame feature correlation in the input data. It can memorize the feature information of data frames corresponding to each time period through various gate structures, and then selectively combine the feature information of the data frame corresponding to the current time period to update the feature information and hidden layer state information of the data frame corresponding to the current time period, thereby performing the aforementioned calculation of inter-frame feature correlation.
[0113] A temporal recursive network model may include at least one Long Short-Term Memory (LSTM) network model, at least one linear fully connected layer, and at least one layer normalization layer.
[0114] like Figure 5 As shown, the temporal model 320 may include a temporal recurrent network model 340 and a dilated convolutional network model 350 connected in sequence. The temporal recurrent network model 340 may include a long short-term memory network model 341, a linear fully connected layer 342, and a normalization layer 343 connected in sequence.
[0115] Sub-step 2041: Input the second feature into the temporal recurrent network model for inter-frame correlation processing to obtain inter-frame correlation sub-features with inter-frame correlation.
[0116] In this embodiment, the target complex spectral feature, the first feature, and the second feature are all composed of data frames corresponding to audio frames at different times. Therefore, the second feature output by the frequency domain module can be input into the time domain recursive network model to process the correlation between the second feature frames, thereby obtaining a correlation sub-feature that can reflect the correlation between the audio frames and the data frames of the second feature. It should be noted that there is a one-to-one correspondence between the data frames of the correlation sub-feature, the second feature, and the first feature, and the correlation sub-feature, the second feature, and the first feature have the same dimension.
[0117] Furthermore, in order to improve the accuracy of the inter-frame correlation sub-features output by the temporal recurrent network model, the corresponding data frames in the first and second features can be fused to obtain fused features. These fused features can then be used as input to the aforementioned temporal recurrent network model to determine the correlation sub-features.
[0118] Sub-step 2042: Input the second feature into the dilated convolutional network model for context information integration processing to obtain context integrated sub-features.
[0119] Dilated convolutional neural network models can increase the receptive field of convolution, enabling the analysis of audio information over a larger temporal range to extract its feature information, thus reducing the loss of information in the audio data during convolution. In a convolutional neural network, the receptive field is the size of the region mapped onto the input data by the pixels in the feature map output by each layer. The size of the receptive field is adjusted by changing the dilation rate in the dilated convolutional network model.
[0120] In this embodiment of the application, the context information refers to the feature information contained in multiple data frames corresponding to different times in the second feature. Since there is a correspondence between the data frames of the second feature and the audio frames, the context information can also reflect the feature information contained in multiple audio frames.
[0121] like Figure 5 As shown, the dilated convolutional network model 350 includes at least one dilated convolutional sub-network model 351, at least one linear fully connected layer 352 connected to the output of the at least one dilated convolutional sub-network model 351, and at least one normalized layer 353 connected to the output of the at least one linear fully connected layer 352.
[0122] Sub-step 2043: Combine the correlation sub-feature and the context sub-feature to obtain the third feature.
[0123] The dilated convolution network model comprises at least one dilated convolution subnetwork model; the dilated convolution subnetwork model comprises at least one first ordinary convolution layer with a nonlinear activation function, at least one first dilated convolution layer with linear activation, at least one second dilated convolution layer with an S-shaped curve activation function, at least one second ordinary convolution layer with a linear activation function, and a submodel activation function for the dilated convolution subnetwork model.
[0124] The at least one first ordinary convolution layer receives the input of the previous layer and outputs the fourth feature to the at least one first dilated convolution layer and the at least one second dilated convolution layer respectively.
[0125] The nonlinear activation function of the first ordinary convolution layer can be a PReLU activation function, and the first ordinary convolution layer can unify the channel dimension of the input (the third feature or the combination of the third feature and the second feature) to a fixed channel number through the nonlinear activation function. Specifically, the channel number of the input feature can be fixed by adjusting the convolution kernel size and the channel number of the nonlinear activation function, to obtain the fourth feature. For example, in the embodiment of the present application, the first ordinary convolution layer can perform one-dimensional convolution on the data frames in the input feature through a 1X1 convolution kernel, to output data frames with 128 channels.
[0126] The convolution kernel of the first dilated convolution layer can be 3X3, the dilation coefficient can be any integer greater than or equal to 1, the channel number is the same as that of the first ordinary convolution layer, and the activation function can be a linear activation function. The convolution kernel of the second dilated convolution layer can be 3X3, the dilation coefficient can be the same as that of the first dilated convolution layer, the channel number is the same as that of the first ordinary convolution layer, and the activation function can be an S-shaped curve activation function (Sigmoid function).
[0127] The first dilated convolution layer can increase 0 elements in the convolution kernel according to the set dilation coefficient, so that the convolution kernel with an original size of 3X3 can perform convolution calculation in a range of 5X5, 7X7 or larger, so that the convolution calculation is not limited to adjacent frames of data. The output of the previous layer (the first ordinary convolution layer) is one-dimensional convolution. In order to increase the perception field of view of the network, the traditional convolutional neural network usually increases the depth of the network or expands the size of the convolution kernel, but this often leads to gradient disappearance or excessive calculation, resulting in a decline in the overall network performance. The first dilated convolution layer in the embodiment of the present application can expand the perception field without changing the size of the convolution kernel, improve the perception field without causing the problem of gradient disappearance, and without significantly increasing the calculation amount.
[0128] The second dilated convolution layer adopts an S-shaped curve activation function as the activation function, so that the output result is normalized to 0-1, and thus the output of the second dilated convolution layer can be used as a weight to correct the output of the first dilated convolution layer.
[0129] The fifth feature output by the at least one first dilated convolution layer and the sixth feature output by the at least one second dilated convolution layer are combined and output to the at least one second normal convolution layer.
[0130] The first dilated convolution layer and the second dilated convolution layer are arranged in parallel and are connected to the first normal convolution layer, taking the fourth feature of the output of the first normal convolution layer as the input, the first dilated convolution layer outputting a fifth feature, and the second dilated convolution layer outputting a sixth feature, and then the fifth feature and the sixth feature are fused and input to the second normal convolution layer. The activation function of the second normal convolution layer can be a linear activation function, and the convolution kernel size and the number of channels can be the same as or different from those of the first normal convolution layer.
[0131] The seventh feature output by the at least one second normal convolution layer and the fourth feature are combined and output.
[0132] In the embodiments of the present application, the combination of the seventh feature and the fourth feature can also be processed by a nonlinear activation function to obtain the final output of the dilated convolution subnetwork model. Specifically, a submodel activation function can be set in the dilated convolution subnetwork model, the combination of the seventh feature and the fourth feature is input to the submodel activation function for processing to obtain the final output of the dilated convolution subnetwork model. The submodel activation function can adopt a nonlinear activation function.
[0133] In the case where the dilated convolution subnetwork model includes at least two, the at least two dilated convolution subnetwork models are combined and output.
[0134] In the embodiments of the present application, the dilated convolution network model can include at least two dilated convolution subnetwork models, and the model structures of the at least two dilated convolution subnetwork models can be the same, and the difference lies in that the dilated convolution layers (including the first dilated convolution layer and the second dilated convolution layer) in different dilated convolution subnetwork models adopt different dilation coefficients, so that different dilated convolution subnetworks can extract features from the input through different sizes of perception fields to obtain multiple feature extraction results extracted under different sizes of perception fields.
[0135] The at least two dilated convolution subnetwork models are arranged in parallel in the dilated convolution network model, the dilated convolution network model fuses the output results of the dilated convolution subnetwork models, and takes the fusion result as the final output result of the dilated convolution network model.
[0136] Referring to Figure 6 ,Figure 6 This illustration shows a schematic diagram of a dilated convolutional network model structure provided in an embodiment of this application, as shown below. Figure 6 As shown, the dilated convolutional network model can contain five dilated convolutional sub-network models 410, 420, 430, 440, and 450, which are set up in parallel. The dilated convolutional sub-network model 410 can contain a first ordinary convolutional layer 411 with a non-linear activation function, a first dilated convolutional layer 412 with linear activation, a second dilated convolutional layer 413 with an sigmoid activation function, a second ordinary convolutional layer 414 with a linear activation function, and a sub-model activation function 415 for the dilated convolutional sub-network model. The first ordinary convolutional layer 411 has a convolutional kernel size of 1x1 (Size-1), 128 output channels, and uses the PReLU activation function; the first dilated convolutional layer 412 ... The kernel size can be 3x3 (Size-3), the number of output channels can be 128, the activation function is a linear activation function, and the dilation coefficient P can be 1; the kernel size of the second dilated convolutional layer 413 can be 3x3 (Size-3), the number of output channels can be 128, the activation function is a sigmoid activation function, and the dilation coefficient P can be 1; the kernel size of the second ordinary convolutional layer 414 can be 1x1 (Size-1), the number of output channels can be 128, and the activation function is a linear activation function; the sub-model activation function 715 can use the PReLU activation function. The model structures of the dilated convolutional sub-network models 420, 430, 440, and 450 can be the same as the model structure of the dilated convolutional sub-network model 410. The expansion coefficient P of the first and second dilated convolutional layers in dilated convolutional sub-network model 420 can be 2; the expansion coefficient P of the first and second dilated convolutional layers in dilated convolutional sub-network model 430 can be 4; the expansion coefficient P of the first and second dilated convolutional layers in dilated convolutional sub-network model 440 can be 8; and the expansion coefficient P of the first and second dilated convolutional layers in dilated convolutional sub-network model 450 can be 16.
[0137] In the embodiments of the present application, the time domain model with the long short-term memory network model is used to extract the time domain features in the audio data, which can not only avoid the problems of gradient disappearance or gradient expansion, but also realize information integration in two sequence directions, so as to fully consider the frequency domain information of the context, so that the extracted time domain features have high accuracy, which helps to improve the speech enhancement effect. In addition, the recursive network and the dilated convolution are used in the time domain module, so as to expand the perception field of the model to the sequence information without causing the problem of gradient disappearance and without significantly increasing the calculation amount, and further improve the extraction quality and efficiency of the features in the audio data. Further, a plurality of dilated convolution sub-network models with different dilated coefficients are set in parallel to simultaneously extract time domain features with different perception field sizes and fuse the extraction results to obtain the final time domain feature extraction result, which further enhances the accuracy of the time domain feature extraction and improves the effect of the model on audio enhancement.
[0138] Step 205, fusing the second feature and the third feature, and inputting the decoder of the target model to obtain a complex mask value.
[0139] This step can refer to step 105, and the embodiments of the present application will not be described again.
[0140] It should be noted that, since the encoder can include N layers of input convolutional layers, and the decoder includes N layers of output convolutional layers corresponding to the N layers of input convolutional layers, in the embodiments of the present application, the output convolutional layers in the decoder can decode the output results of the dilated convolutional network model according to the output results of the corresponding input convolutional layers, so as to restore the low-resolution data obtained after the pre-network processing to the original size (the size of the complex spectral feature), and obtain the complex mask value.
[0141] Step 206, performing enhancement processing on the initial complex spectral feature corresponding to an audio channel according to the complex mask value to obtain an enhanced audio spectral feature.
[0142] Since the complex mask value includes a real part and an imaginary part, and the initial complex spectral feature of each audio channel is obtained by Fourier transform and also includes a real part and an imaginary part, the complex mask value can be multiplied by the initial complex spectral feature of an audio channel to mask the noise features, and an enhanced speech spectral feature is obtained.
[0143] Step 207, performing inverse Fourier transform on the enhanced audio spectral feature to obtain enhanced audio data.
[0144] Optionally, step 207 can further include:
[0145] Sub-step 2071, inverse Fourier transform is performed on the enhanced audio spectrum feature to obtain a time-domain waveform corresponding to each frame.
[0146] Sub-step 2072, the time-domain waveforms corresponding to each frame are synthesized by overlap-add method to obtain enhanced audio data of the entire initial audio data.
[0147] By the overlap-add method, the latter half of the time-domain waveform corresponding to the previous audio frame and the former half of the time-domain waveform corresponding to the next audio frame in the time-domain waveform corresponding to each audio frame are superimposed, and all adjacent audio frames are synthesized by the above method to obtain the enhanced speech data of the entire initial audio data.
[0148] It should be noted that the above steps 202 to 207 can be executed by the trained target model. Specifically, after obtaining the target complex spectrum feature of the initial audio data, the target complex spectrum feature is directly input into the target model to obtain the enhanced speech data output by the target model. Referring to Figure 8 , Figure 8 An audio enhancement flowchart provided by an embodiment of the present application is shown, as shown in Figure 8 The target complex spectrum feature is input into the target model 600 after the initial audio data is processed to obtain the target complex spectrum feature. The target complex spectrum feature is first obtained by the encoder 610 in the target model 600 to obtain the first feature. The first feature is obtained by the frequency-domain-time-domain model 620 in the target model 600 to obtain the fusion result of the second feature and the third feature. The fusion result is obtained by the decoder 630 in the target model 600. The decoder 630 obtains the complex mask value according to the output of the encoder 610 and the output of the frequency-domain-time-domain model 620. The complex mask value is used to enhance the initial complex spectrum feature corresponding to one audio channel to obtain the enhanced speech spectrum feature. The enhanced speech spectrum feature is inverse Fourier transformed to obtain the enhanced speech data.
[0149] In summary, another audio enhancement method provided by the embodiment of the present application includes: obtaining a target complex spectrum feature of initial audio data; the initial audio data includes audio data of at least two audio channels; inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; fusing the second feature and the third feature and inputting the fused second feature and third feature into a decoder of the target model to obtain a complex masking value; and generating enhanced audio data according to the complex masking value. By inputting the multi-channel audio data into the target model, the multi-channel audio data can be converted into multi-channel complex spectrum features, so that the structural features of the speech time domain and frequency domain can be processed separately and then integrated, the causality of the entire system is maintained while the phase influence of the audio data is considered, so that the enhanced audio data obtained finally not only has a good enhancement effect, but also has a low real-time deployment difficulty.
[0150] Referring to Figure 7 , Figure 7 A step flowchart of a model training method provided by the embodiment of the present application is shown as shown in FIG. 3, and specifically includes the following steps: Figure 7
[0151] Step 301: obtaining a target complex spectrum feature of sample audio data; the sample audio data includes audio data of at least two audio channels.
[0152] This step can be referred to step 201, and the embodiment of the present application will not be described here.
[0153] Step 302: inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature.
[0154] This step can be referred to step 201, and the embodiment of the present application will not be described here.
[0155] Step 303: inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature.
[0156] This step can be referred to step 202, and the embodiment of the present application will not be described here.
[0157] Step 304: inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature.
[0158] This step can be referred to step 204, and the embodiment of the present application will not be described here.
[0159] Step 305, fusing the second feature and the third feature, and inputting the decoder of the target model to obtain a complex mask value; the complex mask value is used to generate enhanced audio data.
[0160] This step can refer to step 205, and details are not repeated here.
[0161] Step 306, optimizing the target model according to the complex mask value and a preset loss function to obtain a trained target model.
[0162] In the embodiment of the present application, the sample audio data also corresponds to sample enhanced audio data. The sample enhanced audio data and the enhanced audio data determined according to the complex mask value can be input into the preset loss function to obtain a feature loss between the complex mask value corresponding to the sample audio data and the sample enhanced audio data corresponding to the sample audio data, and the parameters in one or more modules of the target model are adjusted according to the feature loss to obtain a trained audio enhancement model. The preset loss function is used to measure the difference between the enhanced audio data corresponding to the same sample audio data and the sample enhanced audio data.
[0163] Optionally, step 306 can further include:
[0164] Sub-step 3061, generating enhanced audio data according to the complex mask value.
[0165] This step can refer to steps 206 to 207, and details are not repeated here.
[0166] Sub-step 3062, calculating a multi-resolution spectral loss function according to the enhanced audio data and the corresponding noise-free audio data to obtain a spectral loss.
[0167] Multi-resolution spectral loss refers to spectral loss under multiple Fourier transform scales. In an implementation, an initial audio data can be transformed by multiple Fourier transform scales, and the target model is used to determine multiple enhanced audio data corresponding to the initial audio data under multiple Fourier transform scales. A multi-resolution spectral loss function is determined according to the multiple enhanced audio data corresponding to the initial audio data and the corresponding noise-free audio data, and finally the resolution spectral loss functions under different scales are superimposed as the final spectral loss.
[0168] In another embodiment, in order to avoid the efficiency reduction caused by inputting the same initial audio data into the target model for multiple times, the enhanced speech spectrum feature output by the target model can be inversely transformed by Fourier transforms with different analysis parameters (i.e., Fourier transform length, window length, and frame shift) to obtain multi-scale enhanced speech data, and then a multi-resolution spectrum loss function can be calculated according to the multi-scale enhanced speech data and the noise-free speech data, and finally the resolution spectrum loss functions at different scales can be superimposed as the final spectrum loss. For example, in the above different analysis parameters, the Fourier transform length is taken from [512, 1024, 2048], the window length is taken from [240, 600, 1200], and the frame shift is taken from [50, 120, 240].
[0169] In sub-step 3063, a preset speech perceptual quality evaluation loss function is calculated according to the enhanced audio data and the corresponding noise-free audio data to obtain a power loss.
[0170] The speech perceptual quality evaluation loss function is designed based on the perceptual evaluation of the speech quality (PESQ), so that the loss function considers the human auditory masking and threshold effect, and is obtained by calculating the power spectrum of the noise-free speech data and the power spectrum of the enhanced speech data. Specifically, after the two speech data to be compared are subjected to level adjustment, input filter filtering, time alignment and compensation, and auditory transformation, the parameters of the two speech data are extracted, the time-frequency characteristics are integrated, the PESQ score is obtained, and finally the score is mapped to the subjective mean opinion score (MOS). The higher the score, the better the speech quality.
[0171] In sub-step 3064, the target model is optimized according to the spectrum loss and the power loss.
[0172] It should be noted that in the embodiments of the present application, the target model can be optimized according to the spectrum loss and the power loss. Alternatively, only the spectrum loss or the power loss can be obtained, and the target model can be optimized by one of the spectrum loss and the power loss. In addition, other loss functions can be calculated by other loss functions of the enhanced speech data and the corresponding noise-free speech data, and the target model can be optimized according to the other loss functions, and the embodiments of the present application are not limited in detail.
[0173] In the embodiments of the present application, the target model can be optimized by the spectrum loss and the power loss at the same time, so that the finally obtained target model can not only provide good audio enhancement effect, but also make the enhanced audio sound more natural, and further improve the audio enhancement effect.
[0174] In the embodiments of the present application, after the target model is trained, the target model can be published for application. For example, the target model can be published in a smart robot of a mobile terminal, so as to improve the voice recognition of the smart robot.
[0175] For example, when a user uses a mobile terminal to record, the mobile terminal can include at least two microphones, and the mobile terminal can collect sample audio data through the at least two microphones during recording. Each microphone can correspond to an audio channel.
[0176] In summary, the model training method provided in the embodiments of the present application includes: obtaining a target complex spectrum feature of sample audio data; the sample audio data includes audio data of at least two audio channels; inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; fusing the second feature and the third feature and inputting the fused second feature and third feature into a decoder of the target model to obtain a complex mask value; the complex mask value is used to generate enhanced audio data; and the target model is optimized according to the complex mask value and a preset loss function to obtain a trained target model. The target model is trained by using multi-path sample audio data. Since the target model can convert multi-path audio data into multi-path complex spectrum features, the structure features of the voice time domain and the frequency domain can be separately processed and then integrated. The phase influence of the audio data is considered, and the causality of the entire system is maintained. The target model trained finally can not only achieve good enhancement effect, but also has low real-time deployment difficulty.
[0177] The audio enhancement method provided in the embodiments of the present application can be executed by an audio enhancement device. The audio enhancement method is executed by the audio enhancement device in the embodiments of the present application, and the audio enhancement device provided in the embodiments of the present application is described.
[0178] Reference Figure 9 , Figure 9 is a block diagram of an audio enhancement device provided in the embodiments of the present application, as shown in Figure 9 The audio enhancement device includes:
[0179] The acquisition module 801 is configured to obtain a target complex spectrum feature of initial audio data; the initial audio data includes audio data of at least two audio channels.
[0180] The first feature module 802 is configured to input the target complex spectrum feature into an encoder of a target model to obtain a first feature.
[0181] The second feature module 803 is configured to input the first feature into a frequency domain model of the target model, perform frequency domain feature extraction on the first feature, and obtain a second feature.
[0182] The third feature module 804 is configured to input the second feature into a time domain model of the target model, perform time domain feature extraction on the second feature, and obtain a third feature.
[0183] The masking value module 805 is configured to fuse the second feature and the third feature, and input the fused second feature and third feature into a decoder of the target model, and obtain a complex masking value.
[0184] The enhancement module 806 is configured to generate enhanced audio data according to the complex masking value.
[0185] Optionally, the obtaining module comprises:
[0186] The sample obtaining sub-module is configured to obtain audio data of at least two audio channels.
[0187] The Fourier transform sub-module is configured to perform Fourier transform on the audio data of each audio channel respectively, and obtain initial complex spectral features of the channels.
[0188] The splicing sub-module is configured to splice real part features of the complex spectral features respectively, and splice imaginary part features of the complex spectral features respectively, and obtain the target complex spectral features of the multiple channels.
[0189] Optionally, the number of the multiple channels is an integer multiple of the number of the at least two audio channels.
[0190] Optionally, the transform sub-module is further configured to perform short-time Fourier transform on the audio data of each audio channel respectively, and obtain the initial complex spectral features of the channels; the short-time Fourier transform is a transform on the audio data according to a preset window function.
[0191] Optionally, the encoder comprises N layers of input convolution layers, and the decoder comprises N layers of output convolution layers respectively corresponding to the N layers of input convolution layers.
[0192] The first feature module comprises:
[0193] The combined feature sub-module is configured to combine, for each layer of output convolution layer, output features of a previous layer and output features of an input convolution layer corresponding to the output convolution layer, and obtain combined features.
[0194] The combined input sub-module is configured to input the combined features into the output convolution layer.
[0195] Optionally, the frequency domain model comprises a recurrent neural network model, and the second feature module comprises:
[0196] a second feature submodule configured to perform forward recurrent calculation on frequency points of the feature of each audio frame in the target complex frequency spectrum feature in a frequency dimension, to obtain a second feature with frequency correlation.
[0197] Optionally, the frequency domain model comprises a time domain recurrent network model and an inflated convolution network model, and the third feature module comprises:
[0198] a correlation submodule configured to input the second feature into the time domain recurrent network model for inter-frame correlation processing, to obtain an inter-frame correlation sub-feature with inter-frame correlation.
[0199] a context feature submodule configured to input the second feature into the inflated convolution network model for context information integration processing, to obtain a context integration sub-feature.
[0200] a third feature submodule configured to combine the correlation sub-feature and the context sub-feature, to obtain the third feature.
[0201] Optionally, the inflated convolution network model comprises at least one inflated convolution sub-network model; the inflated convolution sub-network model comprises at least one first normal convolution layer with a nonlinear activation function, at least one first inflated convolution layer with linear activation, at least one second inflated convolution layer with an S-shaped curve activation function, at least one second normal convolution layer with a linear activation function, and a sub-model activation function for the inflated convolution sub-network model; wherein at least one first normal convolution layer receives an input of a previous layer and outputs a fourth feature to at least one first inflated convolution layer and at least one second inflated convolution layer respectively; a fifth feature output by at least one first inflated convolution layer and a sixth feature output by at least one second inflated convolution layer are combined and output to at least one second normal convolution layer; a seventh feature output by at least one second normal convolution layer and the fourth feature are combined and output.
[0202] Optionally, the third feature module is further configured to output a combination of at least two inflated convolution sub-network models when the inflated convolution sub-network model comprises at least two.
[0203] Optionally, the enhancement module comprises:
[0204] an enhancement processing submodule configured to perform enhancement processing on an initial complex frequency spectrum feature corresponding to an audio channel according to the complex mask value, to obtain an enhanced audio frequency spectrum feature.
[0205] The enhanced audio acquisition submodule is used to perform an inverse Fourier transform on the enhanced audio spectral features to obtain enhanced audio data.
[0206] Optionally, the enhanced audio acquisition submodule includes:
[0207] The waveform acquisition submodule is used to perform inverse Fourier transform on the enhanced speech spectral features frame by frame to obtain the time-domain waveform corresponding to each frame.
[0208] The waveform synthesis submodule is used to synthesize the time-domain waveforms corresponding to each frame by overlapping and adding them together to obtain the enhanced speech data of the entire initial audio data.
[0209] In summary, the audio enhancement device provided in this application includes: an acquisition module for acquiring target complex spectral features of initial audio data; the initial audio data includes audio data from at least two audio channels; a first feature module for inputting the target complex spectral features into the encoder of a target model to obtain a first feature; a second feature module for inputting the first feature into the frequency domain model of the target model and extracting frequency domain features from the first feature to obtain a second feature; a third feature module for inputting the second feature into the time domain model of the target model and extracting time domain features from the second feature to obtain a third feature; a masking value module for fusing the second feature and the third feature and inputting them into the decoder of the target model to obtain a complex masking value; and an enhancement module for generating enhanced audio data based on the complex masking value. This application, by inputting multiple audio data into the target model, can transform multiple audio data into multiple complex spectral features. Therefore, by separately processing and then uniformly integrating the structural features of the speech in the time and frequency domains, the causality of the entire system is maintained while considering the phase influence of the audio data. This results in enhanced audio data with good enhancement effects and lower real-time deployment difficulty for the overall audio enhancement method.
[0210] The model training method provided in this application can be executed by a model training device. This application uses a model training device to perform model training as an example to illustrate the model training device provided in this application.
[0211] Reference Figure 10 , Figure 10 This is a block diagram of a model training device provided in an embodiment of this application, such as... Figure 10 As shown, the model training device includes:
[0212] The acquisition module 901 is used to acquire the target complex spectral features of the sample audio data; the sample audio data includes audio data from at least two audio channels.
[0213] The first feature module 902 is configured to input the target complex spectrum feature into an encoder of a target model to obtain a first feature.
[0214] The second feature module 903 is configured to input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature.
[0215] The third feature module 904 is configured to input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature to obtain a third feature.
[0216] The masking value module 905 is configured to fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex masking value, and the complex masking value is used to generate enhanced audio data.
[0217] The optimization module 906 is configured to optimize the target model according to the complex masking value and a preset loss function to obtain a trained target model.
[0218] Optionally, the optimization module includes:
[0219] The enhancement sub-module is configured to generate enhanced speech data according to the complex masking value.
[0220] The spectrum loss sub-module is configured to calculate a multi-resolution spectrum loss function according to the enhanced audio data and corresponding noise-free audio data to obtain a spectrum loss.
[0221] The power loss sub-module is configured to calculate a preset audio perceptual quality evaluation loss function according to the enhanced audio data and corresponding noise-free audio data to obtain a power loss.
[0222] The joint optimization sub-module is configured to optimize the target model according to the spectrum loss and the power loss.
[0223] In summary, the model training device provided in the embodiment of the present application comprises: an acquisition module configured to acquire a target complex frequency spectrum feature of sample audio data; the sample audio data comprises audio data of at least two audio channels; a first feature module configured to input the target complex frequency spectrum feature into an encoder of a target model to obtain a first feature; a second feature module configured to input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; a third feature module configured to input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; a masking value module configured to fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex masking value; the complex masking value is used to generate enhanced audio data; and an optimization module configured to optimize the target model according to the complex masking value and a preset loss function to obtain a trained target model. The target model is trained by using multi-channel sample audio data. Since the target model can convert multi-channel audio data into multi-channel complex frequency spectrum features, the structure features of the speech time domain and frequency domain can be processed separately and then integrated. The causality of the entire system is maintained while the phase influence of the audio data is considered. The target model trained finally can not only achieve good enhancement effect, but also has low real-time deployment difficulty.
[0224] The model training device and the audio enhancement device in the embodiment of the present application can be an electronic device or a component in an electronic device, such as an integrated circuit or a chip. The electronic device can be a terminal or other devices except the terminal. For example, the electronic device can be a mobile phone, a tablet computer, a notebook computer, a palm computer, a vehicle-mounted electronic device, a mobile Internet device (MID), an augmented reality (AR) / virtual reality (VR) device, a robot, a wearable device, an ultra-mobile personal computer (UMPC), a netbook, or a personal digital assistant (PDA), etc. The electronic device can also be a server, a network attached storage (NAS), a personal computer (PC), a television (TV), a teller machine, or a self-service machine, etc. The embodiment of the present application is not limited in this regard.
[0225] The model training apparatus and the audio enhancement apparatus in the embodiments of the present application can be apparatuses with an operating system. The operating system can be an Android operating system, can be an ios operating system, and can also be other possible operating systems, and the embodiments of the present application do not make specific limitations.
[0226] The model training apparatus and the audio enhancement apparatus provided in the embodiments of the present application can realize Figures 1 to 7 The method embodiments realize various processes, and to avoid repetition, the various processes are not described herein again.
[0227] Optionally, as shown in Figure 11 The embodiments of the present application also provide an electronic device M00, which includes a processor M01 and a memory M02, and the memory M02 stores programs or instructions that can run on the processor M01. When the programs or instructions are executed by the processor M01, each step of the above-mentioned audio enhancement method or model training method embodiment is realized, and the same technical effect can be achieved. To avoid repetition, each step is not described herein again.
[0228] It should be noted that the electronic device in the embodiments of the present application includes the mobile electronic device and the non-mobile electronic device described above.
[0229] Figure 12 To implement the hardware structure of an electronic device in the embodiments of the present application.
[0230] The electronic device 100 includes but is not limited to the following components: a radio frequency unit 101, a network module 102, an audio output unit 103, an input unit 104, a sensor 105, a display unit 106, a user input unit 107, an interface unit 108, a memory 109, and a processor 110, etc.
[0231] Those skilled in the art can understand that the electronic device 100 can also include a power supply (such as a battery) that supplies power to each component. The power supply can be logically connected to the processor 110 through a power management system, so as to realize functions such as management of charging, discharging, and power consumption management through the power management system. Figure 7 The electronic device structure shown in the embodiments of the present application does not constitute a limitation on the electronic device, and the electronic device can include more or fewer components than those shown, or combine certain components, or different component arrangements, which are not described herein again.
[0232] The processor 110 is configured to obtain a target complex spectrum feature of initial audio data, the initial audio data including audio data of at least two audio channels; input the target complex spectrum feature into an encoder of a target model to obtain a first feature; input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex mask value; and generate enhanced audio data according to the complex mask value.
[0233] The processor 110 is further configured to obtain a target complex spectrum feature of sample audio data, the sample audio data including audio data of at least two audio channels; input the target complex spectrum feature into an encoder of a target model to obtain a first feature; input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; input the second feature into a time domain model to perform time domain feature extraction on the second feature and obtain a third feature; fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex mask value; use the complex mask value to generate enhanced audio data; and optimize the target model according to the complex mask value and a preset loss function to obtain a trained target model.
[0234] In summary, the audio enhancement method provided in the embodiments of the present application includes the following steps: obtaining a target complex spectrum feature of initial audio data, the initial audio data including audio data of at least two audio channels; inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature and obtain a third feature; fusing the second feature and the third feature and inputting the fused second feature and third feature into a decoder of the target model to obtain a complex mask value; and generating enhanced audio data according to the complex mask value. The multi-channel audio data can be converted into multi-channel complex spectrum features by inputting the multi-channel audio data into the target model, so that the structural features of the speech time domain and frequency domain can be processed separately and then integrated. The causality of the entire system is maintained while the phase influence of the audio data is considered, so that the enhanced audio data obtained finally not only has a good enhancement effect, but also has a low real-time deployment difficulty.
[0235] It should be understood that in the embodiments of the present application, the input unit 104 can include a graphics processing unit (GPU) 1041 and a microphone 1042. The graphics processing unit 1041 processes image data of a still picture or a video obtained by an image capture device (such as a camera) in a video capture mode or an image capture mode. The display unit 106 can include a display panel 1061, which can be configured in the form of a liquid crystal display, an organic light-emitting diode, etc. The user input unit 107 includes at least one of a touch panel 1071 and other input devices 1072. The touch panel 1071 is also called a touch screen. The touch panel 1071 can include two parts of a touch detection device and a touch controller. The other input devices 1072 can include, but are not limited to, a physical keyboard, function keys (such as volume control keys, on-off keys, etc.), a trackball, a mouse, a joystick, and the like, which will not be described here.
[0236] The memory 109 can be used to store software programs and various data. The memory 109 can mainly include a first storage area storing programs or instructions and a second storage area storing data, wherein the first storage area can store an operating system, application programs or instructions required by at least one function (such as a sound playing function, an image playing function, etc.), and the like. In addition, the memory 109 can include a volatile memory or a non-volatile memory, or the memory 109 can include both volatile and non-volatile memories. The non-volatile memory can be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically EPROM (EEPROM), or a flash memory. The volatile memory can be a random access memory (RAM), a static random access memory (SRAM), a dynamic random access memory (DRAM), a synchronous dynamic random access memory (SDRAM), a double data rate synchronous dynamic random access memory (DDR SDRAM), an enhanced synchronous dynamic random access memory (ESDRAM), a synch link DRAM (SLDRAM), and a direct memory bus random access memory (Direct Rambus RAM, DRRAM). The memory 109 in the embodiments of the present application includes but is not limited to these and any other suitable types of memory.
[0237] The processor 110 can include one or more processing units; optionally, the processor 110 integrates an application processor and a modem processor, wherein the application processor mainly processes operations related to an operating system, a user interface, and an application program, and the modem processor mainly processes a wireless communication signal, such as a baseband processor. It can be understood that the above-mentioned modem processor can also not be integrated into the processor 110.
[0238] The embodiment of the present application further provides a readable storage medium, and the readable storage medium stores a program or instructions, the program or instructions are executed by a processor to realize various processes of the above-mentioned audio enhancement method embodiment, and the same technical effects can be achieved, and details are not repeated here.
[0239] The processor is the processor in the electronic device in the above-mentioned embodiment. The readable storage medium includes a computer readable storage medium, such as a computer readable only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and the like.
[0240] The embodiment of the present application further provides a chip, and the chip includes a processor and a communication interface, the communication interface is coupled with the processor, and the processor is used to run a program or instructions to realize various processes of the above-mentioned audio enhancement method embodiment, and the same technical effects can be achieved, and details are not repeated here.
[0241] It should be understood that the chip mentioned in the embodiment of the present application can also be referred to as a system level chip, a system chip, a chip system, or a system on chip, and the like.
[0242] The embodiment of the present application provides a computer program product, and the program product is stored in a storage medium, and the program product is executed by at least one processor to realize various processes of the above-mentioned audio enhancement method embodiment, and the same technical effects can be achieved, and details are not repeated here.
[0243] It should be noted that, in the present document, the terms "comprises", "comprising", or any other variations thereof, are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can include other elements not expressly listed or inherent to such process, method, article, or apparatus. An element proceeded by "comprises... a" does not, without more constraints, exclude the presence of additional identical elements in the process, method, article, or apparatus that comprises the element. Furthermore, it is to be understood that the method and apparatus of the present application can be carried out by more than one process, method, article, or apparatus either simultaneously, concurrently, or with intervening action that are carried out at the same time, either in a simultaneous fashion or in a fashion that is interleaved in time. For example, the described methods can be performed in a different order from that described, and / or various steps can be combined or omitted, and / or additional steps can be added, without departing from the scope of the present application. Also, features described with respect to certain examples can be combined in other examples.
[0244] From the above description of the embodiments, it is apparent that the above-mentioned method can be realized by means of software and necessary universal hardware platform, of course, it can also be realized by hardware, but in many cases, the former is a better embodiment. Based on such understanding, the technical solution of the present application can be embodied in the form of computer software product, which is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes a plurality of instructions for making a terminal (which can be a mobile phone, computer, server, or network equipment, etc.) execute the method described in various embodiments of the present application.
[0245] The embodiments of the present application are described above in conjunction with the drawings, but the present application is not limited to the above-described specific embodiments, and the above-described specific embodiments are merely illustrative, rather than limiting, and those skilled in the art can make many forms under the inspiration of the present application without departing from the scope of the present application and the scope protected by the claims.
Claims
1. An audio enhancement method, characterized by, The method comprises: obtaining a target complex spectrum feature of initial audio data; the initial audio data comprises audio data of at least two audio channels; inputting the target complex spectrum feature into an encoder of a target model to obtain a first feature, wherein a feature dimension of the first feature is different from a feature dimension of the target complex spectrum feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature; inputting the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature to obtain a third feature; fusing the second feature and the third feature and inputting the fused second feature and third feature into a decoder of the target model to obtain a complex masking value; generating enhanced audio data according to the complex masking value.
2. The method of claim 1, wherein, The method comprises: obtaining audio data of at least two audio channels; performing Fourier transform on the audio data of each audio channel respectively to obtain initial complex spectrum features of the channels; splicing real part features of the complex spectrum features respectively and splicing imaginary part features of the complex spectrum features respectively to obtain the target complex spectrum features of the multiple channels; the number of the multiple channels is an integer multiple of the number of the at least two audio channels.
3. The method of claim 2, wherein, The method comprises: performing short-time Fourier transform on the audio data of each audio channel respectively to obtain initial complex spectrum features of the channels; the short-time Fourier transform is a transformation on the audio data according to a preset window function.
4. The method of claim 1, wherein, The encoder comprises N layers of input convolution layers, and the decoder comprises N layers of output convolution layers corresponding to the N layers of input convolution layers respectively; The method comprises: combining the output feature of the previous layer and the output feature of the input convolution layer corresponding to the output convolution layer of each layer to obtain a combined feature; inputting the combined feature into the output convolution layer.
5. The method of claim 1, wherein, The frequency domain model comprises a recurrent neural network model, and the method comprises: performing forward recurrent calculation on frequency points of the feature of each audio frame in the target complex spectrum feature along a frequency dimension to obtain a second feature having frequency correlation.
6. The method of claim 1, wherein, The frequency domain model comprises a time domain recurrent network model and an inflated convolution network model, and the method comprises: inputting the second feature into the time domain recurrent network model to perform inter-frame correlation processing to obtain an inter-frame correlation sub-feature having inter-frame correlation; inputting the second feature into the inflated convolution network model to perform context information integration processing to obtain a context integration sub-feature; combining the correlation sub-feature and the context integration sub-feature to obtain the third feature.
7. The method of claim 6, wherein, The dilated convolution network model comprises at least one dilated convolution sub-network model; the dilated convolution sub-network model comprises at least one first ordinary convolution layer with a nonlinear activation function, at least one first dilated convolution layer with linear activation, at least one second dilated convolution layer with an S-shaped curve activation function, at least one second ordinary convolution layer with a linear activation function, and a sub-model activation function for the dilated convolution sub-network model; At least one of the first ordinary convolution layers receives the input of the previous layer and outputs the fourth feature to at least one of the first dilated convolution layers and at least one of the second dilated convolution layers, respectively; The fifth feature output by at least one of the first dilated convolution layers and the sixth feature output by at least one of the second dilated convolution layers are combined and output to at least one of the second ordinary convolution layers; The seventh feature output by at least one of the second ordinary convolution layers and the fourth feature are combined and output.
8. The method of claim 7, wherein, In the case where the dilated convolution sub-network model comprises at least two, the at least two dilated convolution sub-network models are combined and output.
9. The method of claim 1, wherein, The generation of enhanced audio data according to the complex masking value comprises: enhancing the initial complex spectral feature corresponding to an audio channel according to the complex masking value to obtain an enhanced audio spectral feature; performing inverse Fourier transform on the enhanced audio spectral feature to obtain enhanced audio data.
10. The method of claim 9, wherein, The generation of enhanced audio data by performing inverse Fourier transform on the enhanced audio spectral feature comprises: performing inverse Fourier transform on the enhanced audio spectral feature by frame to obtain the time-domain waveform corresponding to each frame; synthesizing the time-domain waveform corresponding to each frame by overlap-add method to obtain the enhanced audio data of the entire initial audio data.
11. A model training method, comprising: The method comprises: obtaining a target complex spectral feature of sample audio data; the sample audio data comprises audio data of at least two audio channels; inputting the target complex spectral feature into an encoder of a target model to obtain a first feature, wherein the feature dimension of the first feature is different from that of the target complex spectral feature; inputting the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature and obtain a second feature; inputting the second feature into a time domain model to perform time domain feature extraction on the second feature and obtain a third feature; fusing the second feature and the third feature and inputting them into a decoder of the target model to obtain a complex masking value; the complex masking value is used to generate enhanced audio data; optimizing the target model according to the complex masking value and a preset loss function to obtain a trained target model.
12. The method of claim 11, wherein, The optimization of the target model according to the complex masking value and the preset loss function comprises: generating enhanced audio data according to the complex masking value; calculating a multi-resolution spectral loss function according to the enhanced audio data and corresponding noise-free audio data to obtain a spectral loss; calculating a preset audio perceptual quality evaluation loss function according to the enhanced audio data and corresponding noise-free audio data to obtain a power loss; According to the spectral loss and the power loss, the target model is optimized.
13. An audio enhancement device, characterized by The device comprises: An acquisition module is configured to acquire a target complex spectral feature of initial audio data; the initial audio data comprises audio data of at least two audio channels; A first feature module is configured to input the target complex spectral feature into an encoder of a target model to obtain a first feature, wherein a feature dimension of the first feature is different from a feature dimension of the target complex spectral feature; A second feature module is configured to input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature; A third feature module is configured to input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature to obtain a third feature; A masking value module is configured to fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex masking value; An enhancement module is configured to generate enhanced audio data according to the complex masking value.
14. The apparatus of claim 13, wherein, The acquisition module comprises: A sample acquisition submodule is configured to acquire audio data of at least two audio channels; A Fourier transform submodule is configured to perform Fourier transform on audio data of each audio channel respectively to obtain initial complex spectral features of the channels; A splicing submodule is configured to splice real part features of the complex spectral features respectively and to splice imaginary part features of the complex spectral features respectively to obtain the target complex spectral feature of the multiple channels.
15. The apparatus of claim 13, wherein, The encoder comprises N layers of input convolution layers, and the decoder comprises N layers of output convolution layers corresponding to the N layers of input convolution layers respectively; the first feature module comprises: A combined feature submodule is configured to combine output features of a previous layer and output features of an input convolution layer corresponding to an output convolution layer of each layer to obtain combined features; A combined input submodule is configured to input the combined features into the output convolution layer.
16. The apparatus of claim 13, wherein, The frequency domain model comprises a time domain recurrent network model and an inflated convolution network model, and the third feature module comprises: A correlation submodule is configured to input the second feature into the time domain recurrent network model to perform inter-frame correlation processing to obtain inter-frame correlation sub-features having inter-frame correlation; A context feature submodule is configured to input the second feature into the inflated convolution network model to perform context information integration processing to obtain context integration sub-features; A third feature submodule is configured to combine the correlation sub-features and the context integration sub-features to obtain the third features.
17. The apparatus of claim 13, wherein, The enhancement module comprises: An enhancement processing submodule is configured to perform enhancement processing on initial complex spectral features corresponding to an audio channel according to the complex masking value to obtain enhanced audio spectral features; An enhanced audio acquisition submodule is configured to perform inverse Fourier transform on the enhanced audio spectral features to obtain enhanced audio data.
18. A model training apparatus, comprising: The device comprises: An acquisition module is configured to acquire a target complex spectral feature of sample audio data; the sample audio data comprises audio data of at least two audio channels; The first feature module is configured to input the target complex spectrum feature into an encoder of a target model to obtain a first feature, wherein a feature dimension of the first feature is different from a feature dimension of the target complex spectrum feature. The second feature module is configured to input the first feature into a frequency domain model of the target model to perform frequency domain feature extraction on the first feature to obtain a second feature. The third feature module is configured to input the second feature into a time domain model of the target model to perform time domain feature extraction on the second feature to obtain a third feature. The masking value module is configured to fuse the second feature and the third feature and input the fused second feature and third feature into a decoder of the target model to obtain a complex masking value, wherein the complex masking value is used to generate enhanced audio data. The optimization module is configured to optimize the target model according to the complex masking value and a preset loss function to obtain a trained target model.
19. The apparatus of claim 18, wherein, The optimization module includes: The enhancement sub-module is configured to generate enhanced speech data according to the complex masking value. The spectrum loss sub-module is configured to calculate a multi-resolution spectrum loss function according to the enhanced audio data and corresponding noise-free audio data to obtain a spectrum loss. The power loss sub-module is configured to calculate a preset audio perceptual quality evaluation loss function according to the enhanced audio data and the corresponding noise-free audio data to obtain a power loss. The joint optimization sub-module is configured to optimize the target model according to the spectrum loss and the power loss.
20. An electronic device, comprising: The electronic device includes a processor and a memory, wherein the memory stores programs or instructions executable on the processor, and the programs or instructions are executed by the processor to implement the audio enhancement method in any one of claims 1 to 10 or the steps of the model training method in any one of claims 11 to 12.
21. A readable storage medium, characterized by, The readable storage medium stores programs or instructions, and the programs or instructions are executed by the processor to implement the audio enhancement method in any one of claims 1 to 10 or the steps of the model training method in any one of claims 11 to 12.
Citation Information
Patent Citations
Feature extraction method and device based on voice signal time domain and frequency domain, and echo cancellation method and device
CN113870888A