Audio noise reduction method, audio noise reduction device and storage medium
By combining the mixed noise reduction technology of time domain waveform and frequency domain spectrum, combined with scene recognition and self-attention mechanism, the sound distortion and real-time problems in the existing audio noise reduction technology are solved, and high-quality and real-time audio noise reduction on mobile devices is achieved.
Patent Information
- Application Number
- CN202410138201.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-31
- Publication Date
- 2025-08-01
AI Technical Summary
The existing audio noise reduction technology is prone to losing noise waveform information during frequency domain feature processing, resulting in sound distortion and unable to achieve real-time noise reduction. Especially at high noise levels, audio quality decreases, high computing requirements, large power consumption, and unstable effect in adapting to different noise environments.
The hybrid noise reduction technology based on time domain waveform and frequency domain spectrum is adopted, combined with the time domain noise reduction network and the frequency domain noise reduction network, feature fusion is performed through scene recognition and self-attention mechanism, and the noise reduction model parameters are adjusted in real time to adapt to different noise scenarios, and unknown scene parameters are optimized through the cloud.
It realizes high-quality and real-time audio noise reduction on mobile devices, ensuring that the target audio is not distorted, reducing the calculation amount, adapting to different noise environments, and improving noise reduction effect and device battery life.
Smart Images

Figure CN120412618A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of audio processing, and in particular, to an audio noise reduction method, an audio noise reduction device, and a storage medium. Background Art
[0002] Audio recorded in real-world scenarios may contain various background noises. For example, audio and video conferencing, automatic speech recognition, and hearing aids. The presence of background noise affects the accurate transmission of target audio information. To solve the above problems, a speech noise reduction technology is needed to eliminate such noises and output a perceptually high-quality speech signal.
[0003] In related technologies, there are technical means for noise reduction based on the frequency-domain characteristics of audio, that is, obtaining the spectrogram of the audio to be processed, using a noise reduction network to process the spectrogram, and converting the processed spectrogram into audio to complete noise reduction. However, this means may cause the loss of target audio and result in distorted audio after noise reduction. Summary of the Invention
[0004] To overcome the problems existing in the related technologies, the present disclosure provides an audio noise reduction method, an audio noise reduction device, and a storage medium.
[0005] According to a first aspect of an embodiment of the present disclosure, an audio noise reduction method is provided, including: obtaining in real time the audio to be processed, and determining the current noise scene according to the audio to be processed; determining the current noise reduction parameter corresponding to the current noise scene according to the correspondence between the noise scene and the noise reduction parameter; adjusting a preset noise reduction model according to the current noise reduction parameter; obtaining the initial audio features of the audio to be processed, using the adjusted preset noise reduction model to perform noise reduction processing on the initial audio features to obtain target audio features, and obtaining the target audio corresponding to the target audio features, where the target audio features are the initial audio features after noise reduction processing, and the initial audio features include the initial frequency-domain features and initial time-domain features of the audio to be processed.
[0006] In one implementation, the determining the current noise scene according to the audio to be processed includes: determining the current noise scene according to the audio to be processed and preset information, where the preset information includes one or more of the following information: application running information of the terminal within a preset time, low-power component running information of the terminal within a preset time, and current location information of the terminal.
[0007] In one implementation, the preset noise reduction model includes a time-domain noise reduction network and a frequency-domain noise reduction network; using the adjusted preset noise reduction model to perform noise reduction processing on the initial audio features to obtain target audio features, including: performing time-domain convolution processing on the initial time-domain features through the time-domain noise reduction network to obtain intermediate time-domain features, and performing frequency-domain convolution processing on the initial frequency-domain features through the frequency-domain noise reduction network to obtain intermediate frequency-domain features; processing the intermediate time-domain features and the intermediate frequency-domain features according to the time-domain noise reduction network and the frequency-domain noise reduction network to obtain the target audio features.
[0008] In one implementation, the processing the intermediate time-domain features and the intermediate frequency-domain features according to the time-domain convolution network and the frequency-domain convolution network to obtain the target audio features includes: processing the intermediate time-domain features based on the self-attention mechanism, and fusing the processed intermediate time-domain features and the intermediate frequency-domain features to obtain an intermediate hybrid feature; performing recursive processing, transposed convolution processing, and predictive modulation mask processing on the intermediate hybrid feature successively through the frequency-domain noise reduction network to obtain target frequency-domain features; performing recursive processing on the target frequency-domain features and the intermediate time-domain features through the time-domain noise reduction network to obtain the target audio features.
[0009] In one implementation, the noise reduction parameters include the model shape parameters of the preset noise reduction model obtained by training under different noise scenarios; the preset noise reduction model is trained in the following manner: simulating different noise scenarios based on a preset audio dataset, and under different noise scenarios, processing the training audio corresponding to the noise scenario through a preset neural network to obtain predicted target frequency-domain features and predicted target time-domain features, the audio dataset includes a noise dataset and a clean audio dataset, and the preset neural network has the same algorithm structure as the preset noise reduction model; obtaining the loss function of the preset neural network according to the predicted target frequency-domain features, the predicted target time-domain features, the time-domain features of the training audio, and the frequency-domain features of the training audio; in response to the loss function satisfying the preset numerical requirements, determining the preset neural network as the preset noise reduction model.
[0010] In one implementation, obtaining the loss function of the preset neural network according to the predicted target frequency-domain feature, the predicted target time-domain feature, the time-domain feature of the training audio, and the frequency-domain feature of the training audio includes: performing an inverse Fourier transform on the predicted target frequency-domain feature to obtain a converted time-domain feature corresponding to the predicted target frequency-domain feature; obtaining a time-domain loss between the predicted target time-domain feature and the time-domain feature of the training audio, obtaining a frequency-domain loss between the converted time-domain feature and the frequency-domain feature of the training audio, and obtaining a time-frequency domain hybrid loss between the predicted target time-domain feature and the converted time-domain feature; determining the loss function according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, a first weight, a second weight, and a third weight, where the first weight corresponds to the time-domain loss, the second weight corresponds to the frequency-domain loss, and the third weight corresponds to the time-frequency domain hybrid loss.
[0011] In one implementation, determining the loss function according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight includes: determining a first product as the product of the time-domain loss and the first weight, determining a second product as the product of the frequency-domain loss and the second weight, and determining a third product as the product of the time-frequency domain hybrid loss and the third weight; determining the sum of the first product, the second product, and the third product as the loss function.
[0012] In one implementation, the method further includes: in response to being unable to determine the current noise scenario based on the audio to be processed, adjusting a preset noise reduction model according to default noise reduction parameters, and uploading the audio data to be processed and preset data collected by the current terminal to the cloud.
[0013] In one implementation, the method further includes: obtaining corresponding data between the noise reduction parameters returned by the cloud and the noise scenario, and updating the corresponding relationship between the locally stored noise scenario and the noise reduction parameters according to the corresponding data.
[0014] According to a second aspect of the embodiments of the present disclosure, there is provided an audio noise reduction device, including: a determination unit, configured to obtain the audio to be processed in real time, determine the current noise scenario according to the audio to be processed, and determine the current noise reduction parameter corresponding to the current noise scenario according to the correspondence between the noise scenario and the noise reduction parameter; an adjustment unit, configured to adjust a preset noise reduction model according to the current noise reduction parameter; and a processing unit, configured to obtain the initial audio feature of the audio to be processed, perform noise reduction processing on the initial audio feature by using the adjusted preset noise reduction model to obtain a target audio feature, and obtain a target audio corresponding to the target audio feature, where the target audio feature is the initial audio feature after noise reduction processing, and the initial audio feature includes the initial frequency domain feature and the initial time domain feature of the audio to be processed.
[0015] In an implementation manner, the determination unit determines the current noise scenario according to the audio to be processed in the following manner: determining the current noise scenario according to the audio to be processed and preset information, where the preset information includes one or more of the following information: the application running information of the terminal within a preset time, the low-power component running information of the terminal within a preset time, and the current location information of the terminal.
[0016] In an implementation manner, the preset noise reduction model includes a time domain noise reduction network and a frequency domain noise reduction network; the processing unit performs noise reduction processing on the initial audio feature by using the adjusted preset noise reduction model to obtain a target audio feature in the following manner: performing time domain convolution processing on the initial time domain feature through the time domain noise reduction network to obtain an intermediate time domain feature, and performing frequency domain convolution processing on the initial frequency domain feature through the frequency domain noise reduction network to obtain an intermediate frequency domain feature; and processing the intermediate time domain feature and the intermediate frequency domain feature according to the time domain noise reduction network and the frequency domain noise reduction network to obtain the target audio feature.
[0017] In an implementation manner, the processing unit processes the intermediate time domain feature and the intermediate frequency domain feature according to the time domain convolution network and the frequency domain convolution network to obtain the target audio feature in the following manner: processing the intermediate time domain feature based on a self-attention mechanism, and fusing the processed intermediate time domain feature and the intermediate frequency domain feature to obtain an intermediate hybrid feature; performing recursive processing, deconvolution processing, and predictive modulation mask processing on the intermediate hybrid feature successively through the frequency domain noise reduction network to obtain a target frequency domain feature; and performing recursive processing on the target frequency domain feature and the intermediate time domain feature through the time domain noise reduction network to obtain the target audio feature.
[0018] In one implementation, the noise reduction parameters include the model form parameters of a preset noise reduction model obtained through training under different noise scenarios; the processing unit trains the preset noise reduction model in the following manner: simulating different noise scenarios based on a preset audio data set, and under different noise scenarios, processing the training audio corresponding to the noise scenario through a preset neural network to obtain the predicted target frequency domain features and the predicted target time domain features. The audio data set includes a noise data set and a clean audio data set, and the preset neural network has the same algorithm structure as the preset noise reduction model; obtaining the loss function of the preset neural network according to the predicted target frequency domain features, the predicted target time domain features, the time domain features of the training audio, and the frequency domain features of the training audio; and in response to the loss function meeting the preset numerical requirements, determining the preset neural network as the preset noise reduction model.
[0019] In one implementation, the processing unit obtains the loss function of the preset neural network according to the predicted target frequency domain features, the predicted target time domain features, the time domain features of the training audio, and the frequency domain features of the training audio in the following manner: performing an inverse Fourier transform process on the predicted target frequency domain features to obtain the converted time domain features corresponding to the predicted target frequency domain features; obtaining the time domain loss between the predicted target time domain features and the time domain features of the training audio, obtaining the frequency domain loss between the converted time domain features and the frequency domain features of the training audio, and obtaining the time-frequency domain hybrid loss between the predicted target time domain features and the converted time domain features; and determining the loss function according to the time domain loss, the frequency domain loss, the time-frequency domain hybrid loss, a first weight, a second weight, and a third weight, where the first weight corresponds to the time domain loss, the second weight corresponds to the frequency domain loss, and the third weight corresponds to the time-frequency domain hybrid loss.
[0020] In one implementation, the processing unit determines the loss function according to the time domain loss, the frequency domain loss, the time-frequency domain hybrid loss, a first weight, a second weight, and the third weight in the following manner: determining the product of the time domain loss and the first weight as a first product, determining the product of the frequency domain loss and the second weight as a second product, and determining the product of the time-frequency domain hybrid loss and the third weight as a third product; and determining the sum of the first product, the second product, and the third product as the loss function.
[0021] In one implementation, the processing unit is further configured to: in response to being unable to determine the current noise scenario based on the audio to be processed, adjusting the preset noise reduction model according to the default noise reduction parameters, and uploading the audio data to be processed and the preset data collected by the current terminal to the cloud.
[0022] In one implementation, the processing unit is further configured to: obtain the corresponding data between the noise reduction parameters and the noise scenarios returned by the cloud, and update the corresponding relationship between the locally stored noise scenarios and the noise reduction parameters according to the corresponding data.
[0023] According to a third aspect of the embodiments of the present disclosure, there is provided an audio noise reduction device, including: a processor; a memory for storing processor-executable instructions; wherein the processor is configured to: execute the audio noise reduction method described in the first aspect or any implementation manner of the first aspect.
[0024] According to a fourth aspect of the embodiments of the present disclosure, there is provided a storage medium storing instructions that, when executed by a processor, enable the processor to execute the audio noise reduction method described in the first aspect or any implementation manner of the first aspect.
[0025] The technical solutions provided by the embodiments of the present disclosure may include the following beneficial effects: When performing audio noise reduction, the audio to be processed is obtained in real time, the current noise scenario is determined according to the audio to be processed, and then the noise reduction parameters corresponding to the current noise scenario are determined, and the preset noise reduction model is adjusted according to the current noise reduction parameters. The adjusted preset noise reduction model is used to perform noise reduction processing on the initial audio features to obtain the target audio features and the target audio corresponding to the target audio features. Through the present disclosure, the noise reduction effect is ensured, the obtained target audio is ensured not to be distorted, the algorithm can be ensured to run in real time on a mobile device, and the noise reduction effect meets the current scenario.
[0026] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and do not limit the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The accompanying drawings herein are incorporated into the specification and form a part of the specification, showing embodiments consistent with the present disclosure, and are used together with the specification to explain the principles of the present disclosure.
[0028] Figure 1 is a flowchart of an audio noise reduction method shown according to an exemplary embodiment.
[0029] Figure 2 is a schematic diagram of an audio noise reduction method shown according to an exemplary embodiment of the present disclosure.
[0030] Figure 3 is a flowchart of a method for determining a current noise scenario according to an audio to be processed shown according to an exemplary embodiment.
[0031] Figure 4 is a flowchart of a method for obtaining target audio features shown according to an exemplary embodiment.
[0032] Figure 5 is a flowchart of a method for obtaining target audio features shown according to an exemplary embodiment.
[0033] Figure 6 is a flowchart of a method for training a preset noise reduction model shown according to an exemplary embodiment.
[0034] Figure 7 is a flowchart of a method for obtaining a loss function of a preset neural network shown according to an exemplary embodiment.
[0035] Figure 8 is a flowchart of a method for determining a loss function shown according to an exemplary embodiment.
[0036] Figure 9 is a flowchart of an audio noise reduction method shown according to an exemplary embodiment.
[0037] Figure 10 is a flowchart of an audio noise reduction method shown according to an exemplary embodiment.
[0038] Figure 11 is a flowchart of an audio noise reduction method shown according to an exemplary embodiment of the present disclosure.
[0039] Figure 12 is a flowchart of a method for training a preset noise reduction model shown according to an exemplary embodiment of the present disclosure.
[0040] Figure 13 is a block diagram of an audio noise reduction device shown according to an exemplary embodiment.
[0041] Figure 14 is a block diagram of a device for audio noise reduction shown according to an exemplary embodiment. Detailed implementation manners
[0042] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with the present disclosure.
[0043] The audio noise reduction method provided by the embodiments of the present disclosure is applied to a scenario where the obtained original audio is subjected to noise reduction to obtain the noise-reduced audio.
[0044] The speech recorded in a real-world scenario may contain various background noises. For example, audio and video conferences, automatic speech recognition, and hearing aids. To solve this problem, speech noise reduction technology aims to eliminate such noises and then output a perceptually high-quality speech signal.
[0045] In addition, on mobile devices such as mobile phones, smart headphones, smart speakers, etc., the requirements for noise reduction algorithms are extremely high, such as noise reduction quality, distortion degree, real-time ability, power consumption, etc., which pose a great challenge to the reliability of noise reduction algorithms.
[0046] In related technologies, speech noise reduction technologies include traditional signal processing methods and machine learning methods. The method based on deep neural network has a large model capacity for processing large-scale training data and has better noise reduction effect and speed. In building a model, speech noise reduction is usually regarded as a regression task, and the network is trained to directly predict clean speech. The noise reduction models in related technologies are mainly divided into two categories: the noise reduction model based on spectrogram and the noise reduction model based on time-domain waveform. The speech noise reduction method in related technologies is generally implemented based on the noise reduction model based on spectrogram. The method of the noise reduction model based on spectrogram first extracts the noise spectrum features (such as spectrogram or the amplitude of complex spectrum). And predicts the modulation mask (such as the ideal ratio mask) or the spectrum features of clean speech. Finally, according to the predicted mask or spectrum features and other information (such as phase) extracted from the noisy speech, the waveform corresponding to the target audio is generated. And the method based on time-domain waveform directly predicts the time-domain waveform representation of clean speech from the noisy waveform. The method based on time-domain waveform in related technologies uses WaveNet or U-Net as the backbone and has different sub-modules, such as dilated convolution, independent WaveNet, Long Short Term Memory (LSTM) and Attention mechanism block, etc. In addition, the traditional sound noise reduction technology based on sparse representation is also a common method. This method uses the characteristics of sparse representation to decompose the mixed sound into multiple sub-signals, and then performs noise reduction processing on the noise sub-signals among them. Common sparse representation methods include the method based on dictionary learning and the method based on compressive sensing, etc.
[0047] However, the noise reduction means in the above related technologies have the following problems: (1) Only using the noise spectrogram as the input is very likely to lose the information in the noise waveform, resulting in sound distortion. (2) The noise reduction network is not causal, resulting in the inability to truly achieve real-time noise reduction. In addition, overly complex processing of the audio signal will also cause a certain processing delay, which may affect real-time performance. (3) The phase estimation of the noisy speech is inaccurate, and obvious noise leakage will occur at high noise levels. (4) At high noise levels, the audio quality of the target audio will decline (that is, the target audio obtained after noise reduction does not sound natural), and expanding the model to a larger network cannot improve the audio quality of the target audio. (5) On miniaturized devices such as smart headphones, the computing power requirements for the device are relatively high, the power consumption is large, and it affects the battery life of the device. (6) The effect is unstable for different noise environments, and there will be situations of incomplete noise reduction or misjudgment.
[0048] In view of this, the present disclosure proposes an audio noise reduction method. When performing audio noise reduction, the audio to be processed is obtained in real time, and the current noise scene is determined according to the audio to be processed. Then, the noise reduction parameters corresponding to the current noise scene are determined, and the preset noise reduction model is adjusted according to the current noise reduction parameters. The adjusted preset noise reduction model is used to perform noise reduction processing on the initial audio features to obtain the target audio features and the target audio corresponding to the target audio features. Through the present disclosure, the noise reduction effect is ensured, the obtained target audio is ensured not to be distorted, the algorithm is ensured to be able to run in real time on mobile devices, and the noise reduction effect conforms to the current scene.
[0049] The present disclosure mainly relates to the technical fields of audio noise reduction, scene analysis, and related artificial intelligence technologies. In particular, it relates to a fast noise reduction technology based on scene recognition and time-frequency domain hybrid features. It can be extended in the technical fields of fast noise reduction on mobile devices and fast adaptation of noise reduction algorithms.
[0050] Figure 1 It is a flowchart of an audio noise reduction method shown according to an exemplary embodiment. As Figure 1 shown, the method includes steps S101 to S104.
[0051] In step S101, the audio to be processed is obtained in real time, and the current noise scene is determined according to the audio to be processed.
[0052] In step S102, according to the correspondence between the noise scene and the noise reduction parameters, the current noise reduction parameters corresponding to the current noise scene are determined.
[0053] In step S103, the preset noise reduction model is adjusted according to the current noise reduction parameters.
[0054] In step S104, the initial audio features of the audio to be processed are obtained, the initial audio features are denoised using the adjusted preset denoising model to obtain the target audio features, and the target audio corresponding to the target audio features is obtained.
[0055] Among them, the target audio features are the initial audio features after denoising, and the initial audio features include the initial frequency domain features and the initial time domain features of the audio to be processed.
[0056] In the embodiments of the present disclosure, the proposed audio denoising method is mainly aimed at low-resource mobile devices (such as mobile terminals, Bluetooth headsets, etc.). It is necessary to establish a voice denoising model algorithm (i.e., the preset denoising model) on the corresponding device for denoising through the voice denoising model algorithm. It can be understood that the denoising requirements are inconsistent in different scenarios. If good denoising effects are to be ensured for various scenarios, the model itself needs to have model parameters corresponding to various scenarios. In this case, there are many model parameters and a large amount of computation during the denoising process, making it impossible to run in real time on mobile devices such as terminals. The present disclosure optimizes for different denoising scenarios, configures different model parameters (denoising parameters) for different scenarios, identifies the current scenario during denoising, selects the model parameters corresponding to the current scenario, adjusts the preset denoising model, and uses the adjusted preset denoising model for denoising. Through the present disclosure, the preset denoising model is adaptively adjusted for different denoising scenarios, the parameters in the model are reduced, and the amount of computation during denoising is reduced, enabling real-time operation on mobile devices.
[0057] In the embodiments of the present disclosure, audio denoising is performed on the time domain features and frequency domain features of the audio to be processed, that is, denoising is performed on the audio to be processed by combining the time domain waveform denoising technology and the frequency domain spectrum denoising technology, combining the advantages of the time domain waveform denoising technology and the frequency domain spectrum denoising technology, ensuring good denoising effects while also ensuring that the target audio obtained after denoising is not distorted.
[0058] In an exemplary embodiment of the present disclosure, the denoising process is represented by a relational expression, that is: \(\hat{x}=f(x_{noisy})\), where \(x_{noisy}\) is the noisy audio (audio to be processed), and \(\hat{x}\) is the processed clean and undistorted speech (target audio), that is, only the target audio (such as human speech) is retained after removing the environmental noise from the noisy audio recorded by the microphone.
[0059] In the embodiments of the present disclosure, a method combining a time-domain waveform noise reduction technique and a frequency-domain spectrum noise reduction technique is adopted to perform noise reduction on the audio to be processed. The main consideration is the main feature that the noise reduction models based on spectrograms and those based on time-domain waveforms are "complementary" under high noise levels. That is, the model based on spectrograms can maintain good speech quality, but there is noise leakage (for example, due to inaccurate phase information extracted from noisy audio). On the other hand, the model based on time-domain waveforms is good at eliminating noise, but may produce speech with degraded quality. In order to combine the advantages of these two models, a hybrid two-stage algorithm framework is adopted, which consists of two main networks: a spectrogram-based noise reduction network and a waveform-based noise reduction network. By means of feature complementary fusion, their respective advantages are fully utilized to improve the final noise reduction quality.
[0060] In an exemplary embodiment of the present disclosure, as Figure 2 shown in the schematic diagram of the audio noise reduction method, the present disclosure realizes audio noise reduction in the following manner: In response to performing noise reduction processing on the audio to be processed, the time-domain features of the audio to be processed (noisy time-domain waveform diagram Noisy) are obtained, and scene recognition is performed (including scenes such as subway, office, high-speed rail, and gymnasium, etc.). After the scene recognition is completed, scene noise reduction adaptation is performed, the preset noise reduction model is adjusted, and the time-domain features of the audio to be processed are subjected to noise reduction processing to obtain the target audio features (cleaned time-domain waveform diagram Clean). Among them, the data required for scene adaptation is obtained from the cloud, and the cloud will optimize the scene recognition algorithm and perform joint debugging of the scene + algorithm based on the scenes encountered by the user, so as to obtain the noise reduction parameters corresponding to different scenes.
[0061] It can be understood that when performing scene recognition, real-time audio information and corresponding information in the device need to be collected to ensure the accuracy of the recognition result. The following embodiments of the present disclosure illustrate the method for determining the current noise scene.
[0062] Figure 3 is a flowchart of a method for determining the current noise scene according to an audio to be processed shown in an exemplary embodiment. As Figure 3 shown, the method includes steps S201 to step S202.
[0063] In step S201, the audio to be processed is obtained in real time.
[0064] In step S202, the current noise scene is determined according to the audio to be processed and the preset information.
[0065] Among them, the preset information includes one or more of the following information: the application running information of the terminal within a preset time, the low-power component running information of the terminal within a preset time, and the current location information of the terminal.
[0066] In the embodiments of the present disclosure, considering the user's habit of using the terminal in daily life, when scene recognition is required, the application data of the terminal is collected and combined with the current audio data to be processed for scene judgment. For example, if it is detected that the traffic card function of the NFC module is activated, the current noise scene can be determined as the subway scene or the bus scene in combination with the current audio data to be processed; further, when the user grants the positioning permission, more accurate positioning can be achieved in combination with the current location information of the terminal.
[0067] In one exemplary embodiment, the audio noise reduction method proposed by the present disclosure can identify the following scenes and perform noise reduction processing:
[0068] Scene 1: Traditional noise reduction scene. 1. Voice communication: In scenarios such as telephone, Internet voice calls, and video conferences, due to limitations of the network environment and device conditions, noise interference often occurs, and fast noise reduction technology can effectively improve the quality of voice communication. 2. Audio processing: In scenarios such as audio recording, audio editing, and audio conversion, due to interference from environmental noise or device noise, the quality of the audio is affected, and fast noise reduction technology can effectively remove the noise and improve the quality of the audio.
[0069] Scene 2: Combining scene recognition technology, that is, the sub-scene noise reduction technology refers to the technology that adopts different noise reduction algorithms and parameters according to different scenes and environments to achieve better noise reduction effects. Its application scenarios are as follows: 1. Conference room: In the conference room, the ratio of human voices to environmental noise is relatively high. The sub-scene noise reduction technology can adopt different noise reduction algorithms and parameters according to the characteristics of human voices and environmental noise to achieve better noise reduction effects. 2. Restaurant: In the restaurant, the environmental noise is relatively complex, including human voices, background music, tableware collision sounds, etc. The sub-scene noise reduction technology can adopt different noise reduction algorithms and parameters according to different noise types to achieve better noise reduction effects. 3. Inside the vehicle: Inside the vehicle, the noise mainly comes from the engine, wheels, and wind noise, etc. The sub-scene noise reduction technology can adopt different noise reduction algorithms and parameters according to different noise sources to achieve better noise reduction effects. 4. Live broadcast site: At the live broadcast site, the noise mainly comes from human voices, music, on-site effects, etc. The sub-scene noise reduction technology can adopt different noise reduction algorithms and parameters according to different noise sources and characteristics to achieve better noise reduction effects.
[0070] Scenario 3: Mobile devices such as mobile phones and headphones. Mobile devices, mobile phones, and headphones noise reduction technology refers to noise reduction technology applied on portable devices such as mobile phones and headphones. At the same time, in order to ensure the application requirements of the end-side devices in low-power scenarios, it is necessary to optimize the noise reduction algorithm, including pruning, compression, etc. However, compression of related models will lead to a decrease in effect. It is estimated that by combining with scene analysis, a single algorithm can target a group of scenes, thus ensuring the stability of its noise reduction, while also being able to compress as much as possible to reduce power consumption. Its application scenarios are as follows: 1. The noise reduction effect on smart noise reduction headphones can be applied to many application scenarios. 2. Noise reduction of smart speaker devices can analyze the current scene and adaptively adjust the adaptation of the noise reduction algorithm and parameter adaptation. The corresponding scenarios of smart noise reduction headphones include: (1) Public transportation: On public transportation, noise mainly comes from engines, wheels, wind noise, etc. Mobile devices, mobile phones, and headphones noise reduction technology can effectively remove noise and improve passengers' riding experience. (2) Offices: In offices, the ratio of human voices to environmental noise is high. Mobile devices, mobile phones, and headphone noise reduction technology can effectively remove noise, improving work efficiency and work quality. (3) Learning places: In learning places, noise mainly comes from human voices and environmental noise. Mobile devices, mobile phones, and headphone noise reduction technology can effectively remove noise, improving learning effects and learning quality. (4) Sports venues: In sports venues, noise mainly comes from human voices and environmental noise. Mobile devices, mobile phones, and headphone noise reduction technology can effectively remove noise, improving exercise experience and exercise effects.
[0071] In the embodiments of the present disclosure, the initial audio features obtained for noise reduction include initial time-domain features and initial frequency-domain features. Based on this, the present disclosure employs a combination of time-domain waveform noise reduction technology and frequency-domain spectrum noise reduction technology to reduce noise in the processed audio. Therefore, the preset noise reduction model in the present disclosure includes a time-domain noise reduction network and a frequency-domain noise reduction network. The following embodiments of the present disclosure illustrate methods for obtaining target audio features.
[0072] Figure 4 FIG. 1 is a flow chart showing a method for obtaining target audio features according to an exemplary embodiment. Figure 4 As shown, the method includes steps S301 to S302.
[0073] In step S301, the initial time domain features are subjected to time domain convolution processing through a time domain denoising network to obtain intermediate time domain features, and the initial frequency domain features are subjected to frequency domain convolution processing through a frequency domain denoising network to obtain intermediate frequency domain features.
[0074] In step S302, the intermediate time domain features and the intermediate frequency domain features are processed according to the time domain noise reduction network and the frequency domain noise reduction network to obtain the target audio features.
[0075] In an embodiment of the present disclosure, when performing noise reduction, the initial time-domain features of the audio to be processed are obtained, and the initial time-domain features are subjected to short-time Fourier feature transformation to obtain initial frequency-domain features. The initial frequency-domain features and the initial membership features are respectively input into a time-domain noise reduction network and a frequency-domain noise reduction network, and after convolutional processing respectively, intermediate time-domain features and intermediate frequency-domain features are obtained, which is convenient for introducing frequency-domain features when using the time-domain noise reduction network to process the intermediate time-domain features and introducing time-domain features when using the frequency-domain noise reduction network to process the intermediate frequency-domain features in subsequent processes. Thus, by combining the advantages of the time-domain noise reduction model and the frequency-domain noise reduction model, while ensuring good noise reduction effect, the target audio is not distorted.
[0076] In an embodiment of the present disclosure, the time-domain noise reduction model mainly processes based on time-domain waveform features, that is, both its input and output are time-domain signals (initial time-domain features). The frequency-domain noise reduction network is mainly processed based on the initial frequency-domain features, and secondly, to establish a streaming audio processing mechanism, reduce the delay of noise reduction processing, and achieve real-time noise reduction.
[0077] In an embodiment of the present disclosure, the frequency-domain features are fed into a hybrid feature Unet network based on the frequency-domain noise reduction network. The hybrid feature Unet network is a Unet composed of three layers of two-dimensional convolution and transposed convolution. Through the Unet network, frequency-domain convolutional processing is performed to obtain intermediate frequency-domain features.
[0078] In an example of the present disclosure, the time-domain noise reduction network is composed of multiple layers of Temporal Convolutional Network (TCN) models, as well as LSTM and convolutional layers of different scales. TCN is an open-source network for time-series modeling in time-domain networks, mainly solving the defect that traditional RNNs cannot be processed in parallel. TCN can adopt a series of arbitrary lengths and output them with the same length. In the case of using a one-dimensional fully convolutional network architecture, causal convolution is used. A key feature is that the output at time t is only convolved with the elements that occurred before t. The main features can be summarized as: applicable sequence model - causal convolution and memory history - dilated convolution, and residual block, ensuring that the model is a causal model. It can be understood that there is a serious problem in the internal design of traditional RNNs: since the network can only process one time step at a time, the next step must wait for the previous step to be processed before it can perform operations. This means that RNNs cannot be processed in large-scale parallel like CNNs, especially when RNN / LSTM processes text bidirectionally. This also means that RNNs are extremely computationally intensive because all intermediate results must be saved before the entire task runs to completion. The present disclosure uses CNN for noise reduction processing of frequency-domain features (spectrograms). When processing images, CNN regards the image as a two-dimensional "block" (a matrix of m*n). When migrated to time series, the sequence can be regarded as a one-dimensional object (a vector of 1*n). Through a multi-layer network structure, a sufficiently large receptive field can be obtained. This approach will make the CNN very deep, but thanks to the advantage of large-scale parallel processing, no matter how deep the network is, it can be processed in parallel, saving a lot of time.
[0079] In the embodiments of the present disclosure, after obtaining the intermediate time-domain features and intermediate frequency-domain features, based on the intermediate time-domain features and intermediate frequency-domain features, when using the time-domain noise reduction network to process the intermediate time-domain features, the frequency-domain features are introduced, and when using the frequency-domain noise reduction network to process the intermediate frequency-domain features, the time-domain features are introduced. The following embodiments of the present disclosure illustrate the method for obtaining the target audio features.
[0080] Figure 5 It is a flowchart of a method for obtaining target audio features shown according to an exemplary embodiment. As Figure 5 shown, the method includes steps S401 to S403.
[0081] In step S401, the intermediate time-domain features are processed based on the self-attention mechanism, and the processed intermediate time-domain features and intermediate frequency-domain features are fused to obtain intermediate hybrid features.
[0082] In step S402, the intermediate mixed features are successively subjected to recursive processing, transposed convolution processing, and predictive modulation mask processing through a frequency-domain noise reduction network to obtain target frequency-domain features.
[0083] In step S403, the target frequency-domain features and the intermediate time-domain features are recursively processed through a time-domain noise reduction network to obtain target audio features.
[0084] In the embodiments of the present disclosure, after obtaining the intermediate time-domain features and the intermediate frequency-domain features, the intermediate frequency-domain features of the time-domain model are introduced into the middle part of the time-domain noise reduction network, and the intermediate time-domain features are processed through a self-attention mechanism block (Self-Attention). The processed intermediate time-domain features are fused with the intermediate frequency-domain features to obtain intermediate mixed features. The intermediate mixed features are successively subjected to recursive processing, transposed convolution processing, and predictive modulation mask processing through a frequency-domain noise reduction network to obtain target frequency-domain features (i.e., the denoised spectrogram). The target frequency-domain features are introduced into the time-domain noise reduction network, and combined with the intermediate time-domain features and the target frequency-domain features to obtain the target time-domain features, i.e., the target audio features, ensuring good noise reduction effect and no distortion of the target audio. In summary, the present disclosure improves the stability of the algorithm in the time dimension by combining the frequency-domain compressed features (intermediate frequency-domain features) with the convolutional features in the time domain dimension, and improves the noise reduction continuity and noise reduction effect of the frequency-domain model.
[0085] In an exemplary embodiment of the present disclosure, for each convolutional layer and transposed convolutional layer, it is essentially composed of a two-dimensional convolution (Conv2d) that preserves channels, a rectified linear unit (ReLU), and a gated linear unit (GLU). The kernel size of each Conv2d is (3, 1), and the stride is 1. For the time-domain features, here the input time-domain features are processed through a self-attention block with 16 heads and 256 dimensions. In addition, the temporal modeling ability is further enhanced by combining LSTM. For the features (intermediate frequency-domain features) after Unet, the denoised spectrum (target frequency-domain features) is obtained by estimating the Mask and multiplying it with the original spectrum.
[0086] In an exemplary embodiment of the present disclosure, for multiple layers of TCN (1, 2, 3), in order to ensure a low time delay, the number of channels of TCN-1 is set to (16, 8, 4), the number of channels of TCN-2 is set to (32, 16, 8), and the number of channels of TCN-3 is set to (64, 32, 16). In addition, after passing through the time-domain TCN module, the frequency-domain features are introduced here, and the original features are mapped to 128 channels using a one-dimensional convolution, and are concatenated and stacked with the frequency-domain dimension in the channel dimension, so that the time-frequency domain features establish a strong correlation matrix, and the strong correlation feature matrix is modeled by sending it into a two-layer LSTM network, and the output features are restored to a time-domain signal.
[0087] In the embodiments of the present disclosure, the noise reduction model based on [ ] involves a time-domain noise reduction network and a frequency-domain noise reduction network. During the training process, it is necessary to consider the loss between the original audio features and the time-domain noise reduction processing results, the loss between the original audio features and the frequency-domain noise reduction processing results, and the loss between the time-domain noise reduction processing results and the frequency-domain noise reduction processing results. The following embodiments of the present disclosure illustrate the training method of the preset noise reduction model.
[0088] Figure 6 FIG. 1 is a flow chart showing a method for training a preset noise reduction model according to an exemplary embodiment. Figure 6 As shown, the method includes steps S501 to S503.
[0089] In step S501, different noise scenes are simulated based on a preset audio data set. In different noise scenes, the training audio in the corresponding noise scene is processed by a preset neural network to obtain predicted target frequency domain features and predicted target time domain features.
[0090] Among them, the audio data set includes a noisy data set and a clean audio data set, and the preset neural network and the preset noise reduction model have the same algorithm structure.
[0091] In step S502, a loss function of a preset neural network is obtained according to the predicted target frequency domain features, the predicted target time domain features, the time domain features of the training audio, and the frequency domain features of the training audio.
[0092] In step S503, in response to the loss function satisfying the preset numerical requirement, the preset neural network is determined as the preset noise reduction model.
[0093] In the embodiment of the present disclosure, various noise scenarios are simulated according to a preset database, and a neural network with the same algorithm structure is trained under different noise scenarios to obtain noise reduction parameters suitable for different noise scenarios, which facilitates parameter selection and model adjustment based on the identified scenarios in practical applications.
[0094] In an exemplary embodiment of the present disclosure, subjective evaluation data is based on the Deep Noise Suppression (DNS Challenge) dataset and data recorded in real-world environments. The dataset contains 400 hours of clean speech and numerous noisy segments, while the training set consists of 200 hours of clean-noise speech pairs, with 20 signal-to-noise (SNR) levels ranging from -8 to 30 dB.
[0095] The following embodiments of the present disclosure further illustrate the training method of the preset noise reduction model.
[0096] Figure 7 FIG. 1 is a flow chart showing a method for obtaining a loss function of a preset neural network according to an exemplary embodiment. Figure 7As shown, the method includes steps S601 to S603.
[0097] In step S601, an inverse Fourier transform process is performed on the predicted target frequency-domain feature to obtain the converted time-domain feature corresponding to the predicted target frequency-domain feature.
[0098] In step S602, the time-domain loss between the predicted target time-domain feature and the time-domain feature of the training audio is obtained, the frequency-domain loss between the converted time-domain feature and the frequency-domain feature of the training audio is obtained, and the time-frequency domain hybrid loss between the predicted target time-domain feature and the converted time-domain feature is obtained.
[0099] In step S603, according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight, a loss function is determined. The first weight corresponds to the time-domain loss, the second weight corresponds to the frequency-domain loss, and the third weight corresponds to the time-frequency domain hybrid loss.
[0100] In an exemplary embodiment of the present disclosure, the domain frequency-domain loss is obtained in the following manner: Assume that s(x; θ)=|STFT(x; θ)| is the magnitude of the linear spectrogram of waveform x, where θ is a set of parameters (frame shift size, window length, and FFT-bin) for calculating STFT. Then, the transformed noisy waveform x is represented by s(x; θspec), that is, the noisy-to-noisy spectrogram y-noisy = s(x-noisy; θspec), and the clean waveform x to the clean spectrogram y = s(x; θspec). Then the frequency-domain loss is: where Ts is the time length of the spectrogram, represents the denoised speech, y represents the original audio, and Ts represents the time frame. Ls is the frequency-domain loss.
[0101] In an exemplary embodiment of the present disclosure, the time-domain loss is obtained in the following manner:
[0102] where Lt is the time-domain loss, s(x) represents the time-domain signal, s(x') represents the time-domain signal after noise reduction processing, and T represents the time length.
[0103] In an embodiment of the present disclosure, during the training process of the two time-domain denoising networks and the frequency-domain denoising network, in order to improve the training effectiveness of the two-branch model and better utilize the mutual assistance ability of features in different modalities, a method of adding time-domain features during the training process of the frequency-domain model and a method of fusing frequency-domain features during the training process of the time-domain model are combined, and the sum of the multi-resolution frequency-domain loss, time-domain loss, and time-domain hybrid loss is used as the final loss function to train and obtain the time-frequency domain hybrid loss. In an exemplary embodiment of the present disclosure, the hybrid loss is: Lmix = MSE(s(x) t ,s(x) s) Among them, Lmix is the time-domain loss, s(x)t represents the time-domain signal after noise reduction processing, and s(x)s represents the frequency-domain signal after noise reduction processing.
[0104] The following embodiments of the present disclosure further illustrate the method for training a preset noise reduction model.
[0105] Figure 8 is a flowchart of a method for determining a loss function shown according to an exemplary embodiment. As Figure 8 shown, the method includes steps S701 to S702.
[0106] In step S701, the product of the time-domain loss and the first weight is determined as the first product, the product of the frequency-domain loss and the second weight is determined as the second product, and the product of the time-frequency domain mixed loss and the third weight is determined as the third product.
[0107] In step S702, the sum of the first product, the second product, and the third product is determined as the loss function.
[0108] In the embodiments of the present disclosure, by configuring a first weight, a second weight, and a third weight for the obtained time-domain loss, frequency-domain loss, and time-frequency domain mixed loss respectively, and combining the three weights to represent the loss function, it is convenient to determine whether the training of the preset noise reduction model is completed based on the loss function.
[0109] In an exemplary embodiment of the present disclosure, the final loss function for model training is: L = α·Lt + β·Ls + λ·Lmix, where Lt is the time-domain loss, Ls is the frequency-domain loss, Lmix is the time-frequency domain mixed loss, α is the first weight, β is the second weight, and λ is the third weight. In one example, the first weight, the second weight, and the third weight are 1, 0.5, and 0.1 respectively. After the model training is completed, the inference model only uses the time-domain model as the main branch and combines the features of the frequency-domain model, that is, the frequency-domain branch no longer outputs the time-domain signal.
[0110] It can be understood that the scenarios and corresponding parameters included in the terminal are limited. When encountering an unrecognizable scenario, it is necessary to feedback to the cloud and use the default method for noise reduction. The following embodiments of the present disclosure further illustrate the audio noise reduction method in the present disclosure.
[0111] Figure 9 is a flowchart of an audio noise reduction method shown according to an exemplary embodiment. As Figure 9 shown, the method includes steps S801 to S802.
[0112] In step S801, the audio to be processed is obtained in real time.
[0113] In step S802, in response to the inability to determine the current noise scenario based on the audio to be processed, the preset noise reduction model is adjusted according to the default noise reduction parameters, and the audio data to be processed collected by the current terminal and the preset data are uploaded to the cloud.
[0114] In the embodiments of the present disclosure, when the current scenario cannot be recognized, the preset noise reduction model is adjusted using the default noise reduction parameters for noise reduction, and then the data in the current scenario is uploaded to the cloud, facilitating developers to generate corresponding noise reduction parameters based on this data and optimizing the local noise reduction algorithm of the terminal subsequently.
[0115] The following embodiments of the present disclosure further illustrate the audio noise reduction method in the present disclosure.
[0116] Figure 10 It is a flowchart of an audio noise reduction method shown according to an exemplary embodiment. As Figure 10 shown, the method includes step S901 to step S902.
[0117] In step S901, the corresponding data between the noise reduction parameters returned by the cloud and the noise scenario is obtained.
[0118] In step S902, the corresponding relationship between the locally stored noise scenario and the noise reduction parameters is updated according to the corresponding data.
[0119] In the embodiments of the present disclosure, improvements are made in terms of the reliability of algorithm adaptation, mainly by combining the optimization of the cloud scene recognition algorithm and the joint debugging of the cloud scene analysis + noise reduction algorithm for scene adaptability optimization. For scenarios that cannot be recognized, they will be uploaded to the cloud for scene recognition training and adaptation, and new noise reduction algorithm training for the new scenario will be added again. After the relevant adaptation is completed, it will be sent to the device side to complete the expansion of the scene and the noise reduction scene. In this way, the scene-specific noise reduction algorithm can be completed for adaptation.
[0120] In an exemplary embodiment of the present disclosure, as Figure 11As shown in the flowchart of the audio noise reduction method, the following method is used for audio noise reduction. In response to noise reduction, the time-domain features of the audio to be processed (the time-domain waveform diagram of the noisy audio) are obtained. The time-domain features are processed by short-time Fourier transform to obtain the frequency-domain features of the audio to be processed (the spectrogram of the noisy audio). The time-domain features are processed by time-domain convolution, and TCN-1, TCN-2, and TCN-3 are traversed to obtain the intermediate time-domain feature Conv1D(32). The frequency-domain features are processed by frequency-domain convolution, and Conv2D(8), Conv2D(16), and Conv2D(32) are traversed to obtain the intermediate frequency-domain feature Conv1D(32). The intermediate frequency-domain features are processed by the self-attention mechanism block Self-Attention, and the processed intermediate time-domain features and intermediate time-frequency domain features are fused to obtain the time-frequency domain hybrid feature Conv1D(32). The time-frequency domain hybrid features are processed by transposed convolution, and ConvTranspose2D(32), ConvTranspose2D(16), and ConvTranspose2D(8) are traversed, and prefabricated mask (MASK) processing is performed to obtain the target frequency-domain features (the denoised spectrogram Spec-ff). The target frequency-domain features and the intermediate time-domain features are recursively processed multiple times by LSTM(128) to obtain the target time-domain features (i.e., the target audio features, the time-domain waveform diagram after noise reduction), and the target audio features are converted into audio to complete the noise reduction.
[0121] In an exemplary embodiment of the present disclosure, Figure 12Flowchart of the training method of the preset noise reduction model. The noise reduction model is trained in the following manner. Obtain the time-domain features (training waveform diagrams) for training. Perform short-time Fourier transform processing on the time-domain features to obtain the frequency-domain features (training spectrograms) for training. Perform time-domain convolution processing on the time-domain features, traverse TCN-1, TCN-2, TCN-3, to obtain the intermediate time-domain features Conv1D(32). Perform frequency-domain convolution processing on the frequency-domain features, traverse Conv2D(8), Conv2D(16), Conv2D(32), to obtain the intermediate frequency-domain features Conv1D(32). Process the intermediate frequency-domain features through the self-attention mechanism block Self-Attention, and fuse the processed intermediate time-domain features and intermediate time-frequency domain features to obtain the time-frequency domain hybrid features Conv1D(32). Perform transposed convolution processing on the time-frequency domain hybrid features, traverse ConvTranspose2D(32), ConvTranspose2D(16), ConvTranspose2D(8), and perform prefabricated mask (MASK) processing to obtain the target frequency-domain features (denoised spectrogram Spec-ff). Perform multiple recursive processing LSTM(128) on the target frequency-domain features and the intermediate time-domain features to obtain the target time-domain features (i.e., target audio features, denoised time-domain waveform diagram). And perform inverse Fourier transform on the target frequency-domain features to obtain the conversion features. Obtain the frequency-domain loss according to the conversion features and the time-domain features for training, obtain the time-domain loss according to the target time-domain features and the time-domain features for training, obtain the hybrid loss according to the target time-domain features and the conversion features, obtain the loss function according to the time-domain loss, frequency-domain loss, and hybrid loss, and complete the model training in response to the loss function meeting the requirements.
[0122] In the embodiments of the present disclosure, by combining the advantages of the time-domain waveform noise reduction and the frequency-domain spectrogram noise reduction modes, the speech noise reduction quality is improved and the distortion degree is reduced. By establishing a combination of two-stage noise reduction algorithms, it is highly possible to prevent noise leakage, and at the same time, good sound quality can be maintained even under high noise levels. Combining with the scene recognition and analysis technology, noise reduction is performed for specific scenes and combined with general noise reduction. The model training method of hybrid features fully integrates the characteristics of the frequency domain and time domain to improve the noise reduction effect. The combination of time-domain TCN and frequency-domain real-time processing, and the comprehensive algorithm can achieve real-time noise reduction. Equipped with a cloud optimization module, the noise reduction algorithm can be adapted and distributed for unknown scenes to establish better environmental adaptability.
[0123] Based on the same concept, the embodiments of the present disclosure also provide an audio noise reduction device 100.
[0124] It can be understood that, in order to implement the above functions, the audio noise reduction device 100 provided in the embodiments of the present disclosure includes the corresponding hardware structures and / or software modules for performing various functions. Combining the units and algorithm steps of the examples disclosed in the embodiments of the present disclosure, the embodiments of the present disclosure can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the way of hardware or computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the technical solution of the embodiments of the present disclosure.
[0125] Figure 13 FIG. is a block diagram of an audio noise reduction device 100 shown according to an exemplary embodiment. Referring to Figure 13 , the device includes a determination unit 101, an adjustment unit 102, and a processing unit 103.
[0126] The determination unit 101 is configured to obtain the audio to be processed in real time, determine the current noise scenario according to the audio to be processed, and determine the current noise reduction parameter corresponding to the current noise scenario according to the correspondence between the noise scenario and the noise reduction parameter.
[0127] The adjustment unit 102 is configured to adjust the preset noise reduction model according to the current noise reduction parameter.
[0128] The processing unit 103 is configured to obtain the initial audio features of the audio to be processed, perform noise reduction processing on the initial audio features using the adjusted preset noise reduction model to obtain the target audio features, and obtain the target audio corresponding to the target audio features.
[0129] Wherein, the target audio features are the initial audio features after noise reduction processing, and the initial audio features include the initial frequency domain features and the initial time domain features of the audio to be processed.
[0130] In one implementation manner, the determination unit 101 determines the current noise scenario according to the audio to be processed in the following manner: determining the current noise scenario according to the audio to be processed and preset information, where the preset information includes one or more of the following information: application running information of the terminal within a preset time, low-power component running information of the terminal within a preset time, and current location information of the terminal.
[0131] In one implementation, the preset noise reduction model includes a time-domain noise reduction network and a frequency-domain noise reduction network. The processing unit 103 uses the adjusted preset noise reduction model to perform noise reduction processing on the initial audio features in the following manner to obtain the target audio features, including: performing time-domain convolution processing on the initial time-domain features through the time-domain noise reduction network to obtain intermediate time-domain features, and performing frequency-domain convolution processing on the initial frequency-domain features through the frequency-domain noise reduction network to obtain intermediate frequency-domain features. According to the time-domain noise reduction network and the frequency-domain noise reduction network, the intermediate time-domain features and the intermediate frequency-domain features are processed to obtain the target audio features.
[0132] In one implementation, the processing unit 103 processes the intermediate time-domain features and the intermediate frequency-domain features according to the time-domain convolution network and the frequency-domain convolution network in the following manner to obtain the target audio features: processing the intermediate time-domain features based on the self-attention mechanism, and fusing the processed intermediate time-domain features and the intermediate frequency-domain features to obtain an intermediate mixed feature. Performing recursive processing, transposed convolution processing, and predictive modulation mask processing on the intermediate mixed feature successively through the frequency-domain noise reduction network to obtain the target frequency-domain feature. Performing recursive processing on the target frequency-domain feature and the intermediate time-domain features through the time-domain noise reduction network to obtain the target audio feature.
[0133] In one implementation, the noise reduction parameters include the model shape parameters of the preset noise reduction model obtained by training in different noise scenarios. The preset noise reduction model processing unit 103 is trained in the following manner: simulating different noise scenarios based on a preset audio data set, and in different noise scenarios, processing the training audio corresponding to the noise scenario through a preset neural network to obtain the predicted target frequency-domain feature and the predicted target time-domain feature. The audio data set includes a noise data set and a clean audio data set, and the preset neural network has the same algorithm structure as the preset noise reduction model. According to the predicted target frequency-domain feature, the predicted target time-domain feature, the time-domain feature of the training audio, and the frequency-domain feature of the training audio, the loss function of the preset neural network is obtained. In response to the loss function meeting the preset numerical requirements, the preset neural network is determined as the preset noise reduction model.
[0134] In one implementation, the processing unit 103 obtains the loss function of the preset neural network according to the predicted target frequency-domain feature, the predicted target time-domain feature, the time-domain feature of the training audio, and the frequency-domain feature of the training audio in the following manner: perform an inverse Fourier transform on the predicted target frequency-domain feature to obtain the converted time-domain feature corresponding to the predicted target frequency-domain feature. Obtain the time-domain loss between the predicted target time-domain feature and the time-domain feature of the training audio, obtain the frequency-domain loss between the converted time-domain feature and the frequency-domain feature of the training audio, and obtain the time-frequency domain hybrid loss between the predicted target time-domain feature and the converted time-domain feature. Determine the loss function according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight, where the first weight corresponds to the time-domain loss, the second weight corresponds to the frequency-domain loss, and the third weight corresponds to the time-frequency domain hybrid loss.
[0135] In one implementation, the processing unit 103 determines the loss function according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight in the following manner: determine the product of the time-domain loss and the first weight as the first product, determine the product of the frequency-domain loss and the second weight as the second product, and determine the product of the time-frequency domain hybrid loss and the third weight as the third product. Determine the sum of the first product, the second product, and the third product as the loss function.
[0136] In one implementation, the processing unit 103 is further configured to: in response to being unable to determine the current noise scenario based on the audio to be processed, adjust the preset noise reduction model according to the default noise reduction parameters, and upload the audio data to be processed and the preset data collected by the current terminal to the cloud.
[0137] In one implementation, the processing unit 103 is further configured to: obtain the corresponding data between the noise reduction parameters returned by the cloud and the noise scenario, and update the corresponding relationship between the noise scenario stored locally and the noise reduction parameters according to the corresponding data.
[0138] Regarding the device in the above embodiments, the specific manners in which each module performs operations have been described in detail in the embodiments related to the method, and will not be elaborated herein.
[0139] Figure 14 It is a block diagram of a device 200 for audio noise reduction shown according to an exemplary embodiment. The device 200 may be provided as a terminal. For example, the device 200 may be a mobile phone, a computer, a digital broadcast terminal, a messaging device, a game console, a tablet device, a medical device, a fitness device, a personal digital assistant, etc.
[0140] Refer to Figure 14, Device 200 may include one or more of the following components: a processing component 202, a memory 204, a power component 206, a multimedia component 208, an audio component 210, an input / output (I / O) interface 212, a sensor component 214, and a communication component 216.
[0141] The processing component 202 generally controls the overall operation of the device 200, such as operations associated with display, telephone calls, data communications, camera operations, and recording operations. The processing component 202 may include one or more processors 220 to execute instructions to complete all or part of the steps of the above methods. In addition, the processing component 202 may include one or more modules to facilitate the interaction between the processing component 202 and other components. For example, the processing component 202 may include a multimedia module to facilitate the interaction between the multimedia component 208 and the processing component 202.
[0142] The memory 204 is configured to store various types of data to support the operation of the device 200. Examples of such data include instructions for any application or method operating on the device 200, contact data, phone book data, messages, pictures, videos, etc. The memory 204 may be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, a magnetic disk, or an optical disk.
[0143] The power component 206 provides power to the various components of the device 200. The power component 206 may include a power management system, one or more power sources, and other components associated with generating, managing, and distributing power for the device 200.
[0144] The multimedia component 208 includes a screen that provides an output interface between the device 200 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touches, swipes, and gestures on the touch panel. The touch sensors can not only sense the boundaries of the touch or swipe actions, but also detect the duration and pressure associated with the touch or swipe operations. In some embodiments, the multimedia component 208 includes a front camera and / or a rear camera. When the device 200 is in an operation mode, such as a shooting mode or a video mode, the front camera and / or the rear camera can receive external multimedia data. Each of the front camera and the rear camera may be a fixed optical lens system or have a focal length and optical zoom capabilities.
[0145] The audio component 210 is configured to output and / or input audio signals. For example, the audio component 210 includes a microphone (MIC) that is configured to receive external audio signals when the device 200 is in an operating mode, such as a call mode, a recording mode, and a voice recognition mode. The received audio signals can be further stored in the memory 204 or transmitted via the communication component 216. In some embodiments, the audio component 210 further includes a speaker for outputting audio signals.
[0146] The I / O interface 212 provides an interface between the processing component 202 and a peripheral interface module, which may be a keyboard, a click wheel, buttons, etc. These buttons can include, but are not limited to: a home button, a volume button, a power button, and a lock button.
[0147] The sensor component 214 includes one or more sensors for providing an assessment of the state of the device 200 in various aspects. For example, the sensor component 214 can detect the on / off state of the device 200, the relative positioning of components, such as the display and keypad of the device 200. The sensor component 214 can also detect a change in the position of the device 200 or a component of the device 200, the presence or absence of user contact with the device 200, the orientation or acceleration / deceleration of the device 200, and the temperature change of the device 200. The sensor component 214 can include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor component 214 can also include a light sensor, such as a CMOS or CCD image sensor, for use in imaging applications. In some embodiments, the sensor component 214 can further include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.
[0148] The communication component 216 is configured to facilitate communication between the device 200 and other devices in a wired or wireless manner. The device 200 can access a wireless network based on a communication standard, such as WiFi, 2G, or 3G, or a combination thereof. In an exemplary embodiment, the communication component 216 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 216 further includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology, and other technologies.
[0149] In an exemplary embodiment, the apparatus 200 may be implemented by one or more application specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field programmable gate arrays (FPGAs), controllers, microcontrollers, microprocessors, or other electronic components for performing the above method.
[0150] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as the memory 204 including instructions, and the above instructions can be executed by the processor 220 of the apparatus 200 to complete the above method. For example, the non-transitory computer-readable storage medium may be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.
[0151] It can be understood that "a plurality of" in the present disclosure means two or more, and other quantifiers are similar thereto. "And / or" describes the association relationship of associated objects and indicates that three relationships may exist. For example, A and / or B may indicate: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates an "or" relationship between the associated objects before and after. The singular forms of "a", "the", and "said" are also intended to include the plural forms unless the context clearly indicates otherwise.
[0152] Furthermore, it can be understood that the terms "first", "second", etc. are used to describe various information, but this information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other, and do not indicate a specific order or importance. In fact, the expressions such as "first" and "second" can be used interchangeably. For example, without departing from the scope of the present disclosure, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information.
[0153] Furthermore, it can be understood that the orientation or positional relationship indicated by the terms "center", "longitudinal", "lateral", "front", "rear", "upper", "lower", "left", "right", "vertical", "horizontal", "top", "bottom", "inner", "outer", etc. is based on the orientation or positional relationship shown in the drawings, and is only for the convenience of describing this embodiment and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation.
[0154] Furthermore, it can be understood that unless otherwise specified, "connection" includes direct connection between the two without other components therebetween, and also includes indirect connection between the two with other elements therebetween.
[0155] It can be further understood that although operations are described in a specific order in the drawings in the embodiments of the present disclosure, it should not be construed as requiring these operations to be performed in the specific order shown or in a serial order, or requiring all the operations shown to obtain the desired result. In certain environments, multitasking and parallel processing may be advantageous.
[0156] Those skilled in the art will readily conceive of other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses, or adaptations of this solution, which follow the general principles of the present disclosure and include well-known knowledge or conventional technical means in the technical field not disclosed in the present disclosure.
[0157] It should be understood that the present disclosure is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present disclosure is only limited by the appended claims.
Claims
1. An audio noise reduction method, characterized in that, Including: Obtain the audio to be processed in real time, and determine the current noise scene according to the audio to be processed; Determine the current noise reduction parameter corresponding to the current noise scene according to the correspondence between the noise scene and the noise reduction parameter; Adjust the preset noise reduction model according to the current noise reduction parameter; Obtain the initial audio features of the audio to be processed, use the adjusted preset noise reduction model to perform noise reduction processing on the initial audio features, obtain the target audio features, and obtain the target audio corresponding to the target audio features. The target audio features are the initial audio features after noise reduction processing, and the initial audio features include the initial frequency domain features and initial time domain features of the audio to be processed.
2. The method according to claim 1, characterized in that The determining the current noise scene according to the audio to be processed includes: Determine the current noise scene according to the audio to be processed and preset information, where the preset information includes one or more of the following information: application running information of the terminal within a preset time, low-power component running information of the terminal within a preset time, and current location information of the terminal.
3. The method according to claim 1, wherein The preset noise reduction model includes a time domain noise reduction network and a frequency domain noise reduction network; The using the adjusted preset noise reduction model to perform noise reduction processing on the initial audio features to obtain the target audio features includes: Perform time domain convolution processing on the initial time domain features through the time domain noise reduction network to obtain intermediate time domain features, and perform frequency domain convolution processing on the initial frequency domain features through the frequency domain noise reduction network to obtain intermediate frequency domain features; Process the intermediate time domain features and the intermediate frequency domain features according to the time domain noise reduction network and the frequency domain noise reduction network to obtain the target audio features.
4. The method according to claim 3, wherein The processing the intermediate time domain features and the intermediate frequency domain features according to the time domain convolution network and the frequency domain convolution network to obtain the target audio features includes: Process the intermediate time domain features based on the self-attention mechanism, and fuse the processed intermediate time domain features and the intermediate frequency domain features to obtain an intermediate mixed feature; Perform recursive processing, transposed convolution processing, and predictive modulation mask processing on the intermediate mixed feature successively through the frequency domain noise reduction network to obtain the target frequency domain features; Perform recursive processing on the target frequency domain features and the intermediate time domain features through the time domain noise reduction network to obtain the target audio features.
5. The method according to any one of claims 1, 3, and 4, characterized in that The noise reduction parameter includes the model shape parameters of the preset noise reduction model obtained by training under different noise scenes; The preset noise reduction model is trained in the following manner: Simulate different noise scenes based on a preset audio dataset. Under different noise scenes, process the training audio corresponding to the noise scene through a preset neural network to obtain predicted target frequency domain features and predicted target time domain features. The audio dataset includes a noise dataset and a clean audio dataset, and the preset neural network has the same algorithm structure as the preset noise reduction model; Obtain the loss function of the preset neural network according to the predicted target frequency domain features, the predicted target time domain features, the time domain features of the training audio, and the frequency domain features of the training audio; In response to the loss function satisfying a preset numerical requirement, determine the preset neural network as the preset noise reduction model.
6. The method according to claim 5, wherein The obtaining of the loss function of the preset neural network according to the predicted target frequency domain feature, the predicted target time domain feature, the time domain feature of the training audio, and the frequency domain feature of the training audio includes: Perform an inverse Fourier transform process on the predicted target frequency domain feature to obtain a converted time domain feature corresponding to the predicted target frequency domain feature; Obtain a time domain loss between the predicted target time domain feature and the time domain feature of the training audio, obtain a frequency domain loss between the converted time domain feature and the frequency domain feature of the training audio, and obtain a time-frequency domain hybrid loss between the predicted target time domain feature and the converted time domain feature; Determine the loss function according to the time domain loss, the frequency domain loss, the time-frequency domain hybrid loss, a first weight, a second weight, and a third weight, where the first weight corresponds to the time domain loss, the second weight corresponds to the frequency domain loss, and the third weight corresponds to the time-frequency domain hybrid loss.
7. The method according to claim 6, wherein The determining of the loss function according to the time domain loss, the frequency domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight includes: Determine a first product as the product of the time domain loss and the first weight, determine a second product as the product of the frequency domain loss and the second weight, and determine a third product as the product of the time-frequency domain hybrid loss and the third weight; Determine the sum of the first product, the second product, and the third product as the loss function.
8. The method according to claim 1, characterized in that The method further includes: In response to being unable to determine the current noise scenario based on the audio to be processed, adjust the preset noise reduction model according to default noise reduction parameters, and upload the audio data to be processed and preset data collected by the current terminal to the cloud.
9. The method according to claim 8, characterized in that The method further includes: Obtain the corresponding data between the noise reduction parameters returned by the cloud and the noise scenario, and update the corresponding relationship between the noise scenario and the noise reduction parameters stored locally according to the corresponding data.
10. An audio noise reduction device, characterized in that, including: [[ID=I3]]A determination unit, configured to obtain the audio to be processed in real time, determine the current noise scenario according to the audio to be processed, and determine the current noise reduction parameter corresponding to the current noise scenario according to the corresponding relationship between the noise scenario and the noise reduction parameter; An adjustment unit, configured to adjust the preset noise reduction model according to the current noise reduction parameter; A processing unit, configured to obtain the initial audio feature of the audio to be processed, perform noise reduction processing on the initial audio feature using the adjusted preset noise reduction model to obtain a target audio feature, and obtain a target audio corresponding to the target audio feature, where the target audio feature is the initial audio feature after noise reduction processing, and the initial audio feature includes the initial frequency domain feature and the initial time domain feature of the audio to be processed.
11. The device according to claim 10, characterized in that, The determining unit determines the current noise scenario according to the audio to be processed in the following manner: Determine a current noise scenario according to the audio to be processed and preset information, where the preset information includes one or more of the following: application running information of the terminal within a preset time, low-power component running information of the terminal within a preset time, and current location information of the terminal.
12. The device according to claim 10, characterized in that, The preset noise reduction model includes a time-domain noise reduction network and a frequency-domain noise reduction network; The processing unit uses the adjusted preset noise reduction model to perform noise reduction processing on the initial audio features in the following manner to obtain target audio features, including: Perform time-domain convolution processing on the initial time-domain features through the time-domain noise reduction network to obtain intermediate time-domain features, and perform frequency-domain convolution processing on the initial frequency-domain features through the frequency-domain noise reduction network to obtain intermediate frequency-domain features; According to the time-domain noise reduction network and the frequency-domain noise reduction network, process the intermediate time-domain features and the intermediate frequency-domain features to obtain the target audio features.
13. The device according to claim 12, wherein The processing unit uses the following method to process the intermediate time-domain features and the intermediate frequency-domain features according to the time-domain convolution network and the frequency-domain convolution network to obtain the target audio features: Process the intermediate time-domain features based on the self-attention mechanism, and fuse the processed intermediate time-domain features and the intermediate frequency-domain features to obtain intermediate hybrid features; Perform recursive processing, transposed convolution processing, and predictive modulation mask processing on the intermediate hybrid features sequentially through the frequency-domain noise reduction network to obtain target frequency-domain features; Perform recursive processing on the target frequency-domain features and the intermediate time-domain features through the time-domain noise reduction network to obtain the target audio features.
14. The device according to any one of claims 10, 12, and 13, characterized in that, The noise reduction parameters include model form parameters of a preset noise reduction model obtained by training under different noise scenarios; The preset noise reduction model is trained by the processing unit in the following manner: Simulate different noise scenarios based on a preset audio dataset. Under different noise scenarios, process training audio corresponding to the noise scenarios through a preset neural network to obtain predicted target frequency-domain features and predicted target time-domain features. The audio dataset includes a noise dataset and a clean audio dataset, and the preset neural network has the same algorithm structure as the preset noise reduction model; Obtain a loss function of the preset neural network according to the predicted target frequency-domain features, the predicted target time-domain features, the time-domain features of the training audio, and the frequency-domain features of the training audio; In response to the loss function satisfying a preset numerical requirement, determine the preset neural network as the preset noise reduction model.
15. The device according to claim 14, characterized in that, The processing unit uses the following method to obtain the loss function of the preset neural network according to the predicted target frequency-domain features, the predicted target time-domain features, the time-domain features of the training audio, and the frequency-domain features of the training audio: Perform inverse Fourier transform processing on the predicted target frequency-domain features to obtain a converted time-domain feature corresponding to the predicted target frequency-domain features; Obtain the time-domain loss between the time-domain features of the prediction target and the time-domain features of the training audio, obtain the frequency-domain loss between the converted time-domain features and the frequency-domain features of the training audio, and obtain the time-frequency domain hybrid loss between the time-domain features of the prediction target and the converted time-domain features; Determine the loss function according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight, where the first weight corresponds to the time-domain loss, the second weight corresponds to the frequency-domain loss, and the third weight corresponds to the time-frequency domain hybrid loss.
16. The device according to claim 15, wherein, The processing unit determines the loss function according to the time-domain loss, the frequency-domain loss, the time-frequency domain hybrid loss, the first weight, the second weight, and the third weight in the following manner: Determine the first product as the product of the time-domain loss and the first weight, determine the second product as the product of the frequency-domain loss and the second weight, and determine the third product as the product of the time-frequency domain hybrid loss and the third weight; Determine the sum of the first product, the second product, and the third product as the loss function.
17. The device according to claim 10, wherein The processing unit is further configured to: In response to being unable to determine the current noise scenario based on the audio to be processed, adjust the preset noise reduction model according to the default noise reduction parameters, and upload the audio data to be processed and the preset data collected by the current terminal to the cloud.
18. The device according to claim 17, characterized in that, The processing unit is further configured to: Obtain the corresponding data between the noise reduction parameters returned by the cloud and the noise scenario, and update the corresponding relationship between the locally stored noise scenario and the noise reduction parameters according to the corresponding data.
19. An audio noise reduction device, characterized in that, Comprising: A processor: A memory for storing instructions executable by the processor; Wherein, the processor is configured to: execute the audio noise reduction method according to any one of claims 1 to 9.
20. A storage medium, characterized in that, Instructions are stored in the storage medium, and when the instructions in the storage medium are executed by the processor, the processor can execute the audio noise reduction method according to any one of claims 1 to 9.