An audio noise reduction method, device, equipment and storage medium
By combining real and complex network models to estimate the amplitude and phase spectrum of audio data, the problem of existing audio noise reduction tools poorly suppressing non-stationary noise is solved, and the audio quality is improved.
Patent Information
- Application Number
- CN202111124158.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-24
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2041-09-24
AI Technical Summary
The existing audio noise reduction tools are not effective, especially for non-stationary noise suppression, resulting in poor audio quality.
The amplitude time-frequency masking and complex time-frequency masking of audio data are estimated respectively by using the preset real-number network model and the preset complex network model. Combining the first-order enhanced amplitude spectrum and complex time-frequency masking, the noise reduction result audio data is determined.
It effectively improves the sound quality of the audio, especially the suppression effect of non-stationary noise, and improves the user experience.
Smart Images

Figure CN115862649B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing, and in particular, to an audio noise reduction method, apparatus, device, and storage medium. Background Art
[0002] During the process of audio recording, due to reasons such as the environment or equipment, there is often noise in the recorded audio, resulting in a poor user experience of the audio.
[0003] Currently, there are very few tools for audio noise reduction, and the noise reduction effects of the only few available tools are not satisfactory.
[0004] Therefore, how to achieve audio noise reduction and thus improve the audio quality is a technical problem that needs to be solved urgently at present. Summary of the Invention
[0005] In order to solve the above technical problems or at least partially solve the above technical problems, embodiments of the present disclosure provide an audio noise reduction method, which can achieve audio noise reduction and thus better improve the audio quality.
[0006] In a first aspect, the present disclosure provides an audio noise reduction method, the method comprising:
[0007] Obtain audio data to be noise-reduced;
[0008] Estimate the magnitude time-frequency mask of the audio data to be noise-reduced by using a preset real number network model; wherein, the magnitude time-frequency mask is used to determine the first-order enhanced magnitude spectrum corresponding to the audio data to be noise-reduced;
[0009] Estimate the complex time-frequency mask of the audio data to be noise-reduced by using a preset complex number network model;
[0010] Determine the noise-reduced result audio data corresponding to the audio data to be noise-reduced based on the first-order enhanced magnitude spectrum corresponding to the audio data to be noise-reduced and the complex time-frequency mask.
[0011] In an optional implementation manner, the estimating the complex time-frequency mask of the audio data to be noise-reduced by using a preset complex number network model includes:
[0012] Determine the complex spectrum to be noise-reduced; wherein, the complex spectrum to be noise-reduced includes a complex spectrum determined based on the first-order enhanced magnitude spectrum corresponding to the audio data to be noise-reduced and the original phase spectrum of the audio data to be noise-reduced, or a complex spectrum determined based on the original spectrum and the original phase spectrum of the audio data to be noise-reduced;
[0013] Input the complex spectrum to be denoised into a preset complex network model, and after being processed by the preset complex network model, output the complex time-frequency mask corresponding to the audio data to be denoised.
[0014] In an alternative embodiment, determining the denoised result audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and the complex time-frequency mask includes:
[0015] Determine the amplitude gain and phase gain based on the complex time-frequency mask;
[0016] Determine the phase enhancement spectrum corresponding to the audio data to be denoised based on the phase gain and the original phase spectrum corresponding to the audio data to be denoised;
[0017] And, determine the second-order enhanced amplitude spectrum corresponding to the audio data to be denoised based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised;
[0018] Determine the denoised result audio data corresponding to the audio data to be denoised based on the second-order enhanced amplitude spectrum and the phase enhancement spectrum.
[0019] In an alternative embodiment, the preset real network model and the preset complex network model are used to form a two-stage time-domain convolutional network (TCN) model.
[0020] In an alternative embodiment, before using the preset real network model to estimate the amplitude time-frequency mask of the audio data to be denoised, it further includes:
[0021] Use audio training samples with a sampling rate higher than a preset sampling rate threshold to train the two-stage TCN model.
[0022] In an alternative embodiment, before using audio training samples with a sampling rate higher than a preset sampling rate threshold to train the two-stage TCN model, it further includes:
[0023] Perform preset data augmentation processing on the audio training samples to obtain augmented audio training samples;
[0024] Correspondingly, using audio training samples with a sampling rate higher than a preset sampling rate threshold to train the two-stage TCN model includes:
[0025] Use the augmented audio training samples to train the two-stage TCN model; wherein, the sampling rate of the augmented audio training samples is higher than the preset sampling rate threshold.
[0026] In a second aspect, the present disclosure provides an audio denoising device, the device includes:
[0027] An acquisition module, configured to acquire audio data to be denoised;
[0028] A first estimation module, configured to estimate the magnitude time-frequency mask of the audio data to be denoised by using a preset real number network model; wherein, the magnitude time-frequency mask is used to determine the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised;
[0029] A second estimation module, configured to estimate the complex time-frequency mask of the audio data to be denoised by using a preset complex number network model;
[0030] A first determination module, configured to determine the denoised result audio data corresponding to the audio data to be denoised based on the first-order enhanced magnitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised.
[0031] In a third aspect, the present disclosure provides a computer-readable storage medium, in which instructions are stored, and when the instructions run on a terminal device, the terminal device is enabled to implement the above method.
[0032] In a fourth aspect, the present disclosure provides a device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the processor executes the computer program, the above method is implemented.
[0033] In a fifth aspect, the present disclosure provides a computer program product, which includes a computer program / instructions, and when the computer program / instructions are executed by a processor, the above method is implemented.
[0034] The technical solution provided by the embodiments of the present disclosure has at least the following advantages compared with the prior art:
[0035] The embodiments of the present disclosure provide an audio denoising method. First, audio data to be denoised is acquired, and then the magnitude time-frequency mask of the audio data to be denoised is estimated by using a preset real number network model, and the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised can be obtained. Furthermore, the complex time-frequency mask of the audio data to be denoised is estimated by using a preset complex number network model, and in combination with the first-order enhanced magnitude spectrum and the complex time-frequency mask, the denoised result audio data corresponding to the audio data to be denoised is determined. The embodiments of the present disclosure enhance the magnitude spectrum of the audio data to be denoised by using a preset real number network model, and enhance both the magnitude spectrum and the phase spectrum of the audio data to be denoised by using a preset complex number network model. It can be seen that the embodiments of the present disclosure can implement the denoising process of the audio data to be denoised, thereby better improving the audio quality. Description of the Drawings
[0036] The accompanying drawings herein are incorporated into and constitute a part of this specification, showing embodiments consistent with the present disclosure and, together with the specification, used to explain the principles of the present disclosure.
[0037] To more clearly illustrate the technical solutions in the embodiments of the present disclosure or in the prior art, the following will briefly introduce the accompanying drawings required for use in the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0038] Figure 1 It is a flowchart of an audio noise reduction method provided by an embodiment of the present disclosure;
[0039] Figure 2 It is a schematic diagram of a two-stage TCN model provided by an embodiment of the present disclosure;
[0040] Figure 3 It is a schematic structural diagram of an audio noise reduction device provided by an embodiment of the present disclosure;
[0041] Figure 4 It is a schematic structural diagram of an audio noise reduction device provided by an embodiment of the present disclosure. Detailed implementation manners
[0042] In order to more clearly understand the above objects, features, and advantages of the present disclosure, the following will further describe the solutions of the present disclosure. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.
[0043] In the following description, many specific details are set forth to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the specification are only a part of the embodiments of the present disclosure, rather than all the embodiments.
[0044] Due to reasons such as the recording environment or equipment, there may be noise in the recorded audio, resulting in poor audio quality and affecting the user experience. Among them, the noise in the audio can be divided into at least two types: stationary noise and non-stationary noise. Stationary noise refers to the noise whose statistical characteristics do not change with time, and common ones include white noise and pink noise, etc.; non-stationary noise refers to the noise whose statistical characteristics change with time, and common ones such as keyboard sounds and mouse click sounds.
[0045] Currently, tools for reducing noise in audio can be implemented using artificial intelligence AI noise reduction models. However, currently, the AI noise reduction models usually have a good inhibitory effect on stationary noise, but a weak inhibitory effect on non-stationary noise, resulting in the current noise reduction tools not being able to guarantee the user experience in terms of the noise reduction effect on audio.
[0046] In practical applications, audio noise reduction tools often use a single network model to achieve audio noise reduction. Although the complexity of the network model is relatively low, it is difficult to ensure the audio noise reduction effect. For example, it is particularly difficult to ensure the suppression effect on non-stationary noise in the audio. Therefore, the embodiments of the present disclosure provide an audio noise reduction method, which uses a preset real number network model and a preset complex number network model to perform noise reduction processing on the audio data to be denoised respectively, and then determines the denoised result audio data corresponding to the audio data to be denoised by combining the noise reduction results of the two. It can be seen that compared with using a single network model to perform audio noise reduction, the embodiments of the present disclosure can have a better suppression effect on non-stationary noise, thereby ensuring the overall audio noise reduction effect and further improving the audio quality better.
[0047] Specifically, the embodiments of the present disclosure obtain the audio data to be denoised, and then use the preset real number network model to estimate the magnitude time-frequency mask of the audio data to be denoised, and a first-order enhanced magnitude spectrum corresponding to the audio data to be denoised can be obtained. Furthermore, use the preset complex number network model to estimate the complex time-frequency mask of the audio data to be denoised, and combine the first-order enhanced magnitude spectrum and the complex time-frequency mask to determine the denoised result audio data corresponding to the audio data to be denoised.
[0048] The embodiments of the present disclosure use the preset real number network model to enhance the magnitude spectrum of the audio data to be denoised, and use the preset complex number network model to enhance both the magnitude spectrum and the phase spectrum of the audio data to be denoised at the same time. It can be seen that the embodiments of the present disclosure can achieve noise reduction processing on the audio data to be denoised, while ensuring the noise reduction effect, and further improving the audio quality better.
[0049] Based on this, the embodiments of the present disclosure provide an audio noise reduction method, refer to Figure 1 , which is a flowchart of an audio noise reduction method provided by the embodiments of the present disclosure. The method includes:
[0050] S101: Obtain the audio data to be denoised.
[0051] The audio data to be denoised in the embodiments of the present disclosure can be any audio segment. Among them, the audio segment can also be an audio segment extracted from a video, etc. The embodiments of the present disclosure do not limit the audio data to be denoised.
[0052] In practical applications, the embodiments of the present disclosure can perform real-time noise reduction processing on the audio data to be denoised during the audio recording stage, or perform noise reduction processing on the audio data to be denoised during the audio editing stage. The embodiments of the present disclosure do not limit the noise reduction scenario.
[0053] S102: Use the preset real number network model to estimate the magnitude time-frequency mask of the audio data to be denoised.
[0054] Among them, the amplitude time-frequency mask is used to determine the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised.
[0055] In the embodiments of the present disclosure, first, a preset real-number network model is trained using audio training samples to obtain a trained preset real-number network model for performing amplitude enhancement processing on the audio data to be denoised. Among them, the preset real-number network model can be implemented based on any AI model. For example, the preset real-number network model can be implemented by a Temporal Convolutional Network (TCN), or can be implemented by a Recurrent Neural Network (RNN), etc.
[0056] In the embodiments of the present disclosure, after the preset real-number network model is trained, the audio data to be denoised can be input into the preset real-number network model for processing, and the amplitude time-frequency mask of the audio data to be denoised is output by the preset real-number network model. Among them, the amplitude time-frequency mask is used to represent the proportional relationship between the enhanced amplitude spectrum and the original amplitude spectrum.
[0057] After the preset real-number network model enhances the amplitude of the spectrum of the audio data to be denoised, the amplitude time-frequency mask of the audio data to be denoised is obtained. Then, by multiplying the amplitude time-frequency mask by the original amplitude spectrum of the audio data to be denoised, the enhanced amplitude spectrum of the audio data to be denoised is obtained as the first-order enhanced amplitude spectrum. Among them, the first-order enhanced amplitude spectrum is the amplitude spectrum after the spectrum of the audio data to be denoised is amplitude-enhanced by the preset real-number network model.
[0058] S103: Estimate the complex time-frequency mask of the audio data to be denoised using a preset complex-number network model.
[0059] In the embodiments of the present disclosure, first, a preset complex-number network model is trained using audio training data to obtain a trained preset complex-number network model for simultaneously enhancing the amplitude and phase of the audio data to be denoised. Among them, the preset complex-number network model can be implemented based on any AI model. For example, the preset complex-number network model can be implemented by a Temporal Convolutional Network (TCN), or can be implemented by a Recurrent Neural Network (RNN), etc.
[0060] In an alternative embodiment, before denoising the audio data to be denoised using a trained preset complex network model, first, the complex spectrum determined based on the original spectrum and the original phase spectrum of the audio data to be denoised is determined as the complex spectrum to be denoised. Then, the complex spectrum to be denoised is input into the preset complex network model for processing, and the complex time-frequency mask corresponding to the audio data to be denoised is output by the preset complex network model. Among them, the complex time-frequency mask is used to represent the proportional relationship between the enhanced spectrum and the original spectrum, and the complex time-frequency mask includes a real part and an imaginary part.
[0061] To improve the denoising effect, in the embodiments of the present disclosure, the complex spectrum determined based on the first-order enhanced amplitude spectrum and the original phase spectrum corresponding to the audio data to be denoised can also be determined as the complex spectrum to be denoised, so that the preset complex network model can further enhance the amplitude and phase of the spectrum of the audio data to be denoised on the basis of the denoising by the preset real network model, thereby further improving the denoising effect.
[0062] Specifically, in an alternative embodiment, first, the original phase spectrum of the audio data to be denoised is obtained, and then the spectrum determined based on the first-order enhanced amplitude spectrum and the original phase spectrum corresponding to the audio data to be denoised is determined as the complex spectrum to be denoised. Furthermore, the complex spectrum to be denoised is input into the preset complex network model for processing, and the complex time-frequency mask corresponding to the audio data to be denoised is output by the preset complex network model.
[0063] S104: Determine the denoised result audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and the complex time-frequency mask.
[0064] In the embodiments of the present disclosure, after the amplitude of the audio data to be denoised is enhanced by the preset real network model and the amplitude and phase of the audio data to be denoised are simultaneously enhanced by the preset complex network model, the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised are obtained respectively. Then, based on the first-order enhanced amplitude spectrum and the complex time-frequency mask corresponding to the audio data to be denoised, the denoised result audio data corresponding to the audio data to be denoised is determined, realizing the denoising process of the audio data to be denoised.
[0065] In an alternative embodiment, first, the amplitude gain and the phase gain are determined based on complex time-frequency masking. The amplitude gain is used to characterize the amplitude enhancement of the spectrum of the audio data to be denoised by a preset complex network model, and the phase gain is used to characterize the phase enhancement of the spectrum of the audio data to be denoised by the preset complex network model. Then, based on the phase gain and the original phase spectrum corresponding to the audio data to be denoised, the phase enhancement spectrum corresponding to the audio data to be denoised is determined. Also, based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised, the second-order enhanced amplitude spectrum corresponding to the audio data to be denoised is determined. The second-order enhanced amplitude spectrum is the amplitude spectrum obtained by amplitude enhancement of the audio data to be denoised through a preset real network model and a preset complex network model. Further, based on the second-order enhanced amplitude spectrum and the phase enhancement spectrum, the enhanced spectrum corresponding to the audio data to be denoised is determined, and based on this enhanced spectrum, the denoised result audio data corresponding to the audio data to be denoised is determined.
[0066] In practical applications, the amplitude gain and the phase gain can be calculated respectively using formulas (1) and (2). The following are formulas (1) and (2):
[0067]
[0068]
[0069] Among them, is used to represent the amplitude gain, is used to represent the real part of the complex time-frequency masking, is used to represent the imaginary part of the complex time-frequency masking, is used to represent the phase gain;
[0070] In addition, the enhanced spectrum corresponding to the audio data to be denoised can be calculated using formula (3). The following is formula (3):
[0071]
[0072] Among them, is used to represent the enhanced spectrum, Y phase is used to represent the original phase spectrum, is used to represent the phase enhancement spectrum, is used to represent the first-order enhanced amplitude spectrum, is used to represent the second-order enhanced amplitude spectrum.
[0073] After obtaining the enhanced spectrum corresponding to the audio data to be denoised, through processing such as inverse Fourier transform, the denoised result audio data corresponding to the audio data to be denoised is obtained.
[0074] It can be seen that in the audio noise reduction method provided by the embodiments of the present disclosure, first, the audio data to be noise-reduced is obtained, and then the amplitude time-frequency mask of the audio data to be noise-reduced is estimated by using a preset real number network model, and the first-order enhanced amplitude spectrum corresponding to the audio data to be noise-reduced can be obtained. Furthermore, the complex time-frequency mask of the audio data to be noise-reduced is estimated by using a preset complex number network model, and in combination with the first-order enhanced amplitude spectrum and the complex time-frequency mask, the noise-reduced result audio data corresponding to the audio data to be noise-reduced is determined. The embodiments of the present disclosure use a preset real number network model to enhance the amplitude spectrum of the audio data to be noise-reduced, and use a preset complex number network model to enhance both the amplitude spectrum and the phase spectrum of the audio data to be noise-reduced at the same time. It can be seen that the embodiments of the present disclosure can implement the noise reduction process of the audio data to be noise-reduced, thereby better improving the sound quality of the audio.
[0075] Since the TCN model has better effects in the field of audio noise reduction compared with other network models, therefore, the embodiments of the present disclosure can implement the preset real number network model and the preset complex number network model based on the TCN model. In addition, in order to further improve the effect of audio noise reduction, the embodiments of the present disclosure can use the two-stage time-domain convolutional network TCN model to perform noise reduction processing on the audio, thereby greatly improving the sound quality of the audio.
[0076] Reference Figure 2 , is a schematic diagram of a two-stage TCN model provided by the embodiments of the present disclosure. Among them, the two-stage TCN model includes a real number TCN model and a complex number TCN model, and Y(n) is used to represent the audio data to be noise-reduced.
[0077] In practical applications, after obtaining Y(n), the short-time Fourier transform STFT and Log|.| are successively performed on Y(n), and the processing results are input into the real number TCN model. After being processed by the real number TCN model, the amplitude time-frequency mask corresponding to Y(n) is output; then the original amplitude spectrum of Y(n) is obtained, and the product of the original amplitude spectrum and the amplitude time-frequency mask is calculated as the first-order enhanced amplitude spectrum corresponding to Y(n). Based on the first-order enhanced amplitude spectrum and the original phase spectrum Y phase The complex number frequency spectrum to be noise-reduced is determined and input into the complex number TCN model. After being processed by the complex number TCN model, the complex time-frequency mask corresponding to Y(n) is output. Among them, the complex time-frequency mask corresponding to Y(n) includes a real part and an imaginary part
[0078] It should be noted that the model architectures, parameters, etc. for implementing the real number TCN model and the complex number TCN model are not limited in the embodiments of the present disclosure.
[0079] In practical applications, before denoising audio using the two-stage TCN model, the two-stage TCN model is first trained. Specifically, audio training data with a sampling rate higher than a preset sampling rate threshold can be used to train the two-stage TCN model so that the trained two-stage TCN model can have a better denoising effect on audio data with a higher sampling rate. Among them, the preset sampling rate threshold can be a value greater than 16K.
[0080] In an alternative implementation, the time-domain loss function SISNR can be used to train the two-stage TCN model. Among them, no further introduction is made to the time-domain loss function SISNR here.
[0081] In addition, in order to improve the robustness of the two-stage TCN model, before training the two-stage TCN model, preset data augmentation processing can be performed on the audio training samples to enrich the diversity of the audio training samples.
[0082] Among them, the preset data augmentation processing can include performing high-pass, low-pass, band-pass, setting different volumes, and / or equalization and other processing operations on the audio training samples with a certain probability.
[0083] In practical applications, after performing preset data augmentation processing on the audio training samples to obtain augmented audio training samples, the augmented audio training samples can be used to train the two-stage TCN model.
[0084] In an alternative implementation, the sampling rate of the augmented audio training samples can be higher than the preset sampling rate threshold to ensure the robustness of the two-stage TCN model for denoising high-sampling-rate audio data.
[0085] The audio denoising method provided by the embodiments of the present disclosure can use the two-stage TCN model to achieve audio denoising, especially has a good suppression effect on non-stationary noise in the audio, further improves the denoising effect, improves the audio quality, and enhances the user experience.
[0086] Based on the above method embodiments, the present disclosure also provides an audio denoising device. Refer to Figure 3 , which is a schematic structural diagram of an audio denoising device provided by the embodiments of the present disclosure. The device includes:
[0087] An acquisition module 301, configured to acquire audio data to be denoised;
[0088] A first estimation module 302, configured to estimate the magnitude time-frequency mask of the audio data to be denoised using a preset real number network model; wherein, the magnitude time-frequency mask is used to determine the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised.
[0089] A second estimation module 303, configured to estimate the complex time-frequency mask of the audio data to be denoised by using a preset complex network model;
[0090] A determination module 304, configured to determine the denoised result audio data corresponding to the audio data to be denoised based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and the complex time-frequency mask.
[0091] In an optional implementation manner, the second estimation module includes:
[0092] A first determination sub-module, configured to determine the complex spectrum to be denoised; wherein, the complex spectrum to be denoised includes a complex spectrum determined based on the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and the original phase spectrum of the audio data to be denoised, or a complex spectrum determined based on the original spectrum and the original phase spectrum of the audio data to be denoised;
[0093] A first processing sub-module, configured to input the complex spectrum to be denoised into a preset complex network model, and after being processed by the preset complex network model, output the complex time-frequency mask corresponding to the audio data to be denoised.
[0094] In an optional implementation manner, the determination module includes:
[0095] A second determination sub-module, configured to determine an amplitude gain and a phase gain based on the complex time-frequency mask;
[0096] A third determination sub-module, configured to determine the phase enhanced spectrum corresponding to the audio data to be denoised based on the phase gain and the original phase spectrum corresponding to the audio data to be denoised;
[0097] A fourth determination sub-module, configured to determine the second-order enhanced amplitude spectrum corresponding to the audio data to be denoised based on the amplitude gain and the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised;
[0098] A fifth determination sub-module, configured to determine the denoised result audio data corresponding to the audio data to be denoised based on the second-order enhanced amplitude spectrum and the phase enhanced spectrum.
[0099] In an optional implementation manner, the preset real network model and the preset complex network model are used to form a two-stage time-domain convolutional network (TCN) model.
[0100] In an optional implementation manner, the apparatus further includes:
[0101] A training module, configured to train the two-stage TCN model by using audio training samples with a sampling rate higher than a preset sampling rate threshold.
[0102] In an alternative embodiment, the apparatus further comprises:
[0103] An augmentation module for performing preset data augmentation processing on the audio training samples to obtain augmented audio training samples;
[0104] Correspondingly, the training module is specifically configured to:
[0105] Train the two-stage TCN model by using the augmented audio training samples; wherein, the sampling rate of the augmented audio training samples is higher than a preset sampling rate threshold.
[0106] In the audio noise reduction apparatus provided by the embodiments of the present disclosure, first, the audio data to be denoised is obtained, and then when the preset real number network model is used to estimate the amplitude time-frequency mask of the audio data to be denoised, the first-order enhanced amplitude spectrum corresponding to the audio data to be denoised can be obtained. Furthermore, the complex time-frequency mask of the audio data to be denoised is estimated by using a preset complex number network model, and in combination with the first-order enhanced amplitude spectrum and the complex time-frequency mask, the denoised result audio data corresponding to the audio data to be denoised is determined. The embodiments of the present disclosure use a preset real number network model to enhance the amplitude spectrum of the audio data to be denoised, and use a preset complex number network model to enhance both the amplitude spectrum and the phase spectrum of the audio data to be denoised at the same time. It can be seen that the embodiments of the present disclosure can implement the noise reduction processing of the audio data to be denoised, thereby better improving the audio quality.
[0107] In addition to the above methods and apparatuses, the embodiments of the present disclosure further provide a computer-readable storage medium. Instructions are stored in the computer-readable storage medium, and when the instructions are run on a terminal device, the terminal device implements the audio noise reduction method described in the embodiments of the present disclosure.
[0108] The embodiments of the present disclosure further provide a computer program product. The computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the audio noise reduction method described in the embodiments of the present disclosure is implemented.
[0109] In addition, the embodiments of the present disclosure further provide an audio noise reduction device. As shown in Figure 4 it may include:
[0110] A processor 401, a memory 402, an input device 403, and an output device 404. The number of processors 401 in the audio noise reduction device may be one or more. Figure 4 Taking one processor as an example. In some embodiments of the present disclosure, the processor 401, the memory 402, the input device 43, and the output device 404 may be connected through a bus or other means. Among them, Figure 4 taking the connection through a bus as an example.
[0111] The memory 402 can be used to store software programs and modules. The processor 401 executes various functional applications and data processing of the audio noise reduction device by running the software programs and modules stored in the memory 402. The memory 402 may mainly include a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function, etc. In addition, the memory 402 may include high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices. The input device 403 can be used to receive input digital or character information, and generate signal inputs related to the user settings and function controls of the audio noise reduction device.
[0112] Specifically, in this embodiment, the processor 401 loads the executable files corresponding to the processes of one or more application programs into the memory 402 according to the following instructions, and the processor 401 runs the application programs stored in the memory 402 to implement various functions of the above audio noise reduction device.
[0113] It should be noted that in this article, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or also includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0114] The above are only specific embodiments of the present disclosure, enabling those skilled in the art to understand or implement the present disclosure. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present disclosure. Therefore, the present disclosure will not be limited to these embodiments described herein, but rather to the broadest scope consistent with the principles and novel features disclosed herein.
Claims
1. An audio noise reduction method, characterized in that The method includes: Obtaining the audio data to be denoised; Estimating the magnitude time-frequency mask of the audio data to be denoised by using a preset real number network model; wherein, the magnitude time-frequency mask is used to determine the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised; Estimating the complex time-frequency mask of the audio data to be denoised by using a preset complex number network model; Determining the denoised result audio data corresponding to the audio data to be denoised based on the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised and the complex time-frequency mask.
2. The method according to claim 1, characterized in that, The estimating the complex time-frequency mask of the audio data to be denoised by using a preset complex number network model includes: Determining the complex spectrum to be denoised; wherein, the complex spectrum to be denoised includes a complex spectrum determined based on the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised and the original phase spectrum of the audio data to be denoised, or a complex spectrum determined based on the original spectrum and the original phase spectrum of the audio data to be denoised; Inputting the complex spectrum to be denoised into the preset complex number network model, and after being processed by the preset complex number network model, outputting the complex time-frequency mask corresponding to the audio data to be denoised.
3. The method according to claim 1 or 2, characterized in that, The determining the denoised result audio data corresponding to the audio data to be denoised based on the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised and the complex time-frequency mask includes: Determining the magnitude gain and the phase gain based on the complex time-frequency mask; Determining the phase enhanced spectrum corresponding to the audio data to be denoised based on the phase gain and the original phase spectrum corresponding to the audio data to be denoised; And determining the second-order enhanced magnitude spectrum corresponding to the audio data to be denoised based on the magnitude gain and the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised; Determining the denoised result audio data corresponding to the audio data to be denoised based on the second-order enhanced magnitude spectrum and the phase enhanced spectrum.
4. The method according to claim 1, wherein The preset real number network model and the preset complex number network model are used to form a two-stage time-domain convolutional network TCN model.
5. The method according to claim 4, wherein Before estimating the magnitude time-frequency mask of the audio data to be denoised by using the preset real number network model, it further includes: Training the two-stage TCN model by using audio training samples with a sampling rate higher than a preset sampling rate threshold.
6. The method according to claim 5, wherein Before training the two-stage TCN model by using audio training samples with a sampling rate higher than a preset sampling rate threshold, it further includes: Performing preset data augmentation processing on the audio training samples to obtain augmented audio training samples; Correspondingly, the training the two-stage TCN model by using audio training samples with a sampling rate higher than a preset sampling rate threshold includes: Training the two-stage TCN model by using the augmented audio training samples; wherein, the sampling rate of the augmented audio training samples is higher than the preset sampling rate threshold.
7. An audio noise reduction device, characterized in that, The device includes: An obtaining module, configured to obtain the audio data to be denoised; A first estimation module, configured to estimate the magnitude time-frequency mask of the audio data to be denoised by using a preset real number network model; wherein, the magnitude time-frequency mask is used to determine the first-order enhanced magnitude spectrum corresponding to the audio data to be denoised; A second estimation module, configured to estimate a complex time-frequency mask of the audio data to be denoised by using a preset complex network model; A determination module, configured to determine denoised result audio data corresponding to the audio data to be denoised based on a first-order enhanced amplitude spectrum corresponding to the audio data to be denoised and the complex time-frequency mask.
8. A computer-readable storage medium, characterized in that, Instructions are stored in the computer-readable storage medium, and when the instructions run on a terminal device, the terminal device implements the method according to any one of claims 1-6.
9. An audio noise reduction device, characterized in that, Comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein when the processor executes the computer program, the method according to any one of claims 1-6 is implemented.
10. A computer program product, characterized in that, The computer program product includes computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method according to any one of claims 1-6 is implemented.
Citation Information
Cited By
Audio denoising method and device, apparatus and storage medium
US12701359B2