An audio noise reduction method, device, equipment, storage medium and product

Through multi-stage noise reduction processing and scene-adaptive algorithms, the problem of background music being damaged in live streaming scenarios has been solved, achieving better audio noise reduction effects and user experience.

CN116469402BActive Publication Date: 2026-04-24BIGO TECH PTE LTD
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BIGO TECH PTE LTD
Filing Date
2023-04-23
Publication Date
2026-04-24

AI Technical Summary

Technical Problem

Existing noise reduction algorithms primarily target human voices in live streaming scenarios, resulting in damage to background music, poor noise reduction performance, and negatively impacting user experience.

Method used

The noise reduction process involves multiple stages, including a first noise reduction model to suppress steady-state noise, a second or third noise reduction process depending on the scene type to preserve music or suppress non-steady-state noise, and a combination of phase noise reduction and enhancement processing. Traditional and AI models are used to process different noise types respectively.

Benefits of technology

It effectively improves audio noise reduction in live streaming scenarios, optimizes user experience, reduces resource waste, and retains background music while reducing noise impact in non-music scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116469402B_ABST
    Figure CN116469402B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides an audio noise reduction method, device, equipment, storage medium and product. The technical scheme provided by the embodiment of the application obtains first noise reduction audio information by performing first noise reduction processing on the noise amplitude spectrum of the to-be-processed audio through a first noise reduction model, performs noise reduction processing on the first noise reduction audio according to a scene type to obtain second noise reduction audio information that retains music and suppresses non-steady-state noise, or third noise reduction audio information that suppresses non-steady-state noise, and performs phase noise reduction and enhancement processing on the second noise reduction audio information or the third noise reduction audio information according to the scene type corresponding to the to-be-processed audio to obtain target noise reduction audio, thereby reducing resource waste caused by independent noise reduction of background music and vocals, weakening the influence of noise reduction ability in a non-music scene while retaining background music, effectively improving the audio noise reduction effect in a live scene, and optimizing user experience.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of audio processing technology, and in particular to an audio noise reduction method, apparatus, device, storage medium, and product. Background Technology

[0002] In online live streaming scenarios, many hosts play background music while speaking. At this time, the audio information captured by the microphone will include human voice, background music and noise. It is necessary to suppress noise while preserving human voice and background music to improve the user's viewing experience.

[0003] Since human voices and background music differ significantly in bandwidth, fundamental frequency, and other characteristics, and the primary target of current noise reduction algorithms is human voices, directly applying noise reduction algorithms designed for human voices to live streaming scenarios would damage the background music, resulting in poor audio noise reduction and negatively impacting the user experience. Summary of the Invention

[0004] This application provides an audio noise reduction method, apparatus, device, storage medium, and product to address the technical problem that noise reduction algorithms in related technologies are mainly designed for single human voice scenarios, resulting in poor noise reduction performance in live streaming scenarios and affecting user experience. This application effectively improves noise reduction performance in live streaming scenarios and optimizes user experience.

[0005] In a first aspect, embodiments of this application provide an audio noise reduction method, comprising:

[0006] Obtain the audio to be processed;

[0007] The audio to be processed is sent to the first noise reduction model, and the first noise reduction model performs a first noise reduction process on the noisy amplitude spectrum of the audio to be processed to obtain the first noise reduction audio information that suppresses steady-state noise.

[0008] Based on the scene type corresponding to the audio to be processed, the first noise reduction frequency information is subjected to a second noise reduction process to obtain a second noise reduction frequency information that retains the music and suppresses non-steady-state noise, or the first noise reduction frequency information is subjected to a third noise reduction process to obtain a third noise reduction frequency information that suppresses non-steady-state noise.

[0009] Based on the scene type corresponding to the audio to be processed, phase noise reduction and enhancement processing are performed on the second noise reduction frequency information or the third noise reduction frequency information to obtain the target noise reduction frequency.

[0010] In a second aspect, embodiments of this application provide an audio noise reduction device, including an audio acquisition module, a first noise reduction module, a second noise reduction module, and an audio enhancement module, wherein:

[0011] The audio acquisition module is configured to acquire the audio to be processed;

[0012] The first noise reduction module is configured to send the audio to be processed to a first noise reduction model, and perform a first noise reduction process on the noisy amplitude spectrum of the audio to be processed through the first noise reduction model to obtain first noise-reduced audio information that suppresses steady-state noise.

[0013] The second noise reduction module is configured to perform a second noise reduction process on the first noise reduction frequency information according to the scene type corresponding to the audio to be processed, to obtain a second noise reduction frequency information that retains the music and suppresses non-steady-state noise, or to perform a third noise reduction process on the first noise reduction frequency information to obtain a third noise reduction frequency information that suppresses non-steady-state noise.

[0014] The audio enhancement module is configured to perform phase noise reduction and enhancement processing on the second noise reduction frequency information or the third noise reduction frequency information according to the scene type corresponding to the audio to be processed, so as to obtain the target noise reduction frequency.

[0015] In a third aspect, embodiments of this application provide an audio noise reduction device, including: a memory and one or more processors;

[0016] The memory is used to store one or more programs;

[0017] When the one or more programs are executed by the one or more processors, the one or more processors implement the audio noise reduction method as described in the first aspect.

[0018] In a fourth aspect, embodiments of this application provide a non-volatile storage medium for storing computer-executable instructions, which, when executed by a computer processor, are used to perform the audio noise reduction method as described in the first aspect.

[0019] In a fifth aspect, embodiments of this application provide a computer program product comprising a computer program stored in a computer-readable storage medium, wherein at least one processor of the device reads from the computer-readable storage medium and executes the computer program, causing the device to perform the audio noise reduction method as described in the first aspect.

[0020] This application embodiment uses a first noise reduction model to perform first noise reduction processing on the noisy amplitude spectrum of the audio to be processed, obtaining first noise-reduced frequency information that suppresses steady-state noise. Based on the scene type, the first noise-reduced frequency is further denoised to obtain second noise-reduced frequency information that preserves music and suppresses non-steady-state noise, or third noise-reduced frequency information that suppresses non-steady-state noise. Then, based on the scene type corresponding to the audio to be processed, phase denoising and enhancement processing are performed on the second or third noise-reduced frequency information to obtain the target noise-reduced frequency. Combining scene type, the complex noise reduction task is divided into multiple serially connected, lower-difficulty sub-tasks. Through multi-stage audio noise reduction tasks, different scene types can share the same noise reduction model, reducing resource waste caused by independent noise reduction of background music and human voices. While preserving background music, the impact on the noise reduction capability of non-music scenes is reduced, effectively improving the audio noise reduction effect in live streaming scenarios and optimizing the user experience. Attached Figure Description

[0021] Figure 1 This is a flowchart of an audio noise reduction method provided in an embodiment of this application;

[0022] Figure 2 This is a flowchart of another audio noise reduction method provided in the embodiments of this application;

[0023] Figure 3 This is a schematic diagram of a first noise reduction processing flow provided in an embodiment of this application;

[0024] Figure 4 This is a schematic diagram of a process for phase denoising and enhancement of second noise-reduced frequency information provided in an embodiment of this application;

[0025] Figure 5 This is a schematic diagram of a process for performing phase denoising and enhancement processing on third noise-reducing frequency information, provided in an embodiment of this application.

[0026] Figure 6 This is a schematic diagram of the structure of an audio noise reduction device provided in an embodiment of this application;

[0027] Figure 7 This is a schematic diagram of the structure of an audio noise reduction device provided in an embodiment of this application. Detailed Implementation

[0028] To make the objectives, technical solutions, and advantages of this application clearer, specific embodiments of this application will be described in further detail below with reference to the accompanying drawings. It should be understood that the specific embodiments described herein are merely for explaining this application and not for limiting it. It should also be noted that, for ease of description, only the parts relevant to this application are shown in the drawings, not all of them. Before discussing exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe operations (or steps) as sequential processes, many of these operations can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but additional steps not included in the drawings may also be present. The above processes can correspond to methods, functions, procedures, subroutines, subroutines, etc.

[0029] The audio denoising method provided in this application can be applied to various audio denoising scenarios, such as denoising audio recorded during live streaming. It aims to suppress steady-state and non-steady-state noise in the audio to be processed using multiple denoising models, depending on the scenario type, to obtain damaged audio information (second or third denoised frequency information). Phase denoising and enhancement processing are then applied to the damaged audio information to obtain the target denoised frequency, effectively improving the audio denoising effect and optimizing the user experience. In live streaming, the host speaks while playing background music. The audio captured by the microphone includes human voice, background music, and noise. To improve the audio experience, the audio denoising algorithm needs to suppress noise while preserving human voice and background music. Traditional denoising algorithms use traditional signal processing methods, which can effectively suppress steady-state noise, but their denoising effect is poor in non-steady-state noise scenarios. AI noise reduction algorithms, employing deep learning methods, can generally remove steady-state and non-steady-state noise, excluding human voices. However, human voices and background music differ significantly in bandwidth, fundamental frequency, and other characteristics. Current noise reduction algorithms primarily target human voices for preservation. If a noise reduction algorithm designed for a single human voice scene is directly applied to a scene with background music, the background music will be damaged, resulting in poor audio noise reduction and negatively impacting user experience. Therefore, this application provides an audio noise reduction method to address the technical problem of poor audio noise reduction performance and negative user experience in current audio noise reduction schemes designed for single human voice scenes.

[0030] Figure 1 A flowchart of an audio noise reduction method provided in an embodiment of this application is given. The audio noise reduction method provided in this embodiment of the application can be executed by an audio noise reduction device, which can be implemented by hardware and / or software and integrated into an audio noise reduction device.

[0031] The following description uses an audio noise reduction device to perform an audio noise reduction method as an example. (Reference) Figure 1 The audio noise reduction method includes:

[0032] S110: Obtain the audio to be processed.

[0033] The audio to be processed provided by this solution can be obtained by recording audio through an external microphone device connected to the audio noise reduction device. For example, the broadcaster can record audio through the microphone configured on the live broadcast host device (such as a computer) to obtain the audio to be processed. Alternatively, the audio can be recorded through the microphone module built into the audio noise reduction device. For example, when the broadcaster is broadcasting live through a mobile terminal such as a mobile phone or tablet, the audio to be processed can be recorded through the microphone module configured on the mobile terminal.

[0034] For example, the audio to be processed, which requires noise reduction, is acquired. Optionally, the audio to be processed can be obtained through real-time recording. Optionally, the audio noise reduction device can be connected to an audio speaker (e.g., a speaker) via wired and / or wireless means, and background music can be played through the audio speaker. In one embodiment, when the broadcaster is speaking without background music playing, the recorded audio to be processed generally includes human voice and noise; when the broadcaster is speaking while playing background music, the recorded audio to be processed generally includes human voice, background music, and noise.

[0035] S120: The audio to be processed is sent to the first noise reduction model. The first noise reduction model performs first noise reduction processing on the noisy amplitude spectrum of the audio to be processed to obtain the first noise-reduced audio information that suppresses steady-state noise.

[0036] This solution allows for the configuration of one or more combinations of a first noise reduction model, a second noise reduction model, a third noise reduction model, and an audio enhancement model in the audio noise reduction device. The first noise reduction model provided by this solution can be used to perform initial noise reduction processing on the amplitude spectrum of the audio, suppressing steady-state noise in the amplitude spectrum. Optionally, the first noise reduction model can be a noise reduction model built based on traditional noise reduction algorithms.

[0037] For example, after obtaining the audio to be processed, the audio is sent to a first noise reduction model. The first noise reduction model performs a first noise reduction process on the noisy amplitude spectrum of the audio to be processed to obtain first noise-reduced frequency information that suppresses steady-state noise. In one embodiment, the first noise-reduced frequency information output by the first noise reduction model retains human voices, background music, and non-steady-state noise.

[0038] Optionally, the audio to be processed can first undergo a short-time Fourier transform to obtain a noisy amplitude spectrum, and then send the noisy amplitude spectrum of the audio to be processed to the first denoising model for the first denoising process. Alternatively, the complete audio to be processed can be sent directly to the first denoising model, which will then perform a short-time Fourier transform on the audio to obtain a noisy amplitude spectrum, and then perform the first denoising process on the noisy amplitude spectrum. The first denoised audio information can be the denoised audio (including amplitude spectrum information and phase spectrum information) of the audio to be processed, where steady-state noise in the amplitude spectrum of the audio is suppressed during the denoising process, and the phase spectrum information in the first denoised audio information is not denoised.

[0039] S130: Based on the scene type corresponding to the audio to be processed, perform a second noise reduction process on the first noise reduction frequency information to obtain a second noise reduction frequency information that retains the music and suppresses non-steady-state noise, or perform a third noise reduction process on the first noise reduction frequency information to obtain a third noise reduction frequency information that suppresses non-steady-state noise.

[0040] The scenario types provided in this solution include music type and non-music type. The music type can be understood as a scenario where the broadcaster plays background music through an external audio speaker, while the non-music type can be understood as a scenario where the broadcaster does not play background music through an external audio speaker.

[0041] In one possible embodiment, the audio noise reduction method provided by this solution performs a second noise reduction process on the first noise-reduced frequency information according to the scene type corresponding to the audio to be processed, to obtain a second noise-reduced frequency information that retains music and suppresses non-steady-state noise, or performs a third noise reduction process on the first noise-reduced frequency information to obtain a third noise-reduced frequency information that suppresses non-steady-state noise. Specifically, if the scene type corresponding to the audio to be processed is music, the first noise-reduced frequency information is subjected to a second noise reduction process to obtain a second noise-reduced frequency information that retains music and suppresses non-steady-state noise; if the scene type corresponding to the audio to be processed is not music, the first noise-reduced frequency information is subjected to a third noise reduction process to obtain a third noise-reduced frequency information that suppresses non-steady-state noise.

[0042] For example, after performing a first noise reduction process on the audio to be processed, residual non-steady-state noise may remain in the first noise-reduced audio information. When the scene type of the audio to be processed is music, such as when a broadcaster plays background music during a live stream, the recorded audio to be processed includes human voice, background music, and noise. After performing the first noise reduction process, a second noise reduction process is applied to the first noise-reduced audio information to obtain a second noise-reduced audio information that retains the music and suppresses non-steady-state noise. When the scene type of the audio to be processed is not music, such as when a broadcaster does not play background music during a live stream, the recorded audio to be processed includes human voice and noise, but does not contain background music. After performing the first noise reduction process, a third noise reduction process is applied to the first noise-reduced audio information to obtain a third noise-reduced audio information that suppresses non-steady-state noise. This solution performs the second and third noise reduction processes on the first noise-reduced audio information according to the music scene and the non-music scene respectively, to obtain noise-reduced audio that retains the audio information of the corresponding scene type and suppresses non-steady-state noise. The two types of scenarios share the same model, which effectively avoids the waste of resources caused by separate noise reduction for background music and human voice. Furthermore, while retaining background music, it reduces the impact on the noise reduction capability of non-music scenes, thereby improving the noise reduction effect in live streaming scenarios.

[0043] In one possible embodiment, when performing a second noise reduction process on the first noise-reduced frequency information to obtain a second noise-reduced frequency information that retains the music and suppresses non-steady-state noise, the solution may send the first noise-reduced frequency information to a second noise reduction model, and then perform noise reduction processing on the first noise-reduced frequency information through the second noise reduction model to obtain a second noise-reduced frequency information that retains the music and suppresses non-steady-state noise.

[0044] Optionally, a pre-trained second noise reduction model can be configured in the audio noise reduction device. The second noise reduction model can be built based on a neural network and is trained with audio information containing non-steady-state noise (which may be the audio information after the first noise reduction process by the first noise reduction model) as input and noise-reduced audio information that preserves music and suppresses non-steady-state noise as output.

[0045] For example, when the scene type corresponding to the audio to be processed is music, the first denoised audio signal output by the first denoising model is sent to the second denoising model. The second denoising model then performs denoising processing on the first denoised audio signal information to obtain second denoised audio signal information that preserves the music while suppressing non-stationary noise. The second denoised audio signal information can be used as damaged audio information under the music scene type for subsequent phase denoising and enhancement processing. This solution uses the second denoising model to perform second denoising processing on the first denoised audio signal information, effectively preserving vocals and background music in the audio while suppressing non-stationary noise.

[0046] In one possible embodiment, the solution performs a third noise reduction process on the first noise-reduced frequency information to obtain third noise-reduced frequency information that suppresses non-steady-state noise. This can be achieved by: sending the first noise-reduced frequency information to a second noise reduction model, performing noise reduction processing on the first noise-reduced frequency information through the second noise reduction model, obtaining second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise; and sending the second noise-reduced frequency information to a third noise reduction model, performing noise reduction processing on the second noise-reduced frequency information through the third noise reduction model, obtaining third noise-reduced frequency information that further suppresses non-steady-state noise.

[0047] Optionally, a pre-trained second and third noise reduction model can be configured in the audio noise reduction device. The second noise reduction model can reuse the previously provided second noise reduction model, meaning it can be used in both music and non-music scenarios. The third noise reduction model can be built based on a neural network and trained using audio information containing non-stationary noise (which may be audio information processed by the second noise reduction model) as input and noise-reduced audio information that suppresses non-stationary noise as output.

[0048] For example, when the scene type corresponding to the audio to be processed is not music, the first noise-reduced frequency signal output by the first noise reduction model is sent to the second noise reduction model. The second noise reduction model performs noise reduction processing on the first noise-reduced frequency signal information to obtain second noise-reduced frequency signal information that preserves music and suppresses non-steady-state noise. Further, the second noise-reduced frequency signal information output by the second noise reduction model is sent to the third noise reduction model. The third noise reduction model performs noise reduction processing on the second noise-reduced frequency signal information to obtain third noise-reduced frequency signal information that suppresses non-steady-state noise. The third noise-reduced frequency signal can be used as damaged audio information in non-music scene types for subsequent phase denoising and enhancement processing. This solution uses the second noise reduction model to perform noise reduction processing on the first noise-reduced frequency signal information, and also uses the third noise reduction model to perform noise reduction processing on the second noise-reduced frequency signal information, effectively preserving human voices in the audio while further suppressing non-steady-state noise. At the same time, this solution can share the second noise reduction model in different scene types, preserving background music and improving noise reduction capabilities in non-music scenes while reducing resource waste caused by independent noise reduction of background music and human voices.

[0049] Optionally, after using the second noise reduction model to denoise the first noise-reduced frequency information, it can be determined whether a third noise reduction model is needed to denoise the second noise-reduced frequency information based on the scene type. Alternatively, the scene type can be determined first, and then it can be determined whether to use only the second noise reduction model to denoise the first noise-reduced frequency information, or to use both the second and third noise reduction models together to denoise the first noise-reduced frequency information.

[0050] The second noise reduction model provided in this solution can be used to reduce noise in the amplitude spectrum of audio, suppressing non-steady-state noise in the amplitude spectrum. Optionally, the second and third noise reduction models provided in this solution can be AI noise reduction models, that is, noise reduction models that can suppress non-steady-state noise in the amplitude spectrum are trained through neural network learning.

[0051] For example, after performing a first noise reduction process on the audio to be processed to obtain a first noise-reduced frequency information, the first noise-reduced frequency information is sent to a trained second noise reduction model. The second noise reduction model performs noise reduction processing on the first noise-reduced frequency information to obtain a second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise.

[0052] Optionally, the first noise-reduced audio information can be sent to the second noise reduction model. The second noise reduction model performs a short-time Fourier transform on the first noise-reduced audio information to obtain a noisy amplitude spectrum, and then performs noise reduction processing on the noisy amplitude spectrum. Alternatively, the first noise-reduced audio information can be subjected to a short-time Fourier transform first to obtain a noisy amplitude spectrum, and then the noisy amplitude spectrum is sent to the second noise reduction model for noise reduction processing. The noise reduction processing of the audio information by the first noise reduction model and / or the second noise reduction model can preserve background music and human voice-related features in the audio information while suppressing noise-related features.

[0053] The second noise-reduced audio information can be the audio (including amplitude spectrum information and phase spectrum information) after noise reduction processing of the first noise-reduced audio information. The non-steady-state noise in the amplitude spectrum of the audio is suppressed in the noise reduction processing, while the phase spectrum information is not noise-reduced.

[0054] The third noise reduction model provided in this solution can be used to denoise the amplitude spectrum of audio, further suppressing non-steady-state noise in the amplitude spectrum. For example, after the second noise reduction model denoises the first noise-reduced frequency information to obtain the second noise-reduced frequency information, the second noise-reduced frequency information is sent to the trained third noise reduction model. The third noise reduction model then denoises the second noise-reduced frequency information to obtain third noise-reduced frequency information that further suppresses non-steady-state noise.

[0055] Optionally, the information sent to the third noise reduction model can be the complete second noise-reduced audio information. The third noise reduction model performs a short-time Fourier transform on the second noise-reduced audio information to obtain a noisy amplitude spectrum, and then performs noise reduction processing on the noisy amplitude spectrum. Alternatively, the second noise-reduced audio information can first undergo a short-time Fourier transform to obtain a noisy amplitude spectrum, and then be sent to the third noise reduction model for noise reduction processing. The third noise reduction model's noise reduction processing of the audio information can preserve human voices in the audio and further suppress non-steady-state noise in the audio.

[0056] The third noise-reduced audio information can be the audio (including amplitude spectrum information and phase spectrum information) after noise reduction processing of the second noise-reduced audio information. The non-steady-state noise in the amplitude spectrum of the audio is further suppressed in the noise reduction process, while the phase spectrum information is not noise-reduced.

[0057] S140: Based on the scene type corresponding to the audio to be processed, perform phase noise reduction and enhancement processing on the second or third noise reduction frequency information to obtain the target noise reduction frequency.

[0058] For example, after denoising the first noise-reduced frequency information to obtain the second noise-reduced frequency information or the third noise-reduced frequency information, the scene type corresponding to the audio to be processed is determined, and phase denoising and enhancement processing are performed on the second noise-reduced frequency information or the third noise-reduced frequency information according to the scene type corresponding to the audio to be processed to obtain the target noise-reduced frequency.

[0059] The phase denoising process for the second or third noise-reduced frequency information can be understood as denoising the phase spectrum corresponding to the second or third noise-reduced frequency information. Optionally, the enhancement process for the second noise-reduced frequency information can be the recovery process of the fundamental frequency and harmonics of the amplitude spectrum of the second noise-reduced frequency information, or the enhancement process for the third noise-reduced frequency information can be the recovery process of the fundamental frequency and harmonics of the amplitude spectrum of the third noise-reduced frequency information.

[0060] Optionally, the corresponding phase denoising and enhancement processing methods can be determined for the sounds that need to be preserved and the noise that needs to be suppressed in different scene types. For example, for music, phase denoising and enhancement processing can be performed on the second noise-reduced frequency information to obtain the target noise-reduced frequency, while for non-music types, phase denoising and enhancement processing can be performed on the third noise-reduced frequency information to obtain the target noise-reduced frequency.

[0061] In one possible embodiment, a trained audio enhancement model can be used to perform phase denoising and enhancement processing on the second or third noise-reduced frequency information. Based on this, when this solution performs phase denoising and enhancement processing on the second or third noise-reduced frequency information according to the scene type corresponding to the audio to be processed, to obtain the target noise-reduced frequency, if the scene type corresponding to the audio to be processed is music, the second noise-reduced frequency information is sent to the audio enhancement model, and the audio enhancement model performs phase denoising and enhancement processing on the second noise-reduced frequency information to obtain the target noise-reduced frequency; if the scene type corresponding to the audio to be processed is not music, the third noise-reduced frequency information is sent to the audio enhancement model, and the audio enhancement model performs phase denoising and enhancement processing on the third noise-reduced frequency information to obtain the target noise-reduced frequency.

[0062] The above describes a process where a first noise reduction model is used to perform noise reduction on the noisy amplitude spectrum of the audio to be processed, resulting in first noise-reduced frequency information that suppresses steady-state noise. Depending on the scene type, the first noise-reduced frequency information is further denoised to obtain second noise-reduced frequency information that preserves music while suppressing non-steady-state noise, or third noise-reduced frequency information that suppresses non-steady-state noise. Then, based on the scene type corresponding to the audio to be processed, phase denoising and enhancement processing are performed on the second or third noise-reduced frequency information to obtain the target noise-reduced frequency. By combining scene types, the complex noise reduction task is divided into multiple interconnected, less difficult sub-tasks. Through multi-stage audio noise reduction tasks, different scene types can share the same noise reduction model, reducing resource waste caused by independent noise reduction of background music and human voices. While preserving background music, the impact on the noise reduction capability of non-music scenes is reduced, effectively improving the audio noise reduction effect in live streaming scenarios and optimizing the user experience.

[0063] Based on the above embodiments, Figure 2 A flowchart of another audio noise reduction method provided in an embodiment of this application is given, which is a concretization of the above-described audio noise reduction method. (Reference) Figure 2 The audio noise reduction method includes:

[0064] S210: Obtain the audio to be processed.

[0065] S220: The audio to be processed is sent to the first noise reduction model. The first noise reduction model performs first noise reduction processing on the noisy amplitude spectrum of the audio to be processed to obtain the first noise-reduced audio information that suppresses steady-state noise.

[0066] In one possible embodiment, such as Figure 3 As shown in the schematic diagram of a first noise reduction process, the first noise reduction model provided in this solution performs first noise reduction processing on the noisy amplitude spectrum of the audio to be processed to obtain first noise-reduced audio information that suppresses steady-state noise, including steps S221-S224:

[0067] S221: Perform a short-time Fourier transform on the audio to be processed to obtain the first noisy amplitude spectrum and the first noisy phase spectrum of the audio to be processed.

[0068] S222: Perform noise estimation and gain estimation on the first noisy amplitude spectrum to obtain the gain estimate.

[0069] S223: Based on the gain estimate, the first noisy amplitude spectrum is denoised to obtain the first denoised amplitude spectrum.

[0070] S224: Perform inverse short-time Fourier transform on the first noise reduction amplitude spectrum and the first noisy phase spectrum to obtain the first noise reduction frequency information.

[0071] For example, after obtaining the audio to be processed, a short-time Fourier transform is performed on the audio to be processed to obtain the first noisy amplitude spectrum and the first noisy phase spectrum of the audio to be processed. Optionally, the short-time Fourier transform of the audio to be processed is performed outside the first denoising model, that is, the first noisy amplitude spectrum and the first noisy phase spectrum are obtained by performing a short-time Fourier transform on the audio to be processed, and then the first noisy amplitude spectrum and / or the first noisy phase spectrum are submitted to the first denoising model for the first denoising process.

[0072] In one embodiment, the first noise reduction model performs noise estimation and gain estimation on the first noisy amplitude spectrum to obtain a gain estimate, and then performs noise reduction on the first noisy amplitude spectrum based on the gain estimate to obtain a first noise-reduced amplitude spectrum.

[0073] For example, the first noise reduction model uses methods such as quantile noise estimation to obtain a noise estimate of the first noisy amplitude spectrum, and uses this noise estimate to calculate the speech probability of the first noisy amplitude spectrum or the audio to be processed, and further obtains an accurate noise estimate of the first noisy amplitude spectrum or the current frame of the audio to be processed, and based on this noise estimate, uses gain estimation methods such as Wiener filtering to determine the gain estimate.

[0074] In one embodiment, the first noise reduction model performs noise reduction processing on the first noisy amplitude spectrum based on the determined gain estimate to obtain the first noise-reduced amplitude spectrum. For example, the gain estimate is multiplied by the first noisy amplitude spectrum to obtain the first noise-reduced amplitude spectrum, thus achieving noise reduction processing on the first noisy amplitude spectrum. After obtaining the first noise-reduced amplitude spectrum, an inverse short-time Fourier transform is performed on the first noise-reduced amplitude spectrum and the first noisy phase spectrum to obtain the first noise-reduced audio information. At this time, the first noise-reduced audio information records the corresponding information of the second noisy amplitude spectrum (i.e., the first noise-reduced amplitude spectrum) and the second noisy phase spectrum. The first noise reduction model does not perform noise reduction processing on the first noisy phase spectrum of the audio to be processed. The second noisy phase spectrum is consistent with the first noisy phase spectrum. The second noisy amplitude spectrum suppresses steady-state noise relative to the first noisy amplitude spectrum, but there is still a possibility of non-steady-state noise.

[0075] Optionally, the inverse short-time Fourier transform of the first denoised amplitude spectrum and the first noisy phase spectrum can be performed through the first denoising model, that is, the first denoised frequency information is directly output by the first denoising model, or the inverse short-time Fourier transform of the first denoised amplitude spectrum and the first noisy phase spectrum can be performed outside the first denoising model after the first denoising model has denoised the first noisy amplitude spectrum and output the first denoised amplitude spectrum.

[0076] This scheme obtains a first noisy amplitude spectrum and a first noisy phase spectrum by performing a short-time Fourier transform on the audio to be processed. Noise estimation and gain estimation are then performed on the first noisy amplitude spectrum to obtain a gain estimate. Based on the gain estimate, the first noisy amplitude spectrum is further denoised to obtain a first denoised amplitude spectrum that suppresses steady-state noise. After performing an inverse short-time Fourier transform on the first denoised amplitude spectrum and the first noisy phase spectrum, the first denoised frequency information that suppresses steady-state noise can be obtained, thus achieving suppression of steady-state noise in the audio to be processed and effectively improving the signal-to-noise ratio of the audio. Optionally, the first denoising model provided in this scheme can adopt a traditional denoising model, reducing the training cost, memory cost, and computational resource cost for the denoising model.

[0077] S230: The first noise reduction frequency information is sent to the second noise reduction model. The second noise reduction model performs noise reduction processing on the first noise reduction frequency information to obtain the second noise reduction frequency information that preserves the music and suppresses non-steady-state noise.

[0078] In one possible embodiment, the second noise reduction model provided by this solution, when processing the first noise-reduced audio information to obtain second noise-reduced audio information that preserves music and suppresses non-steady-state noise, may involve: acquiring the second noisy amplitude spectrum of the first noise-reduced audio information; inputting the first noise-reduced audio information into a trained first noise reduction neural network; and processing the second noisy amplitude spectrum through the first noise reduction neural network to obtain a second noise-reduced amplitude spectrum that suppresses non-steady-state noise. The inputs to the first and second noise reduction neural networks can be complete audio information, and the processing object for the audio information is the amplitude spectrum corresponding to the audio information.

[0079] The second denoising model provided in this solution can be built and trained based on the first denoising neural network. The first denoising neural network can be trained based on a first noisy sample audio and a first clean sample audio containing human voice and music (i.e., background music). For example, the first noisy sample audio can be used as the input to the first denoising neural network, and the first clean sample audio can be used as the output to train the first denoising neural network. Optionally, before inputting the first noisy sample audio into the first denoising neural network, feature extraction can be performed on the first noisy sample audio to improve the network's denoising performance. The extracted features can be amplitude spectrum, BFCC (Bark-Frequency Cepstral Coefficients) features, fundamental frequency features, etc., and the amplitude spectrum of the first clean sample audio containing human voice and background music can be used as the label for the network output. The loss value used in the training process is used as the basis for updating the network parameters. The loss function can be MSE (Mean Square Error), SDR (Signal to Distortion Ratio), etc. This scheme trains a second noise reduction model using a first noisy sample audio and a first clean sample audio containing human voices and music. The second noise reduction model effectively preserves human voices and background music in the audio while suppressing non-steady-state noise.

[0080] In one possible embodiment, when the second noise reduction model provided by this solution performs noise reduction processing on the first noise reduction frequency information to obtain the second noise reduction frequency information that retains the music and suppresses non-steady-state noise, it may be as follows: obtain the second noisy amplitude spectrum of the first noise reduction frequency information, perform noise reduction processing on the second noisy amplitude spectrum, and obtain the second noise reduction frequency information that retains the music and suppresses non-steady-state noise.

[0081] For example, the second denoising model acquires the second noisy amplitude spectrum of the first denoised frequency information, inputs the first denoised frequency information into the trained second denoising model, and the second denoising model performs denoising processing on the second noisy amplitude spectrum to obtain a second denoised amplitude spectrum that suppresses non-steady-state noise. Optionally, the input to the second denoising model can be the complete first denoised frequency information, or the first denoised frequency information can be first subjected to a short-time Fourier transform to obtain the second noisy amplitude spectrum and the second noisy phase spectrum of the first denoised frequency information, and then the second noisy amplitude spectrum is input into the second denoising model for denoising processing.

[0082] Optionally, after obtaining the second noise reduction amplitude spectrum, a second noise-reduced frequency spectrum that preserves music and suppresses non-steady-state noise can be obtained based on the second noise reduction amplitude spectrum and the second noisy phase spectrum. At this time, the second noise-reduced frequency spectrum records a third noisy amplitude spectrum (i.e., the second noise reduction amplitude spectrum) and a third noisy phase spectrum. The second noise reduction model does not perform noise reduction processing on the second noisy phase spectrum of the first noise-reduced frequency spectrum. The third noisy phase spectrum is consistent with the second noisy phase spectrum. The third noisy amplitude spectrum suppresses non-steady-state noise compared to the second noisy amplitude spectrum. However, in non-musical scenarios, the third noisy amplitude spectrum may still contain noise similar to background music.

[0083] For example, after obtaining the second noise-reduced frequency information, the scene type corresponding to the audio to be processed is further determined. Optionally, the scene type corresponding to the audio to be processed can be determined based on whether background music is enabled on the corresponding broadcaster terminal (e.g., audio noise reduction device) or the live broadcast room, or it can be determined by the scene classification module configured in the live broadcast application, wherein the scene classification model can be trained based on the classification model. The scene types corresponding to the audio to be processed provided by this solution include music type and non-music type. Based on this, after obtaining the second noise-reduced frequency information, if the scene type corresponding to the audio to be processed is non-music type, the solution jumps to step S240; if the scene type corresponding to the audio to be processed is music type, the solution jumps to step S250.

[0084] S240: When the scene type corresponding to the audio to be processed is not music, the second noise reduction frequency information is sent to the third noise reduction model. The third noise reduction model performs noise reduction processing on the second noise reduction frequency information to obtain the third noise reduction frequency information that further suppresses non-steady-state noise.

[0085] For example, when the scene type corresponding to the audio to be processed is not music, there may be noise in the second noise reduction frequency information that is similar to the background music. It is necessary to send the second noise reduction frequency information to the trained third noise reduction model, and the third noise reduction model performs noise reduction processing on the second noise reduction frequency information to obtain the third noise reduction frequency information, which further suppresses the noise in the audio information that is similar to the background music in the non-music scene.

[0086] In one possible embodiment, when the third noise reduction model provided by this solution performs noise reduction processing on the second noise reduction frequency information to obtain the third noise reduction frequency information, it may be as follows: obtain the third noisy amplitude spectrum of the second noise reduction frequency information, input the second noise reduction frequency information into the trained second noise reduction neural network, and perform noise reduction processing on the third noisy amplitude spectrum through the second noise reduction neural network to obtain the third noise reduction amplitude spectrum that retains the human voice.

[0087] The third noise reduction model provided in this solution can be built and trained based on the second noise reduction neural network. The second noise reduction neural network can be trained based on the first noise-reduced sample audio and the second clean sample audio containing human voices. The first noise-reduced sample audio provided in this solution can be obtained by denoising the second noisy sample audio using the second noise reduction model. For example, the second noise reduction neural network can be trained using the first noise-reduced sample audio as input and the second clean sample audio as output. Optionally, feature extraction can be performed on the first noise-reduced sample audio before inputting it into the second noise reduction model to improve the network's noise reduction performance. This solution trains the third noise reduction model using the first noise-reduced sample audio and the second clean sample audio, effectively reducing residual noise in the second noise-reduced audio information in non-music scenes.

[0088] For example, the third denoising model obtains the third noisy amplitude spectrum of the second denoised frequency information, inputs the second denoised frequency information into a trained second denoising neural network, and the second denoising neural network performs denoising processing on the third noisy amplitude spectrum to obtain a third denoised amplitude spectrum that preserves the human voice. Optionally, the input to the third denoising model can be the complete second denoised frequency information, or the second denoised frequency information can be first subjected to a short-time Fourier transform to obtain the third noisy amplitude spectrum and the third noisy phase spectrum of the second denoised frequency information, and then the third noisy amplitude spectrum is input into the third denoising model for denoising processing.

[0089] In one embodiment, the third noise-reduced audio information can be determined based on the third noise-reduced amplitude spectrum and the third noisy phase spectrum. The third noise-reduced audio information records a fourth noisy amplitude spectrum (i.e., the third noise-reduced amplitude spectrum) and a fourth noisy phase spectrum. The third noise reduction model does not perform noise reduction processing on the third noisy phase spectrum of the second noise-reduced audio information; the fourth noisy phase spectrum is consistent with the third noisy phase spectrum, and the fourth noisy amplitude spectrum suppresses noise similar to background music compared to the third noisy amplitude spectrum. This solution uses the third noise reduction model to perform noise reduction processing on the second noise-reduced audio information (excluding music genres), effectively suppressing residual noise similar to background music in the second noise-reduced audio and improving the audio noise reduction effect.

[0090] S250: Based on the scene type corresponding to the audio to be processed, perform phase noise reduction and enhancement processing on the second or third noise reduction frequency information to obtain the target noise reduction frequency.

[0091] In one possible embodiment, when this solution performs phase denoising and enhancement processing on the second or third noise-reduced frequency information based on the scene type corresponding to the audio to be processed to obtain the target noise-reduced frequency, it may be as follows: if the scene type corresponding to the audio to be processed is music, the second noise-reduced frequency information is sent to the audio enhancement model, and the audio enhancement model performs phase denoising and enhancement processing on the second noise-reduced frequency information to obtain the target noise-reduced frequency; if the scene type corresponding to the audio to be processed is not music, the third noise-reduced frequency information is sent to the audio enhancement model, and the audio enhancement model performs phase denoising and enhancement processing on the third noise-reduced frequency information to obtain the target noise-reduced frequency.

[0092] In one embodiment, the audio enhancement model provided by this solution, when performing phase denoising and enhancement processing on the second or third noise-reduced audio information to obtain the target noise-reduced audio, may do the following: send the damaged noisy amplitude spectrum and damaged noisy phase spectrum of the damaged audio information to a trained complex network; perform fundamental frequency and harmonic recovery processing on the damaged noisy amplitude spectrum and denoising processing on the damaged noisy phase spectrum through the complex network to obtain a complex mask, wherein the damaged audio information is the second or third noise-reduced audio information; enhance the damaged audio information using the complex mask to obtain enhanced audio information; and perform inverse short-time Fourier transform on the enhanced audio information to obtain the target noise-reduced audio.

[0093] Specifically, when the scene type corresponding to the audio to be processed is music, the corresponding damaged audio information is the second noise-reduced frequency information, and the corresponding damaged noise amplitude spectrum and damaged noise phase spectrum are the third noise amplitude spectrum and the third noise phase spectrum. When the scene type corresponding to the audio to be processed is not music, the corresponding damaged audio information is the third noise-reduced frequency information, and the corresponding damaged noise amplitude spectrum and damaged noise phase spectrum are the fourth noise amplitude spectrum and the fourth noise phase spectrum.

[0094] For example, the damaged audio amplitude spectrum and damaged audio phase spectrum are sent to a trained complex network. The complex network performs fundamental frequency and harmonic recovery processing on the damaged audio amplitude spectrum and noise reduction processing on the damaged audio phase spectrum to obtain a complex mask.

[0095] Furthermore, the audio enhancement model utilizes the aforementioned determined complex mask to enhance the amplitude spectrum and phase spectrum of the damaged noisy signal. For example, the complex mask is used to multiply the amplitude spectrum and add the phase spectrum. Further, the enhanced amplitude and phase spectra are subjected to inverse short-time Fourier transform to obtain the target noise-reduced frequency.

[0096] In one embodiment, when the scene type corresponding to the audio to be processed is music, since the background music in the first noise-reduced frequency information is preserved when the second noise-reduced frequency information is obtained by denoising the first noise-reduced frequency information through the second noise-reduction model, the main features of the human voice and background music are retained. Therefore, the audio enhancement model can be directly used to perform phase denoising and enhancement processing on the second noise-reduced frequency information to recover the fundamental frequency and harmonics in the second noise-reduced frequency information, which helps to improve the estimation of audio phase and improve audio quality. This solution, when the scene type corresponding to the audio to be processed is music, obtains the target noise-reduced frequency by performing phase denoising and enhancement processing on the second noise-reduced frequency information through the audio enhancement model, effectively recovering the fundamental frequency and harmonics of the human voice and background music, and improving audio quality.

[0097] In one possible embodiment, such as Figure 4 The provided schematic diagram illustrates a process for phase denoising and enhancement of the second noise-reduced frequency information. The audio enhancement model provided in this solution, when performing phase denoising and enhancement on the second noise-reduced frequency information to obtain the target noise-reduced frequency, includes steps S251-S253:

[0098] S251: Send the third noisy amplitude spectrum and the third noisy phase spectrum of the second noise-reduced frequency information to the trained complex network. The complex network performs fundamental frequency and harmonic recovery processing on the third noisy amplitude spectrum and noise reduction processing on the third noisy phase spectrum to obtain the first complex mask.

[0099] S252: Enhance the third noisy amplitude spectrum and the third noisy phase spectrum using the first complex mask.

[0100] S253: Perform inverse short-time Fourier transform on the enhanced third noisy amplitude spectrum and the third noisy phase spectrum to obtain the target noise-reduced frequency.

[0101] For example, the third noisy amplitude spectrum and the third noisy phase spectrum of the second noise-reduced frequency information are sent to the trained complex network. The complex network performs fundamental frequency and harmonic recovery processing on the third noisy amplitude spectrum and noise reduction processing on the third noisy phase spectrum to obtain the first complex mask.

[0102] The audio enhancement model provided in this solution is built and trained based on a complex network. The complex network can obtain the frequency domain correlation of audio in the time and frequency domain by using two-dimensional convolution or self-attention mechanism, which can effectively recover the fundamental frequency and harmonics of the amplitude spectrum, help estimate the clean phase, and improve audio quality.

[0103] Furthermore, the audio enhancement model utilizes the aforementioned determined first complex mask to enhance the third noisy amplitude spectrum and the third noisy phase spectrum. For example, the first complex mask is used to multiply the third noisy amplitude spectrum and add the third noisy phase spectrum. Further, an inverse short-time Fourier transform is performed on the enhanced third noisy amplitude spectrum and the third noisy phase spectrum to obtain the target noise-reduced frequency.

[0104] Optionally, the inverse short-time Fourier transform of the third noisy amplitude spectrum and the third noisy phase spectrum can be performed through the audio enhancement model, i.e., the target noise-reduced frequency is directly output by the audio enhancement model, or the inverse short-time Fourier transform of the third noisy amplitude spectrum and the third noisy phase spectrum can be performed outside the audio enhancement model after the audio enhancement model outputs the third noisy amplitude spectrum and the third noisy phase spectrum.

[0105] In one embodiment, when the scene type corresponding to the audio to be processed is not music, after the third noise reduction information is obtained by denoising the second noise reduction information using a third noise reduction model, the third noise reduction information suppresses noise similar to the background music. Then, the audio enhancement model is used to perform phase denoising and enhancement processing on the third noise reduction information to recover the fundamental frequency and harmonics in the third noise reduction information, which helps to improve the estimation of the audio phase and improve audio quality. In this solution, when the scene type corresponding to the audio to be processed is music, the third noise reduction information is phase-denoised and enhanced using an audio enhancement model to obtain the target noise reduction, effectively recovering the fundamental frequency and harmonics of the human voice and background music, thus improving audio quality.

[0106] In one possible embodiment, such as Figure 5 A flowchart illustrating phase denoising and enhancement processing of third-order noise-reduced frequency information is provided. The audio enhancement model provided in this solution, when performing phase denoising and enhancement processing on the third-order noise-reduced frequency information to obtain the target noise-reduced frequency, includes steps S271-S273:

[0107] S254: Send the fourth noisy amplitude spectrum and the fourth noisy phase spectrum of the third noise reduction frequency information to the trained complex network. The complex network performs fundamental frequency and harmonic recovery processing on the fourth noisy amplitude spectrum and noise reduction processing on the fourth noisy phase spectrum to obtain the second complex mask.

[0108] S255: Enhance the fourth noisy amplitude spectrum and the fourth noisy phase spectrum using the second complex mask.

[0109] S256: Perform inverse short-time Fourier transform on the enhanced fourth-band noisy amplitude spectrum and the fourth-band noisy phase spectrum to obtain the target noise-reduced frequency.

[0110] For example, the fourth noisy amplitude spectrum and the fourth noisy phase spectrum of the third noise-reduced frequency information are sent to the trained complex network. The complex network performs fundamental frequency and harmonic recovery processing on the fourth noisy amplitude spectrum and noise reduction processing on the fourth noisy phase spectrum to obtain the second complex mask.

[0111] The audio enhancement model provided in this solution is built and trained based on a complex network. The complex network can obtain the frequency domain correlation of audio in the time and frequency domain by using two-dimensional convolution or self-attention mechanism, which can effectively recover the fundamental frequency and harmonics of the amplitude spectrum, help estimate the clean phase, and improve audio quality.

[0112] Furthermore, the audio enhancement model utilizes the aforementioned determined second complex mask to enhance the fourth noisy amplitude spectrum and the fourth noisy phase spectrum. For example, the second complex mask is used to multiply the fourth noisy amplitude spectrum and add the fourth noisy phase spectrum. Further, an inverse short-time Fourier transform is performed on the enhanced fourth noisy amplitude spectrum and the fourth noisy phase spectrum to obtain the target noise-reduced frequency.

[0113] Optionally, the inverse short-time Fourier transform of the fourth noisy amplitude spectrum and the fourth noisy phase spectrum can be performed through the audio enhancement model, i.e., the target noise-reduced frequency is directly output by the audio enhancement model, or the inverse short-time Fourier transform of the fourth noisy amplitude spectrum and the fourth noisy phase spectrum can be performed outside the audio enhancement model after the audio enhancement model outputs the fourth noisy amplitude spectrum and the fourth noisy phase spectrum.

[0114] In one embodiment, the computation of a complex network can be represented as follows:

[0115] W*h=(A+iB)*(x+iy)=(A*xB*t)+i(B*x+A*y)

[0116] Where h = x + iy is the complex input vector (audio information), W = + iB is the complex weight matrix, and A and B are the training parameters of the complex network. The complex network can effectively maintain the coupling relationship between the amplitude spectrum and the phase spectrum while simultaneously achieving noise reduction of both, exhibiting good noise reduction performance. The output of the complex network can be a complex mask. Complex masks can be used to enhance the amplitude spectrum of noise reduction |X t,f |and noisy phase spectrum θ Xt, ,Right now: Where t and f represent the number of time-domain frames and the number of frequency-domain points, respectively, they can be used to... The target noise reduction frequency is obtained by performing an inverse short-time Fourier transform.

[0117] In one possible embodiment, the audio enhancement model provided by this solution can be trained based on a second denoised sample audio, a third denoised sample audio, and a third clean sample audio containing human voice and / or music. The second denoised sample audio is obtained by denoising the third noisy sample audio using the second denoising model, and the third denoised sample audio is obtained by denoising the fourth noisy sample audio using the third denoising model. For example, the complex network can be trained using the second and third denoised sample audio as inputs and the third clean sample audio as output. Optionally, feature extraction can be performed on the second and third denoised sample audio before inputting them into the complex network to improve the network's enhancement performance. This solution trains the third denoising model using the second, third, and third clean sample audio, and trains the audio enhancement model using the corresponding clean audio. Using the audio denoising amplitude spectrum and noisy phase spectrum as inputs, it utilizes frequency domain correlation to repair the fundamental frequency and harmonics while effectively estimating the clean phase, thereby improving audio quality.

[0118] As described above, the first noise reduction model performs first noise reduction processing on the noisy amplitude spectrum of the audio to be processed to obtain first noise-reduced frequency information that suppresses steady-state noise. Based on the scene type, the first noise-reduced frequency is further denoised to obtain second noise-reduced frequency information that preserves the music while suppressing non-steady-state noise, or third noise-reduced frequency information that suppresses non-steady-state noise. Then, based on the scene type corresponding to the audio to be processed, phase denoising and enhancement processing are performed on the second or third noise-reduced frequency information to obtain the target noise-reduced frequency. Combining scene type, the complex noise reduction task is divided into several serially connected, less difficult sub-tasks. First, the first noise reduction model suppresses steady-state noise; then, scene classification labels are introduced, and the second and third noise reduction models jointly complete the suppression of non-steady-state noise; finally, the audio enhancement model performs audio enhancement. In music-related scenarios, the second noise reduction model can suppress non-steady-state noise while preserving the music. When the scenario is determined to be non-music-related, the third noise reduction model further processes the noise reduction results of the second noise reduction model, filters out residual noise, and reduces the impact on the noise reduction effect in non-music scenarios. Parts of the model can be shared between non-music and music scenarios, reducing the waste of resources caused by independent noise reduction of background music and human voices. While preserving background music, it also weakens the impact on the noise reduction capability of non-music scenarios, which can effectively improve the audio noise reduction effect and optimize the user experience.

[0119] Figure 6 This is a schematic diagram of the structure of an audio noise reduction device provided in an embodiment of this application. (Reference) Figure 6 The audio noise reduction device includes an audio acquisition module 61, a first noise reduction module 62, a second noise reduction module 63, and an audio enhancement module 64.

[0120] The audio acquisition module 61 is configured to acquire the audio to be processed; the first noise reduction module 62 is configured to send the audio to be processed to the first noise reduction model, and perform first noise reduction processing on the noisy amplitude spectrum of the audio to be processed through the first noise reduction model to obtain first noise-reduced frequency information that suppresses steady-state noise; the second noise reduction module 63 is configured to perform second noise reduction processing on the first noise-reduced frequency information according to the scene type corresponding to the audio to be processed to obtain second noise-reduced frequency information that preserves music and suppresses non-steady-state noise, or perform third noise reduction processing on the first noise-reduced frequency information to obtain third noise-reduced frequency information that suppresses non-steady-state noise; and the audio enhancement module 64 is configured to perform phase noise reduction and enhancement processing on the second noise-reduced frequency information or the third noise-reduced frequency information according to the scene type corresponding to the audio to be processed to obtain the target noise-reduced frequency.

[0121] The above describes a process where a first noise reduction model is used to perform noise reduction on the noisy amplitude spectrum of the audio to be processed, resulting in first noise-reduced frequency information that suppresses steady-state noise. Depending on the scene type, the first noise-reduced frequency information is further denoised to obtain second noise-reduced frequency information that preserves music while suppressing non-steady-state noise, or third noise-reduced frequency information that suppresses non-steady-state noise. Then, based on the scene type corresponding to the audio to be processed, phase denoising and enhancement processing are performed on the second or third noise-reduced frequency information to obtain the target noise-reduced frequency. By combining scene types, the complex noise reduction task is divided into multiple interconnected, less difficult sub-tasks. Through multi-stage audio noise reduction tasks, different scene types can share the same noise reduction model, reducing resource waste caused by independent noise reduction of background music and human voices. While preserving background music, the impact on the noise reduction capability of non-music scenes is reduced, effectively improving the audio noise reduction effect in live streaming scenarios and optimizing the user experience.

[0122] In one possible embodiment, the second noise reduction module 63 is configured as follows:

[0123] When the scene type corresponding to the audio to be processed is music, the first noise reduction frequency information is subjected to the second noise reduction process to obtain the second noise reduction frequency information that retains the music and suppresses non-steady-state noise.

[0124] When the scene type corresponding to the audio to be processed is not music, the first noise reduction frequency information is subjected to the third noise reduction process to obtain the third noise reduction frequency information that suppresses non-steady-state noise.

[0125] In one possible embodiment, when the second noise reduction module 63 performs a second noise reduction process on the first noise-reduced frequency information to obtain second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise, it is configured as follows:

[0126] The first noise-reduced frequency information is sent to the second noise reduction model. The second noise reduction model performs noise reduction processing on the first noise-reduced frequency information to obtain the second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise.

[0127] In one possible embodiment, when the second noise reduction module 63 performs a third noise reduction process on the first noise-reduced frequency information to obtain third noise-reduced frequency information that suppresses non-steady-state noise, it is configured as follows:

[0128] The first noise-reduced frequency information is sent to the second noise reduction model. The second noise reduction model performs noise reduction processing on the first noise-reduced frequency information to obtain the second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise.

[0129] The second noise reduction frequency information is sent to the third noise reduction model, and the second noise reduction frequency information is processed by the third noise reduction model to obtain the third noise reduction frequency information that further suppresses non-steady-state noise.

[0130] In one possible embodiment, when the second noise reduction model performs noise reduction processing on the first noise reduction frequency information to obtain the second noise reduction frequency information that retains the music and suppresses non-steady-state noise, it is configured to acquire the second noisy amplitude spectrum of the first noise reduction frequency information, perform noise reduction processing on the second noisy amplitude spectrum, and obtain the second noise reduction frequency information that retains the music and suppresses non-steady-state noise.

[0131] In one possible embodiment, when the third noise reduction model performs noise reduction processing on the second noise reduction frequency information to obtain the third noise reduction frequency information that suppresses non-steady-state noise, it is configured to acquire the third noisy amplitude spectrum of the second noise reduction frequency information, perform noise reduction processing on the third noisy amplitude spectrum, and obtain the third noise reduction amplitude spectrum that preserves human voice.

[0132] In one possible embodiment, the audio enhancement module 64 is configured as follows:

[0133] When the scene type corresponding to the audio to be processed is music, the second noise reduction frequency information is sent to the audio enhancement model. The audio enhancement model performs phase noise reduction and enhancement processing on the second noise reduction frequency information to obtain the target noise reduction frequency.

[0134] If the scene type corresponding to the audio to be processed is not music, the third noise reduction frequency information is sent to the audio enhancement model. The audio enhancement model performs phase noise reduction and enhancement processing on the third noise reduction frequency information to obtain the target noise reduction frequency.

[0135] In one possible embodiment, when the audio enhancement model performs phase denoising and enhancement processing on the second or third noise-reduced frequency information to obtain the target noise-reduced frequency, it is configured as follows:

[0136] The damaged audio information with noise amplitude spectrum and damaged audio with noise phase spectrum are sent to a trained complex network. The complex network performs fundamental frequency and harmonic recovery processing on the damaged audio amplitude spectrum and noise reduction processing on the damaged audio phase spectrum to obtain a complex mask. The damaged audio information is the second noise reduction frequency information or the third noise reduction frequency information.

[0137] The damaged audio information is enhanced by using a complex mask to obtain the enhanced audio information;

[0138] The target noise-reduced frequency is obtained by performing an inverse short-time Fourier transform on the enhanced audio information.

[0139] It is worth noting that in the above-described embodiments of the audio noise reduction device, the various units and modules included are only divided according to functional logic, but are not limited to the above division, as long as the corresponding functions can be achieved; in addition, the specific names of each functional unit are only for easy differentiation and are not used to limit the protection scope of the embodiments of the present invention.

[0140] This application also provides an audio noise reduction device, which can integrate the audio noise reduction apparatus provided in this application. Figure 7 This is a schematic diagram of the structure of an audio noise reduction device provided in an embodiment of this application. (Reference) Figure 7 The audio noise reduction device includes: an input device 73, an output device 74, a memory 72, and one or more processors 71; the memory 72 is used to store one or more programs; when one or more programs are executed by one or more processors 71, the one or more processors 71 implement the audio noise reduction method provided in the above embodiments. The audio noise reduction device, apparatus, and computer provided above can be used to execute the audio noise reduction method provided in any of the above embodiments, and have corresponding functions and beneficial effects.

[0141] This application also provides a non-volatile storage medium storing computer-executable instructions, which, when executed by a computer processor, are used to perform the audio noise reduction method provided in the above embodiments. Of course, the computer-executable instructions provided in this application are not limited to the audio noise reduction method provided above; they can also perform related operations in the audio noise reduction method provided in any embodiment of this application. The audio noise reduction apparatus, device, and storage medium provided in the above embodiments can perform the audio noise reduction method provided in any embodiment of this application. Technical details not described in detail in the above embodiments can be found in the audio noise reduction method provided in any embodiment of this application.

[0142] Based on the above embodiments, this application also provides a computer program product. The technical solution of this application, in essence or in other words, the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. The computer program product is stored in a storage medium and includes several instructions to cause a computer device, mobile terminal, or processor therein to execute all or part of the steps of the audio noise reduction method provided in the various embodiments of this application.

Claims

1. An audio noise reduction method, characterized in that, include: Obtain the audio to be processed; The audio to be processed is sent to the first noise reduction model, and the first noise reduction model performs a first noise reduction process on the noisy amplitude spectrum of the audio to be processed to obtain the first noise reduction audio information that suppresses steady-state noise. When the scene type corresponding to the audio to be processed is music, the first noise reduction frequency information is subjected to a second noise reduction process to obtain a second noise reduction frequency information that retains music and suppresses non-steady-state noise; when the scene type corresponding to the audio to be processed is not music, the first noise reduction frequency information is subjected to a third noise reduction process to obtain a third noise reduction frequency information that suppresses non-steady-state noise. Based on the scene type corresponding to the audio to be processed, phase noise reduction and enhancement processing are performed on the second noise reduction frequency information or the third noise reduction frequency information to obtain the target noise reduction frequency.

2. The audio noise reduction method according to claim 1, characterized in that, The second noise reduction process performed on the first noise-reduced frequency information to obtain second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise includes: The first noise-reduced frequency information is sent to the second noise reduction model, and the second noise reduction model performs noise reduction processing on the first noise-reduced frequency information to obtain the second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise.

3. The audio noise reduction method according to claim 1, characterized in that, The third noise reduction process performed on the first noise-reduced frequency information to obtain third noise-reduced frequency information that suppresses non-steady-state noise includes: The first noise-reduced frequency information is sent to the second noise reduction model, and the second noise reduction model performs noise reduction processing on the first noise-reduced frequency information to obtain the second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise. The second noise-reduced frequency information is sent to the third noise reduction model, and the second noise-reduced frequency information is processed by the third noise reduction model to obtain the third noise-reduced frequency information that further suppresses non-steady-state noise.

4. The audio noise reduction method according to claim 2 or 3, characterized in that, The second noise reduction model, when processing the first noise-reduced frequency information to obtain second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise, includes: The second noisy amplitude spectrum of the first noise-reduced frequency information is obtained, and the second noisy amplitude spectrum is subjected to noise reduction processing to obtain the second noise-reduced frequency information that preserves the music and suppresses non-steady-state noise.

5. The audio noise reduction method according to claim 3, characterized in that, The third noise reduction model, when performing noise reduction processing on the second noise-reduced frequency information to obtain third noise-reduced frequency information that suppresses non-steady-state noise, includes: The third noise-band amplitude spectrum of the second noise-reduced frequency information is obtained, and the third noise-band amplitude spectrum is subjected to noise reduction processing to obtain the third noise-reduced amplitude spectrum that preserves the human voice.

6. The audio noise reduction method according to claim 1, characterized in that, The step of performing phase denoising and enhancement processing on the second or third noise-reduced frequency information based on the scene type corresponding to the audio to be processed to obtain the target noise-reduced frequency includes: When the scene type corresponding to the audio to be processed is music, the second noise reduction frequency information is sent to the audio enhancement model. The audio enhancement model performs phase noise reduction and enhancement processing on the second noise reduction frequency information to obtain the target noise reduction frequency. If the scene type corresponding to the audio to be processed is not music, the third noise reduction frequency information is sent to the audio enhancement model. The audio enhancement model performs phase noise reduction and enhancement processing on the third noise reduction frequency information to obtain the target noise reduction frequency.

7. The audio noise reduction method according to claim 6, characterized in that, The audio enhancement model, when performing phase denoising and enhancement processing on the second or third noise-reduced frequency information to obtain the target noise-reduced frequency, includes: The damaged audio information is sent to a trained complex network to perform fundamental frequency and harmonic recovery processing on the damaged audio information and noise reduction processing on the damaged audio information and the damaged audio information and the noise reduction processing on the damaged audio information and the noise reduction processing on the damaged audio information and the noise reduction processing on the damaged audio information and the noise reduction processing on the damaged audio information and the noise reduction processing on the damaged audio information and the noise reduction processing on the damaged audio information. The damaged audio information is enhanced using the complex mask to obtain enhanced audio information. The enhanced audio information is subjected to inverse short-time Fourier transform to obtain the target noise-reduced frequency.

8. An audio noise reduction device, characterized in that, It includes an audio acquisition module, a first noise reduction module, a second noise reduction module, and an audio enhancement module, wherein: The audio acquisition module is configured to acquire the audio to be processed; The first noise reduction module is configured to send the audio to be processed to a first noise reduction model, and perform a first noise reduction process on the noisy amplitude spectrum of the audio to be processed through the first noise reduction model to obtain first noise-reduced audio information that suppresses steady-state noise. The second noise reduction module is configured to perform a second noise reduction process on the first noise reduction frequency information when the scene type corresponding to the audio to be processed is music, to obtain a second noise reduction frequency information that retains music and suppresses non-steady-state noise; and to perform a third noise reduction process on the first noise reduction frequency information when the scene type corresponding to the audio to be processed is not music, to obtain a third noise reduction frequency information that suppresses non-steady-state noise. The audio enhancement module is configured to perform phase noise reduction and enhancement processing on the second noise reduction frequency information or the third noise reduction frequency information according to the scene type corresponding to the audio to be processed, so as to obtain the target noise reduction frequency.

9. An audio noise reduction device, characterized in that, include: Memory and one or more processors; The memory is used to store one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the audio noise reduction method as described in any one of claims 1-7.

10. A non-volatile storage medium for storing computer-executable instructions, characterized in that, The computer-executable instructions, when executed by a computer processor, are used to perform the audio noise reduction method as described in any one of claims 1-7.

11. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the audio noise reduction method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Voice enhancement method and system based on phase compensation

    CN108735213A

  • Speech signal processing method and device

    CN110875049A

  • Audio denoising method and device, electronic equipment and storage medium

    CN112908352A

  • Audio denoising method and device, apparatus and storage medium

    US20240284100A1

  • Audio noise reduction method and apparatus, device, storage medium and product

    WO2024222373A1