Scenario-Based Audio Noise Reduction for Music-Preserving Live Streams
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional noise reduction algorithms designed for human voice scenarios adversely affect background music in live streaming scenarios, leading to poor noise reduction effects and user experience.
Innovation Solution
An audio noise reduction method that employs multiple noise reduction models based on scenario type, suppressing stationary and non-stationary noise, and performing phase noise reduction and enhancement processing to preserve human voice and background music.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional noise reduction algorithms designed for human voice scenarios are applied to live streaming scenarios, then human voice can be preserved, but background music is damaged and noise reduction effect deteriorates
Solution Approach 1:
The patent divides the noise reduction task into multiple specialized models: a first noise reduction model for human voice scenarios and a second noise reduction model for live streaming scenarios with background music. Each model is trained on specific data corresponding to its target scenario, allowing optimal performance for each case without compromise.
Solution Approach 2:
The system dynamically selects which noise reduction model to apply based on the detected scenario type. The determination module identifies whether the audio contains background music and switches between models accordingly, making the system adaptable to different operating conditions rather than using a fixed approach.
2Device complexity
If a single noise reduction model is used for all scenarios, then device complexity is reduced, but noise reduction performance deteriorates in specific scenarios like live streaming
Solution Approach 1:
Instead of one complex universal model, the patent uses multiple specialized models with simpler, scenario-specific architectures. The first model handles human voice scenarios while the second handles live streaming with music, reducing the complexity burden on each individual model while improving overall performance.
Solution Approach 2:
The determination module acts as a universal controller that manages multiple specialized models, selecting the appropriate one based on scenario detection. This allows the system to maintain multiple functions through a coordinated architecture rather than requiring each model to be universally capable.
3Object-generated harmful factors
If noise reduction algorithms suppress all noise components, then noise is reduced, but background music is also suppressed and audio quality deteriorates
Solution Approach 1:
The second noise reduction model is specifically trained to differentiate between noise components and music components in the audio spectrum. It applies different suppression strategies to different frequency regions and time segments, preserving music while removing noise based on the local characteristics of each audio segment.
Solution Approach 2:
The patent changes the training parameters and target objectives of the noise reduction model based on the scenario. For live streaming with music, the model is trained to preserve music parameters while suppressing noise parameters, whereas for human voice scenarios, different parameter optimization is applied to preserve voice quality.
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Provided in the embodiments of the present application are an audio noise reduction method and apparatus, a device, a storage medium and a product. The technical solution provided by the embodiments of the present application comprises: by means of a first noise reduction model, performing first noise reduction processing on a noisy amplitude spectrum of an audio to be processed, so as to obtain first noise-reduced audio information in which steady noise has been suppressed; performing noise reduction processing on the first noise-reduced audio according to a scenario type, so as to obtain second noise-reduced audio information in which music has been retained and unsteady noise has been suppressed or third noise-reduced audio information in which unsteady noise has been suppressed; and performing phase noise reduction and enhancement processing on the second noise-reduced audio information or the third noise-reduced audio information according to the scenario type corresponding to the audio to be processed, so as to obtain a target noise-reduced audio. Thus, the resource waste caused by independent noise reduction of background music and vocals is reduced, the influence on the noise reduction capability in a non-music scenario is weakened while the background music is retained, the audio noise reduction effect in a live broadcast scenario can be effectively improved, and the user experience is optimized.