Audio Processing for Immersive Media Using Artificial Sound Masking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing techniques for removing unwanted sounds, such as wind noise, often degrade the perceptual quality of multimedia content and can introduce inconsistencies between audio and video components, particularly in scenarios where immersive media like VR, AR, and MR are involved.
Innovation Solution
An apparatus and method that determine the location and intensity of unwanted sounds within multimedia data, perform audio processing to remove unwanted sounds, and add artificial sounds synchronized in time and space to maintain perceptual quality and consistency, using a combination of audio and video analysis to identify regions of interest and adjust sound removal and addition accordingly.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If wind noise removal is applied using known audio processing techniques, then unwanted sound is removed, but perceptual quality of the audio is reduced
Solution Approach 1:
The patent applies wind noise removal techniques to eliminate unwanted wind noise from audio recordings. The harmful wind noise is converted into a benefit by systematically removing it while using artificial wind noise generation to compensate for the removal artifacts, thereby maintaining natural sound characteristics and perceptual quality.
Solution Approach 2:
The system dynamically adjusts processing parameters based on detected wind noise characteristics. By analyzing audio signals and identifying temporal-spatial locations of wind noise, the system applies variable processing strength and artificial noise addition parameters to maintain optimal perceptual quality while effectively removing unwanted sound.
2Object-affected harmful factors
If wind noise removal is applied, then unwanted sound is removed, but consistency between audio and video components is reduced
Solution Approach 1:
The patent converts the harmful effect of audio-video inconsistency into a benefit by using video analysis to guide audio processing. The visual information about wind presence and characteristics is used to generate artificial wind noise that matches the video scene, thereby restoring and enhancing audio-video consistency.
Solution Approach 2:
The system introduces an intermediary approach by using video data as a mediator between audio processing and the final output. Video analysis of wind effects serves as a reference to generate artificial audio wind noise that is consistent with the visual scene, bridging the gap between audio and video modalities.
3Reliability
If audio processing is performed to remove unwanted sound, then sound quality is improved, but artifacts are introduced
Solution Approach 1:
The patent converts the harmful processing artifacts into a benefit by using artificial wind noise generation to mask these artifacts. The artifacts introduced by wind noise removal are counteracted by adding synthesized wind noise that is imperceptibly different from the original, thereby improving overall sound quality.
Solution Approach 2:
The system dynamically adjusts the parameters of artificial wind noise addition based on the detected original wind noise characteristics. By varying the intensity, temporal pattern, and spectral characteristics of the artificial noise, the system effectively masks processing artifacts while maintaining natural sound quality.
Data Source
AI summary
An apparatus, method and computer program is disclosed. The apparatus may comprise a means comprising at least one processor and at least one memory including computer program code, the at least one memory and computer program code configured to, with the at least one processor, to receive multimedia data representing a scene, the multimedia data comprising at least audio data representing an audio component of the scene. Another operation may comprise determining a location of unwanted sound in the scene. Another operation may comprise performing first audio processing to remove at least part of the unwanted sound from the determined location. Another operation may comprise performing second audio processing to add artificial sound associated to the unwanted sound at the determined location.


