Terminal Sound Processing Using Interaural Level Difference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing sound processing technologies face challenges in accurately picking up target sound sources during video shooting due to severe aliasing noise and low accuracy, especially in low signal-to-noise ratios, where directional interference noise from behind the camera is misidentified as the target sound, leading to poor video quality.
Innovation Solution
A sound processing method and apparatus that utilize two microphones positioned at the front and rear of a terminal to calculate the interaural level difference, determine if a rear sound signal is present, and filter it out, ensuring that only sound within the camera's field of view is captured, thereby improving noise suppression and accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional sound source localization algorithms (beamforming, delay difference) are used, then targeted sound pickup is achieved, but aliasing noise occurs in low signal-to-noise ratio scenarios causing low accuracy
Solution Approach 1:
The patent introduces an intermediary verification mechanism using a second algorithm to check whether the sound source direction determined by the first algorithm falls within the camera's shooting range. This intermediary step acts as a mediator to identify and filter out aliased noise sources that the first algorithm incorrectly identifies as target sound sources.
Solution Approach 2:
The patent implements a feedback mechanism where the sound source direction determination result is fed back to verify consistency with camera shooting range information. If the determined direction falls outside the shooting range, the system recognizes it as aliasing noise and adjusts the sound pickup accordingly, creating a closed-loop feedback system to improve accuracy.
2Reliability
If fixed beam or adaptive beam method is used to reduce out-of-beam interference, then targeted sound pickup is improved, but noise from directions outside beam coverage is not effectively suppressed
Solution Approach 1:
The patent extends the noise suppression approach from traditional spatial beamforming to a new dimension by incorporating camera shooting range information as a constraint. This adds a geometric dimension (camera field of view) to the sound processing, allowing suppression of noise not just based on directional beams but also based on whether the sound source direction falls within the camera's observable range.
3Measurement precision
If delay difference based sound source localization is used, then azimuth determination is achieved, but severe aliasing noise occurs in low signal-to-noise ratio causing misidentification of noise as target
Solution Approach 1:
The patent performs preliminary verification of the sound source direction against camera shooting range information before finalizing the target sound source identification. This preliminary action prevents misidentification of aliased noise as target by checking consistency with the camera's actual field of view before committing to the sound pickup decision.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively filters out noise beyond the camera's range, enhancing the quality of the video by improving user experience through better noise suppression and accurate sound source localization.
Implementation Method 1
calculating an interaural level difference between the two microphones based on collected sound signals according to a preset first algorithm
Data Source
Figure 1
Figure 2A~2B
Figure 2C
AI summary
The present invention discloses a sound processing method and apparatus. The method is applied to a terminal equipped with two microphones at the top of the terminal, where the two microphones are located respectively in the front and at the back of the terminal, and the method is applied to a non-video-call scenario. The method includes: when it is detected that a camera of the terminal is in a shooting state, collecting a sound signal by using the two microphones; calculating an interaural level difference between the two microphones based on collected sound signals according to a preset first algorithm; determining whether the interaural level difference meets a sound source direction determining condition; if the determining condition is met, determining, based on the interaural level difference, whether the sound signal includes a rear sound signal, where the rear sound signal is a sound signal whose sound source is located behind the camera; and if it is determined that the sound signal includes a rear sound signal, filtering out the rear sound signal from the sound signal. In this way, in a low signal-to-noise ratio scenario, sound source localization is performed based on an interaural level difference, thereby improving accuracy in picking up a sound source within a shooting range.