Video Chat Audio Beamforming for Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In video chats, especially multi-party conferences, omnidirectional recording often picks up excessive background noise and noise from other participants, severely affecting voice quality due to the inability to selectively focus on specific speakers.
Innovation Solution
A method where a first terminal divides the video calling screen into multiple angular domains, determines beam configuration information for each domain, and sends this information to a second terminal to perform beamforming processing, enhancing the signal strength of the target domain while attenuating other domains, thereby improving voice quality by reducing background noise and noise from multiple participants.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If omnidirectional recording is used to pick up voices of all participants, then the coverage of audio recording is improved, but background noise and noise from other participants increases
Solution Approach 1:
The patent segments the audio recording space into multiple angular domains, each corresponding to a specific participant or region. Instead of uniform omnidirectional recording, the system divides the 360-degree space into discrete angular sectors and applies independent beamforming control to each sector, allowing selective enhancement of desired speakers while suppressing others.
Solution Approach 2:
The patent implements local quality by applying different beamforming characteristics to different angular domains. Each angular domain has customized beam configuration information (beam direction, beam width, gain) tailored to the specific participant in that region. This allows the system to optimize audio quality locally for each speaker while maintaining overall multi-party coverage.
2Measurement precision
If beamforming processing is applied to enhance target speaker signal, then voice quality of specific participant is improved, but signal strength of other participants is attenuated
Solution Approach 1:
The patent implements dynamic beamforming where the beam configuration information is not fixed but can be adjusted in real-time based on which participant is currently speaking or needs to be highlighted. The system dynamically switches between different angular domain configurations, enhancing the active speaker's signal while temporarily attenuating others, then switching to capture different speakers as they take turns speaking.
Solution Approach 2:
The system employs periodic scanning or switching through different angular domains, systematically cycling through each participant's designated angular sector. This periodic action ensures that over time, all participants receive adequate attention and their signals are enhanced during their respective turns, while maintaining the ability to suppress background noise from non-active participants.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively reduces background noise and noise from other participants, enhancing voice quality in video chats by selectively focusing on specific speakers, thus improving communication clarity.
Implementation Method 1
performing, by the second terminal according to the beam configuration information corresponding to the target angular domain, beamforming processing on an audio signal obtained through recording, so as to enhance signal strength of an audio signal of the target angular domain, and attenuate signal strength of an audio signal of another angular domain
Data Source
Figure 1
Figure 2
Figure 3~4
AI summary
Embodiments of the present invention disclose a method for recording in a video chat, and a terminal, to reduce background noise and noise of multiple persons in a video chat process, and improve voice quality of a video chat. A first terminal divides a video calling screen into multiple angular domains. After beam configuration information of each angular domain is determined, beam configuration information of a target angular domain of the first terminal is sent to a second terminal. The second terminal performs, according to the beam configuration information, beamforming processing on an audio signal obtained through recording, so that signal strength of an audio signal of the target angular domain is enhanced, and signal strength of an audio signal of another angular domain is attenuated.