Audio Pickup Interference Reduction via Face-Source Localization

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Video conferencing environments often suffer from unwanted background noise interference, which can be distracting and difficult to manage, especially in open spaces where noise from non-speakers can disrupt the audio feed.

Innovation Solution

A system utilizing a microphone array to calculate the pan, tilt, and distance of sound sources, coupled with face detection and motion analysis, to selectively mute or attenuate audio signals not originating from a human speaker, ensuring only targeted speech is picked up and transmitted.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a microphone array is used to capture audio in a video conference, then the audio coverage is improved, but background noise interference increases

Engineering Contradiction:
Improveaudio coverageVSAvoidbackground noise interference
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and isolates the desired audio signal from the sound source location (detected via microphone array and face detector) while separating it from unwanted background noise. The system extracts only the relevant speech signal by comparing audio source location with video face location, effectively taking out the useful signal from the noisy environment.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The system uses feedback by continuously monitoring the location of audio sources via microphone array and comparing it with the location of detected faces in video frames. This feedback loop allows the system to adjust audio pickup in real-time, muting or attenuating signals that do not correspond to a detected face, thereby reducing background noise while maintaining reliable audio coverage.

Inventive Principle:
Principle #23Feedback

2Loss of information

If audio from all directions is captured, then no speech is missed, but unwanted noise from non-speakers is picked up

Engineering Contradiction:
Improvespeech detection completenessVSAvoidnoise from non-speakers
Core Design Contradiction:
Loss of informationVSObject-generated harmful factors

Solution Approach 1:

The patent applies local quality by directing audio pickup focus to specific locations in space rather than uniformly capturing all sounds. The system identifies the precise location of the speaker using the microphone array and face detector, then concentrates audio capture on that local region while attenuating sounds from other directions, ensuring complete speech detection from the target speaker while excluding noise from non-speakers.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The system extracts the speech signal from the specific spatial location where a face is detected, separating it from sounds originating at other locations. By comparing audio source coordinates with face detector coordinates, the system takes out only the relevant speech signal while leaving unwanted noise from other directions excluded.

Inventive Principle:
Principle #2Taking out (Extraction)

3Loss of information

If the microphone array captures all audio signals, then the audio feed is complete, but background distractions increase

Engineering Contradiction:
Improveaudio completenessVSAvoidbackground distractions
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The system employs feedback by continuously comparing the spatial location of detected audio sources with the spatial location of detected faces. This feedback mechanism allows the system to maintain complete audio information from the target speaker while dynamically filtering out background distractions that do not correspond to a detected face, thus resolving the contradiction between audio completeness and background distraction.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent extracts the desired speech signal from the complete audio field by identifying and isolating sounds originating from the same spatial location as a detected face. This extraction process maintains audio completeness for the speaker of interest while removing background distractions from other sources.

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentUS10490202B2Interference-free audio pickup in a video conference
Publication Date: 2019.11.26 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US10490202B2 patent drawing
  • US10490202B2 patent drawing
  • US10490202B2 patent drawing

AI summary

A videoconference apparatus at a first location detects audio from a location and determines whether the sound should be included in an audio-video stream sent to a second location, or excluded as an interfering noise. Determining whether to include the audio involves using a face detector to see if there is a face at the source of the sound. If a face is present, the audio data from the location will be transmitted to the second location. If a face is not present, additional motion checks are performed to determine whether the sound corresponds to a person talking, (such as a presenter at a meeting), or whether the sound is instead unwanted noise.