Stereo Sound Generation Using Microphone and Face Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In videoconferences, directional audio fails to accurately distinguish between speakers, especially when participants are moving or multiple people are in the same frame, leading to decreased intelligibility and difficulty in identifying who is speaking.

Innovation Solution

Implementing a system that uses a microphone array and face detection to generate stereo sound signals by weighting the amplitude of audio channels based on the direction of arrival and location of speakers, allowing for more natural sound representation and improved speaker differentiation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If directional audio is used to distinguish speakers in video conferences, then audio directionality is improved, but speaker differentiation fails when multiple people are in the same frame or when speakers are moving

Engineering Contradiction:
Improveaudio directionalityVSAvoidspeaker identification
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent transitions from 2D video frame-based audio positioning to 3D spatial audio positioning by incorporating depth information from face detection. This allows the system to differentiate speakers not just by their position in the video frame but by their actual spatial location in three-dimensional space, resolving the contradiction by adding a dimensional layer to audio directionality.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Solution Approach 2:

The patent introduces face detection technology as an intermediary between the microphone array and the audio output system. The face detection module provides additional spatial information about speaker locations that complements the directional audio data, enabling more accurate speaker differentiation when traditional directional audio alone is insufficient.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Ease of operation

If traditional directional audio is used, then audio processing is simple, but intelligibility decreases when speakers are moving or in the same frame

Engineering Contradiction:
Improveaudio processing simplicityVSAvoidaudio intelligibility
Core Design Contradiction:
Ease of operationVSReliability

Solution Approach 1:

The patent makes the audio system multi-functional by integrating it with face detection capabilities. The same audio processing system that handles directional audio also utilizes face detection data to enhance speaker differentiation, allowing the system to adapt to various scenarios (single speaker, multiple speakers, moving speakers) without requiring completely separate processing paths.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent implements dynamic audio processing that adapts to changing speaker positions and configurations. By continuously updating face detection data and recalculating audio positioning based on current speaker locations, the system maintains high intelligibility even when speakers are moving or entering/exiting the frame, transforming static audio processing into a dynamic responsive system.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS12047739B2Stereo sound generation using microphone and/or face detection
Publication Date: 2024.07.23 CISCO TECHNOLOGY INC
  • US12047739B2 patent drawing
  • US12047739B2 patent drawing
  • US12047739B2 patent drawing

AI summary

Presented herein are techniques to generate a stereo sound signal based on direction of arrival of sound signals. A method includes receiving sound signals at a microphone, outputting a mono sound signal from the microphone, determining a direction of arrival of the sound signals with respect to the microphone, generating, from the mono audio signal, a stereo sound signal having a first channel and a second channel, wherein an amplitude of the first channel and an amplitude of the second channel are weighted, respectively, based on the direction of arrival of the sound signals at the microphone.