Acoustic Speaker Zoom via Microphone Array Signal Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing videoconferencing systems face challenges in efficiently zooming in on active speakers using a single camera, as they often require additional equipment and struggle with detecting lip movements from a distance, leading to suboptimal robustness and increased costs.

Innovation Solution

A method involving a wide-angle camera and a surround microphone with multiple sensors to capture sound fields, processing sound data to determine the direction of sound propagation, and generating a signal to control video zooming around the active speaker, using techniques like source separation and beamforming to accurately identify and focus on the active speaker.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If two cameras are used (one to zoom and another to detect lip movements), then the ability to zoom on active speakers is improved, but the device complexity and cost increase

Engineering Contradiction:
Improvespeaker detection accuracyVSAvoidnumber of cameras
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The single camera system performs multiple functions: capturing the visual scene and enabling digital zoom operations. The microphone array performs dual functions of sound capture and speaker direction detection through acoustic processing. This multi-functionality eliminates the need for separate dedicated zoom cameras and lip-detection cameras.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces mechanical/optical zoom mechanisms with digital zoom processing. Instead of using a second camera with optical zoom capability, the system uses software-based digital zoom on the single camera feed, controlled by acoustic speaker detection. This substitution reduces hardware complexity while maintaining zoom functionality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If two cameras are used to achieve zoom on active speakers, then the zoom functionality is improved, but the cost increases

Engineering Contradiction:
Improvespeaker detection accuracyVSAvoidsystem cost
Core Design Contradiction:
Measurement precisionVSEase of manufacture

Solution Approach 1:

The single camera system performs multiple functions: capturing the visual scene and enabling digital zoom operations. The microphone array performs dual functions of sound capture and speaker direction detection through acoustic processing. This multi-functionality eliminates the need for separate dedicated zoom cameras and lip-detection cameras.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent replaces mechanical/optical zoom mechanisms with digital zoom processing. Instead of using a second camera with optical zoom capability, the system uses software-based digital zoom on the single camera feed, controlled by acoustic speaker detection. This substitution reduces hardware complexity while maintaining zoom functionality.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If lip movement detection is used to identify active speakers, then speaker identification is improved, but the reliability decreases when faces are far from the camera

Engineering Contradiction:
Improvespeaker identification accuracyVSAvoiddetection robustness
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent introduces acoustic signal processing as an intermediary method for speaker detection. Instead of relying directly on visual lip movement detection which fails at distances, the system uses microphone arrays to detect sound sources and their directions. This acoustic intermediary provides reliable speaker identification regardless of distance, which then controls the visual zoom.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces visual lip movement detection with acoustic sound source detection. The microphone array with beamforming or ambisonic processing substitutes for the camera-based lip detection, providing distance-independent speaker identification. This substitution maintains reliability across varying distances while achieving the same speaker detection goal.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Applied Scientific Principles

This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.

Function Achieved in This Case

This solution allows for real-time, cost-effective zooming on active speakers without the need for additional cameras, improving robustness and reducing bandwidth requirements by accurately determining the direction of sound propagation and applying digital zoom accordingly.

Implementation Method 1

A sound acquisition from a microphone comprising a plurality of sensors to capture a sound field

Methodology Applied
Scientific EffectSound propagation: Sound

Data Source

PatentEP3721248B1Processing of data of a video sequence in order to zoom on a speaker detected in the sequence
Publication Date: 2022.08.03 ORANGE SA
  • EP3721248B1 patent drawingFigure 1A~8
  • EP3721248B1 patent drawingFigure 2A~2B
  • EP3721248B1 patent drawingFigure 3

AI summary

The invention relates to processing a video sequence containing a succession of images of one or more speakers, acquired by a wide-angle camera (102), the method including: - audio acquisition from a microphone (103) including a plurality of sensors for sensing an audio field; - processing the audio data acquired by the microphone (103) in order to determine at least one direction (ANG) of origin of sound coming from a speaker (LA), in relation to an optical axis (AO) of the wide-angle camera; - generating a signal (304) containing data (ANG) of said direction of origin of the sound in relation to the optical axis (AO) of the camera, for the purpose of utilizing said signal in rendering of the acquired images by applying a zoom to a zone around the speaker (LA) emitting the sound whose direction of origin corresponds to said data of the signal.