Audio-Video Talker Localization Using Motion Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing videoconferencing solutions for localizing an active talker often rely solely on audio information, leading to reduced accuracy and increased hardware and computational requirements, especially when the speaker is facing away from the device.

Innovation Solution

A method that combines audio and motion information analysis, using a unique algorithm to weight lower frequencies and detect motion at candidate angles to accurately identify the active talker, allowing for high-definition display with fewer cameras and less computational resources.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If only audio information is used to localize an active talker, then hardware requirements are reduced, but localization accuracy decreases and the system becomes more cumbersome

Engineering Contradiction:
Improvelocalization accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines audio information processing with motion detection from video feed to localize the active talker. The audio module identifies candidate talker positions while the video module detects motion, and their results are integrated to determine the final talker location, achieving high accuracy without requiring multiple cameras

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The video feed serves multiple functions: it provides motion detection data for talker localization and simultaneously serves as the display output for the conference participants. This multi-functionality reduces hardware requirements while maintaining localization accuracy

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If multiple cameras are used to improve talker localization accuracy, then measurement precision increases, but device complexity and cost increase

Engineering Contradiction:
Improvetalker localization accuracyVSAvoidnumber of cameras
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system merges audio-based candidate position identification with video-based motion detection to achieve accurate talker localization using only a single camera. The audio module provides directional information while the video module confirms motion at the predicted location, eliminating the need for multiple cameras

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The audio processing module acts as an intermediary that predicts talker position from audio signals, allowing the single camera to focus on detecting motion at the predicted location rather than scanning the entire field of view. This intermediary step enables accurate localization with minimal video processing

Inventive Principle:
Principle #24Intermediary (Mediator)

3Reliability

If audio-only processing is used, then computational resources are reduced, but localization accuracy and reliability decrease

Engineering Contradiction:
Improvelocalization reliabilityVSAvoidcomputational resource usage
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The audio processing module performs preliminary action by identifying candidate talker positions and time intervals before the video motion detection is executed. This preliminary filtering allows the system to focus computational resources on verifying motion only at the most likely candidate positions, improving reliability while controlling computational load

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial action by using audio processing to identify candidate positions and then applying video motion detection only to those specific candidates rather than analyzing the entire video feed. This selective approach achieves reliable localization with reduced computational requirements

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10122972B2System and method for localizing a talker using audio and video information
Publication Date: 2018.11.06 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US10122972B2 patent drawing
  • US10122972B2 patent drawing
  • US10122972B2 patent drawing

AI summary

A videoconferencing endpoint includes at least one processor a number of microphones and at least one camera. The endpoint can receive audio information and visual motion information during a teleconferencing session. The audio information includes one or more angles with respect to the microphone from a location of a teleconferencing session. The system evaluates the audio information is evaluated to determine at least one candidate angle corresponding to a possible location of an active talker. The candidate angle can be analyzed further with respect to the motion information to determine whether the candidate angle correctly corresponds to person who is speaking during the teleconferencing session. The person's face can then be framed within a frame view.