Spatial Voice Tracking for Multi-User Conferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current noise reduction techniques in video conferencing fail to ensure that all in-room users' voices are loud and clear for far-end users, as they are designed for single users and do not account for varying physical locations, leading to inconsistent voice detection and difficulty in focusing on a specific user's voice.

Innovation Solution

A tracking system that adjusts voice levels based on in-room users' physical locations, allowing for equalization and muting of voices from different directions, using metadata and face IDs to determine user positions and enhance the voice of the focused user, without requiring additional hardware.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If noise reduction techniques are applied to video conferencing, then background noise is reduced, but all in-room users' voices cannot be made loud and clear for far-end users

Engineering Contradiction:
Improvebackground noiseVSAvoidvoice detection consistency
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent segments the audio processing by creating separate voice channels for different spatial locations. Each microphone array channel is processed independently to capture and enhance voices from specific directions, allowing selective noise reduction and voice enhancement for multiple users simultaneously rather than treating all audio as a single mixed signal

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality enhancement by adjusting audio characteristics based on spatial location. Different gain levels, noise reduction parameters, and voice enhancement settings are applied to voices detected from different directions and distances, ensuring each user's voice is optimized according to their specific position in the room

Inventive Principle:
Principle #3Local quality

2Measurement precision

If single-user noise reduction techniques are used, then individual user voice quality is improved, but voices from multiple users at different physical locations become inconsistent

Engineering Contradiction:
Improvevoice qualityVSAvoidmulti-user location adaptation
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent implements dynamic adaptability by continuously tracking the positions of multiple users and adjusting audio processing parameters in real-time. The system dynamically updates gain levels, beamforming directions, and noise reduction settings based on current user locations detected through metadata and face ID tracking, allowing the system to adapt to changing spatial configurations

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent creates a universal audio processing system that handles multiple users simultaneously with different spatial positions. The same processing pipeline is applied to all detected voices, with parameters automatically adjusted based on each user's location, making the system versatile for various meeting configurations without requiring separate processing for each user

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Reliability

If voice levels are adjusted based on physical location, then all users' voices become loud and clear, but the system complexity increases

Engineering Contradiction:
Improvevoice clarityVSAvoidtracking system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent uses metadata and face ID tracking as intermediary systems to bridge the gap between physical user positions and audio processing adjustments. These intermediaries provide the necessary spatial information without requiring complex direct measurement systems, simplifying the overall architecture while enabling location-based voice enhancement

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent creates a virtual model of user positions and spatial relationships based on metadata and camera data. This virtual copy of the physical environment is then used to drive audio processing decisions, avoiding the need for complex physical sensors and measurement systems while achieving accurate location-based voice adjustment

Inventive Principle:
Principle #26Copying

4Ease of operation

If far-end users need to focus on a specific user's voice, then voice selection capability is improved, but automatic tracking of user positions is required

Engineering Contradiction:
Improvevoice focusing capabilityVSAvoiduser position tracking
Core Design Contradiction:
Ease of operationVSExtent of automation

Solution Approach 1:

The patent implements self-service automation where the system automatically tracks user positions and adjusts voice enhancement without requiring manual intervention. The metadata and face ID systems continuously monitor user locations and automatically update audio processing parameters, allowing far-end users to focus on specific speakers simply by indicating interest while the system handles all tracking and adjustment operations autonomously

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20240194215A1Equalizing and tracking speaker voices in spatial conferencing
Publication Date: 2024.06.13 INTEL CORP
  • US20240194215A1 patent drawing
  • US20240194215A1 patent drawing
  • US20240194215A1 patent drawing

AI summary

This disclosure describes systems, methods, and devices related to user tracking. A device may identify metadata comprising depth sensing information and camera information received from an in-room device located at a first location having a first camera. The device may perform face recognition on one or more in-room users. The device may calculate a distance of a first in-room user based on the metadata and a first number of pixels across the face of the first in-room user. The device may calculate a distance between the first in-room user and a second in-room user based on the metadata and the first number of pixels across the face of the first in-room user and a number of pixels across the face of the second in-room user.