Video Frame Occupancy Analysis for Audio Intimacy Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current video conferencing systems fail to accurately adjust audio output based on the location of the near-end user within the video frame, leading to inconsistent intimacy and social characteristics in audio reproduction.

Innovation Solution

A video conferencing system that analyzes video frames and camera settings to determine the user's position, adjusting audio parameters such as loudness, directivity, reverberation, and equalization to mimic the user's intended intimacy level, by transmitting these parameters to the far-end system for precise audio reproduction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If audio output is transmitted without adjustment based on user location, then the system is simple and easy to operate, but the audio reproduction lacks accuracy in representing intimacy and social characteristics

Engineering Contradiction:
Improveaudio reproduction accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The system performs preliminary analysis of video frames to determine user position and occupancy ratio before audio transmission. Audio parameters are pre-adjusted based on the determined user location, allowing the far-end system to receive already-optimized audio without requiring complex real-time processing capabilities.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary processing stage where video frame analysis determines user position, which then serves as input for audio parameter adjustment. This intermediary mechanism (video analysis → position determination → audio parameter mapping) bridges the gap between simple video capture and accurate audio reproduction without requiring direct complex interaction between all system components.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If audio parameters are adjusted based on video frame analysis, then audio intimacy and social characteristics are accurately reproduced, but the processing time and computational resources increase

Engineering Contradiction:
Improveaudio reproduction fidelityVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system analyzes only essential video frame characteristics (user position and occupancy ratio) rather than performing comprehensive video analysis. This partial action approach focuses computational resources on the specific parameters needed for audio adjustment, achieving reliable audio reproduction without excessive processing time.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The patent transforms spatial parameters from video analysis (user position, frame occupancy ratio) into audio parameters (loudness, directivity, reverberation, equalization). This parameter transformation allows the system to achieve accurate audio reproduction by mapping visual spatial information to acoustic characteristics without requiring complex audio processing.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If the system analyzes video frames to determine user position, then audio output accurately reflects user location, but the device complexity and processing requirements increase

Engineering Contradiction:
Improveaudio adaptation to user positionVSAvoidvideo analysis capability
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The video frame analysis mechanism serves multiple functions: it determines user position for audio adjustment, calculates occupancy ratio for intimacy control, and provides spatial context for overall scene understanding. This multi-functionality allows the system to achieve adaptable audio output without adding separate dedicated components for each function.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The existing video capture infrastructure is utilized to perform the analysis needed for audio adjustment. The video stream, already being captured and transmitted, is repurposed to provide spatial information for audio parameter control, allowing the system to achieve position-based audio adaptation without requiring additional sensors or complex external analysis systems.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS9883140B2Using the location of a near-end user in a video stream to adjust audio settings of a far-end system
Publication Date: 2018.01.30 APPLE INC
  • US9883140B2 patent drawing
  • US9883140B2 patent drawing
  • US9883140B2 patent drawing

AI summary

A video conferencing system is described that includes a near-end and a far-end system. The near-end system records both audio and video of one or more users proximate to the near-end system. This recorded audio and video is transmitted to the far-end system through the data connection. The video stream and/or one or more settings of the recording camera are analyzed to determine the amount of a video frame occupied by the recorded user(s). The video conferencing system may directly analyze the video frames themselves and/or a zoom setting of the recording camera to determine a ratio or percentage of the video frame occupied by the recorded user(s). By analyzing video frames associated with an audio stream, the video conferencing system may drive a speaker array of the far-end system to more accurately reproduce sound content based on the position of the recorded user in a video frame.