Cascaded Camera View Selection via Audio Subject Tracking

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing videoconferencing systems fail to automatically select the optimal view from multiple non-homogeneous camera devices, leading to unsatisfactory participant capture during videoconferences.

Innovation Solution

A method involving a cascaded network of cameras and microphones, where the location of a subject is determined using audio data from microphones, and this information is used to estimate and select the optimal view from multiple camera perspectives, ensuring the subject is aligned more closely with the camera capturing the preferred feed.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple non-homogeneous camera devices are used to capture views from different angles, then the coverage and variety of views are improved, but the automatic selection of optimal view becomes difficult and unsatisfactory

Engineering Contradiction:
Improveview coverageVSAvoidcamera system
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent introduces an intermediary processing system that receives video feeds from multiple non-homogeneous cameras and uses audio data as a mediator to determine subject location. This intermediary layer automatically selects the optimal view by comparing subject positions across camera feeds, resolving the complexity of manual view selection while maintaining versatile coverage from multiple camera angles.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system implements feedback by continuously monitoring audio signals to track subject location and using this information to dynamically select the best camera view. The audio-based location detection provides real-time feedback that guides the automatic view selection process, ensuring the optimal perspective is always chosen despite the complexity of multiple non-homogeneous cameras.

Inventive Principle:
Principle #23Feedback

2Reliability

If the system automatically selects views based on subject location, then the videoconferencing quality is improved, but the complexity of processing multiple camera feeds increases

Engineering Contradiction:
Improvesubject capture accuracyVSAvoidprocessing system
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent replaces complex mechanical view selection processes with an audio-based detection system. Instead of relying on complex image processing and subject tracking algorithms across multiple camera feeds, the system uses audio signals to determine subject location, which then guides simple view selection logic. This substitution reduces processing complexity while maintaining reliable subject capture accuracy.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If audio data is used to detect subject location, then the accuracy of view selection is improved, but the dependency on audio quality increases

Engineering Contradiction:
Improvesubject location detectionVSAvoidaudio quality sensitivity
Core Design Contradiction:
Measurement precisionVSObject-affected harmful factors

Solution Approach 1:

The patent makes the audio system multi-functional by using it not only for communication but also for subject location detection and view selection. This universal use of audio data allows the system to achieve precise measurement of subject location without requiring separate sensing mechanisms, though it does create sensitivity to audio quality variations.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11803984B2Optimal view selection in a teleconferencing system with cascaded cameras
Publication Date: 2023.10.31 HEWLETT PACKARD DEVELOPMENT COMPANY LP
  • US11803984B2 patent drawing
  • US11803984B2 patent drawing
  • US11803984B2 patent drawing

AI summary

A method (1000) for operating cameras (202) in a cascaded network (100), comprising: capturing a first view (1200) with a first lens (326) having a first focal point (328) and a first centroid (352), the first view (1200) depicting a subject (1106); capturing a second view (1202) with a second lens (326) having a second focal point (328) and a second centroid (352); detecting a first location of the subject (1106), relative the first lens (326), wherein detecting the first location of the subject (1106), relative the first lens (326), is based on audio captured by a plurality of microphones (204); estimating a second location of the subject (1106), relative the second lens (326), based on the first location of the subject (1106) relative the first lens (326); selecting a portion (1206) of the second view (1202) as depicting the subject (1106) based on the estimate of the second location of the subject (1106) relative the second lens (326).