Cascaded Camera View Selection via Audio Subject Tracking
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing videoconferencing systems fail to automatically select the optimal view from multiple non-homogeneous camera devices, leading to unsatisfactory participant capture during videoconferences.
Innovation Solution
A method involving a cascaded network of cameras and microphones, where the location of a subject is determined using audio data from microphones, and this information is used to estimate and select the optimal view from multiple camera perspectives, ensuring the subject is aligned more closely with the camera capturing the preferred feed.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple non-homogeneous camera devices are used to capture views from different angles, then the coverage and variety of views are improved, but the automatic selection of optimal view becomes difficult and unsatisfactory
Solution Approach 1:
The patent introduces an intermediary processing system that receives video feeds from multiple non-homogeneous cameras and uses audio data as a mediator to determine subject location. This intermediary layer automatically selects the optimal view by comparing subject positions across camera feeds, resolving the complexity of manual view selection while maintaining versatile coverage from multiple camera angles.
Solution Approach 2:
The system implements feedback by continuously monitoring audio signals to track subject location and using this information to dynamically select the best camera view. The audio-based location detection provides real-time feedback that guides the automatic view selection process, ensuring the optimal perspective is always chosen despite the complexity of multiple non-homogeneous cameras.
2Reliability
If the system automatically selects views based on subject location, then the videoconferencing quality is improved, but the complexity of processing multiple camera feeds increases
Solution Approach 1:
The patent replaces complex mechanical view selection processes with an audio-based detection system. Instead of relying on complex image processing and subject tracking algorithms across multiple camera feeds, the system uses audio signals to determine subject location, which then guides simple view selection logic. This substitution reduces processing complexity while maintaining reliable subject capture accuracy.
3Measurement precision
If audio data is used to detect subject location, then the accuracy of view selection is improved, but the dependency on audio quality increases
Solution Approach 1:
The patent makes the audio system multi-functional by using it not only for communication but also for subject location detection and view selection. This universal use of audio data allows the system to achieve precise measurement of subject location without requiring separate sensing mechanisms, though it does create sensitivity to audio quality variations.
Data Source
AI summary
A method (1000) for operating cameras (202) in a cascaded network (100), comprising: capturing a first view (1200) with a first lens (326) having a first focal point (328) and a first centroid (352), the first view (1200) depicting a subject (1106); capturing a second view (1202) with a second lens (326) having a second focal point (328) and a second centroid (352); detecting a first location of the subject (1106), relative the first lens (326), wherein detecting the first location of the subject (1106), relative the first lens (326), is based on audio captured by a plurality of microphones (204); estimating a second location of the subject (1106), relative the second lens (326), based on the first location of the subject (1106) relative the first lens (326); selecting a portion (1206) of the second view (1202) as depicting the subject (1106) based on the estimate of the second location of the subject (1106) relative the second lens (326).


