Videoconference Display Controller Active Speaker Highlighting
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large-scale videoconferences, it is difficult for viewers to identify the speaker's site when multiple sites are connected, as the display area for each site is reduced, making it hard to distinguish the speech site from others.
Innovation Solution
A videoconference device that includes a video input unit, voice input unit, communication controller, and display controller, which synthesizes video data from multiple sites based on screen layout and detects voice levels to highlight the speech site, ensuring the display area of the speech site is larger than others.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If video data from each site is displayed in regions having the same area, then all sites can be displayed simultaneously, but the display area of each video data is reduced when the number of sites is large, making it hard to understand the speech site
Solution Approach 1:
The patent applies local quality by giving different display areas to different video regions based on their importance. The speech site (active speaker) is assigned a larger display area while other sites maintain smaller equal areas. This resolves the contradiction by allowing multiple sites to be displayed simultaneously while ensuring the speech site receives sufficient attention for viewer understanding.
Solution Approach 2:
The patent implements dynamics by making the display layout adaptive rather than static. The system dynamically adjusts which site receives the larger display area based on real-time detection of the active speaker. This allows the display configuration to change automatically according to conference conditions, maintaining effectiveness across different numbers of sites and speaking patterns.
2Quantity of substance
If the number of connected sites increases, then more participants can join the conference, but the display area for each site becomes smaller, reducing viewer understanding
Solution Approach 1:
The patent resolves this contradiction by applying local quality to the display layout, where the speech site is granted a larger display area while other sites maintain smaller equal areas. This ensures that even as the number of connected sites increases, the critical information (speech site) remains prominently displayed, preventing loss of viewer understanding.
Solution Approach 2:
The system uses feedback from voice level detection to automatically adjust the display layout. The level detector continuously monitors voice levels at each site, and this feedback drives the display controller to highlight the appropriate speech site. This closed-loop approach ensures that viewer understanding is maintained regardless of the number of connected sites.
Data Source
AI summary
A videoconference device displays video data from a speech site such that a viewer can easily understand even in a case where the number of sites is large. A communication controller receives each piece of video data and voice data from conference terminal devices of a plurality of other sites. A video and voice synthesizer determines a screen layout depending on the number of sites participating in a videoconference, and generates synthesized video data obtained by synthesizing video data of each site according to the screen layout. At this time, the video and voice synthesizer generates the synthesized video data such that display of the video data of each site where a level of voice data is higher than or equal to a threshold is highlighted more than display of video data of the other sites. A video and voice output controller displays the synthesized video data on a screen of a display device.


