Video Image Composition for Priority Speaker Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-person video conferences, moderators and assigned speakers often get ignored due to background noise and the challenge of promptly focusing the camera on the speaking participant.

Innovation Solution

A video image composition method that utilizes a priority level list to determine the display levels of participants based on their identities, and detects speaking status to prioritize and display the main video image accordingly, ensuring the moderator is prominently displayed when speaking.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of time

If the camera follows the participant after they start to give a talk, then the participant can be captured in the video, but the moderator's message may be ignored by participants due to delay in camera switching

Engineering Contradiction:
Improvetime delay in camera switchingVSAvoidreliability of message delivery
Core Design Contradiction:
Loss of timeVSReliability

Solution Approach 1:

The system performs preliminary identification of the speaker before they actually start speaking by detecting voice activation or motion cues in advance. This allows the camera to switch to the speaker proactively rather than reactively, eliminating the time delay and ensuring participants notice the speaker immediately when they begin their message.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system continuously monitors audio and video inputs to detect when a participant begins to speak, providing real-time feedback to the camera control system. This feedback mechanism enables dynamic camera switching based on actual speaking status, ensuring the moderator or active speaker is always prominently displayed without manual intervention.

Inventive Principle:
Principle #23Feedback

2Loss of information

If voice-only notification is used to assign participants to speak, then the assignment can be communicated, but participants may ignore the voice due to background noises from conversations or other situations

Engineering Contradiction:
Improveinformation delivery effectivenessVSAvoidbackground noise interference
Core Design Contradiction:
Loss of informationVSObject-affected harmful factors

Solution Approach 1:

The system introduces visual indicators as an intermediary channel to complement voice notifications. When a participant is assigned to speak or when someone begins speaking, the system displays visual cues such as name tags, priority indicators, or highlighted frames around the speaker's video feed. This dual-channel approach (visual + audio) ensures message delivery even when audio alone is obscured by background noise.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system uses color changes or visual highlighting to indicate speaking status or message importance. For example, the speaker's video frame may be highlighted with a colored border, or their name may appear in a distinct color on the display. These visual signals cut through background noise interference and immediately draw participants' attention to the active speaker or assigned participant.

Inventive Principle:
Principle #32Color changes

Data Source

PatentUS12277799B2Video image composition method and electronic device
Publication Date: 2025.04.15 AMTRAN TECHNOLOGY CO LTD
  • US12277799B2 patent drawing
  • US12277799B2 patent drawing
  • US12277799B2 patent drawing

AI summary

The present disclosure provides a video image composition method including the following steps. A priority level list is obtained, and the priority level list includes multiple priority levels of multiple person identities. Multiple video streams are received. Multiple identity labels corresponding to human face frame images from the video streams are determined. The multiple display levels of the human face frame images are determined according to the identity labels and priority level list. A part of the human face frame images being in speaking status are detected. At least one of the part of the human face frame images being in speaking status is constituted as a main display area of a video image, according to the display levels.