Composite Video Stream Reframing for Unified Speaker and Content Viewing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conference platforms require participants to follow two separate video streams for viewing a speaker and shared digital content, which can be inefficient and resource-intensive.

Innovation Solution

A composite video stream is generated that integrates digital content with participant video by cropping, reframing, and compositing layers to create a unified video output, allowing seamless presentation of both within a single region.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If separate video streams are used for speaker and digital content, then participants can view both, but the system becomes resource-intensive and inefficient

Engineering Contradiction:
Improvevideo conference efficiencyVSAvoidprocessing resources
Core Design Contradiction:
ProductivityVSLoss of energy

Solution Approach 1:

The patent merges the speaker video stream and digital content into a single composite video stream. The system composites the speaker's video feed with digital content (such as slides or documents) into one unified stream, eliminating the need for participants to switch between or follow separate streams. This reduces processing resource consumption and improves overall system efficiency.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The composite video stream serves multiple functions simultaneously - it displays both the speaker and the digital content in a single stream. This multi-functional approach allows the system to replace what would traditionally require two separate video streams, reducing the computational burden on the system while maintaining comprehensive information delivery.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Ease of operation

If separate video streams are used for speaker and digital content, then both can be displayed, but participants must follow multiple streams

Engineering Contradiction:
Improveparticipant viewing experienceVSAvoidvideo stream management
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system combines multiple video streams into a single composite stream that simultaneously presents both the speaker and digital content. This eliminates the complexity of managing and switching between multiple streams, simplifying the participant experience while maintaining all necessary visual information in one unified feed.

Inventive Principle:
Principle #5Merging (Combining)

3Stability of the object's composition

If digital content is overlaid on participant video, then integration is improved, but visual clarity and contrast may be compromised

Engineering Contradiction:
Improvevideo stream integrationVSAvoidvisual contrast
Core Design Contradiction:
Stability of the object's compositionVSIllumination intensity

Solution Approach 1:

The system applies different processing characteristics to different regions of the composite video stream. The speaker's video feed maintains its original quality and contrast in certain regions, while digital content is overlaid in specific areas. This localized quality maintenance ensures that both the speaker and content remain visually distinct and clear, preventing degradation of visual contrast while achieving integration.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS12401756B2Generating a composite video stream having digital content and a participant video for real-time presentation in a user interface of a video conference system
Publication Date: 2025.08.26 GOOGLE LLC
  • US12401756B2 patent drawing
  • US12401756B2 patent drawing
  • US12401756B2 patent drawing

AI summary

A first event associated with a first client device of multiple client devices of participants of a video conference is identified. The first event indicates a request to present content in a user interface (UI) including regions each corresponding to one of multiple video streams from the client devices. A first video layer that includes a reframed visual representation of the first participant is generated from the first video segment of a first video stream from the first client device. A first content layer is generated. The first video layer and the first content layer are composited into the first composite video segment in which the reframed visual representation of the first participant is positioned adjacent to the first part of the content. The first composite video segment is provided to the client devices as a real-time video stream for presentation.