Independently Configurable Annotation Layers for Live Video Compositing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video communication systems struggle to effectively present multiple composite videos with annotation layers during live sessions, leading to limited instruction and collaboration, particularly in hybrid classrooms, due to technological intensity and hardware requirements.

Innovation Solution

A system that generates composite videos with media backgrounds and allows multiple presenters to overlay live annotations, enabling seamless presentation and collaboration by determining user boundaries, providing media backgrounds, and managing annotation permissions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Device complexity

If one video shows both the instructor and the whiteboard, then the setup is simple, but the instructor must turn her back to the students and can obscure the writing with her physical presence

Engineering Contradiction:
Improvevideo setup complexityVSAvoidinstructor visibility and whiteboard visibility
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The system segments the video content into multiple independent video feeds (instructor video, whiteboard video, student presentation video) that can be separately captured, processed, and composed. This allows the instructor to face the students while the whiteboard is captured by a separate camera, eliminating the need to turn around or risk obscuring the board.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system transitions from a single two-dimensional video view to a multi-layered composite video structure with multiple dimensions of content (instructor layer, whiteboard layer, annotation layer). This dimensional expansion allows all elements to be simultaneously visible without physical obstruction.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Loss of information

If two cameras are used trained on the instructor and the whiteboard, then both can be shown, but it is difficult to share the whiteboard meaningfully in addition to the instructor's own video

Engineering Contradiction:
Improvewhiteboard and instructor visibilityVSAvoidvideo switching complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system merges multiple separate video feeds (instructor camera, whiteboard camera, student presentation camera) into a single composite video stream. This integration allows the instructor's video and whiteboard video to be displayed simultaneously in a unified view, eliminating the need for manual switching and making it easy to share both elements meaningfully.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The composite video system serves multiple functions simultaneously: it displays the instructor's face, shows the whiteboard content, incorporates student presentations, and accepts live annotations. This multi-functionality replaces the need for separate video switching operations with a single universal video composition system.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Loss of information

If far-end students are catered to with specialized hardware, then instruction quality improves, but near-end participants' instruction is limited

Engineering Contradiction:
Improvefar-end student instruction qualityVSAvoidinstruction accessibility for all participants
Core Design Contradiction:
Loss of informationVSAdaptability or versatility

Solution Approach 1:

The annotation and composite video system is designed to be universally accessible to all participants regardless of their location (near-end or far-end). The same technology that enhances far-end student experience also benefits near-end students, as everyone can view the composite video with annotations and interact with the same tools, eliminating the instruction gap between different participant groups.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Adaptability or versatility

If students present before the class, then student participation increases, but they are unable to simultaneously present materials while showing themselves

Engineering Contradiction:
Improvestudent participationVSAvoidpresenter visibility and material visibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The system merges the student presenter's video feed with their presentation materials into a single composite view. This allows students to simultaneously show themselves and their presentation materials without requiring separate devices or complex switching, thereby increasing participation while maintaining visibility of both elements.

Inventive Principle:
Principle #5Merging (Combining)

5Loss of information

If a teacher marks up student presentations, then instruction quality improves, but the teacher is unable to do so with existing systems

Engineering Contradiction:
Improveinstructional feedback capabilityVSAvoidannotation system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system introduces an annotation layer as an intermediary between the teacher and the student presentation. This layer allows the teacher to draw, write, and mark up the presentation in real-time without directly interacting with the presentation software itself, simplifying the process while enabling rich instructional feedback.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12407538B2Independently configurable annotation layers
Publication Date: 2025.09.02 ZOOM COMMUNICATIONS INC
  • US12407538B2 patent drawing
  • US12407538B2 patent drawing
  • US12407538B2 patent drawing

AI summary

Methods, systems, and apparatus, including computer programs encoded on computer storage media relate to a method for providing multi-point video presentations with live annotations within a communication platform. First, the system receives video feeds depicting imagery of a number of users. The system then determines a boundary about each user in the video feeds, with the boundaries each having an interior portion and an exterior portion. The system provides a media background for the exterior portions, then generates a composite video for each of the feeds. The system then determines that one or more client devices have annotation permissions, and receives one or more annotation inputs corresponding to at least one of the composite videos. Finally, the system updates at least one of the composite videos to additionally depict the annotation inputs within a third layer.