3D Video Compositing With Scenario Templates for Immersive Meetings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video conferencing technologies present user videos in a fixed, two-dimensional format, lacking immersion and naturalness, failing to provide a compelling meeting experience.

Innovation Solution

A video generating device and method that create three-dimensional portrait models from real-time images, select a suitable three-dimensional scenario template based on user quantity and position, and composite these models within a spatially labeled template to generate immersive videos.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional fixed two-dimensional video format is used, then device complexity is low, but immersion and naturalness are poor

Engineering Contradiction:
Improveimmersion and naturalnessVSAvoiddevice complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent transforms conventional two-dimensional video presentation into three-dimensional virtual scenario presentations. By constructing 3D virtual scenarios with spatial labels and compositing user images into these three-dimensional spaces, the system provides immersive and natural meeting experiences, directly resolving the contradiction between simplicity and immersion through dimensional transformation.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Adaptability or versatility

If fixed dimensional arrangement is used, then ease of operation is high, but adaptability to different user quantities is poor

Engineering Contradiction:
Improveadaptability to user quantityVSAvoidease of operation
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements dynamic adaptation by allowing the video presentation format to automatically adjust based on the number of users. The system dynamically selects appropriate three-dimensional scenario templates and spatial arrangements according to user quantity, enabling flexible adaptation without requiring manual configuration, thus resolving the contradiction between ease of operation and adaptability.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes key parameters including the number of spatial labels, spatial arrangement configurations, and scenario template selections based on user quantity. By dynamically adjusting these parameters, the system adapts to different meeting scales while maintaining automated operation, effectively resolving the contradiction between operational simplicity and adaptability.

Inventive Principle:
Principle #35Parameter changes

3Adaptability or versatility

If traditional background display is used, then manufacturing precision requirements are low, but meeting experience quality is poor

Engineering Contradiction:
Improvemeeting experience qualityVSAvoidimage compositing precision
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent introduces spatial labels as intermediary elements that facilitate precise compositing of user images into three-dimensional scenarios. These spatial labels serve as reference markers that guide the compositing process, enabling high-precision integration of multiple images while maintaining automated operation, thus resolving the contradiction between experience quality and precision requirements.

Inventive Principle:
Principle #24Intermediary (Mediator)

Data Source

PatentUS12499609B2Video generating device and method
Publication Date: 2025.12.16 INSTITUTE FOR INFORMATION INDUSTRY
  • US12499609B2 patent drawing
  • US12499609B2 patent drawing
  • US12499609B2 patent drawing

AI summary

A video generating device and method are provided. The device analyzes a plurality of real-time images corresponding to a plurality of users to segment a target image from each of the real-time images. The device generates a three-dimensional portrait model corresponding to each of the users based on the target image of each of the real-time images. The device determines a first three-dimensional scenario template from the three-dimensional scenario templates based on a user quantity of the users and the position quantity corresponding to each of the three-dimensional scenario templates. The device composites the three-dimensional portrait models to the spatial label position of the first three-dimensional scenario template to generate a video corresponding to the users.