Context-Aware Video Stream Compositing for Object Segmentation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing video conferencing software struggles to segment and include objects from a user's environment in the video stream, leading to distortions or incomplete representation of objects when users attempt to describe or show objects during a call.

Innovation Solution

The implementation of multimodal analysis to determine the interactive context of the user, identifying and tracking real-world objects, and dynamically compositing the video stream to include both the user and relevant objects, even when they are not directly in front of the user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional virtual background segmentation is used, then the user outline can be segmented, but objects in the user's environment cannot be properly segmented or included in the video stream

Engineering Contradiction:
Improveobject segmentation capabilityVSAvoidsegmentation accuracy
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The patent segments the video processing task into multiple components: user detection, object detection, interaction analysis, and composite stream generation. By dividing the segmentation task into these independent modules, the system can accurately identify both users and objects separately, then combine them appropriately in the video stream.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an intermediary analysis layer that examines video data, audio data, and screen content to determine interaction context. This intermediary layer acts as a mediator between raw input data and the final segmentation decisions, enabling accurate identification of objects that users are interacting with.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Adaptability or versatility

If the system tries to include all objects in the user's environment, then object representation improves, but video stream complexity and processing requirements increase

Engineering Contradiction:
Improveenvironmental object inclusionVSAvoidvideo processing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent performs preliminary detection and classification of objects in the user's environment before final video stream composition. By pre-identifying potential objects of interest and their relevance to user interaction, the system reduces the complexity of real-time processing while maintaining comprehensive object inclusion.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent applies partial action by selectively including only those objects that are relevant to user interaction in the video stream, rather than including all detected objects. This selective approach balances comprehensive environmental representation with manageable processing complexity.

Inventive Principle:
Principle #16Partial or excessive action

3Device complexity

If the system segments only the user outline, then processing remains simple, but objects described or shown by users appear distorted or incomplete

Engineering Contradiction:
Improvesegmentation processing simplicityVSAvoidobject representation completeness
Core Design Contradiction:
Device complexityVSLoss of information

Solution Approach 1:

The patent merges the user segmentation results with object segmentation results to create a composite video stream. By combining these segmentation layers and applying appropriate transparency and positioning, the system preserves complete object representation while maintaining processing efficiency through reusable segmentation components.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250203037A1Context-aware object interaction for video conference stream compositing
Publication Date: 2025.06.19 INTEL CORP
  • US20250203037A1 patent drawing
  • US20250203037A1 patent drawing
  • US20250203037A1 patent drawing

AI summary

Various aspects of context aware object interaction prediction and segmentation, including video stream segmentation of a virtual background during a video call, are discussed. An example method of segmentation includes: receiving video data that depicts a human user and an object in a scene; receiving context data from another other data source, which is related to an interaction of the human user with the object; analyzing the context data to determine a shape of the object and a type of the interaction of the human user with the object; and generating a video stream that includes a virtual background overlaid on the video data. The virtual background can be segmented based on at least one outline of the human user, and the virtual background can be further segmented based on the shape of the object and the type of the interaction of the human user with the object.