Unstructured Video Stream Semantic Labeling for XR

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional video streams lack semantic labeling, making it difficult for machine systems to identify and manipulate objects represented by pixel values, as they only process images independently of semantic content.

Innovation Solution

A method involving a first electronic device with image sensors and processors that generates pixel characterization vectors, determines instance label values, and adds semantic label values to identify objects within an unstructured video stream, enabling semantic understanding and display of extended reality content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional video stream processing is used, then processing speed and simplicity are maintained, but semantic understanding and object identification capability are lost

Engineering Contradiction:
Improvesemantic informationVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system performs preliminary semantic labeling and object identification on video frames before further processing. By pre-computing pixel characterization vectors, instance labels, and semantic labels during the initial processing stage, the system ensures semantic information is captured and preserved for subsequent operations without requiring complex real-time analysis later.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The video processing pipeline is segmented into distinct functional modules: pixel characterization vector generation, instance segmentation labeling, semantic labeling, and object identification. Each module handles a specific aspect of semantic understanding, allowing the system to process complex semantic information through a series of simpler, specialized steps rather than a single complex operation.

Inventive Principle:
Principle #1Segmentation

2Adaptability or versatility

If semantic labeling is added to video streams, then object identification and manipulation capability are improved, but data processing complexity increases

Engineering Contradiction:
Improveobject manipulation capabilityVSAvoidprocessing complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The semantic labeling system is designed to be universally applicable to various video processing tasks and different types of objects. The pixel characterization vectors and semantic labels generated can be used for multiple purposes including object identification, manipulation, tracking, and analysis, making the added processing complexity beneficial across multiple functions rather than serving a single purpose.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If pixel characterization vectors and semantic labels are generated for all pixels, then object identification accuracy is improved, but processing time and computational load increase

Engineering Contradiction:
Improveobject identification accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

Rather than processing every pixel uniformly, the system focuses computational resources on pixels that are more likely to be part of objects of interest. The pixel characterization vector generation and semantic labeling are applied selectively based on detected features and patterns, performing partial processing on the full pixel set while maintaining high identification accuracy for relevant regions.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20220012283A1Capturing Objects in an Unstructured Video Stream
Publication Date: 2022.01.13 APPLE INC
  • US20220012283A1 patent drawing
  • US20220012283A1 patent drawing
  • US20220012283A1 patent drawing

AI summary

A method includes obtaining a first unstructured video stream that provides pixel values for a plurality of pixels and corresponds to a portion of a second unstructured video stream being displayed on a second electronic device different from the first electronic device. Obtaining the first unstructured video stream includes obtaining pass-through image data including the portion of a second unstructured video stream. The method includes generating respective pixel characterization vectors for a first portion of the plurality of pixels. Generating each of the respective pixel characterization vectors includes determining a respective instance label value. The method includes identifying a first object within the first portion of the plurality of pixels associated with a particular instance label value. The method includes generating respective semantic label values corresponding to pixels associated with the first object. The respective semantic label values are added to pixel characterization vectors associated with the first object.