Video Encapsulation Metadata for Primary-Auxiliary Sequence Alignment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing standards like ISO/IEC 14496-12 and ISO/IEC 23002-3 fail to encapsulate primary and auxiliary video sequences with different characteristics beyond spatial and temporal resolutions, projection, and geometry due to varying sensor manufacturing methods, leading to misalignment and inefficiencies in processing.

Innovation Solution

The method involves writing and parsing encapsulated data that includes alignment status, resampling data, and calibration information to align and process primary and auxiliary video sequences with different characteristics, allowing efficient storage and transmission while enabling real-time alignment at the receiving end.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If existing standards (ISO/IEC 14496-12 and ISO/IEC 23002-3) are used to encapsulate video sequences, then storage and transmission are simplified, but video sequences with different characteristics beyond spatial and temporal resolutions cannot be properly encapsulated, leading to misalignment

Engineering Contradiction:
Improvecapability to encapsulate video sequences with different characteristicsVSAvoidalignment precision of visual contents
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The encapsulation structure is segmented into distinct components: primary video sequence data, auxiliary video sequence data, and separate alignment status metadata. This segmentation allows each component to be independently processed and stored with its specific characteristics, enabling support for diverse video types while maintaining proper alignment through dedicated metadata fields.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The invention adds a new dimension to the encapsulation format by introducing explicit alignment status metadata beyond the traditional spatial and temporal resolution parameters. This additional dimensional information (alignment status indicating projection and geometry characteristics) enables the system to handle video sequences with different sensor manufacturing characteristics while maintaining compatibility with existing standards.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Productivity

If video sequences from sensors with different manufacturing characteristics are encapsulated without alignment status information, then storage efficiency is maintained, but visual contents misalignment occurs during playback

Engineering Contradiction:
Improvestorage and transmission efficiencyVSAvoidalignment accuracy of visual contents
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The alignment status is determined and encoded during the initial encapsulation process at the sender side. By performing this alignment analysis in advance and storing the results in metadata, the system avoids the need for complex real-time alignment computations during playback, thereby maintaining storage efficiency while ensuring reliable visual content alignment.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The alignment status metadata acts as an intermediary between the primary and auxiliary video sequences. This intermediate data structure carries information about projection and geometry characteristics, enabling the receiver to properly align visual contents from sensors with different manufacturing characteristics without requiring direct complex transformations between the video streams.

Inventive Principle:
Principle #24Intermediary (Mediator)

3Manufacturing precision

If alignment processing is performed at the receiving end in real-time, then alignment accuracy is improved, but computational overhead increases

Engineering Contradiction:
Improvealignment precision of visual contentsVSAvoidcomputational energy consumption
Core Design Contradiction:
Manufacturing precisionVSUse of energy by moving object

Solution Approach 1:

The alignment status is pre-determined and encoded in metadata during the encapsulation phase. This preliminary action shifts the computational workload from the receiving end to the sending end, where the alignment characteristics can be analyzed once and stored. The receiver then only needs to read and apply this pre-computed alignment information, dramatically reducing real-time computational energy consumption while maintaining high alignment precision.

Inventive Principle:
Principle #10Preliminary action

4Adaptability or versatility

If encapsulation format supports only spatial and temporal resolutions, then device compatibility is maintained, but device agnosticism is reduced

Engineering Contradiction:
Improvedevice agnosticismVSAvoidencapsulation format complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The encapsulation format is designed with multi-functionality, maintaining compatibility with existing standards for spatial and temporal resolution while adding support for projection and geometry characteristics. The alignment status metadata serves multiple purposes: it enables device agnosticism by accommodating different sensor manufacturing characteristics, maintains backward compatibility with legacy systems, and provides a foundation for future extensions without requiring fundamental format changes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250373881A1Method and apparatus of signaling encapsulated data representing primary video sequence associated with auxiliary video sequence, and method and apparatus of parsing encapsulated data representing primary video sequence associated with auxiliary video sequence
Publication Date: 2025.12.04 BEIJING XIAOMI MOBILE SOFTWARE CO LTD
  • US20250373881A1 patent drawing
  • US20250373881A1 patent drawing
  • US20250373881A1 patent drawing

AI summary

A method of signaling encapsulated data representing a primary video sequence associated with an auxiliary video sequence, the primary and auxiliary video sequences resulting from image projections methods applied on signals captured by sensors, includes: writing visual contents of the primary and auxiliary video sequences as encapsulated data; and writing an alignment status as encapsulated data, the alignment status indicating whether the visual contents of the primary and auxiliary video sequences are aligned.