Cross-Modal Data Reconstruction and Compression for Missing Modalities

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing technologies struggle to effectively utilize and manage multimodal data, leading to issues such as missing or erroneous data and inefficient data compression across different data modalities.

Innovation Solution

Utilizing cross-modal machine learning models to analyze and link different data modalities, enabling the reconstruction of missing data and efficient compression by identifying calibration points and leveraging relationships between data types.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If traditional single-modality data processing is used, then data management is simple, but data completeness and accuracy deteriorate when data is missing or erroneous

Engineering Contradiction:
Improvedata completenessVSAvoiddata processing complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent combines multiple data modalities (audio, video, text) into a unified processing framework. Cross-modal machine learning models integrate these different modalities to互相 supplement each other, allowing the system to reconstruct missing or erroneous data by leveraging relationships across modalities, thereby improving data completeness without requiring separate processing systems for each modality.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces cross-modal machine learning models as intermediary components that mediate between different data modalities. These models learn and exploit the relationships between modalities (e.g., between audio and video, or between text and video), enabling the system to infer missing information from available modalities while maintaining manageable processing complexity through standardized intermediary processing steps.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If all multimodal data is stored and transmitted, then data completeness is maintained, but storage requirements and transmission bandwidth increase significantly

Engineering Contradiction:
Improvedata integrityVSAvoiddata volume
Core Design Contradiction:
Loss of informationVSQuantity of substance

Solution Approach 1:

The patent extracts and stores only the essential calibration points and relationship information from complete multimodal datasets. Instead of storing all raw data, the system identifies key reference points that capture the essential cross-modal relationships, allowing for efficient storage and reconstruction of complete data when needed, thereby reducing data volume while maintaining data integrity.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary analysis to identify calibration points and establish cross-modal relationships before actual data storage or transmission. By pre-processing the data to extract essential relationships and calibration information, the system reduces the amount of data that needs to be stored or transmitted while ensuring that complete information can be reconstructed when required, thus preventing information loss with reduced data volume.

Inventive Principle:
Principle #10Preliminary action

3Productivity

If cross-modal analysis is performed to identify relationships between modalities, then data compression efficiency improves, but computational complexity increases

Engineering Contradiction:
Improvecompression efficiencyVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent segments the cross-modal analysis process into distinct stages: calibration point identification, relationship learning, and compression application. This segmentation allows the system to perform complex cross-modal analysis only when needed for calibration and relationship establishment, then apply these pre-computed relationships for efficient compression, thereby improving compression efficiency while managing computational complexity through divided processing steps.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12524505B2Cross-modal data completion and compression
Publication Date: 2026.01.13 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12524505B2 patent drawing
  • US12524505B2 patent drawing
  • US12524505B2 patent drawing

AI summary

Aspects relate to analyzing multimodal datasets using one or more cross-modal machine learning models. The machine learning models are operable to generate analysis data related to the different data modalities. The analysis data can be used to identify related portions of data in the different modalities. Once these relationships between the different modalities of a data are identified, the relationships can be leveraged to perform various different processes. For example, a first portion of data having a first modality can be used to reconstruct missing or erroneous data from a second modality. The relationship between content stored in the different modalities can further be leveraged to perform compression on multimodal data sets.