Multimodal Document Augmentation Using Image and Text Embeddings

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Traditional technologies struggle to effectively understand and characterize multimodal unstructured data in unstructured documents, leading to difficulties in establishing knowledge bases and performing data augmentation, which limits the ability to search and retrieve valuable information.

Innovation Solution

Generate image and text embeddings using pre-trained multimodal deep learning neural networks, acquire descriptive information from a storage library based on these embeddings, and add it to the unstructured document to enrich its content.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If traditional technologies are used to process unstructured documents, then the processing method is simple, but the ability to understand and characterize multimodal data is insufficient

Engineering Contradiction:
Improveability to understand and characterize multimodal dataVSAvoidprocessing method complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent segments the unstructured document into multiple modalities (text, images, tables, formulas) and processes each modality separately through dedicated processing modules. This segmentation enables the system to understand and characterize each type of data with appropriate methods while maintaining overall system manageability.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces embedding vectors as intermediary representations that bridge different modalities. The text embedding module, image embedding module, and other components convert diverse data types into unified vector representations, enabling cross-modality understanding and characterization without requiring complex direct processing between different data types.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Quantity of substance

If data augmentation is performed on unstructured documents, then the amount and diversity of data increase, but the complexity of data processing increases

Engineering Contradiction:
Improveamount and diversity of dataVSAvoiddata processing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The patent performs preliminary processing by generating embedding vectors for text, images, and other modalities before actual data augmentation. This preliminary embedding step organizes the data in a standardized format, making subsequent augmentation operations more efficient and manageable despite the increased data volume and diversity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the representation parameters of unstructured data by converting various modalities into unified embedding vectors. This parameter transformation enables the system to handle diverse data types through consistent mathematical operations, reducing processing complexity while maintaining data amount and diversity through augmentation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12619814B2Method, device, and computer program product for data augmentation
Publication Date: 2026.05.05 DELL PROD LP
  • US12619814B2 patent drawing
  • US12619814B2 patent drawing
  • US12619814B2 patent drawing

AI summary

Embodiments of the present disclosure relate to a method, a device, and a computer program product for data augmentation. The method includes generating an image embedding based on an image in an unstructured document, and generating a text embedding based on text in the unstructured document and associated with the image. The method further includes acquiring descriptive information from a storage library based on the generated image embedding and text embedding. The method further includes adding the acquired descriptive information into the unstructured document. In this way, it can be possible not only to understand and analyze the unstructured document across modalities, but also to enrich it with a characterization of multimodal data in the unstructured document, thus increasing the amount and diversity of data.