Multimodal Document Augmentation Using Image and Text Embeddings
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional technologies struggle to effectively understand and characterize multimodal unstructured data in unstructured documents, leading to difficulties in establishing knowledge bases and performing data augmentation, which limits the ability to search and retrieve valuable information.
Innovation Solution
Generate image and text embeddings using pre-trained multimodal deep learning neural networks, acquire descriptive information from a storage library based on these embeddings, and add it to the unstructured document to enrich its content.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional technologies are used to process unstructured documents, then the processing method is simple, but the ability to understand and characterize multimodal data is insufficient
Solution Approach 1:
The patent segments the unstructured document into multiple modalities (text, images, tables, formulas) and processes each modality separately through dedicated processing modules. This segmentation enables the system to understand and characterize each type of data with appropriate methods while maintaining overall system manageability.
Solution Approach 2:
The patent introduces embedding vectors as intermediary representations that bridge different modalities. The text embedding module, image embedding module, and other components convert diverse data types into unified vector representations, enabling cross-modality understanding and characterization without requiring complex direct processing between different data types.
2Quantity of substance
If data augmentation is performed on unstructured documents, then the amount and diversity of data increase, but the complexity of data processing increases
Solution Approach 1:
The patent performs preliminary processing by generating embedding vectors for text, images, and other modalities before actual data augmentation. This preliminary embedding step organizes the data in a standardized format, making subsequent augmentation operations more efficient and manageable despite the increased data volume and diversity.
Solution Approach 2:
The patent changes the representation parameters of unstructured data by converting various modalities into unified embedding vectors. This parameter transformation enables the system to handle diverse data types through consistent mathematical operations, reducing processing complexity while maintaining data amount and diversity through augmentation.
Data Source
AI summary
Embodiments of the present disclosure relate to a method, a device, and a computer program product for data augmentation. The method includes generating an image embedding based on an image in an unstructured document, and generating a text embedding based on text in the unstructured document and associated with the image. The method further includes acquiring descriptive information from a storage library based on the generated image embedding and text embedding. The method further includes adding the acquired descriptive information into the unstructured document. In this way, it can be possible not only to understand and analyze the unstructured document across modalities, but also to enrich it with a characterization of multimodal data in the unstructured document, thus increasing the amount and diversity of data.


