Multimodal Data Enhancement via Knowledge Graph Inference
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data enhancement methods for AI subtasks, particularly in multimodal training data, face limitations as they often rely on single technologies, failing to effectively supplement information between different modalities, which can lead to semantic coherence and correctness issues in natural language processing.
Innovation Solution
A data enhancement method that involves obtaining sub-data from multiple modalities, determining entity objects matching the data type, and performing inference on these objects based on entity relationship information in a knowledge graph to generate new, enhanced data, thereby bridging information gaps between modalities.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If single technology is used to process information in multimodal training data, then processing simplicity is maintained, but information supplement between different modalities is wasted
Solution Approach 1:
The patent merges multiple processing technologies (image processing, text processing, knowledge graph inference) into a unified multimodal processing system. The system combines image data, text data, and knowledge graph data through integrated processing steps including entity recognition, relationship extraction, and cross-modal reasoning, thereby充分利用 the complementary information between different modalities without losing valuable information supplement.
Solution Approach 2:
The processing system is designed with multi-functionality to handle diverse data types simultaneously. It can process images, text, and knowledge graph structures through a unified framework that performs entity recognition, relationship extraction, and inference across all modalities, making the system versatile and capable of utilizing information from any modality according to its specific functions.
2Adaptability or versatility
If natural language processing is enhanced by changing or deleting phrases, then data diversity is improved, but semantic coherence and correctness are affected
Solution Approach 1:
The patent introduces knowledge graphs as an intermediary structure to mediate between diverse data inputs and semantic coherence requirements. The knowledge graph provides structured entity relationships and common sense knowledge that guide the processing of image and text data, ensuring that enhancements maintain semantic correctness while allowing data diversity through multiple processing paths and reasoning approaches.
Solution Approach 2:
The system implements feedback mechanisms where inference results from knowledge graph reasoning are used to validate and refine processed data. The cross-modal reasoning process continuously checks whether enhanced data maintains consistency with established knowledge, providing feedback that preserves semantic coherence while allowing controlled diversity in data representation.
Data Source
AI summary
A data enhancement method includes obtaining first data including sub-data of a plurality of modalities. The sub-data of one modality corresponds to one data type and the data types of different modalities are different. The method further includes determining, in the sub-data of each modality, an entity object matching the data type of the sub-data, and performing inference on the entity objects corresponding to different modalities based on entity relationship information in a knowledge graph to obtain second data different from the first data.


