Multimodal Vector Correction for Image-Document Integration
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing techniques for solving problems using multiple modal information, such as ViLBERT, BERT, VideoBERT, and MCAN, suffer from poor accuracy when integrating information from images and documents, as they either refer to vectors without explicit correction or fail to distinguish between modal types effectively.
Innovation Solution
A method that corrects vectors based on correlations between different modal types, generates aggregate vectors through self-attention layers, and combines these vectors with a predetermined vector to improve accuracy in multimodal problem-solving, using a Co-Attention Network that updates models based on training data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing techniques (ViLBERT, BERT, VideoBERT, MCAN) are used to integrate information from images and documents, then the system can process multiple modal inputs, but the accuracy of solutions is poor due to inadequate vector correction and modal distinction
Solution Approach 1:
The patent segments the vector processing into distinct stages: initial vector generation from each modal, separate correction processes for image and document vectors, and final aggregation. This segmentation allows each modal to be processed independently with modal-specific correction, improving both accuracy and reliability of multimodal integration
Solution Approach 2:
The patent applies parameter changes by correcting vectors based on correlation metrics between different modal vectors. The correction process modifies vector parameters (direction, magnitude) dynamically based on inter-modal relationships, enabling adaptive integration that improves solution accuracy while maintaining reliable modal distinctions
2Ease of manufacture
If vectors from different modal types are aggregated without explicit correction, then the processing is simpler, but the accuracy of integrating image and document information deteriorates
Solution Approach 1:
The patent performs preliminary correction actions on vectors from each modal before aggregation. By correcting image vectors and document vectors separately based on their correlations with other modal vectors, the system prepares high-quality input vectors that maintain simplicity in the aggregation step while achieving high integration accuracy
Solution Approach 2:
The correction mechanism acts as an intermediary between raw modal vectors and the final aggregated vector. This intermediary process refines individual modal representations by considering cross-modal correlations, enabling accurate integration without complicating the overall processing architecture
Data Source
AI summary
A method includes: correcting a vector of a first modal by using a correlation between the vector of the first modal and a vector of a second modal different from the first modal; correcting the vector of the second modal by using the correlation between the vector of the first modal and the vector of the second modal; generating a first vector by using a correlation of two different types of vectors obtained from the corrected vector of the first modal; generating a second vector by using the correlation of the two different types of vectors obtained from the corrected vector of the second modal; generating a third vector in which the first and second vectors are aggregated by using the correlation of the two different types of vectors obtained from a combined vector including a predetermined vector, the generated first and second vectors; and outputting the generated third vector.


