Image-Text Feature Fusion for Accurate Object Association

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current image-based data processing methods have low accuracy in learning the association relationship between text and objects in images, leading to incorrect text processing.

Innovation Solution

A method that extracts features of objects and text, fuses them based on a matching degree, and processes the text using a visual question answer system to improve accuracy.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the current method extracts low-level features of image and text separately using two different underlying representation systems, then the processing can be performed, but the association relationship between text and objects has low accuracy

Engineering Contradiction:
Improveassociation relationship accuracyVSAvoidfeature extraction complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges the separate image feature extraction and text feature extraction processes into a unified feature extraction framework. By using a single underlying representation system to extract features from both image and text, the method eliminates the need for two different representation systems while improving the accuracy of association relationships between text and objects.

Inventive Principle:
Principle #5Merging (Combining)

2Measurement precision

If the method learns high-level features and associates them through an associated learning module, then text processing can be performed, but the association relationship accuracy remains low

Engineering Contradiction:
Improveassociation relationship accuracyVSAvoidtext processing efficiency
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent applies preliminary action by extracting high-level features directly from low-level features in advance, before the association learning process. This pre-extraction of meaningful high-level features enables more accurate association relationships to be learned subsequently, improving both accuracy and efficiency of text processing.

Inventive Principle:
Principle #10Preliminary action

3Reliability

If the method processes text based on learned association relationships, then text processing is completed, but incorrect processing occurs due to low association accuracy

Engineering Contradiction:
Improvetext processing correctnessVSAvoidassociation relationship accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms where the learned association relationships are continuously refined based on processing results. The system uses the outcomes of text processing to adjust and improve the association relationships between text and object features, thereby increasing both the accuracy of associations and the correctness of text processing over time.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP3696729B1Method, apparatus, device and readable storage medium for image-based data processing
Publication Date: 2026.01.21 BEIJING BAIDU NETCOM SCI & TECH CO LTD
  • EP3696729B1 patent drawingFigure 1a~1b
  • EP3696729B1 patent drawingFigure 1c
  • EP3696729B1 patent drawingFigure 2a~2b

AI summary

Embodiments of the present disclosure disclose a method, apparatus, device, and readable storage medium for image-based data processing. The method comprises: acquiring an image and a to-be-processed text; extracting features of a plurality of objects in the image, and extracting a feature of the text; fusing the features of the plurality of objects into a fused feature of the image based on a matching degree between the feature of the text and a feature of each object of the plurality of objects; and processing the text based on the fused feature of the image and the feature of the text. Embodiments of the present disclosure can accurately learn an association relationship between a text and each object in an image, and improve the processing accuracy.