Multimodal Object Recognition Using Text-Guided Region Matching

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing object recognition technologies struggle to accurately identify objects in media content due to the limitations of relying solely on image information, often leading to inaccuracies and inefficiencies.

Innovation Solution

A method and apparatus that integrates image and text information to determine candidate object regions, utilizing visual and text features for precise object recognition, involving modules for candidate region determination, object region filtering, and feature-based matching.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If object recognition relies solely on image information, then the processing is simple and fast, but the recognition accuracy is low

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines image information and text information into a unified object recognition system. The image processing module extracts visual features from media content while the text processing module extracts textual features from captions or descriptions. These two feature sets are then integrated through a matching module that correlates visual and textual representations to identify objects, thereby improving recognition accuracy by leveraging multimodal data fusion.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system is designed to handle multiple types of input data (images and text) through a single unified processing framework. The same architectural structure can process both visual and textual modalities, making the system versatile and adaptable to different media content types while maintaining high recognition accuracy across diverse object categories.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If only image information is used for object recognition, then the processing speed is fast, but the identification precision is insufficient

Engineering Contradiction:
Improveobject identification precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs preliminary extraction of both visual features from images and textual features from captions in parallel before the matching stage. This preliminary processing prepares feature representations in advance, allowing the subsequent matching operation to proceed efficiently without time-consuming real-time analysis, thus reducing overall processing time while maintaining high precision.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent implements continuous feature extraction and matching operations where image processing and text processing occur simultaneously and continuously rather than sequentially. The matching module continuously correlates visual and textual features as new data becomes available, maintaining steady processing throughput while improving identification precision through comprehensive multimodal analysis.

Inventive Principle:
Principle #20Continuity of useful action

3Reliability

If existing object recognition technologies are used, then the system is simple, but the recognition accuracy is low due to reliance on single modality

Engineering Contradiction:
Improveobject recognition reliabilityVSAvoidsystem structure complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges image processing and text processing into a unified recognition system. The image processing module extracts visual features while the text processing module extracts textual features from captions or descriptions. These feature sets are then integrated through a matching module that correlates visual and textual representations, thereby improving reliability by cross-validating object identification across multiple modalities.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system incorporates feedback mechanisms where the matching module continuously refines object identification based on the correlation between visual and textual features. If discrepancies are detected between modalities, the system can re-evaluate and adjust its recognition results, improving reliability through iterative refinement and cross-modal validation.

Inventive Principle:
Principle #23Feedback

Data Source

PatentUS20250371842A1Method, apparatus, device and storage medium for object recognition
Publication Date: 2025.12.04 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250371842A1 patent drawing
  • US20250371842A1 patent drawing
  • US20250371842A1 patent drawing

AI summary

The embodiment of the disclosure provides a method, apparatus, device, and storage medium for object recognition. The method includes: determining a set of first candidate object regions based on image information of a media content; determining, based on text information associated with the media content, an object region from the set of first candidate object regions; and determining an object matching the object region based on a visual feature of the object region and a text feature, the text feature being determined based on the text information. Based on the manner, disclosure may recognize an object in the media content for multimodal information of the image information of the media content and text information associated with the media content, which may effectively improve the accuracy of the object recognition.