Multimodal Image Segmentation Using Response Heat Maps

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Image segmentation under language indication faces challenges due to a semantic gap between images and linguistic descriptions, making it difficult to accurately segment specified objects.

Innovation Solution

Fusing visual features from an original image with text features from a description language to create a multimodal feature, determining a visual region using a response heat map, and employing an image segmentation model to refine the segmentation result.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If image segmentation is performed under language indication, then the ability to segment specified objects is improved, but the semantic gap between images and linguistic descriptions causes segmentation accuracy to deteriorate

Engineering Contradiction:
Improvelanguage indication capabilityVSAvoidsegmentation accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent introduces a response heat map as an intermediary between the input image and the segmentation output. The heat map visually represents the model's attention distribution across different regions, serving as a mediator that bridges the semantic gap between language indications and image segmentation, thereby improving segmentation accuracy while maintaining language indication capability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent implements a feedback mechanism where the response heat map is generated based on the interaction between visual features and text features, and this heat map is then used to guide the final segmentation result. The feedback loop allows the system to adjust and refine segmentation based on the attention distribution, resolving the contradiction between adaptability and precision

Inventive Principle:
Principle #23Feedback

2Loss of information

If visual features and text features are fused to create multimodal features, then the elimination of semantic gap is improved, but the complexity of the processing system increases

Engineering Contradiction:
Improvesemantic gap eliminationVSAvoidprocessing system complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent merges visual features extracted from the image with text features from the language indication to create unified multimodal features. This combining approach allows the system to process both visual and linguistic information in an integrated manner, effectively eliminating the semantic gap while managing complexity through feature-level fusion rather than system-level complexity

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If a response heat map is generated to determine visual regions, then the accuracy of target object identification is improved, but the computational time and resources increase

Engineering Contradiction:
Improvetarget object identification accuracyVSAvoidcomputational time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent generates the response heat map as a preliminary step before final segmentation, pre-processing the attention distribution information in advance. This preliminary action allows the subsequent segmentation to be more efficient and accurate, as the heat map already highlights the relevant regions, reducing the computational burden during the final segmentation phase

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS12608813B2Image segmentation method for a target object, device, and storage medium
Publication Date: 2026.04.21 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US12608813B2 patent drawing
  • US12608813B2 patent drawing
  • US12608813B2 patent drawing

AI summary

Provided are an image segmentation method and apparatus, a device, and a storage medium. The image segmentation method includes: fusing a visual feature corresponding to an original image with a text feature corresponding to a description language to obtain a multimodal feature, where the description language is used for specifying a target object to be segmented in the original image; determining a visual region of the target object according to an image corresponding to the multimodal feature and recording an image corresponding to the visual region as a response heat map; and determining a segmentation result of the target object according to the image corresponding to the multimodal feature and the response heat map.