Image Metadata Association via Region Segmentation and Object Relationship Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing methods fail to accurately add metadata to objects in images, especially when the relevant text information is not nearby, leading to difficulties in search and retrieval due to the lack of consideration for object relationships and importance.

Innovation Solution

An information processing apparatus and method that divides input image data into regions, identifies objects, and adds metadata based on object types, with specific objects associating with other regions, allowing for the inclusion of relevance and importance metadata, even when objects are not adjacent, by using an MFP system that performs OCR and vectorization processes.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If text objects are used to express content of photograph or graphic objects, then the content can be searched, but the text object is not necessarily in the vicinity of the photograph or graphic object, leading to inaccurate metadata addition

Engineering Contradiction:
Improvemetadata addition accuracyVSAvoidobject relationship analysis complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the image into multiple regions and identifies objects within each region. It then establishes relationships between objects across different regions by analyzing spatial positions, distances, and semantic connections. This segmentation approach allows the system to accurately associate text objects with photograph or graphic objects even when they are not in close proximity, resolving the contradiction between metadata accuracy and analysis complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces an object relationship analysis mechanism that acts as an intermediary between text objects and photograph/graphic objects. This intermediary analyzes various factors including spatial relationships, semantic relevance, and contextual information to establish accurate associations. By using this intermediary analysis layer, the system can accurately add metadata without requiring text objects to be physically adjacent to the objects they describe.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Loss of information

If conventional OCR process is used to add metadata, then character encoding information can be obtained, but importance and relevance information of objects is not considered

Engineering Contradiction:
Improveobject importance and relevance informationVSAvoidmetadata processing efficiency
Core Design Contradiction:
Loss of informationVSProductivity

Solution Approach 1:

The patent performs preliminary object identification and relationship analysis before adding metadata. It first identifies all objects in the image, determines their types (text, photograph, graphic), and establishes relationships between them. This preliminary action allows the system to capture importance and relevance information upfront, which is then used during the metadata addition phase. This approach prevents loss of critical object relationship information while maintaining processing efficiency through structured preliminary analysis.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent creates a universal object relationship analysis framework that handles multiple object types (text, photograph, graphic) and multiple relationship types (spatial, semantic, hierarchical) through a single integrated system. This multi-functional approach allows the system to extract various types of information including importance, relevance, and contextual relationships in one unified process, rather than requiring separate processing for each information type, thus maintaining productivity while reducing information loss.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Measurement precision

If text from adjacent zones is used for image zone annotation, then useful keywords can be generated, but the method does not consider object relationships and importance

Engineering Contradiction:
Improvekeyword annotation accuracyVSAvoidobject relationship and importance information
Core Design Contradiction:
Measurement precisionVSLoss of information

Solution Approach 1:

The patent implements a feedback mechanism where the object relationship analysis results are continuously refined based on metadata addition outcomes. The system analyzes spatial relationships, semantic connections, and importance levels, uses this information to add appropriate metadata, and then refines the relationship models based on the results. This feedback loop ensures that keyword annotations are not only based on adjacent text but also enhanced by understood object relationships and importance, improving accuracy without losing critical relationship information.

Inventive Principle:
Principle #23Feedback

Data Source

PatentEP2166467B1Information processing apparatus, control method thereof, computer program, and storage medium
Publication Date: 2012.06.06 CANON KK
  • EP2166467B1 patent drawingFigure 1
  • EP2166467B1 patent drawingFigure 2~3
  • EP2166467B1 patent drawingFigure 4~5

AI summary

An information processing apparatus (100) is provided. The apparatus (100) comprises an association unit (301) configured to divide input image data into regions and to associate each region with one or more types of objects; an addition unit (302) configured to add metadata to each object based on the type of each object; and a determination unit (304) configured to determine whether or not a specific object that associates a first one of the regions with a second one of the regions different from the first one is present among the objects. In the case where the determination unit (304) has determined that the specific object is present, the addition unit (302) is configured to further add, to a first object that is present in the first one of the regions, metadata for associating the second one of the regions with the first one of the regions.