Object Recognition Interpretation with Fixed Segmentation for Small Objects

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Superpixels often include both the object and its peripheral area, leading to inappropriate contribution calculation in object recognition results, especially for small objects.

Innovation Solution

Interpreting object recognition results in units of segments that geometrically divide the image, using a model interpreting unit to generate explanatory diagrams.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If superpixels are used to interpret object recognition results, then the interpretation can be performed in a simplified manner, but for small objects the superpixel may include both the object and peripheral area leading to inaccurate contribution calculation

Engineering Contradiction:
Improveease of interpretationVSAvoidcontribution calculation accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent divides the image into multiple segments (e.g., grid-based segments) instead of using superpixels. Each segment is a fixed geometric region that can be independently analyzed. This segmentation approach allows for more precise localization of small objects since the segment boundaries are fixed and known, enabling accurate attribution of recognition contributions to specific spatial regions without the uncertainty of superpixel boundaries.

Inventive Principle:
Principle #1Segmentation

2Device complexity

If superpixels are used for interpretation, then processing can be simplified, but the superpixel boundaries may not align with object boundaries causing mixed contributions

Engineering Contradiction:
Improveinterpretation complexityVSAvoidrecognition result reliability
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent employs fixed geometric segmentation (such as grid division) of the image into multiple segments. Each segment has predetermined boundaries that are independent of object boundaries. This approach simplifies the interpretation process by providing a regular, structured framework for analyzing recognition results, while the fine-grained segment structure ensures that even small objects can be localized within specific segments, maintaining reliability.

Inventive Principle:
Principle #1Segmentation

3Productivity

If larger segmentation units are used, then the number of segments decreases simplifying processing, but the precision of locating small objects deteriorates

Engineering Contradiction:
Improveprocessing efficiencyVSAvoidobject location precision
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent uses fine-grained segmentation by dividing the image into multiple small fixed segments (e.g., using a grid pattern with sufficient resolution). This creates a large number of segments that can precisely locate small objects while maintaining systematic processing. The regular structure of fixed segments allows for efficient computation despite the increased segment count, as each segment can be processed independently using standardized procedures.

Inventive Principle:
Principle #1Segmentation

Data Source

PatentUS12387361B2Information processing to appropriately interpret a recognition result of an object recognition model
Publication Date: 2025.08.12 SONY SEMICON SOLUTIONS CORP
  • US12387361B2 patent drawing
  • US12387361B2 patent drawing
  • US12387361B2 patent drawing

AI summary

The present technique relates to an information processing device and an information processing method that enable a recognition result of an object recognition model to be appropriately interpreted. The information processing device includes an interpreting unit that performs interpretation of a recognition result of an object recognition model in units of segments which geometrically divide an image. For example, the present technique is applied to a device which interprets and explains an object recognition model that performs object recognition in front of a vehicle.