Gaze-Based Mask Generation for Complex Object Labeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing image processing systems struggle to accurately distinguish and label specific objects, especially artistic or complex shapes, in two-dimensional images, leading to difficulties in object recognition and three-dimensional reconstruction.

Innovation Solution

An image processing system and method that utilize a processor and gaze detector to identify view hotspots corresponding to the user's gaze direction, combined with an indicator signal, to generate accurate mask blocks for object blocks in two-dimensional images, enabling precise labeling of specific objects.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If conventional image segmentation models (CNN) are used to divide objects into two-dimensional images, then the segmentation process can be completed, but the system cannot accurately distinguish or label specific objects with complex or artistic shapes

Engineering Contradiction:
Improveobject recognition accuracyVSAvoidsystem complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent introduces view hotspots and indicator signals as intermediary elements between the image segmentation model and the final object labeling. These intermediaries provide additional spatial and contextual information that helps distinguish specific objects, particularly those with complex or artistic shapes that conventional models struggle to identify accurately.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent adds a new dimension of information by incorporating view hotspots (spatial attention points) and indicator signals (user intent markers) into the traditional two-dimensional image segmentation process. This multi-dimensional approach enables more accurate object identification by combining visual data with spatial and user interaction data.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

2Measurement precision

If the system attempts to accurately label all objects in a two-dimensional image, then object recognition precision improves, but the processing time and computational resources increase

Engineering Contradiction:
Improveobject labeling accuracyVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent performs preliminary segmentation and identification of view hotspots before final object labeling. By pre-processing the image to identify key spatial points and potential objects, the system reduces the computational burden during the final labeling stage, thereby decreasing overall processing time while maintaining accuracy.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent focuses computational resources on regions of interest identified by view hotspots and indicator signals rather than processing the entire image uniformly. This selective approach allows the system to achieve high labeling accuracy for specific objects while reducing unnecessary computational expenditure on less relevant areas.

Inventive Principle:
Principle #16Partial or excessive action

3Measurement precision

If conventional segmentation methods are used without additional user interaction, then the system operation is simple, but the system cannot correctly identify specific objects or their shapes

Engineering Contradiction:
Improveobject identification accuracyVSAvoiduser interaction complexity
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent implements a feedback mechanism where indicator signals from user interaction (such as gaze detection or explicit selection) are fed back into the segmentation process. This feedback loop allows the system to adjust its focus and improve object identification accuracy based on user intent, while maintaining ease of operation through natural interaction modalities.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent enables the system to automatically process and interpret user interactions (such as gaze direction or selection signals) without requiring complex manual input. The system self-adjusts its segmentation and labeling based on these interactions, improving object identification while keeping the user interface simple and intuitive.

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS11615549B2Image processing system and image processing method
Publication Date: 2023.03.28 HTC CORP
  • US11615549B2 patent drawing
  • US11615549B2 patent drawing
  • US11615549B2 patent drawing

AI summary

An image processing method includes the following steps: dividing an object block into a two-dimensional image; identifying at least one view hotspot in a viewing field corresponding to pupil gaze direction; receiving the view hotspot and an indicator signal; wherein the indicator signal is used to remark the object block; and generating a mask block that corresponds to the object block according to the view hotspot; wherein the indicator signal determines the label of the mask block.