Learning Device Mask Generation for Still Image Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models face reduced recognition accuracy for objects with complex textures in still images due to difficulties in identifying object boundaries.

Innovation Solution

A learning device with a mask generation mechanism that adjusts parameters based on the difference between object masks in still and moving images, using a loss calculation to refine the object mask generation process.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If a machine learning model is trained to acquire object representations for individual objects in still images, then object recognition capability is achieved, but recognition accuracy deteriorates for objects with complicated textures due to difficulty in boundary recognition

Engineering Contradiction:
Improveobject recognition capabilityVSAvoidboundary recognition accuracy
Core Design Contradiction:
ReliabilityVSMeasurement precision

Solution Approach 1:

The patent introduces an intermediary training mechanism using moving images as a mediator to improve still image object recognition. By training the model with moving images that provide temporal context and motion information, the system generates more accurate object masks for still images, thereby resolving the boundary recognition problem for complex textured objects without sacrificing general recognition capability

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-training the machine learning model with moving images before deploying it for still image analysis. This preliminary training phase allows the model to learn robust object boundary representations from temporal sequences, which then improves its performance when analyzing static images with complicated textures

Inventive Principle:
Principle #10Preliminary action

2Manufacturing precision

If object masks are generated for each individual object in still images, then object segmentation is achieved, but recognition accuracy is lowered when objects have complicated textures

Engineering Contradiction:
Improveobject mask generation accuracyVSAvoidobject recognition accuracy
Core Design Contradiction:
Manufacturing precisionVSMeasurement precision

Solution Approach 1:

The patent uses moving images as an intermediary training data source to improve the quality of object masks generated for still images. The temporal information and motion cues from moving images help the model distinguish object boundaries more accurately, even for complex textured objects, thereby improving both mask generation quality and subsequent recognition accuracy

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent changes the training parameters by incorporating moving image data with temporal dynamics into the training process. This parameter change allows the model to learn more robust features for boundary detection, improving its ability to generate accurate object masks for still images with complicated textures

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240153255A1Learning device, parameter adjustment method and recording medium
Publication Date: 2024.05.09 NEC CORP
  • US20240153255A1 patent drawing
  • US20240153255A1 patent drawing
  • US20240153255A1 patent drawing

AI summary

A learning device includes a learning model for a still image. The learning model includes a mask generation means configured to generate a first object mask identifying an area in which an object exists in a still image, for each individual object. A first parameter including at least one parameter used for processing of generating the first object mask is adjusted based on a first loss. The first loss indicates a difference of the first object mask with respect to a second object mask identifying an area in which an object exists in a moving image including the still image, for each individual object.