Learning Device Mask Generation for Still Image Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models face reduced recognition accuracy for objects with complex textures in still images due to difficulties in identifying object boundaries.
Innovation Solution
A learning device with a mask generation mechanism that adjusts parameters based on the difference between object masks in still and moving images, using a loss calculation to refine the object mask generation process.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If a machine learning model is trained to acquire object representations for individual objects in still images, then object recognition capability is achieved, but recognition accuracy deteriorates for objects with complicated textures due to difficulty in boundary recognition
Solution Approach 1:
The patent introduces an intermediary training mechanism using moving images as a mediator to improve still image object recognition. By training the model with moving images that provide temporal context and motion information, the system generates more accurate object masks for still images, thereby resolving the boundary recognition problem for complex textured objects without sacrificing general recognition capability
Solution Approach 2:
The patent applies preliminary action by pre-training the machine learning model with moving images before deploying it for still image analysis. This preliminary training phase allows the model to learn robust object boundary representations from temporal sequences, which then improves its performance when analyzing static images with complicated textures
2Manufacturing precision
If object masks are generated for each individual object in still images, then object segmentation is achieved, but recognition accuracy is lowered when objects have complicated textures
Solution Approach 1:
The patent uses moving images as an intermediary training data source to improve the quality of object masks generated for still images. The temporal information and motion cues from moving images help the model distinguish object boundaries more accurately, even for complex textured objects, thereby improving both mask generation quality and subsequent recognition accuracy
Solution Approach 2:
The patent changes the training parameters by incorporating moving image data with temporal dynamics into the training process. This parameter change allows the model to learn more robust features for boundary detection, improving its ability to generate accurate object masks for still images with complicated textures
Data Source
AI summary
A learning device includes a learning model for a still image. The learning model includes a mask generation means configured to generate a first object mask identifying an area in which an object exists in a still image, for each individual object. A first parameter including at least one parameter used for processing of generating the first object mask is adjusted based on a first loss. The first loss indicates a difference of the first object mask with respect to a second object mask identifying an area in which an object exists in a moving image including the still image, for each individual object.


