Visual Marker Dataset Generation for Object Recognition

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating training datasets for object recognition and position/posture estimation in the food and logistics industries are inefficient, as they require multiple sensors and 3D models, and are prone to issues like marker reflection and concealment, leading to inaccurate object recognition and position estimation.

Innovation Solution

A method using visual markers to generate training datasets by acquiring images of objects with markers, concealing the marker area, and correlating posture and position information to create accurate training datasets for machine learning, without requiring multiple sensors or 3D models, allowing for efficient object recognition and position estimation.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Extent of automation

If a marker is attached to an object for automated data collection, then the automation extent is improved, but the marker may be reflected in the target object or hidden by it, reducing measurement precision

Engineering Contradiction:
Improveautomation of data collectionVSAvoidaccuracy of object recognition
Core Design Contradiction:
Extent of automationVSMeasurement precision

Solution Approach 1:

The patent extracts and removes the marker from the captured image through image processing techniques. The marker detection unit identifies marker positions, and the image processing unit removes these markers from the training data images, preventing marker reflection issues while maintaining automation benefits

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent introduces an intermediary processing system between marker attachment and data collection. The marker serves as an intermediary tool for automation, but through image processing, its harmful effects are eliminated while its useful function (enabling automated detection) is preserved

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If multiple sensors and 3D models are used for accurate position estimation, then measurement precision is improved, but the device complexity increases

Engineering Contradiction:
Improveaccuracy of position estimationVSAvoidnumber of sensors and processing systems
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent merges multiple functions into a single integrated system. The object recognition unit, position estimation unit, and marker removal functions are combined in one system that processes images directly, eliminating the need for separate multiple sensors and 3D model processing systems

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent replaces complex mechanical/multi-sensor systems with an optical/image processing-based system. Instead of using multiple physical sensors and 3D models, the system uses 2D image capture and computational image processing to achieve accurate position and posture estimation

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Measurement precision

If manual data input is performed for object position and posture, then measurement precision is maintained, but the productivity decreases

Engineering Contradiction:
Improveaccuracy of position dataVSAvoidspeed of data collection
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The system performs self-service automated data collection. The object recognition unit automatically detects objects and extracts position/posture information from images without human intervention, while the trained model automatically processes the data, achieving both high speed and maintained precision through automated intelligent processing

Inventive Principle:
Principle #25Self-service

4Device complexity

If a single RGB camera is used for object recognition, then the device complexity is reduced, but the measurement precision of position and posture estimation deteriorates

Engineering Contradiction:
Improvesimplicity of camera systemVSAvoidaccuracy of position and posture
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent compensates for the limitations of 2D RGB camera data by introducing depth information through marker-based measurements. The position estimation unit calculates 3D position and posture by combining 2D image coordinates with known marker dimensions and geometric relationships, effectively adding a depth dimension through computational methods

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS11276194B2Learning dataset creation method and device
Publication Date: 2022.03.15 NARA INSTITUTE OF SCIENCE AND TECHNOLOGY
  • US11276194B2 patent drawing
  • US11276194B2 patent drawing
  • US11276194B2 patent drawing

AI summary

Provided are a method and a device that can efficiently generate a training dataset. Object information is associated with a visual marker, a training dataset generation jig that is configured from a base part and a marker is used, said base part being provided with an area that serves as a guide for positioning a target object and said marker being fixed on the base part, the target object is positioned using the area as a guide and in this condition an image group of the entire object including the marker is acquired, the object information that was associated with the visual marker is acquired from the acquired image group, a reconfigured image group is generated from this image group by performing a concealment process on a region corresponding to the visual marker or the training dataset generation jig, a bounding box is set in the reconfigured image group on the basis of the acquired object information, information relating to the bounding box, the object information, and estimated target object position information and posture information are associated with a captured image, and a training dataset for performing object recognition and position/posture estimation for the target object is generated.