Visual Marker Dataset Generation for Object Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating training datasets for object recognition and position/posture estimation in the food and logistics industries are inefficient, as they require multiple sensors and 3D models, and are prone to issues like marker reflection and concealment, leading to inaccurate object recognition and position estimation.
Innovation Solution
A method using visual markers to generate training datasets by acquiring images of objects with markers, concealing the marker area, and correlating posture and position information to create accurate training datasets for machine learning, without requiring multiple sensors or 3D models, allowing for efficient object recognition and position estimation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Extent of automation
If a marker is attached to an object for automated data collection, then the automation extent is improved, but the marker may be reflected in the target object or hidden by it, reducing measurement precision
Solution Approach 1:
The patent extracts and removes the marker from the captured image through image processing techniques. The marker detection unit identifies marker positions, and the image processing unit removes these markers from the training data images, preventing marker reflection issues while maintaining automation benefits
Solution Approach 2:
The patent introduces an intermediary processing system between marker attachment and data collection. The marker serves as an intermediary tool for automation, but through image processing, its harmful effects are eliminated while its useful function (enabling automated detection) is preserved
2Measurement precision
If multiple sensors and 3D models are used for accurate position estimation, then measurement precision is improved, but the device complexity increases
Solution Approach 1:
The patent merges multiple functions into a single integrated system. The object recognition unit, position estimation unit, and marker removal functions are combined in one system that processes images directly, eliminating the need for separate multiple sensors and 3D model processing systems
Solution Approach 2:
The patent replaces complex mechanical/multi-sensor systems with an optical/image processing-based system. Instead of using multiple physical sensors and 3D models, the system uses 2D image capture and computational image processing to achieve accurate position and posture estimation
3Measurement precision
If manual data input is performed for object position and posture, then measurement precision is maintained, but the productivity decreases
Solution Approach 1:
The system performs self-service automated data collection. The object recognition unit automatically detects objects and extracts position/posture information from images without human intervention, while the trained model automatically processes the data, achieving both high speed and maintained precision through automated intelligent processing
4Device complexity
If a single RGB camera is used for object recognition, then the device complexity is reduced, but the measurement precision of position and posture estimation deteriorates
Solution Approach 1:
The patent compensates for the limitations of 2D RGB camera data by introducing depth information through marker-based measurements. The position estimation unit calculates 3D position and posture by combining 2D image coordinates with known marker dimensions and geometric relationships, effectively adding a depth dimension through computational methods
Data Source
AI summary
Provided are a method and a device that can efficiently generate a training dataset. Object information is associated with a visual marker, a training dataset generation jig that is configured from a base part and a marker is used, said base part being provided with an area that serves as a guide for positioning a target object and said marker being fixed on the base part, the target object is positioned using the area as a guide and in this condition an image group of the entire object including the marker is acquired, the object information that was associated with the visual marker is acquired from the acquired image group, a reconfigured image group is generated from this image group by performing a concealment process on a region corresponding to the visual marker or the training dataset generation jig, a bounding box is set in the reconfigured image group on the basis of the acquired object information, information relating to the bounding box, the object information, and estimated target object position information and posture information are associated with a captured image, and a training dataset for performing object recognition and position/posture estimation for the target object is generated.


