Data Creation Apparatus for Selecting Training Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Selecting appropriate image data for machine learning training from vast datasets is labor-intensive and time-consuming, particularly when multiple subjects are captured in images, increasing the cost of creating training data.
Innovation Solution
A data creation apparatus that sets and applies conditions related to identification and image quality information for selecting and creating training data from image data with multiple subjects, including resolution, brightness, and noise levels, and suggests additional conditions to enhance data accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If image data with multiple subjects is manually selected and annotated for training data creation, then the accuracy and relevance of training data can be ensured, but the time consumption and labor cost increase significantly
Solution Approach 1:
The system enables self-service by automatically selecting and annotating image data without human intervention. The image data itself contains metadata (exif information) that the system uses to autonomously determine selection criteria and annotate multiple subjects, allowing the data creation process to serve itself rather than requiring manual labor for each step
Solution Approach 2:
The system changes parameters by utilizing existing metadata parameters (exif information) embedded in image data to make selection decisions. By extracting and evaluating parameters such as拍摄时间、location、camera settings already present in the image files, the system transforms unstructured image collections into structured training data based on these parameter values
2Reliability
If comprehensive annotation of all subjects in image data is performed, then the quality and completeness of training data improves, but the processing complexity and cost increase
Solution Approach 1:
The system achieves universality by using a single automated process that handles multiple functions: selecting images based on criteria, detecting multiple subjects within each image, and annotating all detected subjects. This multi-functional approach replaces multiple separate manual operations (selection, detection, annotation) with one integrated system that processes all aspects simultaneously
Solution Approach 2:
The system substitutes mechanical manual operations with automated computational processes. Instead of human operators manually reviewing and annotating images, the system uses algorithmic processing to automatically select images based on predefined criteria and detect/annotate subjects using image processing techniques, replacing the mechanical nature of manual work with automated digital processing
3Measurement precision
If selective criteria are applied to filter image data, then the relevance of training data increases, but the amount of usable data may decrease
Solution Approach 1:
The system applies segmentation by dividing the image data selection process into distinct segments or stages: first filtering images based on exif metadata criteria (time, location, camera settings), then separately processing subject detection and annotation. This segmented approach allows comprehensive filtering while maintaining data volume by processing different aspects in separate manageable stages
Solution Approach 2:
The system performs preliminary action by pre-filtering image data using exif information before the actual training data creation process. By establishing selection criteria based on metadata in advance and pre-identifying suitable images, the system prepares the data set beforehand, ensuring high relevance while efficiently managing the volume of data that proceeds to the annotation stage
Data Source
AI summary
Image data to be used for creating training data is appropriately selected from a plurality of pieces of image data in each of which an image in which a plurality of subjects are captured is recorded.One embodiment of the present invention is a data creation apparatus that creates training data used in machine learning from image data in which accessory information is recorded in an image in which a plurality of subjects are captured, the data creation apparatus being configured to execute setting processing of setting any setting condition related to identification information and to image quality information with respect to a plurality of pieces of image data in which the accessory information including a plurality of pieces of the identification information assigned in association with the plurality of subjects and a plurality of pieces of the image quality information assigned in association with the plurality of subjects is recorded, and creation processing of creating the training data based on selection image data in which the identification information and the image quality information satisfying the setting condition are recorded.


