Image Data Preprocessing for ML Tuning Under Data Drift
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing machine learning models face precision deterioration due to data drift during practical use, especially in irregular or small-scale applications, where insufficient data for tuning leads to reduced classification accuracy.
Innovation Solution
An information processing apparatus that collects and preprocesses image data to meet tuning conditions, dynamically adjusts preprocessing methods, and re-trains the model to adapt to data drift, ensuring adequate data availability and maintaining high inference precision.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If the DL model is tuned using collected image data, then the model adapts to data drift and maintains classification precision, but insufficient data quantity leads to inadequate tuning and reduced accuracy
Solution Approach 1:
The system generates synthetic image data that copies and replicates real object characteristics through computer-generated imagery. This synthetic data is created to supplement insufficient collected images, providing adequate training samples for model tuning without requiring additional physical imaging, thereby resolving the contradiction between data quantity and classification precision.
Solution Approach 2:
The system varies parameters such as lighting conditions, camera angles, and object positions when generating synthetic image data. By changing these parameters to create diverse synthetic samples, the system ensures adequate data quantity and variety for effective model tuning, maintaining classification precision even when real collected data is insufficient.
2Quantity of substance
If image data is collected and preprocessed to meet tuning conditions, then adequate data is available for model tuning, but the preprocessing complexity increases
Solution Approach 1:
The system automatically determines whether collected data meets tuning conditions and self-adjusts by generating synthetic data only when necessary. This self-service approach simplifies preprocessing complexity by eliminating manual data assessment and generation decisions, while ensuring adequate data availability through automated synthetic data creation when real data is insufficient.
Solution Approach 2:
The system performs preliminary assessment of collected image data to determine if tuning conditions are met before proceeding with model tuning. This preliminary action prevents unnecessary synthetic data generation and simplifies the overall preprocessing workflow, while ensuring data availability is adequately verified before tuning begins.
3Measurement precision
If the model is tuned frequently to adapt to data drift, then classification precision is maintained, but the time and computational resources required increase
Solution Approach 1:
The system monitors classification precision and uses this feedback to determine when model tuning is necessary. By implementing feedback-based tuning triggers rather than frequent scheduled tuning, the system maintains inference precision while reducing unnecessary tuning operations, thereby minimizing time and computational resource loss.
Solution Approach 2:
The system performs model tuning periodically based on monitored performance degradation rather than continuously. This periodic action approach maintains inference precision by tuning only when necessary, reducing the time and computational resources spent on frequent unnecessary tuning operations while still adapting to data drift effectively.
Data Source
AI summary
An information processing apparatus detects a machine learning model that outputs, based on plural sets of image data resulting from imaging of objects, inference results for the plural sets of image data, the inference results having been in practical use already, collects sets of image data resulting from imaging of objects by means of at least one or more cameras, the objects being related to a target that inference results from the machine learning model detected are to be practically used for, determines whether or not the sets of image data collected satisfy a tuning condition for the machine learning model detected, specifies, based on a result of a determination on whether or not the tuning condition is satisfied, a kind of preprocessing to be executed on the sets of image data collected, and generates sets of image data that have been subjected to the kind of preprocessing specified.


