Image Data Preprocessing for ML Tuning Under Data Drift

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing machine learning models face precision deterioration due to data drift during practical use, especially in irregular or small-scale applications, where insufficient data for tuning leads to reduced classification accuracy.

Innovation Solution

An information processing apparatus that collects and preprocesses image data to meet tuning conditions, dynamically adjusts preprocessing methods, and re-trains the model to adapt to data drift, ensuring adequate data availability and maintaining high inference precision.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If the DL model is tuned using collected image data, then the model adapts to data drift and maintains classification precision, but insufficient data quantity leads to inadequate tuning and reduced accuracy

Engineering Contradiction:
Improveclassification precisionVSAvoiddata quantity
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system generates synthetic image data that copies and replicates real object characteristics through computer-generated imagery. This synthetic data is created to supplement insufficient collected images, providing adequate training samples for model tuning without requiring additional physical imaging, thereby resolving the contradiction between data quantity and classification precision.

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The system varies parameters such as lighting conditions, camera angles, and object positions when generating synthetic image data. By changing these parameters to create diverse synthetic samples, the system ensures adequate data quantity and variety for effective model tuning, maintaining classification precision even when real collected data is insufficient.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If image data is collected and preprocessed to meet tuning conditions, then adequate data is available for model tuning, but the preprocessing complexity increases

Engineering Contradiction:
Improvedata availabilityVSAvoidpreprocessing complexity
Core Design Contradiction:
Quantity of substanceVSDevice complexity

Solution Approach 1:

The system automatically determines whether collected data meets tuning conditions and self-adjusts by generating synthetic data only when necessary. This self-service approach simplifies preprocessing complexity by eliminating manual data assessment and generation decisions, while ensuring adequate data availability through automated synthetic data creation when real data is insufficient.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs preliminary assessment of collected image data to determine if tuning conditions are met before proceeding with model tuning. This preliminary action prevents unnecessary synthetic data generation and simplifies the overall preprocessing workflow, while ensuring data availability is adequately verified before tuning begins.

Inventive Principle:
Principle #10Preliminary action

3Measurement precision

If the model is tuned frequently to adapt to data drift, then classification precision is maintained, but the time and computational resources required increase

Engineering Contradiction:
Improveinference precisionVSAvoidtuning time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system monitors classification precision and uses this feedback to determine when model tuning is necessary. By implementing feedback-based tuning triggers rather than frequent scheduled tuning, the system maintains inference precision while reducing unnecessary tuning operations, thereby minimizing time and computational resource loss.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system performs model tuning periodically based on monitored performance degradation rather than continuously. This periodic action approach maintains inference precision by tuning only when necessary, reducing the time and computational resources spent on frequent unnecessary tuning operations while still adapting to data drift effectively.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20260051155A1Non-transitory computer-readable recording medium, information processing apparatus, and information processing method
Publication Date: 2026.02.19 FUJITSU LTD
  • US20260051155A1 patent drawing
  • US20260051155A1 patent drawing
  • US20260051155A1 patent drawing

AI summary

An information processing apparatus detects a machine learning model that outputs, based on plural sets of image data resulting from imaging of objects, inference results for the plural sets of image data, the inference results having been in practical use already, collects sets of image data resulting from imaging of objects by means of at least one or more cameras, the objects being related to a target that inference results from the machine learning model detected are to be practically used for, determines whether or not the sets of image data collected satisfy a tuning condition for the machine learning model detected, specifies, based on a result of a determination on whether or not the tuning condition is satisfied, a kind of preprocessing to be executed on the sets of image data collected, and generates sets of image data that have been subjected to the kind of preprocessing specified.