Balanced Learning Data Generation for Error Diagnosis Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In the context of multi-function peripherals, existing methods for estimating error occurrence based on operating history face challenges due to imbalanced learning data ratios, leading to reduced accuracy in identifying devices that require preventive maintenance.
Innovation Solution
The system extracts specific error entries as positive examples, selects a portion of negative examples, and generates learning data to create a diagnosis model that estimates the likelihood of error occurrence, ensuring a balanced ratio between positive and negative examples to improve accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If all error history data is used as learning data, then the model can learn from comprehensive data, but the ratio between positive examples and negative examples becomes unbalanced, reducing model accuracy
Solution Approach 1:
The patent extracts only the necessary portion of negative examples (error-free operation data) to create a balanced learning dataset. Specifically, it selects negative examples corresponding to devices that did not exhibit the target error, balancing their quantity with positive examples (devices that exhibited the error). This extraction approach maintains sufficient learning data while achieving class balance for accurate model training.
Solution Approach 2:
The patent changes the parameter of learning data composition by dynamically adjusting the ratio of positive to negative examples. Instead of using all available data with fixed imbalance, it modifies the dataset parameters to achieve a balanced 1:1 ratio between positive and negative examples, thereby optimizing model accuracy for error occurrence prediction.
2Measurement precision
If a balanced ratio of positive and negative examples is created, then model accuracy improves, but storage and computational resources are reduced
Solution Approach 1:
The patent extracts and retains only the essential portion of learning data needed for balanced training. By selecting a specific quantity of negative examples that matches the number of positive examples, it maintains adequate data volume for accurate model training while minimizing unnecessary data storage, thus optimizing the balance between accuracy and resource efficiency.
3Adaptability or versatility
If comprehensive error history data is processed, then all error types can be analyzed, but computational resources and processing time increase
Solution Approach 1:
The patent extracts and processes only the relevant subset of error history data needed for each specific error type analysis. For each target error type, it identifies corresponding positive examples (devices with that error) and matched negative examples (devices without that error), processing only this focused subset rather than all error data, thereby reducing computational overhead while maintaining analysis effectiveness.
Data Source
AI summary
According to one embodiment, information processing device includes a storage and a processing circuit. The storage is configured to store configuration data of a plurality of devices and error history data including a plurality of entries each with a type of error. The processing circuit is configured to: extract the entries with a specific type of error which are positive examples; extract the remaining entries which are negative examples; select a part of the negative examples with each type of error; generate learning data based on the positive examples, the selected negative examples and the configuration data at a time when the specific type of error occurred; and generate a diagnosis model for estimation of possibility the specific type of error occurs by using the learning data.


