Data Generation Device for Balancing Class Distribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data oversampling techniques, such as FSMOTE, cause overlap between different classes during machine learning model training, leading to decreased classification accuracy and fairness.
Innovation Solution
A data generation device that selects origin data from minority clusters based on the distribution of majority clusters, generating new composite data at positions corresponding to the majority cluster's distribution to avoid overlap and balance class sizes, thereby suppressing overlap and improving classification fairness.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If data oversampling technique is used to generate new data for minority classes, then the quantity of minority class data is improved, but overlap between different classes occurs leading to decreased classification accuracy
Solution Approach 1:
The patent applies local quality by generating synthetic data with different levels of noise addition based on the specific characteristics of each data point. Data points closer to decision boundaries receive different noise treatment compared to those clearly within their class region, thereby generating diverse synthetic samples while maintaining class separability and avoiding overlap that would harm classification accuracy.
Solution Approach 2:
The patent changes parameters by systematically varying noise levels and adding perturbations to generated data. By controlling the magnitude and type of noise added during synthetic data generation, the method creates diverse samples that expand minority class representation without causing overlap with majority class regions, thus resolving the contradiction between quantity and accuracy.
2Quantity of substance
If data oversampling technique is used to generate new data for minority classes, then the quantity of minority class data is improved, but fairness in classification deteriorates
Solution Approach 1:
The patent applies local quality by differentiating the treatment of data points based on their local density and distance to decision boundaries. Minority class data points in sparse regions receive targeted synthetic generation with appropriate noise levels, while avoiding generation in regions where overlap would occur. This localized approach ensures fair representation without compromising the reliability and fairness of classification outcomes.
Solution Approach 2:
The patent incorporates feedback mechanisms by evaluating the quality and distribution of generated data, and adjusting the generation process accordingly. By monitoring whether generated samples cause overlap or improve representation, the system iteratively refines the oversampling strategy to achieve both quantity improvement and fairness maintenance in classification.
3Device complexity
If existing oversampling methods generate data based on minority cluster alone, then the generation process is simple, but the generated data overlaps with majority cluster reducing effectiveness
Solution Approach 1:
The patent merges the consideration of both minority and majority clusters in the data generation process. By simultaneously analyzing the distribution characteristics of both clusters and generating data that respects the boundaries between them, the method achieves higher generation precision without excessive complexity. The unified approach considers inter-cluster relationships while maintaining computational efficiency.
Solution Approach 2:
The patent applies preliminary action by first analyzing the distribution and boundaries of both minority and majority clusters before generating synthetic data. This preliminary assessment of cluster characteristics enables the generation process to target appropriate regions and avoid overlap, improving data generation precision while keeping the overall process manageable through structured preliminary analysis.
Data Source
AI summary
A non-transitory computer-readable recording medium storing a data generation program for causing a computer to execute processing including: selecting, based on first distribution of data included in a first data group in which a value of a first attribute is a first value among a plurality of data groups obtained by classifying a plurality of pieces of data based on an attribute, first data from a second data group in which the value of the first attribute is a second value among the plurality of data groups; and generating new data in which the value of the first attribute is the second value based on the first data.


