Data Generation Device for Balancing Class Distribution

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data oversampling techniques, such as FSMOTE, cause overlap between different classes during machine learning model training, leading to decreased classification accuracy and fairness.

Innovation Solution

A data generation device that selects origin data from minority clusters based on the distribution of majority clusters, generating new composite data at positions corresponding to the majority cluster's distribution to avoid overlap and balance class sizes, thereby suppressing overlap and improving classification fairness.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If data oversampling technique is used to generate new data for minority classes, then the quantity of minority class data is improved, but overlap between different classes occurs leading to decreased classification accuracy

Engineering Contradiction:
Improvequantity of minority class dataVSAvoidclassification accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent applies local quality by generating synthetic data with different levels of noise addition based on the specific characteristics of each data point. Data points closer to decision boundaries receive different noise treatment compared to those clearly within their class region, thereby generating diverse synthetic samples while maintaining class separability and avoiding overlap that would harm classification accuracy.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes parameters by systematically varying noise levels and adding perturbations to generated data. By controlling the magnitude and type of noise added during synthetic data generation, the method creates diverse samples that expand minority class representation without causing overlap with majority class regions, thus resolving the contradiction between quantity and accuracy.

Inventive Principle:
Principle #35Parameter changes

2Quantity of substance

If data oversampling technique is used to generate new data for minority classes, then the quantity of minority class data is improved, but fairness in classification deteriorates

Engineering Contradiction:
Improvequantity of minority class dataVSAvoidclassification fairness
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent applies local quality by differentiating the treatment of data points based on their local density and distance to decision boundaries. Minority class data points in sparse regions receive targeted synthetic generation with appropriate noise levels, while avoiding generation in regions where overlap would occur. This localized approach ensures fair representation without compromising the reliability and fairness of classification outcomes.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent incorporates feedback mechanisms by evaluating the quality and distribution of generated data, and adjusting the generation process accordingly. By monitoring whether generated samples cause overlap or improve representation, the system iteratively refines the oversampling strategy to achieve both quantity improvement and fairness maintenance in classification.

Inventive Principle:
Principle #23Feedback

3Device complexity

If existing oversampling methods generate data based on minority cluster alone, then the generation process is simple, but the generated data overlaps with majority cluster reducing effectiveness

Engineering Contradiction:
Improvedata generation process complexityVSAvoiddata generation precision
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent merges the consideration of both minority and majority clusters in the data generation process. By simultaneously analyzing the distribution characteristics of both clusters and generating data that respects the boundaries between them, the method achieves higher generation precision without excessive complexity. The unified approach considers inter-cluster relationships while maintaining computational efficiency.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent applies preliminary action by first analyzing the distribution and boundaries of both minority and majority clusters before generating synthetic data. This preliminary assessment of cluster characteristics enables the generation process to target appropriate regions and avoid overlap, improving data generation precision while keeping the overall process manageable through structured preliminary analysis.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240161011A1Computer-readable recording medium storing data generation program, data generation method, and data generation device
Publication Date: 2024.05.16 FUJITSU LTD
  • US20240161011A1 patent drawing
  • US20240161011A1 patent drawing
  • US20240161011A1 patent drawing

AI summary

A non-transitory computer-readable recording medium storing a data generation program for causing a computer to execute processing including: selecting, based on first distribution of data included in a first data group in which a value of a first attribute is a first value among a plurality of data groups obtained by classifying a plurality of pieces of data based on an attribute, first data from a second data group in which the value of the first attribute is a second value among the plurality of data groups; and generating new data in which the value of the first attribute is the second value based on the first data.