Mixed Sampling Classification Training for Imbalanced Data

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing deep learning classification models face limitations in performance improvement, particularly due to the accuracy of target data and label factors, when applied to actual data sets.

Innovation Solution

The method involves dividing data sets into training and testing data, applying different sampling techniques based on data characteristics, and reconstructing training data using mixed sampling to enhance classification performance through Nth and N+1th classification models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If standard deep learning classification models are trained on raw data sets, then the model structure can be simple and training can be straightforward, but the classification accuracy is limited due to data quality issues and class imbalance

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata preprocessing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent segments the data set into multiple type groups based on classification results, then applies different sampling techniques to each group. This segmentation allows targeted processing of imbalanced classes while maintaining overall system effectiveness

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Different sampling techniques (under-sampling, re-labeling, data augmentation) are applied to specific type groups based on their characteristics. This local quality approach ensures that each data group receives the most appropriate processing method for its specific needs

Inventive Principle:
Principle #3Local quality

2Reliability

If the entire data set is used for training without sampling, then all available data is utilized, but class imbalance and data quality issues limit model performance

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent performs preliminary classification and sampling operations before the main training process. By pre-processing the data to create balanced type groups, the subsequent training becomes more efficient and reliable

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent changes the composition and distribution parameters of training data through various sampling techniques. This transforms the raw data parameters into optimized training set parameters that improve model reliability

Inventive Principle:
Principle #35Parameter changes

3Measurement precision

If sampling techniques are applied to balance class distribution, then classification accuracy improves, but the data processing time and computational overhead increase

Engineering Contradiction:
Improveclassification accuracyVSAvoiddata processing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent applies sampling techniques selectively to specific type groups rather than the entire data set. This partial action approach focuses computational resources on the most critical imbalanced classes, reducing overall processing time while maintaining accuracy improvements

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS12481918B2Method and apparatus for improving performance of classification on the basis of mixed sampling
Publication Date: 2025.11.25 ELECTRONICS & TELECOMM RES INST
  • US12481918B2 patent drawing
  • US12481918B2 patent drawing
  • US12481918B2 patent drawing

AI summary

A method of improving performance of classification on the basis of mixed sampling is applied. The present invention is directed to providing a method and apparatus for improving the performance of classification on the basis of mixed sampling that are capable of, when learning a model that classifies types using deep learning by dividing the entire data set and using the divided data set for training, validation, and testing, applying different sampling techniques by types of data according to the characteristics of the training data in order to improve the classification performance.