Machine Learning Training Data Sampling via Importance Correction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current machine learning systems require significant resources to process large datasets, leading to high costs and diminished training performance due to random sampling methods that discard valuable data.
Innovation Solution
A method and system for training machine learning systems by dividing sample data into sets based on sampling time periods, setting sampling rates, determining importance values, and correcting sample data using these values to input optimized data for training.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If random sampling is used to reduce data amount, then processing cost is reduced, but training performance deteriorates due to loss of useful data
Solution Approach 1:
The patent applies local quality by assigning different sampling rates to different sample sets based on their time periods and importance values. Instead of uniform random sampling, the system identifies valuable samples through importance scoring and applies higher sampling rates to high-importance samples while using lower rates for low-importance samples. This localized differentiation preserves useful data while reducing overall data volume, resolving the contradiction between data reduction and training performance maintenance.
2Measurement precision
If large scale data is processed to maintain training performance, then prediction precision is improved, but machine resource consumption increases
Solution Approach 1:
The patent changes the parameter of sampling rate from a fixed uniform value to a dynamic variable that depends on sample importance and time period. By calculating importance values for different sample sets and adjusting sampling rates accordingly, the system processes fewer high-value samples and more low-value samples, reducing overall machine resource consumption while maintaining prediction precision through selective preservation of important data.
3Reliability
If sampling rate is increased to preserve data quality, then training effectiveness is improved, but processing time increases
Solution Approach 1:
The patent applies partial action by using different sampling rates for different sample sets rather than uniformly high sampling across all data. High-importance sample sets receive higher sampling rates to preserve training effectiveness, while low-importance sample sets receive lower sampling rates to reduce processing time. This selective partial sampling maintains training effectiveness for critical data while minimizing overall processing time through reduced sampling of less critical data.
Data Source
AI summary
The present disclosure provides a method and a system for training a machine learning system. Multiple pieces of sample data are used for training the machine learning system. The method includes acquiring multiple sample sets, each sample set including sample data in a corresponding sampling time period; setting a sampling rate for each sample set according to the corresponding sampling time period; acquiring multiple sample sets sampled according to set sampling rates; determining importance values of the multiple sampled sample sets; correcting each piece of sample data in the multiple sampled sample sets by using a corresponding importance value to obtain corrected sample data; and inputting the corrected sample data into the machine learning system to train the machine learning system.


