ML Data Sampling Control for Accuracy-Cost Balance
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing methods for managing metadata and log information do not effectively reduce the data amount of article data used for machine learning, leading to increased management and sending costs without ensuring the accuracy of the machine learning model.
Innovation Solution
An information processing method that acquires and stores article data, calculates the accuracy of a machine learning model, and determines reduction in data items or sampling rate to maintain reference accuracy, then transmits control data to adjust the data items and sampling rate for efficient data collection.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of energy
If the data amount of article data is reduced to lower management and sending costs, then the costs are reduced, but the accuracy of the machine learning model may deteriorate
Solution Approach 1:
The patent changes parameters of the article data including data items, sampling rate, and data amount to optimize the balance between cost reduction and model accuracy. By systematically adjusting these parameters and evaluating their impact on machine learning model accuracy, the system identifies optimal parameter settings that reduce data amount while maintaining acceptable accuracy levels
Solution Approach 2:
The patent implements a feedback mechanism where the accuracy of the machine learning model is calculated and evaluated based on the reduced article data. This accuracy information feeds back into the data reduction process, allowing the system to iteratively adjust data items, sampling rates, and data amounts to maintain accuracy thresholds while achieving cost reduction goals
2Measurement precision
If all data items and sampling rates are maintained to ensure machine learning accuracy, then the accuracy is preserved, but the data amount increases leading to higher management and sending costs
Solution Approach 1:
The patent systematically adjusts parameters including data items, sampling rates, and data amount to find the optimal configuration. By changing these parameters in a controlled manner and evaluating their effect on model accuracy, the system reduces the data amount (quantity of substance) while maintaining accuracy through intelligent parameter selection
Solution Approach 2:
The patent applies different quality levels to different data items by selectively reducing certain data items while maintaining others. Instead of uniformly reducing all data, the system identifies and preserves critical data items that contribute most to model accuracy while reducing or eliminating less important data items, achieving local optimization of data quality
Data Source
AI summary
A server includes: an acquisition part that acquires and stores article data in a memory; a determination part that calculates an accuracy of a machine learning model which performs machine learning by using the article data stored in the memory, and determines at least one of reduction in data item and reduction in sampling rate so that the calculated accuracy satisfies a reference accuracy; and a transmission part that transmits, to an article, control data for controlling the article to send the article data by using at least one of a data item after the reduction and a sampling rate after the reduction.


