Task-Specific Error Correction Data Determination for Machine Learning
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing data cleansing technologies fail to provide appropriate error correction tailored to specific analysis tasks in machine learning, as the type of error correction required varies with the analysis task.
Innovation Solution
An information processing apparatus and method that calculates the degree of influence of errors on a machine learning model's evaluation index, determining which data to correct based on these influences, allowing for tailored error correction.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional data cleansing technology is used, then data errors can be corrected, but the correction cannot be tailored to specific analysis tasks in machine learning
Solution Approach 1:
The patent applies local quality by calculating the degree of influence for each error type separately and determining correction priorities based on task-specific requirements. Different error types (missing values, format deviations, anomalous values) are evaluated independently to identify which errors most impact the specific machine learning task at hand, allowing targeted correction of the most critical errors rather than uniform treatment of all errors
Solution Approach 2:
The patent changes the parameter of error evaluation from generic to task-specific by introducing an evaluation index that reflects the specific machine learning task requirements. The system calculates how each error type affects this task-specific evaluation index, thereby adapting the data cleansing process to the particular analysis task being performed
2Reliability
If all errors in target data are corrected, then data quality improves, but the work量和 resources required increase significantly
Solution Approach 1:
The patent applies partial action by calculating the degree of influence for each error type and selectively correcting only those errors that have the greatest impact on the machine learning task. Rather than correcting all errors uniformly, the system identifies and prioritizes correction of errors with high influence on the evaluation index, achieving significant data quality improvement with reduced processing effort
Solution Approach 2:
The system performs self-service by automatically calculating the degree of influence for each error type and determining correction priorities without requiring manual assessment. The machine learning model itself provides feedback on which errors most affect its performance, allowing the data cleansing process to be automatically optimized for the specific task
Data Source
AI summary
To enable appropriate error correction that suits an analysis task, an information processing apparatus (1) includes: an acquisition unit (11) that acquires target data; a calculation unit (12) that calculates, for respective ones of a plurality of errors included in the target data or for respective ones of attributes of the plurality of errors, corresponding degrees of influence that the respective ones of the plurality of errors exert on an evaluation index of a machine learning model; and a determination unit (13) that determines data to be corrected in the target data on the basis of the degrees of influence calculated by the calculation unit (12).


