Task-Specific Error Correction Data Determination for Machine Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing data cleansing technologies fail to provide appropriate error correction tailored to specific analysis tasks in machine learning, as the type of error correction required varies with the analysis task.

Innovation Solution

An information processing apparatus and method that calculates the degree of influence of errors on a machine learning model's evaluation index, determining which data to correct based on these influences, allowing for tailored error correction.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If conventional data cleansing technology is used, then data errors can be corrected, but the correction cannot be tailored to specific analysis tasks in machine learning

Engineering Contradiction:
Improvetask-specific error correction capabilityVSAvoiddata analysis accuracy
Core Design Contradiction:
Adaptability or versatilityVSMeasurement precision

Solution Approach 1:

The patent applies local quality by calculating the degree of influence for each error type separately and determining correction priorities based on task-specific requirements. Different error types (missing values, format deviations, anomalous values) are evaluated independently to identify which errors most impact the specific machine learning task at hand, allowing targeted correction of the most critical errors rather than uniform treatment of all errors

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent changes the parameter of error evaluation from generic to task-specific by introducing an evaluation index that reflects the specific machine learning task requirements. The system calculates how each error type affects this task-specific evaluation index, thereby adapting the data cleansing process to the particular analysis task being performed

Inventive Principle:
Principle #35Parameter changes

2Reliability

If all errors in target data are corrected, then data quality improves, but the work量和 resources required increase significantly

Engineering Contradiction:
Improvedata qualityVSAvoiddata cleansing efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent applies partial action by calculating the degree of influence for each error type and selectively correcting only those errors that have the greatest impact on the machine learning task. Rather than correcting all errors uniformly, the system identifies and prioritizes correction of errors with high influence on the evaluation index, achieving significant data quality improvement with reduced processing effort

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system performs self-service by automatically calculating the degree of influence for each error type and determining correction priorities without requiring manual assessment. The machine learning model itself provides feedback on which errors most affect its performance, allowing the data cleansing process to be automatically optimized for the specific task

Inventive Principle:
Principle #25Self-service

Data Source

PatentUS20250335291A1Correction data determination apparatus, correction data determination method, and storage medium
Publication Date: 2025.10.30 NEC CORP
  • US20250335291A1 patent drawing
  • US20250335291A1 patent drawing
  • US20250335291A1 patent drawing

AI summary

To enable appropriate error correction that suits an analysis task, an information processing apparatus (1) includes: an acquisition unit (11) that acquires target data; a calculation unit (12) that calculates, for respective ones of a plurality of errors included in the target data or for respective ones of attributes of the plurality of errors, corresponding degrees of influence that the respective ones of the plurality of errors exert on an evaluation index of a machine learning model; and a determination unit (13) that determines data to be corrected in the target data on the basis of the degrees of influence calculated by the calculation unit (12).