Labeling Reliability Scoring for Machine Learning Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing labeling techniques for data in supervised machine learning face challenges with human variation leading to false labels and lack of error countermeasures, particularly in methods relying on manpower or automated ranking of distances.
Innovation Solution
An information processing apparatus that determines labels for learning data by evaluating the reliability of both the label and the user applying it, using a scoring system based on confidence degrees and user reliability to select the most accurate label, thereby reducing false label effects.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If labeling is performed using a large amount of manpower, then labeling coverage is improved, but reliability deteriorates due to human variation and false labels
Solution Approach 1:
The patent introduces an automated labeling system that acts as an intermediary between human labelers and the final labeled dataset. This system calculates reliability scores for both users and labels, and uses these scores to automatically select and prioritize labels for review and acceptance, reducing the impact of human errors while maintaining high labeling coverage
Solution Approach 2:
The system implements feedback mechanisms where reliability information is continuously calculated and updated based on user performance and label quality. This feedback is then used to adjust the labeling process, prioritizing reviews of labels from users with lower reliability scores and improving overall labeling accuracy through iterative refinement
2Productivity
If automated labeling is performed, then productivity is improved, but reliability deteriorates due to lack of error countermeasures
Solution Approach 1:
The patent introduces an automated reliability calculation system that acts as an intermediary between automated labeling processes and final label acceptance. This system calculates reliability scores for automated labels and uses these scores to determine which labels require human review and which can be accepted automatically, maintaining both high productivity and reliability
Solution Approach 2:
The system dynamically adjusts the threshold for automatic label acceptance based on calculated reliability scores. When reliability is high, more labels are accepted automatically to maintain productivity; when reliability is low, more labels are flagged for human review to maintain accuracy, creating a flexible balance between speed and quality
3Ease of operation
If all labels are trusted without verification, then ease of operation is improved, but reliability deteriorates due to undetected false labels
Solution Approach 1:
The system implements self-service mechanisms where the automated reliability calculation and label selection process operates independently without requiring manual intervention for each label. The system automatically identifies high-reliability labels for acceptance and low-reliability labels for review, maintaining ease of operation while improving reliability through intelligent automation
Data Source
AI summary
An information processing apparatus includes an obtaining unit configured to obtain information of a plurality of labels applied to the learning data by a plurality of users, information regarding reliability of each applied label itself, and information regarding reliability of a user who applies the relevant label, wherein the information of the label is information regarding a result to be recognized in a case where the predetermined recognition is performed on the learning data and a determination unit configured to determine a label to the learning data from among the plurality of labels based on the reliability of the label itself and the reliability of the user who applies the relevant label.


