Adversarial Attack System for Mislabel Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing methods for identifying mislabeled data samples in machine learning are inefficient due to low confidence in mislabel identification, requiring costly and time-consuming manual verification by expert annotators.

Innovation Solution

A processor-implemented method using adversarial attacks to identify mislabeled data samples by training a data-driven model, computing logit scores, and performing adversarial attacks to sort samples based on adversarial perturbation strength, recommending candidate mislabeled samples below a predefined threshold.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If manual verification by expert annotators is used for mislabel identification, then identification accuracy can be maintained, but the process becomes costly and time-consuming

Engineering Contradiction:
Improvemislabel identification accuracyVSAvoidverification time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system uses the trained data-driven model to automatically verify labels by computing logit scores and performing adversarial attacks, allowing the system to self-verify without requiring manual intervention from expert annotators for each sample

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the manual mechanical verification process with an automated computational system that uses adversarial attacks and logit score analysis to identify mislabeled samples, substituting human expert judgment with algorithmic decision-making

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Measurement precision

If manual verification by expert annotators is used for mislabel identification, then identification accuracy can be maintained, but the process becomes costly

Engineering Contradiction:
Improvemislabel identification accuracyVSAvoidresource consumption
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The system automatically performs verification using its own trained model, eliminating the need to consume human expert annotator time and resources for each verification task

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent replaces the resource-intensive manual verification process with a computationally efficient automated system that uses adversarial attacks and logit score comparisons to identify mislabeled samples

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

3Productivity

If existing mislabel identification methods are used, then some mislabeled data can be identified, but confidence in identification is low

Engineering Contradiction:
Improvemislabel identification capabilityVSAvoididentification confidence
Core Design Contradiction:
ProductivityVSReliability

Solution Approach 1:

The system uses feedback from adversarial attacks by computing logit scores before and after perturbation, and by analyzing the magnitude of changes in predictions to confidently identify mislabeled samples based on the model's own confidence metrics

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces low-confidence manual identification with a high-confidence automated system that uses adversarial attacks and logit score analysis to objectively determine mislabeling with measurable confidence metrics

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS20220335335A1Method and system for identifying mislabeled data samples using adversarial attacks
Publication Date: 2022.10.20 TATA CONSULTANCY SERVICES LTD
  • US20220335335A1 patent drawing
  • US20220335335A1 patent drawing
  • US20220335335A1 patent drawing

AI summary

Mislabeled data when used for various applications such as training of Machine Learning (ML) models, can cause erroneous results. The state-of-the-art systems performs the mislabel identification with low confidence, and some require manual intervention. The disclosure herein generally relates to data processing, and, more particularly, to a method and system for identifying mislabeled samples using adversarial attacks. The mislabeled sample may refer to a) a data sample that is tagged with a wrong/incorrect label, or b) a distorted/confusing data sample having similarity with multiple classes. The system performs adversarial attack on training data using varying values of adversarial perturbations, and then identifies, for each of the misguided data samples, least value of adversarial perturbation that was required to misguide each of the data samples. Further, the data samples which were misguided by small values of adversarial perturbation, are identified as candidate mislabeled data samples.