Wobbliness Measurement for ML Model Overfitting and Security

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Deep learning models are susceptible to overfitting and security threats, such as backdoors and adversarial attacks, due to the increased availability of training data and computation power, which can lead to misclassification and performance degradation, and existing methods do not effectively measure overfitting to individual input samples or detect backdoors without labeled test datasets.

Innovation Solution

A method that uses the 'wobbliness' measurement to quantify the stability of a machine learning model's decision surface by sampling around data points, calculating variance, entropy, and area, allowing for the identification of overfitting and potential security threats, including backdoors and adversarial examples, without requiring labeled test datasets.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If deep learning models are trained with increased availability of training data and computation power, then model accuracy and performance are improved, but susceptibility to overfitting and security threats increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidsusceptibility to overfitting and security threats
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent applies preliminary action by performing overfitting measurements and security threat assessments during the model training process, before deployment. This allows detection of overfitting and potential backdoor attacks early, enabling corrective actions to be taken before the model causes harm in production environments.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent introduces an intermediary measurement mechanism called 'overfitting measurement' that acts as a mediator between the training process and deployment. This measurement system evaluates model behavior on unlabeled datasets and provides feedback about overfitting and security vulnerabilities without requiring labeled test data, thus bridging the gap between training and production safety.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If existing methods are used to detect overfitting and security threats, then model safety can be assessed, but labeled test datasets are required which are often unavailable

Engineering Contradiction:
Improvemodel safety assessmentVSAvoidrequirement for labeled test datasets
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts the essential measurement capability from the requirement for labeled data. By focusing on measuring overfitting through unlabeled datasets and using the differences between training and measurement performances, the method removes the dependency on labeled test datasets while maintaining safety assessment capability.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent uses inexpensive, easily obtainable unlabeled datasets for measurement purposes instead of expensive, hard-to-obtain labeled test datasets. These unlabeled datasets can be freely collected and used for overfitting measurement without the need for expert annotation, making safety assessment accessible and scalable.

Inventive Principle:
Principle #27Cheap short-living objects (Disposable)

3Reliability

If overfitting measurements are performed to detect security threats, then model vulnerabilities can be identified, but computational resources and time are consumed

Engineering Contradiction:
Improvedetection of model vulnerabilitiesVSAvoidtime for overfitting measurement
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent applies partial action by performing overfitting measurements on a subset of unlabeled data rather than requiring comprehensive labeled test datasets. This partial measurement approach provides sufficient safety assessment without consuming excessive computational resources or time, enabling efficient detection of overfitting and security threats.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11494496B2Measuring overfitting of machine learning computer model and susceptibility to security threats
Publication Date: 2022.11.08 INTERNATIONAL BUSINESS MACHINE CORPORATION
  • US11494496B2 patent drawing
  • US11494496B2 patent drawing
  • US11494496B2 patent drawing

AI summary

Mechanisms are provided to determine a susceptibility of a trained machine learning model to a cybersecurity threat. The mechanisms execute a trained machine learning model on a test dataset to generate test results output data, and determine an overfit measure of the trained machine learning model based on the generated test results output data. The overfit measure quantifies an amount of overfitting of the trained machine learning model to a specific sub-portion of the test dataset. The mechanisms apply analytics to the overfit measure to determine a susceptibility probability that indicates a likelihood that the trained machine learning model is susceptible to a cybersecurity threat based on the determined amount of overfitting of the trained machine learning model. The mechanisms perform a corrective action based on the determined susceptibility probability.