AI Reasoning Model Pruning for Anomalous Data Detection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Computer-based reasoning systems face challenges in balancing the need for broad training data coverage with the requirement of reducing model size, leading to inefficiencies in computational and memory resources due to large data sets and unnecessary features.

Innovation Solution

The use of entropy-based techniques to assess and reduce the number of data elements by calculating information gain and surprisal, retaining only those that provide significant informational value, and directing training towards areas with high surprisal, while flagging and correcting anomalous data elements.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If broad training data coverage is obtained, then model coverage and reliability are improved, but model size and computational resources increase

Engineering Contradiction:
Improvemodel coverageVSAvoidmodel size
Core Design Contradiction:
ReliabilityVSQuantity of substance

Solution Approach 1:

The patent extracts and removes redundant, anomalous, and low-value data elements from the training dataset. By calculating information gain and surprisal metrics, the system identifies and eliminates data points that do not contribute meaningfully to model coverage, thereby reducing model size while preserving essential coverage characteristics.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent transforms the training data by applying parameter-based filtering using information gain and surprisal thresholds. Data elements are selectively retained or removed based on their calculated information metrics, changing the composition parameters of the dataset to achieve an optimal balance between coverage and size.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If large data sets are used, then model coverage is improved, but computational efficiency deteriorates

Engineering Contradiction:
ImprovecoverageVSAvoidcomputational efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent extracts computationally expensive redundant data elements from the training set. By removing duplicate and anomalous entries through information gain analysis, the system reduces the computational burden during training and inference while maintaining the essential coverage of the data.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies partial action by selectively processing only the most informative data elements. Rather than uniformly processing all data points, the system focuses computational resources on high-value elements identified through surprisal calculations, achieving efficient coverage with reduced computational effort.

Inventive Principle:
Principle #16Partial or excessive action

3Reliability

If more features are included in the model, then coverage is improved, but model complexity and inefficiency increase

Engineering Contradiction:
ImprovecoverageVSAvoidmodel complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent applies local quality by selectively including features based on their local information value. Different features are evaluated individually using information gain metrics, and only those with significant local contribution to coverage are retained, creating a non-uniform feature set optimized for efficiency.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the feature set into distinct categories based on information gain thresholds. Features are divided into retained, removed, and borderline groups, allowing systematic management of model complexity while preserving essential coverage capabilities.

Inventive Principle:
Principle #1Segmentation

4Quantity of substance

If data elements with low information gain are removed, then model size is reduced, but information loss may occur

Engineering Contradiction:
Improvemodel sizeVSAvoidinformational value
Core Design Contradiction:
Quantity of substanceVSLoss of information

Solution Approach 1:

The patent implements feedback through iterative calculation of information gain and surprisal metrics. The system continuously evaluates the impact of removing data elements on overall model coverage, adjusting the selection process to ensure that only truly redundant elements are removed while preserving informational value.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent replaces simple size-based filtering with an information-theoretic evaluation system. Instead of mechanically removing a fixed percentage of data, the system uses surprisal and information gain calculations to intelligently identify elements for removal, substituting computational analysis for brute-force reduction.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

Data Source

PatentUS11262742B2Anomalous data detection in computer based reasoning and artificial intelligence systems
Publication Date: 2022.03.01 HOWSO INC
  • US11262742B2 patent drawing
  • US11262742B2 patent drawing
  • US11262742B2 patent drawing

AI summary

Techniques are provided herein for creating well-balanced computer-based reasoning systems and using those to control systems. The techniques include receiving a request to determine whether to use one or more particular data elements, features, cases, etc. in a computer-based reasoning model (e.g., as data elements, cases or features are being added, or as part of pruning existing features or cases). Conviction measures are determined and inclusivity conditions are tested. The result of comparing the conviction measure can be used to determine whether to include or exclude the feature, case, etc. in the model and/or whether there are anomalies in the model. A controllable system may then be controlled using the computer-based reasoning model. Examples controllable systems include self-driving cars, image labeling systems, manufacturing and assembly controls, federated systems, smart voice controls, automated control of experiments, energy transfer systems, health care systems, cybersecurity systems, and the like.