AI Reasoning Model Reduction Using Entropy-Based Feature Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer-based reasoning systems face challenges in balancing the need for broad coverage with the requirement of reducing model size, as large training datasets increase computational and memory resources, often including unnecessary data elements and inefficient parameters.
Innovation Solution
The use of entropy-based techniques to assess the informational value of data elements, such as surprisal and information gain, to selectively include or exclude data points, thereby reducing the model size while maintaining broad coverage by retaining only the most informative data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If large training datasets are used to achieve broad coverage, then the model's coverage and completeness improve, but computational resources and memory usage increase
Solution Approach 1:
The patent extracts and removes redundant or low-value data elements from the training dataset. By applying entropy-based filtering, the system identifies and excludes data points that contribute minimally to model coverage, retaining only the most informative subset. This extraction process reduces dataset size while preserving essential coverage, thereby lowering computational resource requirements without significantly compromising model adaptability.
Solution Approach 2:
The patent changes the parameter of data element selection from uniform inclusion to entropy-based selective inclusion. By calculating entropy values for each data element and applying threshold-based filtering, the system transforms the training dataset composition to include only high-information-density elements. This parameter change optimizes the balance between coverage breadth and computational efficiency.
2Adaptability or versatility
If large training datasets are used to achieve broad coverage, then the model's coverage and completeness improve, but memory usage increases
Solution Approach 1:
The patent extracts and removes redundant or low-value data elements from the training dataset. By applying entropy-based filtering, the system identifies and excludes data points that contribute minimally to model coverage, retaining only the most informative subset. This extraction process reduces dataset size while preserving essential coverage, thereby lowering memory usage without significantly compromising model adaptability.
Solution Approach 2:
The patent changes the parameter of data element selection from uniform inclusion to entropy-based selective inclusion. By calculating entropy values for each data element and applying threshold-based filtering, the system transforms the training dataset composition to include only high-information-density elements. This parameter change optimizes the balance between coverage breadth and memory efficiency.
3Adaptability or versatility
If comprehensive data elements and features are included in the model, then model coverage improves, but model size and processing complexity increase
Solution Approach 1:
The patent extracts and removes redundant or low-value data elements from the training dataset. By applying entropy-based filtering, the system identifies and excludes data points that contribute minimally to model coverage, retaining only the most informative subset. This extraction process reduces dataset size while preserving essential coverage, thereby lowering computational resource requirements without significantly compromising model adaptability.
Solution Approach 2:
The patent changes the parameter of data element selection from uniform inclusion to entropy-based selective inclusion. By calculating entropy values for each data element and applying threshold-based filtering, the system transforms the training dataset composition to include only high-information-density elements. This parameter change optimizes the balance between coverage breadth and memory efficiency.
Data Source
AI summary
Techniques are provided herein for creating well-balanced computer-based reasoning systems and using those to control systems. The techniques include receiving a request to determine whether to use one or more particular data elements, features, cases, etc. in a computer-based reasoning model (e.g., as data elements, cases or features are being added, or as part of pruning existing features or cases). Conviction measures are determined and inclusivity conditions are tested. The result of comparing the conviction measure can be used to determine whether to include or exclude the feature, case, etc. in the model and/or whether there are anomalies in the model. A controllable system may then be controlled using the computer-based reasoning model. Examples controllable systems include self-driving cars, image labeling systems, manufacturing and assembly controls, federated systems, smart voice controls, automated control of experiments, energy transfer systems, health care systems, cybersecurity systems, and the like.


