Entropy-Based Feature Selection for Computer Reasoning Models
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer-based reasoning systems face challenges in balancing the need for broad training data coverage with the requirement of reducing model size, leading to inefficiencies in computational and memory resources due to large data sets and unnecessary features.
Innovation Solution
The use of entropy-based techniques to assess and reduce the number of data elements by calculating information gain and surprisal, focusing on retaining only the most informative data points and features, and dynamically directing training to areas with high informational value.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If broad training data coverage is obtained, then model coverage and reliability are improved, but model size and computational resource requirements increase
Solution Approach 1:
The patent extracts and removes redundant, duplicate, and low-value data elements from the training dataset. By identifying and eliminating unnecessary features and data points that do not contribute meaningfully to model performance, the system reduces model size while preserving essential coverage. This extraction process directly addresses the contradiction by separating useful information from wasteful data accumulation.
Solution Approach 2:
The patent applies parameter changes by adjusting data element selection criteria based on entropy-based metrics such as information gain and surprisal. By dynamically changing which data elements are included or excluded based on their informational value, the system optimizes the balance between coverage and size. This allows the model to maintain reliability with a more compact representation of training data.
2Reliability
If large data sets are used, then model coverage is improved, but computational efficiency and memory resource usage deteriorate
Solution Approach 1:
The patent extracts and eliminates computationally expensive redundant data elements while preserving those with high informational value. By removing duplicate contexts, redundant features, and low-surprisal data points, the system reduces the computational burden during training and inference while maintaining model coverage. This extraction directly improves computational efficiency without sacrificing reliability.
Solution Approach 2:
Instead of starting with comprehensive data collection and then filtering, the patent inverts the approach by selectively including only high-value data elements from the beginning. By using entropy-based criteria to pre-filter training data before model training, the system avoids wasting computational resources on processing low-value data, thereby improving overall computational efficiency while maintaining coverage.
3Reliability
If comprehensive features are included, then model coverage is improved, but model complexity and resource requirements increase
Solution Approach 1:
The patent extracts and removes unnecessary features and data elements that increase model complexity without contributing to coverage. By identifying and eliminating redundant features across context-action pairs, the system reduces model complexity while preserving the essential feature set needed for comprehensive coverage. This selective extraction directly addresses the contradiction between coverage and complexity.
Solution Approach 2:
The patent applies parameter changes by adjusting feature selection based on entropy-based metrics. Features are dynamically included or excluded based on their information gain and surprisal values, optimizing the feature set to balance coverage and complexity. This parameter-driven feature selection ensures that only necessary features are retained, reducing model complexity while maintaining reliability.
4Loss of information
If more data elements are retained, then informational value is improved, but model size and resource usage increase
Solution Approach 1:
The patent applies parameter changes by using entropy-based metrics (information gain, surprisal) to dynamically determine which data elements to retain. By changing the selection parameters from uniform inclusion to value-based inclusion, the system maximizes informational value while minimizing data element count. This parameter-driven selection ensures that each retained data element contributes maximally to informational value.
Solution Approach 2:
The patent extracts and removes data elements with low informational value, such as those with low surprisal or high redundancy. By systematically identifying and eliminating these low-value elements, the system reduces the total data element count while preserving the informational value contained in high-surprisal, high-information-gain data points. This extraction process directly resolves the contradiction between informational value and data quantity.
Data Source
AI summary
Techniques are provided herein for creating well-balanced computer-based reasoning systems and using those to control systems. The techniques include receiving a request to determine whether to use one or more particular features, cases, etc. in a computer-based reasoning model (e.g., as cases or features are being added, or as part of pruning existing features or cases). Conviction measures (such as targeted or untargeted conviction, contribution, surprisal, etc.) are determined and inclusivity conditions are tested. The result of comparing the conviction measure can be used to determine whether to include or exclude the feature, case, etc. in the computer-based reasoning model. A controllable system may then be controlled using the computer-based reasoning model. Examples controllable systems include self-driving cars, image labeling systems, manufacturing and assembly controls, federated systems, smart voice controls, automated control of experiments, energy transfer systems, and the like.


