Entropy-Based Reasoning Model Pruning for Broad Data Coverage
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Computer-based reasoning systems face challenges in balancing the need for broad training data coverage with the requirement of reducing model size, leading to inefficiencies due to large memory usage and processing costs.
Innovation Solution
The use of entropy-based techniques to assess the informational value of data elements, allowing for the reduction of data sets by retaining only those with high surprisal values and removing or correcting elements with low surprisal, thereby controlling model size while maintaining informational breadth.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If broad training data coverage is obtained, then model coverage and reliability are improved, but model size and computing resource usage increase
Solution Approach 1:
The patent extracts and removes data elements with low surprisal values from the training dataset. By calculating surprisal for each data element and selectively removing those below a threshold, the system reduces model size while preserving the most informative data, thereby resolving the contradiction between model coverage and model size.
Solution Approach 2:
The patent changes the parameter of data element selection from uniform inclusion to surprisal-based filtering. By introducing surprisal calculation as a selection criterion and adjusting the surprisal threshold parameter, the system optimizes the balance between maintaining model coverage and reducing model size.
2Loss of information
If all collected data elements are included in the model, then informational breadth is maximized, but processing time and computational resources increase
Solution Approach 1:
The patent extracts only the most informative data elements by removing those with low surprisal values. This selective extraction reduces the total number of data elements that need to be processed during model training and inference, thereby reducing processing time while maintaining informational breadth through preservation of high-surprisal elements.
Solution Approach 2:
The patent applies partial action by including only a subset of data elements (those with high surprisal) rather than processing all collected data. This partial inclusion strategy reduces computational workload and processing time while still achieving sufficient informational breadth for effective model performance.
3Quantity of substance
If model size is reduced, then computing resource usage decreases, but coverage and informational value may be compromised
Solution Approach 1:
The patent changes the selection parameter from uniform data inclusion to surprisal-based filtering. By using surprisal values as a criterion and adjusting the threshold parameter, the system identifies and retains data elements with the highest informational value, ensuring that model size reduction does not compromise informational value.
Solution Approach 2:
The patent applies local quality by treating data elements differently based on their individual surprisal characteristics. Instead of uniform inclusion or exclusion, each data element is evaluated and retained or removed based on its local informational value, ensuring that high-value elements are preserved while reducing overall model size.
Data Source
AI summary
Techniques are provided herein for creating well-balanced computer-based reasoning systems and using those to control systems. The techniques include receiving a request to determine whether to include one or more particular data elements in a computer-based reasoning model and determining two probability density or mass functions (“PDMFs”), one for the data set including the one or more particular data elements, once for the data set excluding it. Surprisal is determined based on those two PDMFs, and inclusion in the computer-based reasoning model is determined based on surprisal. A system is later controlled using the computer-based reasoning model.


