Model Compression Rate Determination via Importance Value Turning Point
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for compressing machine learning models to fit devices with limited resources require time-consuming training and adjustment processes, leading to inefficient resource consumption and potential accuracy loss due to manual setting of pruning rates.
Innovation Solution
Determining a near-zero importance value subset and a target importance value within the subset to calculate a model compression rate, allowing for optimal compression without compromising model performance, thereby reducing the need for repeated training and resource-intensive adjustments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If manual setting of pruning rates is used for model compression, then model size is reduced, but time-consuming training and adjustment processes are required leading to inefficient resource consumption
Solution Approach 1:
The patent performs preliminary analysis of importance values before actual pruning to determine the optimal compression rate. By pre-calculating which parameters have near-zero importance values and determining the turning point in advance, the system avoids time-consuming trial-and-error training adjustments, directly achieving model compression with minimal retraining time.
Solution Approach 2:
The system automatically determines the optimal compression rate by analyzing the model's own importance value distribution without requiring manual intervention. The method self-identifies the turning point where pruning begins to significantly impact accuracy, enabling autonomous optimization that eliminates the need for manual pruning rate setting and repeated training adjustments.
2Quantity of substance
If higher compression rates are applied to reduce model size, then resource consumption is reduced, but model accuracy may be compromised
Solution Approach 1:
The patent replaces manual trial-and-error adjustment of pruning rates with an automated analytical system based on importance value distribution. By substituting the mechanical process of repeated training and accuracy testing with an algorithmic analysis of parameter importance, the system objectively determines the optimal compression rate that maintains accuracy while maximizing size reduction.
Solution Approach 2:
The method analyzes changes in parameter importance values to identify the optimal compression point. By examining the distribution of importance values and detecting the turning point where parameters transition from near-zero to significant importance, the system dynamically determines the compression rate that preserves model accuracy while achieving maximum compression.
3Reliability
If repeated training and adjustment processes are performed to optimize compression, then model performance is maintained, but overall resource consumption increases
Solution Approach 1:
The patent performs preliminary analysis of importance value distribution before committing to a compression rate. By pre-identifying parameters with near-zero importance and determining the turning point in advance, the system avoids multiple rounds of training and adjustment, significantly reducing computational resource consumption while maintaining model performance.
4Adaptability or versatility
If manual determination of compression rate is used, then flexibility in optimization is maintained, but processing time and complexity increase
Solution Approach 1:
The system automatically determines the optimal compression rate by analyzing its own importance value distribution without requiring manual input or intervention. The method self-identifies the turning point and determines the optimal compression rate algorithmically, maintaining adaptability to different models while eliminating time-consuming manual determination processes.
Data Source
AI summary
A method for determining a model compression rate comprises determining a near-zero importance value subset from an importance value set associated with a machine learning model, a corresponding importance value in the importance value set indicating an importance degree of a corresponding input of a processing layer of the machine learning model, importance values in the near-zero importance value subset being closer to zero than other importance values in the importance value set; determining a target importance value from the near-zero importance value subset, the target importance value corresponding to a turning point of a magnitude of the importance values in the near-zero importance value subset; determining a proportion of importance values less than the target importance value in the importance value set in the importance value set; and determining the compression rate for the machine learning model based on the determined proportion.


