Portion-Specific Model Compression for ML Accuracy

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional compression schemes for machine-learned models often result in a substantial reduction of their accuracy, limiting their utility in resource-constrained devices and requiring excessive computing resources.

Innovation Solution

A method for portion-specific compression and optimization of machine-learned models, where a computing system evaluates cost functions to select candidate compression schemes for specific model portions, applying them to retain accuracy while meeting latency and memory constraints, enabling distillation training for optimized performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If conventional compression schemes are applied to machine-learned models, then model size is reduced, but accuracy is substantially reduced

Engineering Contradiction:
Improvemodel sizeVSAvoidmodel accuracy
Core Design Contradiction:
Quantity of substanceVSManufacturing precision

Solution Approach 1:

The patent divides the machine-learned model into multiple portions or layers, allowing different compression schemes to be applied to different segments. This segmentation enables selective compression where less critical portions are compressed more aggressively while critical portions maintain higher precision, thus reducing overall model size without substantially reducing accuracy.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies different compression schemes to different portions of the model based on their specific characteristics and importance. Each portion receives a customized compression treatment tailored to its local requirements, ensuring that critical portions maintain high accuracy while less critical portions contribute more to size reduction.

Inventive Principle:
Principle #3Local quality

2Quantity of substance

If model size is reduced through compression, then storage requirements are reduced, but computational performance deteriorates

Engineering Contradiction:
Improvestorage requirementsVSAvoidcomputational performance
Core Design Contradiction:
Quantity of substanceVSProductivity

Solution Approach 1:

The patent introduces dynamic elements into the compression approach, allowing the system to adapt compression levels based on performance requirements. The compression scheme can be adjusted dynamically to balance between storage efficiency and computational performance depending on the specific use case and resource constraints.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent changes key parameters of the compression scheme, such as precision levels and compression ratios, to optimize the balance between model size and computational performance. By carefully selecting and adjusting these parameters, the system achieves significant size reduction while maintaining acceptable performance levels.

Inventive Principle:
Principle #35Parameter changes

3Device complexity

If uniform compression is applied to all model portions, then implementation is simplified, but accuracy loss increases

Engineering Contradiction:
Improvecompression implementation complexityVSAvoidmodel accuracy
Core Design Contradiction:
Device complexityVSManufacturing precision

Solution Approach 1:

The patent segments the model into portions that can be processed independently with different compression schemes. This segmentation maintains implementation simplicity by allowing each portion to be compressed separately using straightforward algorithms, while the variety of schemes applied to different segments prevents uniform accuracy loss.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies compression selectively to only certain portions of the model rather than uniformly to all portions. This partial action approach simplifies implementation by focusing compression efforts where they are most beneficial while leaving critical portions uncompressed or lightly compressed, thereby maintaining accuracy.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS20240232686A1Portion-Specific Model Compression for Optimization of Machine-Learned Models
Publication Date: 2024.07.11 GOOGLE LLC
  • US20240232686A1 patent drawing
  • US20240232686A1 patent drawing
  • US20240232686A1 patent drawing

AI summary

Systems and methods of the present disclosure are directed to portion-specific compression and optimization of machine-learned models. For example, a method for portion-specific compression and optimization of machine-learned models includes obtaining data descriptive of one or more respective sets of compression schemes for one or more model portions of a plurality of model portions of a machine-learned model. The method includes evaluating a cost function to respectively select one or more candidate compression schemes from the one or more sets of compression schemes. The method includes respectively applying the one or more candidate compression schemes to the one or more model portions to obtain a compressed machine-learned model comprising one or more compressed model portions that correspond to the one or more model portions.