Machine Learning Model Compression for Mobile Deployment

Overview of Technical Issues:

The compression processing module faces conflicting functional requirements where insufficient compression fails to reduce the trained model structure adequately, causing memory overflow and deployment failure on mobile devices, while excessive compression produces a harmful effect by degrading inference accuracy below acceptable performance thresholds; the goal is to achieve successful mobile deployment while maintaining model accuracy within acceptable bounds.

Solution directions generated for this problem

Problem Direction 1 :

ImproveCompression ratio intensity
VS
ConstraintInference accuracy retention

Inspiration 1 : Cross-domain reference

Application Principle: #26 Copying
Cross-domain applicability Assess applicability
Moving picture decoding device, moving picture decoding method, and moving picture decoding program
Innovative Solution Refine solution

Multi-tier model variant deployment with device-adaptive selection

Create multiple compressed model variants for deployment
How to solve :
  • Generate three model variants from the trained model: Tier-1 (50-80MB, 4-bit quantization, 60% pruning), Tier-2 (100-130MB, 8-bit quantization, 40% pruning), Tier-3 (180-200MB, mixed 8/16-bit, 20% pruning)
  • each variant targets different device memory capacities
  • Implement device capability detection module (5KB overhead) that measures available RAM, CPU cores, and GPU availability at app launch, then selects the highest-tier model the device can support — detection completes within 200ms using standard Android/iOS APIs
  • Deploy all three variants in a single installation package using differential compression (shared base weights + tier-specific deltas), reducing total package size to 220-250MB versus 430MB for separate models
  • only the selected variant loads into runtime memory
Expected Effect : Accuracy degradation ≤1.5% on 95% devices, memory fit rate 100%, package size reduction 45%
Risk Control :
  • variant synchronization during model updates
  • detection algorithm misjudging device capability
  • delta reconstruction introducing numerical errors

Problem Direction 2 :

ImproveModel information preservation capability
VS
ConstraintMust not deteriorate

Inspiration 1 : Cross-domain reference

Application Principle: #10 Preliminary action
Cross-domain applicability Assess applicability
Massively parallel single cell analysis
Innovative Solution Refine solution

Pre-compression knowledge extraction and runtime compensation for mobile model deployment

Extract critical model knowledge before compression then restore during inference
How to solve :
  • Before compression, extract and store decision boundary maps, attention weight matrices, and critical feature activations from the full-precision model as lightweight reference files (5-15MB)
  • compress these references using lossless encoding and embed as lookup tables in the deployment package
  • Apply aggressive 4-bit quantization and 60% structured pruning to reduce the main model to 50-120MB, accepting temporary accuracy loss during this compression phase
  • During mobile inference, the compressed model generates predictions with confidence scoring
  • when confidence falls below 0.85 threshold, retrieve corresponding pre-extracted knowledge references to recalibrate layer outputs, restoring accuracy to within 1.5% of the original model
Expected Effect : Model size 50-120MB, accuracy degradation ≤1.5%, inference latency +8-12ms for low-confidence cases only
Risk Control :
  • knowledge extraction completeness insufficient
  • confidence threshold calibration inaccurate
  • lookup table retrieval overhead excessive
Patsnap Eureka Solution