Machine Learning Model Compression for Mobile Deployment
Overview of Technical Issues:
The compression processing module faces conflicting functional requirements where insufficient compression fails to reduce the trained model structure adequately, causing memory overflow and deployment failure on mobile devices, while excessive compression produces a harmful effect by degrading inference accuracy below acceptable performance thresholds; the goal is to achieve successful mobile deployment while maintaining model accuracy within acceptable bounds.
Solution directions generated for this problem
Problem Direction 1 :
ImproveCompression ratio intensity
VSConstraintInference accuracy retention
Inspiration 1 : Cross-domain reference
Application Principle: #26 Copying
Cross-domain applicability
Moving picture decoding device, moving picture decoding method, and moving picture decoding program
Innovative Solution Refine solution
Multi-tier model variant deployment with device-adaptive selection
Create multiple compressed model variants for deployment
How to solve :
- Generate three model variants from the trained model: Tier-1 (50-80MB, 4-bit quantization, 60% pruning), Tier-2 (100-130MB, 8-bit quantization, 40% pruning), Tier-3 (180-200MB, mixed 8/16-bit, 20% pruning)
- each variant targets different device memory capacities
- Implement device capability detection module (5KB overhead) that measures available RAM, CPU cores, and GPU availability at app launch, then selects the highest-tier model the device can support — detection completes within 200ms using standard Android/iOS APIs
- Deploy all three variants in a single installation package using differential compression (shared base weights + tier-specific deltas), reducing total package size to 220-250MB versus 430MB for separate models
- only the selected variant loads into runtime memory
Expected Effect : Accuracy degradation ≤1.5% on 95% devices, memory fit rate 100%, package size reduction 45%
Risk Control :
- variant synchronization during model updates
- detection algorithm misjudging device capability
- delta reconstruction introducing numerical errors
Problem Direction 2 :
ImproveModel information preservation capability
VSConstraintMust not deteriorate
Inspiration 1 : Cross-domain reference
Application Principle: #10 Preliminary action
Cross-domain applicability
Massively parallel single cell analysis
Innovative Solution Refine solution
Pre-compression knowledge extraction and runtime compensation for mobile model deployment
Extract critical model knowledge before compression then restore during inference
How to solve :
- Before compression, extract and store decision boundary maps, attention weight matrices, and critical feature activations from the full-precision model as lightweight reference files (5-15MB)
- compress these references using lossless encoding and embed as lookup tables in the deployment package
- Apply aggressive 4-bit quantization and 60% structured pruning to reduce the main model to 50-120MB, accepting temporary accuracy loss during this compression phase
- During mobile inference, the compressed model generates predictions with confidence scoring
- when confidence falls below 0.85 threshold, retrieve corresponding pre-extracted knowledge references to recalibrate layer outputs, restoring accuracy to within 1.5% of the original model
Expected Effect : Model size 50-120MB, accuracy degradation ≤1.5%, inference latency +8-12ms for low-confidence cases only
Risk Control :
- knowledge extraction completeness insufficient
- confidence threshold calibration inaccurate
- lookup table retrieval overhead excessive
