Iterative Channel Pruning and Knowledge Distillation for Neural Network Compression

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing model compression techniques, such as knowledge distillation and channel pruning alone, result in compressed models with worse performance due to inefficient resource utilization and convergence issues.

Innovation Solution

A method combining iterative channel pruning and knowledge distillation, where channel pruning reduces model scale and knowledge distillation adjusts weight coefficients to improve training results and convergence, achieving step-by-step compression and better performance.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Volume of stationary object

If channel pruning is applied to reduce model scale, then model size is reduced, but model performance deteriorates

Engineering Contradiction:
Improvemodel sizeVSAvoidmodel performance
Core Design Contradiction:
Volume of stationary objectVSReliability

Solution Approach 1:

The model compression process is divided into multiple iterative stages, where each stage performs channel pruning followed by knowledge distillation. This segmentation allows gradual optimization rather than single-step aggressive pruning, maintaining performance while reducing size.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The teacher model serves as an intermediary that transfers knowledge to the student model during distillation. This intermediary mechanism allows the pruned student model to learn from the full teacher model, compensating for performance loss while maintaining the reduced size.

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If knowledge distillation is used to train compressed model, then model performance is improved, but training time increases

Engineering Contradiction:
Improvemodel performanceVSAvoidtraining time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The training process is segmented into multiple iterative rounds, each with controlled pruning and distillation steps. This segmentation prevents any single training phase from becoming excessively time-consuming while achieving cumulative performance improvement.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of performing full knowledge distillation from scratch, the method performs partial distillation in each iterative round, focusing on the specific changes made in that round. This partial action reduces overall training time while still achieving performance gains.

Inventive Principle:
Principle #16Partial or excessive action

3Productivity

If iterative pruning and distillation are performed, then model compression ratio is improved, but computational complexity increases

Engineering Contradiction:
Improvecompression ratioVSAvoidcomputational complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The complex compression task is segmented into repeated simple cycles of pruning and distillation. Each cycle is computationally manageable, but their iteration achieves high overall compression ratios that would be difficult to achieve in a single complex operation.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The method employs periodic alternating actions of pruning and distillation. This periodic pattern creates a rhythm of compression and recovery that manages computational complexity by alternating between aggressive pruning phases and gentler distillation phases.

Inventive Principle:
Principle #19Periodic action

Data Source

PatentUS20240362486A1Model training method and apparatus, and readable storage medium
Publication Date: 2024.10.31 BEIJING ZITIAO NETWORK TECH CO LTD
  • US20240362486A1 patent drawing
  • US20240362486A1 patent drawing
  • US20240362486A1 patent drawing

AI summary

A model training method includes: acquiring a sample data set corresponding to a target task, a teacher model and an ith initial student model; performing an ith time of channel pruning on the ith initial student model, to acquire a student model subjected to the ith time of channel pruning; performing knowledge distillation according to the sample data set, the teacher model and the student model subjected to the ith time of channel pruning, to acquire an (i+1)th initial student model, wherein a compression ratio of the (i+1)th initial student model to the ith initial student model is equal to a preset ith compression ratio; and updating i to be i+1, and returning to the step of performing the ith time of channel pruning on the ith initial student model, until the updated i is greater than a threshold value N, to acquire a target student model.