Neural Network Compression Profile Search for Resource-Limited Deployment
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
The increasing scale and complexity of neural networks require more computational resources, which can limit their application in addressing new problems or providing solutions in different ways.
Innovation Solution
The development of techniques to search for and apply compression profiles and policies to trained neural networks, allowing for the reduction of network size while maintaining accuracy, thereby enabling their implementation across various systems with different resource limitations.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If neural networks are scaled up to address more complex problems, then problem-solving capability is improved, but computational resource requirements increase
Solution Approach 1:
The patent extracts and removes unnecessary or redundant neurons, channels, and filters from the neural network through automated pruning techniques. This extraction process reduces the network size and computational resource requirements while preserving the essential functionality needed to solve complex problems, directly addressing the contradiction between capability and resource consumption.
Solution Approach 2:
The patent changes key parameters of the neural network including reducing the number of neurons per layer, adjusting filter sizes, modifying activation functions, and changing network depth. These parameter modifications allow the network to maintain problem-solving capability while significantly reducing computational resource requirements through systematic optimization.
2Measurement precision
If neural network size is increased to improve accuracy, then model performance is improved, but memory and power requirements increase
Solution Approach 1:
The patent applies automated pruning to extract and remove redundant neurons, channels, and filters that do not contribute significantly to model accuracy. This extraction reduces memory requirements by eliminating unnecessary model parameters while preserving the structural elements essential for maintaining high accuracy performance.
Solution Approach 2:
The patent performs preliminary training of the neural network to full size before applying compression techniques. This preliminary action ensures the network achieves optimal accuracy with the complete architecture, allowing subsequent pruning to remove only truly redundant elements while preserving accuracy-critical components.
3Adaptability or versatility
If neural network complexity is increased to address new problems, then functionality is improved, but ease of deployment is worsened
Solution Approach 1:
The patent implements self-service through automated compression profiles that can be applied to neural networks without manual intervention. The system automatically analyzes the network architecture, identifies compression opportunities, applies appropriate pruning techniques, and validates performance, making deployment of complex networks straightforward even for users without deep expertise in neural network optimization.
Solution Approach 2:
The patent creates universal compression profiles that can be applied across different neural network architectures and problem domains. These multi-functional profiles enable the same compression techniques to be used for various network types and applications, simplifying deployment processes while maintaining adaptability to different functionality requirements.
Data Source
AI summary
Compression profiles may be searched for trained neural networks. An iterative compression profile search may be performed response to a search request. Different prospective compression profiles may be generated for trained neural networks according to a search policy. Performance of compressed versions of the trained neural networks according to the compression profiles may be tracked. The search policy may be updated according to an evaluation of the performance of the compression profiles for the compressed versions of the trained neural networks using compression performance criteria. When a search criteria is satisfied, a result for the compression profile search may be provided.


