Neural Network Compression Profile Search for Resource-Limited Deployment

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The increasing scale and complexity of neural networks require more computational resources, which can limit their application in addressing new problems or providing solutions in different ways.

Innovation Solution

The development of techniques to search for and apply compression profiles and policies to trained neural networks, allowing for the reduction of network size while maintaining accuracy, thereby enabling their implementation across various systems with different resource limitations.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If neural networks are scaled up to address more complex problems, then problem-solving capability is improved, but computational resource requirements increase

Engineering Contradiction:
Improveproblem-solving capabilityVSAvoidcomputational resource requirements
Core Design Contradiction:
Adaptability or versatilityVSUse of energy by moving object

Solution Approach 1:

The patent extracts and removes unnecessary or redundant neurons, channels, and filters from the neural network through automated pruning techniques. This extraction process reduces the network size and computational resource requirements while preserving the essential functionality needed to solve complex problems, directly addressing the contradiction between capability and resource consumption.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent changes key parameters of the neural network including reducing the number of neurons per layer, adjusting filter sizes, modifying activation functions, and changing network depth. These parameter modifications allow the network to maintain problem-solving capability while significantly reducing computational resource requirements through systematic optimization.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If neural network size is increased to improve accuracy, then model performance is improved, but memory and power requirements increase

Engineering Contradiction:
Improvemodel accuracyVSAvoidmemory requirements
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent applies automated pruning to extract and remove redundant neurons, channels, and filters that do not contribute significantly to model accuracy. This extraction reduces memory requirements by eliminating unnecessary model parameters while preserving the structural elements essential for maintaining high accuracy performance.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent performs preliminary training of the neural network to full size before applying compression techniques. This preliminary action ensures the network achieves optimal accuracy with the complete architecture, allowing subsequent pruning to remove only truly redundant elements while preserving accuracy-critical components.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If neural network complexity is increased to address new problems, then functionality is improved, but ease of deployment is worsened

Engineering Contradiction:
ImprovefunctionalityVSAvoidease of deployment
Core Design Contradiction:
Adaptability or versatilityVSEase of operation

Solution Approach 1:

The patent implements self-service through automated compression profiles that can be applied to neural networks without manual intervention. The system automatically analyzes the network architecture, identifies compression opportunities, applies appropriate pruning techniques, and validates performance, making deployment of complex networks straightforward even for users without deep expertise in neural network optimization.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent creates universal compression profiles that can be applied across different neural network architectures and problem domains. These multi-functional profiles enable the same compression techniques to be used for various network types and applications, simplifying deployment processes while maintaining adaptability to different functionality requirements.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12314277B2Searching compression profiles for trained neural networks
Publication Date: 2025.05.27 AMAZON TECH INC
  • US12314277B2 patent drawing
  • US12314277B2 patent drawing
  • US12314277B2 patent drawing

AI summary

Compression profiles may be searched for trained neural networks. An iterative compression profile search may be performed response to a search request. Different prospective compression profiles may be generated for trained neural networks according to a search policy. Performance of compressed versions of the trained neural networks according to the compression profiles may be tracked. The search policy may be updated according to an evaluation of the performance of the compression profiles for the compressed versions of the trained neural networks using compression performance criteria. When a search criteria is satisfied, a result for the compression profile search may be provided.