Neural Network Compression via Reinforcement Learning

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current techniques for compressing neural networks often compromise accuracy for improved performance, lacking efficient methods to achieve both desired accuracy and performance simultaneously.

Innovation Solution

A system utilizing a reinforcement learning model processes performance and accuracy policies to determine a combination strategy for compressing neural networks, employing techniques like pruning, quantization, and sparsity to achieve target metrics on various processing units.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If neural networks are compressed using traditional techniques, then performance is improved, but accuracy deteriorates

Engineering Contradiction:
ImproveperformanceVSAvoidaccuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies parameter changes by systematically varying compression parameters (pruning ratio, quantization bits, sparsity level) to find optimal configurations that simultaneously satisfy both performance and accuracy requirements. The reinforcement learning model learns the relationship between compression parameters and model quality metrics, enabling parameter optimization that resolves the contradiction between compression performance and model accuracy.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements dynamics by using a reinforcement learning model that adaptively determines compression strategies based on target performance and accuracy metrics. Rather than applying static compression rules, the system dynamically adjusts compression configurations through iterative learning, allowing the model to navigate the trade-off between performance improvement and accuracy preservation in a flexible, goal-directed manner.

Inventive Principle:
Principle #15Dynamics

2Quantity of substance

If compression techniques are applied to neural networks, then resource usage is reduced, but model quality deteriorates

Engineering Contradiction:
Improvemodel sizeVSAvoidmodel quality
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent uses parameter changes to control model size while preserving quality by optimizing compression parameters such as pruning ratios and quantization precision. The reinforcement learning model learns to select parameter configurations that achieve target model sizes without excessive quality loss, resolving the contradiction between reducing model quantity and maintaining reliability.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent implements feedback mechanisms where the reinforcement learning model receives feedback on model quality metrics and resource usage, then adjusts compression strategies accordingly. This closed-loop control enables the system to iterate toward solutions that simultaneously achieve desired model size reduction and quality preservation, rather than accepting quality deterioration as an inevitable trade-off.

Inventive Principle:
Principle #23Feedback

3Productivity

If aggressive compression is applied, then deployment efficiency is improved, but accuracy is compromised

Engineering Contradiction:
Improvedeployment efficiencyVSAvoidaccuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent applies dynamics by using reinforcement learning to adaptively determine the optimal level of compression aggression based on target accuracy and performance requirements. The system dynamically adjusts compression intensity, transitioning from aggressive to conservative strategies as needed, rather than applying fixed aggressive compression rules that inevitably compromise accuracy.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent uses parameter changes to control compression aggression by optimizing parameters such as pruning ratio, quantization precision, and sparsity level. The reinforcement learning model learns to select parameter configurations that achieve deployment efficiency improvements without excessive accuracy loss, resolving the contradiction between aggressive compression benefits and accuracy preservation.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20240160905A1Techniques for compressing neural networks
Publication Date: 2024.05.16 NVIDIA CORP
  • US20240160905A1 patent drawing
  • US20240160905A1 patent drawing
  • US20240160905A1 patent drawing

AI summary

Apparatuses, systems, and techniques to compress neural networks. In at least one embodiment, one or more first neural networks are used to cause one or more compressed neural networks to be selected based, at least in part, on accuracy and performance of the one or more compressed neural networks.