Neural Network Inference Compression for Edge AI Speed-Accuracy Balance

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The complexity and size of deep neural network models pose challenges for real-time applications on edge nodes, requiring efficient compression and acceleration methods to meet intelligent system demands.

Innovation Solution

A neural network inference acceleration method involving model compression, graph optimization, and deployment optimization, utilizing techniques like model quantification, pruning, and distillation, along with platform-specific optimization strategies and tools like MNN, OpenVINO, TensorRT, and TVM, to enhance inference efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If deep neural networks are used in edge nodes, then intelligence and real-time performance are improved, but model complexity and data requirements increase

Engineering Contradiction:
Improveinference speedVSAvoidmodel complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent extracts and removes redundant or less important features from the neural network model through feature selection and dimensionality reduction techniques. This extraction process eliminates unnecessary computational complexity while preserving the essential intelligence and real-time performance capabilities needed for edge node deployment.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different optimization strategies to different parts of the neural network model. By identifying and optimizing specific layers or components locally rather than uniformly across the entire model, the system achieves better inference speed and reduced complexity while maintaining overall model effectiveness for edge computing applications.

Inventive Principle:
Principle #3Local quality

2Speed

If model compression is applied, then inference speed is improved, but model accuracy may deteriorate

Engineering Contradiction:
Improveinference speedVSAvoidmodel accuracy
Core Design Contradiction:
SpeedVSMeasurement precision

Solution Approach 1:

The patent implements feedback mechanisms during the model compression process where the system continuously monitors accuracy metrics and adjusts compression parameters accordingly. This feedback loop allows the model to maintain acceptable accuracy levels while achieving the desired inference speed improvements for real-time edge node operations.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The patent systematically changes key model parameters such as quantization precision, pruning ratios, and compression factors to find the optimal balance between inference speed and accuracy. By carefully adjusting these parameters, the system achieves faster inference while minimizing accuracy deterioration through controlled parameter transformations.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS12633094B2Neural network inference acceleration method, target detection method, device, and storage medium
Publication Date: 2026.05.19 BOE TECHNOLOGY GROUP CO LTD
  • US12633094B2 patent drawing
  • US12633094B2 patent drawing
  • US12633094B2 patent drawing

AI summary

A neural network inference acceleration method includes: acquiring a neural network model to be accelerated and an accelerated data set; automatically performing accelerating process on the neural network model to be accelerated by using the accelerated data set to obtain the accelerated neural network model, wherein the accelerating process includes at least one of the following: model compression, graph optimization and deployment optimization, wherein the model compression includes at least one of the following: model quantification, model pruning and model distillation, wherein the graph optimization is the optimization for the directed graph of the neural network model to be accelerated, and the deployment optimization is the optimization for the deployment platform of the neural network model to be accelerated; and performing inference evaluation on the accelerated neural network model.