Neural Network Quantization Device with Optimized Division Position

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

The accuracy of neural network calculations decreases when using quantized variables, leading to reduced recognition performance, as the transition from floating-point to fixed-point numbers compromises both computation efficiency and accuracy.

Innovation Solution

An information processing device optimizes division positions for quantization to minimize quantization errors, allowing for the conversion of floating-point data into fixed-point data while maintaining calculation accuracy, by setting thresholds that reduce differences between pre- and post-quantization values, thereby improving recognition rates and computation efficiency.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If quantization is applied to convert floating-point data to fixed-point data for neural network calculations, then computation efficiency is improved, but calculation accuracy decreases

Engineering Contradiction:
Improvecomputation efficiencyVSAvoidcalculation accuracy
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent changes the parameter of division positions in quantization from fixed to dynamically optimized values. By adjusting the division positions based on the actual data distribution and neural network requirements, the quantization process achieves better precision while maintaining computational efficiency. This parameter optimization allows the system to find the best trade-off point between speed and accuracy for each specific calculation scenario.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces dynamic adjustment of quantization division positions during the neural network training and inference process. Instead of using static quantization parameters, the system adaptively modifies division positions based on data characteristics and performance requirements, enabling the quantization scheme to evolve and optimize itself for different computational contexts and accuracy demands.

Inventive Principle:
Principle #15Dynamics

2Use of energy by moving object

If quantization is applied to reduce computational resources and power consumption, then energy efficiency is improved, but recognition performance decreases

Engineering Contradiction:
Improvepower consumptionVSAvoidrecognition performance
Core Design Contradiction:
Use of energy by moving objectVSReliability

Solution Approach 1:

The patent optimizes quantization parameters including division positions and precision levels based on the specific requirements of different neural network layers and tasks. By dynamically adjusting these parameters, the system achieves significant energy savings while maintaining recognition performance above acceptable thresholds, effectively decoupling energy consumption from performance degradation.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent applies different quantization strategies and precision levels to different parts of the neural network based on their importance and sensitivity. Critical layers maintain higher precision with finer division positions, while less critical layers use coarser quantization, creating a localized quality distribution that optimizes the overall energy-performance trade-off across the entire system.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS11675567B2Quantization device, quantization method, and recording medium
Publication Date: 2023.06.13 FUJITSU LTD
  • US11675567B2 patent drawing
  • US11675567B2 patent drawing
  • US11675567B2 patent drawing

AI summary

An information processing device that executes calculation of a neural network, includes a memory; and a processor coupled to the memory and the processor configured to: set a division position for quantization of a variable to be used for the calculation so that a quantization error based on a difference between the variable before the quantization and the variable after the quantization is reduced; and quantize the variable based on the division position set.