Mixed-Precision Neural Network Computing System

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computing systems for image rendering using neural networks face challenges in handling computation complexity and memory usage, which limits their ability to perform image rendering within a timely manner due to high resource requirements.

Innovation Solution

A mixed-precision computing system is implemented, comprising high-precision and low-precision computation units that process neural network layers, with a controller managing data transmission between them, allowing for efficient computation by adjusting bit widths based on precision needs, thereby reducing complexity and memory usage.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If high-precision computation units are used to process neural network layers, then image rendering quality is improved, but computation complexity and memory usage increase

Engineering Contradiction:
Improveimage rendering qualityVSAvoidcomputation complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent applies different precision levels (first bit width for high precision, second bit width for low precision) to different computation units based on the specific requirements of neural network layers. This allows high-precision processing only where necessary while using low-precision processing elsewhere, thereby maintaining image rendering quality while reducing overall computation complexity and memory usage.

Inventive Principle:
Principle #3Local quality

2Measurement precision

If high-precision computation units are used to process neural network layers, then image rendering quality is improved, but memory usage increases

Engineering Contradiction:
Improveimage rendering qualityVSAvoidmemory usage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent implements a mixed-precision architecture where different computation units use different bit widths (first bit width for high precision, second bit width for low precision) based on the specific needs of each neural network layer. This localized precision approach reduces the overall memory footprint compared to using high precision throughout, while still maintaining the necessary image rendering quality.

Inventive Principle:
Principle #3Local quality

3Device complexity

If low-precision computation units are used to process neural network layers, then computation complexity is reduced, but image rendering quality deteriorates

Engineering Contradiction:
Improvecomputation complexityVSAvoidimage rendering quality
Core Design Contradiction:
Device complexityVSMeasurement precision

Solution Approach 1:

The patent strategically assigns different precision levels to different computation units based on the specific requirements of each neural network layer. Critical layers that require high image rendering quality are processed by high-precision computation units, while less critical layers use low-precision units to reduce computation complexity. This selective approach resolves the contradiction by applying precision locally where needed rather than uniformly across the entire system.

Inventive Principle:
Principle #3Local quality

4Measurement precision

If uniform high-precision processing is applied to all neural network layers, then image rendering quality is maintained, but processing speed decreases

Engineering Contradiction:
Improveimage rendering qualityVSAvoidprocessing speed
Core Design Contradiction:
Measurement precisionVSProductivity

Solution Approach 1:

The patent implements a mixed-precision neural network architecture where different computation units process different layers at different precision levels (first bit width vs. second bit width). This allows the system to maintain high image rendering quality for critical layers while accelerating processing for less critical layers through low-precision computation, thereby improving overall processing speed without sacrificing necessary quality.

Inventive Principle:
Principle #3Local quality

5Productivity

If uniform low-precision processing is applied to all neural network layers, then processing speed is improved, but image rendering quality deteriorates

Engineering Contradiction:
Improveprocessing speedVSAvoidimage rendering quality
Core Design Contradiction:
ProductivityVSMeasurement precision

Solution Approach 1:

The patent creates a heterogeneous computation system where different computation units use different bit widths (first bit width for high precision, second bit width for low precision) based on the specific requirements of each neural network layer. This allows the system to achieve high processing speeds through low-precision processing for appropriate layers while maintaining high image rendering quality for critical layers through high-precision processing, thus resolving the speed-quality tradeoff.

Inventive Principle:
Principle #3Local quality

Data Source

PatentUS20240296308A1Mixed-precision Neural Network Systems
Publication Date: 2024.09.05 SHANGHAI TECH UNIV
  • US20240296308A1 patent drawing
  • US20240296308A1 patent drawing
  • US20240296308A1 patent drawing

AI summary

A computing system for encoding a machine learning model comprises a plurality of layers and a plurality of computation units. A first set of computation units are configured to process data at a first bit width. A second set of computation units are configured to process at a second bit width. The first bit width is higher than the second bit width. A memory is coupled to the computation units. A controller is coupled to the computation units and the memory. The controller is configured to provide instructions for encoding the machine learning model. The first set of computation units are configured to compute a first set of layers and the second set of computation units are configured to compute a second set of layers.