Heterogeneous Deep Learning Accelerator with Mixed-Precision Compute Tiles

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Small-scale users, such as businesses and individuals, face challenges in affording dedicated artificial intelligence hardware accelerators and are hesitant to use cloud-based services due to concerns over data privacy, security, and configurability, while existing hardware accelerators often waste computing power on transfer learning tasks that do not require high precision.

Innovation Solution

The development of heterogeneous hardware acceleration devices and systems that include both low-precision and high-precision components, allowing low-precision components to handle inference tasks and high-precision components to handle training tasks, with the ability to convert data formats using de-quantizers, enabling efficient performance of transfer learning tasks.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Power

If homogeneous high-precision hardware accelerators are used, then computing power and flexibility are improved, but cost increases significantly

Engineering Contradiction:
Improvecomputing powerVSAvoidcost
Core Design Contradiction:
PowerVSEase of manufacture

Solution Approach 1:

The patent applies local quality by creating heterogeneous hardware accelerators with different precision components in specific locations. Low-precision components handle inference tasks while high-precision components handle training tasks, placing each type of component where it is most needed for the specific workload, rather than using uniform high-precision components throughout the entire system.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent segments the hardware accelerator into distinct low-precision and high-precision components. This segmentation allows the system to divide computational workloads accordingly, with inference operations assigned to low-precision components and training operations assigned to high-precision components, optimizing both cost and performance.

Inventive Principle:
Principle #1Segmentation

2Power

If cloud-based hardware accelerator services are used, then access to high-performance computing is improved, but data privacy and security concerns worsen

Engineering Contradiction:
Improvecomputing powerVSAvoiddata privacy risk
Core Design Contradiction:
PowerVSObject-affected harmful factors

Solution Approach 1:

The patent enables self-service by providing organizations with dedicated heterogeneous hardware accelerators that can be configured and operated locally. This allows organizations to maintain control over their data and computing resources while still accessing the benefits of specialized AI hardware, eliminating the need to rely on external cloud services for sensitive workloads.

Inventive Principle:
Principle #25Self-service

3Measurement precision

If high-precision hardware accelerators are used for transfer learning, then model accuracy is improved, but computing resource waste increases

Engineering Contradiction:
Improvemodel accuracyVSAvoidcomputing resource waste
Core Design Contradiction:
Measurement precisionVSLoss of energy

Solution Approach 1:

The patent applies local quality by matching precision levels to task requirements. Low-precision components are used for inference tasks where high accuracy is not critical, while high-precision components are reserved for training tasks where model accuracy is paramount. This localized assignment of precision levels eliminates the waste of using high-precision resources for tasks that do not require them.

Inventive Principle:
Principle #3Local quality

Solution Approach 2:

The patent applies partial action by using only the necessary precision level for each specific task. Instead of always using high-precision components for all operations, the system uses low-precision components for inference where full precision is not needed, and only activates high-precision components when actually required for training operations.

Inventive Principle:
Principle #16Partial or excessive action

4Power

If dedicated AI hardware accelerators are purchased, then computing performance is improved, but configurability and flexibility worsen

Engineering Contradiction:
Improvecomputing performanceVSAvoidconfigurability
Core Design Contradiction:
PowerVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing heterogeneous hardware accelerators that can perform multiple functions through different precision components. The same hardware platform can be configured to handle inference tasks using low-precision components, training tasks using high-precision components, or a combination of both, providing versatility without requiring multiple dedicated systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12067479B2Heterogeneous deep learning accelerator
Publication Date: 2024.08.20 T-HEAD (SHANGHAI) SEMICON CO LTD
  • US12067479B2 patent drawing
  • US12067479B2 patent drawing
  • US12067479B2 patent drawing

AI summary

Systems and methods for heterogenous hardware acceleration are disclosed. The systems and methods can include a neural network processing unit comprising compute tiles. Each of a first set of the compute tiles can include a first tensor array configured to support operations in a first number format. Each of a second set of the compute tiles can include a second tensor array configured to support operations in a second number format, the second number format supporting a greater range or a greater precision than the first number format, and a de-quantizer configured to convert data in the first number format to data in the second number format. The systems and methods can include neural network processing units, multi-chip hardware accelerators and distributed hardware accelerators including low-precision components for performing interference tasks and high-precision components for performing training tasks. Transfer learning tasks can be performed using low-precision components and high-precision components.