AI Integrated Circuit with Dynamic Data Prefetching and Compression
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current deep-learning accelerators, such as GPUs, face inefficiencies in performance/power ratio and lack specialized features for AI operations, with long adaptation periods for new algorithms and inadequate buffer pre-load techniques and data compression methods, making them less suitable for both training and inference phases.
Innovation Solution
An AI integrated circuit with a command processor, processing elements, task constructor, L1 and L2 caches, and deep-learning accelerators that perform arithmetic operations, hardware multiplication-addition, activation functions, and pooling, along with dynamic data prefetching and compression/decompression mechanisms to optimize matrix operations and reduce bandwidth requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If GPU is used for deep learning acceleration, then programmability and software environment are improved, but performance/power ratio deteriorates due to redundant graphics processing functions
Solution Approach 1:
The patent extracts and removes redundant graphics processing functions (task management, storage buffers, rasterization, rendering output units) from the GPU architecture, retaining only the essential parallel computing components needed for deep learning operations. This extraction eliminates approximately one-third of the GPU area that was dedicated to graphics-specific functions, thereby reducing power consumption while preserving programmability and software environment capabilities.
Solution Approach 2:
The patent applies local quality by modifying specific regions of the GPU architecture to have AI-optimized properties. The processing elements are reconfigured with AI-specific features such as specialized activation function units, pooling units, and data compression mechanisms in particular areas, while other areas maintain general-purpose computing capabilities. This allows the system to achieve better performance/power ratio for AI workloads without completely sacrificing versatility.
2Productivity
If ASIC/FPGA is used for deep learning acceleration, then calculation efficiency is improved, but adaptability to new algorithms deteriorates due to fixed architecture
Solution Approach 1:
The patent introduces dynamic configurability to the ASIC/FPGA architecture through programmable processing elements that can be reconfigured for different deep learning operations. The processing elements include configurable activation function units and pooling units that can adapt to various algorithms. Additionally, the architecture supports dynamic data compression and pre-loading strategies that can be adjusted based on the specific algorithm being executed, enabling high calculation efficiency while maintaining adaptability to new algorithms.
3Adaptability or versatility
If GPU is used for deep learning, then software ecosystem is improved, but performance/power ratio deteriorates due to lack of dedicated AI features
Solution Approach 1:
The patent implements preliminary action through data pre-loading and pre-compression mechanisms. The system pre-loads data into on-chip buffers before processing and applies compression techniques to reduce data bandwidth requirements. This preliminary preparation of data reduces the need for continuous high-power memory access during computation, thereby lowering overall power consumption while maintaining the GPU's software ecosystem capabilities.
4Productivity
If GPU is used for deep learning, then parallel processing capability is improved, but performance/power ratio deteriorates due to excessive graphics processing elements
Solution Approach 1:
The patent extracts and removes excessive graphics processing elements that are not needed for deep learning operations. Specifically, it eliminates 3D graphics rendering modules and other graphics-specific processing units that consume significant power but provide no benefit for AI workloads. The remaining processing elements are optimized for parallel matrix operations and tensor computations, maintaining high parallel processing capability while reducing power consumption.
Solution Approach 2:
The patent applies parameter changes by modifying the operational parameters of the processing elements to optimize for AI workloads. The processing elements are configured with parameters optimized for matrix multiplication, convolution operations, and activation functions rather than graphics rendering parameters. This parameter optimization enables efficient parallel processing for deep learning while reducing the power required per operation.
Data Source
AI summary
An artificial intelligence integrated circuit is provided. The artificial intelligence integrated circuit includes a flash memory, a dynamic random access memory (DRAM), and a memory controller. The flash memory is configured to store a logical-to-physical mapping (L2P) table that is divided into a plurality of group-mapping (G2P) tables. The memory controller includes a first processing core and a second processing core. The first processing core receives a host access command from a host. When a specific G2P table corresponding to a specific logical address in the host access command is not stored in the DRAM, the first processing core determines whether the second processing core has loaded the specific G2P table from the flash memory to the DRAM according to the values in a first column in a first bit map and in a second column of a second bit map.


