AI Chip Data Conveying and Multiplexed Arbitration for Tensor Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current data processing methods for artificial intelligence, particularly in deep learning, face challenges with complex data access and storage paths using general-purpose processors, and specialized hardware devices lack flexibility in handling multi-dimensional tensor data.
Innovation Solution
An apparatus for data processing comprising an input memory, data conveying components, multiplexed arbitration components, and an output memory, which processes tensor data by parsing instructions to determine read and write addresses and commands, enabling flexible data conveying and transposition without hardware modification, and utilizing on-chip memories to enhance bandwidth and reduce access and storage delays.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If a general purpose processor (CPU, GPU, or DSP) is used for data processing, then the system has high flexibility and adaptability, but the data access and storage path becomes complex and is limited by access bandwidth
Solution Approach 1:
The system is divided into distinct functional modules: input memory for data storage, data conveying components for data transmission, multiplexed arbitration components for access control, and output memory for result storage. This segmentation simplifies the data access path by dedicating specific components to specific functions, reducing the complexity inherent in general-purpose processors while maintaining flexibility through programmable control.
2Productivity
If a special purpose hardware device (ASIC or FPGA) is used for data processing, then the data conveying and data transposition efficiency is improved, but the flexibility is reduced
Solution Approach 1:
The multiplexed arbitration component serves multiple functions: it manages data conveying between memory and processing units, handles data transposition operations, and controls access to both input and output memories. This multi-functional design allows the hardware to achieve specialized processing efficiency while maintaining flexibility through software-configurable operation modes, resolving the contradiction between specialization and adaptability.
3Quantity of substance
If external memory is used for data storage, then the system has high capacity, but the access bandwidth is limited and access delays increase
Solution Approach 1:
The system employs a hierarchical memory structure where input memory and output memory are nested within the data processing apparatus, providing fast access for active data. This nested memory architecture allows the system to maintain high-capacity external memory storage while providing a high-speed access path through the integrated input/output memories, effectively resolving the bandwidth and delay limitations of external memory access.
Data Source
AI summary
Disclosed are an apparatus for data processing, an artificial intelligence chip, and an electronic device. The apparatus for data processing includes: at least one input memory, at least one data conveying component, at least one multiplexed arbitration component, and at least one output memory. The input memory is connected to the data conveying component, the data conveying component is connected to the multiplexed arbitration component, and the multiplexed arbitration component is connected to the output memory.


