AI Chip Data Conveying and Multiplexed Arbitration for Tensor Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current data processing methods for artificial intelligence, particularly in deep learning, face challenges with complex data access and storage paths using general-purpose processors, and specialized hardware devices lack flexibility in handling multi-dimensional tensor data.

Innovation Solution

An apparatus for data processing comprising an input memory, data conveying components, multiplexed arbitration components, and an output memory, which processes tensor data by parsing instructions to determine read and write addresses and commands, enabling flexible data conveying and transposition without hardware modification, and utilizing on-chip memories to enhance bandwidth and reduce access and storage delays.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a general purpose processor (CPU, GPU, or DSP) is used for data processing, then the system has high flexibility and adaptability, but the data access and storage path becomes complex and is limited by access bandwidth

Engineering Contradiction:
ImproveflexibilityVSAvoiddata access and storage path complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The system is divided into distinct functional modules: input memory for data storage, data conveying components for data transmission, multiplexed arbitration components for access control, and output memory for result storage. This segmentation simplifies the data access path by dedicating specific components to specific functions, reducing the complexity inherent in general-purpose processors while maintaining flexibility through programmable control.

Inventive Principle:
Principle #1Segmentation

2Productivity

If a special purpose hardware device (ASIC or FPGA) is used for data processing, then the data conveying and data transposition efficiency is improved, but the flexibility is reduced

Engineering Contradiction:
Improvedata conveying efficiencyVSAvoidflexibility
Core Design Contradiction:
ProductivityVSAdaptability or versatility

Solution Approach 1:

The multiplexed arbitration component serves multiple functions: it manages data conveying between memory and processing units, handles data transposition operations, and controls access to both input and output memories. This multi-functional design allows the hardware to achieve specialized processing efficiency while maintaining flexibility through software-configurable operation modes, resolving the contradiction between specialization and adaptability.

Inventive Principle:
Principle #6Universality (Multi-functionality)

3Quantity of substance

If external memory is used for data storage, then the system has high capacity, but the access bandwidth is limited and access delays increase

Engineering Contradiction:
Improvedata storage capacityVSAvoiddata access speed
Core Design Contradiction:
Quantity of substanceVSSpeed

Solution Approach 1:

The system employs a hierarchical memory structure where input memory and output memory are nested within the data processing apparatus, providing fast access for active data. This nested memory architecture allows the system to maintain high-capacity external memory storage while providing a high-speed access path through the integrated input/output memories, effectively resolving the bandwidth and delay limitations of external memory access.

Inventive Principle:
Principle #7Nested doll (Nesting)

Data Source

PatentUS11023391B2Apparatus for data processing, artificial intelligence chip and electronic device
Publication Date: 2021.06.01 KUNLUNXIN TECHNOLOGY (BEIJING) CO LTD
  • US11023391B2 patent drawing
  • US11023391B2 patent drawing
  • US11023391B2 patent drawing

AI summary

Disclosed are an apparatus for data processing, an artificial intelligence chip, and an electronic device. The apparatus for data processing includes: at least one input memory, at least one data conveying component, at least one multiplexed arbitration component, and at least one output memory. The input memory is connected to the data conveying component, the data conveying component is connected to the multiplexed arbitration component, and the multiplexed arbitration component is connected to the output memory.