Fused Array Instructions for Low-Power AI Coprocessing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional computing devices require significant power and resources to perform array, matrix, and tensor operations, making it inefficient for AI applications, especially in battery-powered mobile devices, and lead to data security and privacy concerns due to reliance on power-hungry datacenters.
Innovation Solution
Integration of an array coprocessor with a main processor core, sharing a unified Instruction Set Architecture (ISA), to accelerate computations with energy efficiency and reduced hardware costs, using architectures like RVA23, and advanced packaging technologies, enabling seamless communication and resource sharing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If conventional processors (CPU/GPU) are used to perform array, matrix, and tensor operations, then computation capability is achieved, but power consumption increases significantly
Solution Approach 1:
The patent segments the processing system by introducing a dedicated array processing unit (APU) that works in conjunction with the main CPU. The APU is specifically designed to handle array, matrix, and tensor operations, while the CPU handles control and other computational tasks. This segmentation allows the system to achieve high computation capability for AI workloads without requiring the entire system to consume high power, as the APU can be optimized for energy-efficient array operations.
Solution Approach 2:
The patent introduces an intermediary structure - a unified memory architecture that sits between the CPU and APU, allowing efficient data sharing and reducing data movement overhead. This intermediary memory system enables the APU to access data efficiently without requiring constant high-power communication with the CPU, thus maintaining high productivity while reducing overall power consumption.
2Reliability
If data is processed locally on mobile devices, then data security and privacy are improved, but computational resources and power are limited
Solution Approach 1:
The patent implements a dynamic processing architecture where the APU can be activated or deactivated based on workload requirements. When AI computations are needed, the APU becomes active to provide enhanced computation capability locally. When not needed, the system operates in a lower-power state. This dynamic approach allows mobile devices to achieve high computation capability when required while maintaining energy efficiency for data security and privacy preservation.
Solution Approach 2:
The patent employs parameter changes by allowing the APU to operate at different performance levels and precision modes. For less computationally intensive tasks, the APU can operate in lower-power modes or reduced precision, while still providing sufficient computation capability for local AI processing. This enables the system to balance between computation capability and power consumption, making local processing viable for mobile devices.
3Productivity
If array operations are performed using conventional processors, then computational tasks are completed, but hardware costs and resource consumption increase
Solution Approach 1:
The patent designs the APU with a unified architecture that can handle multiple types of operations - array operations, matrix operations, and tensor operations - all within a single processing unit. This multi-functionality eliminates the need for separate hardware components for each operation type, reducing overall device complexity while maintaining high productivity for various AI computational tasks.
Data Source
AI summary
Systems and methods are directed to fusing array operations associated with an integrated circuit. An integrated circuit receives fused instruction that include a first operand and a second operand that each comprises an array. The fused instruction is executed as a single instruction rather than as a combination of at least two instructions. In response to receiving the fused instruction, the integrated circuit generates an output based on the first operand and the second operand. The generating includes generating, by multipliers and adders of the integrated circuit, at least one sum based on the first operand and the second operand, and generating, by a generator of the integrated circuit, an output based on the at least one sum.


