Compact Vector Arithmetic Accelerator for IoT Energy Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional data processing devices for IoT and wearable devices face challenges in achieving low power consumption and small form factor due to high energy usage and large silicon area, making them unsuitable for power-sensitive applications like IoT and wearable devices.
Innovation Solution
A compact, all-in-one signal processing, linear, and non-linear vector arithmetic accelerator with a single programmable compute engine and configurable internal memory, optimized for machine learning and deep learning models, which minimizes system area and energy usage by performing vector operations efficiently.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Use of energy by moving object
If conventional data processing devices are used for IoT and wearable devices, then they can perform general data processing tasks, but they consume high energy and occupy large silicon area
Solution Approach 1:
The patent segments the data processing function into two distinct parts: a general-purpose processor for control and a dedicated neural processing engine (NPE) for vector operations. This segmentation allows the NPE to handle energy-intensive computations efficiently while the main processor remains lightweight, resolving the contradiction between low energy consumption and adequate data processing capability
Solution Approach 2:
The NPE acts as an intermediary between the general-purpose processor and the external world, handling all vector operations and neural network computations. This intermediary approach offloads energy-intensive tasks from the main processor to a specialized unit designed for efficiency, enabling low energy consumption while maintaining high productivity for AI workloads
2Area of stationary object
If conventional data processing devices are used for IoT and wearable devices, then they can perform general data processing tasks, but they occupy large silicon area
Solution Approach 1:
The patent segments the processing architecture into a compact NPE with dedicated vector units and a minimal main processor. The NPE integrates tightly-c coupled memory and compute units in a space-efficient configuration, reducing overall silicon area while maintaining strong data processing capability for neural network operations
Solution Approach 2:
The NPE is designed as a universal engine that can perform multiple types of vector operations (dot products, element-wise operations, matrix multiplications) and support various neural network architectures. This multi-functionality eliminates the need for separate dedicated hardware for each operation type, significantly reducing silicon area while preserving comprehensive data processing capability
3Area of stationary object
If a single programmable compute engine is used, then system area and energy usage are minimized, but the device must handle diverse vector operations
Solution Approach 1:
The compute engine is designed with dynamic configurability through programmable control, allowing the same hardware to adapt its behavior for different vector operations. The engine can be programmed to perform dot products, element-wise additions, multiplications, and other operations by changing control parameters rather than requiring separate fixed-function hardware for each operation type
Solution Approach 2:
The single compute engine incorporates universal vector processing units that can execute multiple types of operations on different data types (integers, floating-point). This universality is achieved through configurable arithmetic logic units and flexible data paths, enabling one engine to replace multiple specialized units while maintaining comprehensive vector operation capability
Data Source
AI summary
Disclosed are methods, devices and systems for all-in-one signal processing, linear and non-linear vector arithmetic accelerator. The accelerator, which in some implementations can operate as a companion co-processor and accelerator to a main system, can be configured to perform various linear and non-linear arithmetic operations, and is customized to provide shorter execution times and fewer task operations for corresponding arithmetic vector operation, thereby providing an overall energy saving. The compact accelerator can be implemented in devices in which energy consumption and footprint of the electronic circuits are important, such as in Internet of Things (IoT) devices, in sensors and as part of artificial intelligence systems.


