Reconfigurable Streaming Clusters for CNN Layer Adaptation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current convolutional accelerators face challenges in efficiently managing the varied memory and compute requirements across different layers of convolutional neural networks (CNNs), particularly in resource-constrained hardware, which hampers real-time performance and throughput for applications like image recognition.

Innovation Solution

The proposed solution involves a hardware accelerator with a reconfigurable crossbar switch and stream engines that enable flexible data streaming and mode control for processing elements, allowing for efficient data management and operation mode switching between compute and memory modes, thereby optimizing the processing of convolutional layers.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If a fixed architecture is used in convolutional accelerators, then the hardware design is simple, but it cannot adapt to varied memory and compute requirements across different CNN layers

Engineering Contradiction:
Improveadaptability to varied memory and compute requirementsVSAvoidhardware architecture complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The patent implements dynamic reconfiguration of the crossbar switch connectivity based on the specific computational requirements of different CNN layers. The crossbar switch can be reconfigured to connect processing elements in different topologies (e.g., fully connected, partially connected, or disconnected) depending on whether a layer requires high compute intensity, high memory intensity, or a balance thereof. This dynamic adaptability resolves the contradiction by allowing a single hardware architecture to serve multiple computational needs without requiring separate fixed architectures for each scenario.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The processing elements are designed as universal units that can function in multiple roles: as compute units, as memory storage units, or as both simultaneously. The same processing element can be configured to perform different operations based on the reconfigured crossbar switch connectivity. This multi-functionality allows the hardware to adapt to varied requirements across CNN layers without requiring specialized dedicated hardware for each function, thereby improving adaptability while controlling overall device complexity.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Productivity

If resource-constrained hardware is used, then the hardware is more efficient and cost-effective, but real-time performance and throughput are hampered

Engineering Contradiction:
Improvereal-time processing throughputVSAvoidhardware resource constraints
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The hardware accelerator is segmented into multiple processing elements that can be independently configured and connected through the reconfigurable crossbar switch. This segmentation allows the system to dynamically allocate computational resources based on the specific needs of each CNN layer, maximizing the utilization of limited hardware resources. By segmenting the processing workload and resources, the system achieves high throughput on resource-constrained hardware without requiring excessive hardware complexity.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system changes operational parameters by dynamically reconfiguring the crossbar switch connectivity and adjusting the modes of processing elements (compute mode, memory mode, or both) based on the computational characteristics of each CNN layer. This parameter change approach allows the hardware to optimize its performance for specific tasks within the constraints of limited resources, thereby improving real-time throughput without requiring proportional increases in hardware complexity or resource consumption.

Inventive Principle:
Principle #35Parameter changes

3Productivity

If data is streamed through a fixed routing path, then the data streaming is simple, but it cannot optimize for different computational patterns of CNN layers

Engineering Contradiction:
Improvedata streaming efficiencyVSAvoiddata routing complexity
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The data routing paths are made dynamic through the reconfigurable crossbar switch, which can establish different connectivity patterns between processing elements based on the computational requirements of the current CNN layer. For layers requiring high data movement efficiency, the crossbar can be configured to create optimized data paths. This dynamic routing capability improves data streaming efficiency for different computational patterns while managing routing complexity through a systematic reconfiguration approach rather than requiring complex fixed routing infrastructure.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentUS20240281397A1Reconfigurable, streaming-based clusters of processing elements, and multi-modal use thereof
Publication Date: 2024.08.22 STMICROELECTRONICS INT NV
  • US20240281397A1 patent drawing
  • US20240281397A1 patent drawing
  • US20240281397A1 patent drawing

AI summary

A hardware accelerator includes processing elements of a neural network, each processing element having a memory; a stream switch; stream engines coupled to functional circuits via the stream switch, wherein the stream engines, in operation, generate data streaming requests to stream data to and from functional circuits of the plurality of functional circuits; a first system bus interface coupled to the stream engines; a second system bus interface coupled to the processing elements; and mode control circuitry, which, in operation, sets respective modes of operation for the plurality of processing elements. The modes of operation include: a compute mode of operation in which the processing element performs computing operations using the memory associated with the processing element; and a memory mode of operation in which the memory associated with the processing element performs memory operations, bypassing the stream switch, via the second system bus interface.