Cascade Connected Data Processing Engine Cores
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Integrated circuits (ICs) with programmable circuitry face limitations in flexibility and communication efficiency due to restricted connectivity between data processing engines (DPEs), which hampers the formation of varied clusters and concurrent data processing.
Innovation Solution
Implementing a cascade connection architecture between DPEs, allowing each core to send data directly to multiple target cores and receive data from multiple source cores, with programmable inputs and outputs to enable flexible communication and cluster formation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a single processor is used in the IC, then the device complexity is reduced, but the productivity and communication efficiency between data processing engines deteriorate
Solution Approach 1:
The IC is divided into multiple data processing engines (DPEs), each containing its own processor core. This segmentation allows concurrent execution of multiple user applications across different DPEs, thereby improving productivity while maintaining manageable complexity through modular architecture.
Solution Approach 2:
Multiple DPEs are combined within a single IC with shared resources including L2 cache memory, interconnect fabric, and configuration memory. This merging enables efficient resource utilization and high-speed communication between cores while preserving the benefits of parallel processing.
2Device complexity
If restricted connectivity is used between DPEs, then the device complexity is reduced, but the adaptability and flexibility in forming clusters deteriorate
Solution Approach 1:
The interconnect fabric provides universal connectivity that can be dynamically configured to support various cluster topologies and communication patterns. Each core can be programmatically connected to any other core or memory resource, enabling flexible cluster formation for different application requirements while using a single unified interconnect structure.
Solution Approach 2:
The connectivity between DPEs is dynamically reconfigurable through programmable interconnect settings. This allows the system to adapt its communication architecture at runtime based on the specific computational tasks and data flow requirements, providing versatility without requiring multiple fixed connectivity options.
3Device complexity
If traditional inter-DPE communication is used, then the device complexity is reduced, but the speed of data transfer between cores deteriorates
Solution Approach 1:
Each DPE is equipped with local L2 cache memory that can be directly accessed by its associated core, providing fast data access for frequently used data. This local caching strategy reduces the need for slower remote memory accesses while maintaining a relatively simple communication architecture.
Solution Approach 2:
The interconnect fabric acts as an intermediary that provides high-speed direct communication paths between DPEs. This dedicated interconnect structure enables efficient data transfer between cores without burdening the main system bus, achieving fast communication while keeping the overall architecture manageable.
Data Source
AI summary
An integrated circuit includes a plurality of data processing engines (DPEs) DPEs. Each DPE may include a core configured to perform computations. A first DPE of the plurality of DPEs includes a first core coupled to an input cascade connection of the first core. The input cascade connection is directly coupled to a plurality of source cores of the plurality of DPEs. The input cascade connection includes a plurality of inputs, wherein each of the plurality of inputs is connected to a cascade output of a different one of the plurality of source cores. The input cascade connection is programmable to enable a selected one of the plurality of inputs.


