3D Dataflow Architecture With RAPCs for Programmable Routing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multiprocessor architectures face challenges in achieving fast computing with flexibility, minimizing unintended interactions, and optimizing data flow routing while handling complex algorithms efficiently.

Innovation Solution

A 3D dataflow architecture using Reconfigurable Algorithmic Pipeline Cores (RAPCs) with a pipelined pseudo 3D dataflow routing structure, allowing simultaneous data processing and reducing bottlenecks through a simplified hardware design that eliminates unnecessary CPU components and enables flexible algorithm execution.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If traditional multiprocessor architectures are used, then computing speed can be achieved, but flexibility and unintended interactions between program elements increase

Engineering Contradiction:
Improvecomputing speedVSAvoidflexibility
Core Design Contradiction:
SpeedVSAdaptability or versatility

Solution Approach 1:

The system segments the multiprocessor architecture into distinct computational elements (CEs) organized in a grid, where each CE is independent and can be individually configured. This segmentation allows parallel processing for speed while maintaining clear boundaries that reduce unintended interactions, and the modular nature enables flexible reconfiguration for different algorithms.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The architecture implements dynamic reconfigurability where computational elements and their interconnections can be programmatically adjusted during operation. This allows the system to adapt to different computational tasks and algorithms, providing flexibility while maintaining optimized performance paths for each specific computation.

Inventive Principle:
Principle #15Dynamics

2Adaptability or versatility

If complex routing structures are used to handle data flow, then algorithm flexibility increases, but data flow bottlenecks and latency increase

Engineering Contradiction:
Improvealgorithm flexibilityVSAvoiddata flow latency
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Each computational element in the grid has standardized local interfaces and data flow paths, creating uniform local quality throughout the architecture. This standardization enables predictable data flow timing while the overall grid structure provides flexibility for different algorithms by combining these uniform local units in various configurations.

Inventive Principle:
Principle #3Local quality

3Adaptability or versatility

If more CPU components are added to handle complex algorithms, then algorithm capability increases, but hardware size and complexity increase

Engineering Contradiction:
Improvealgorithm capabilityVSAvoidhardware complexity
Core Design Contradiction:
Adaptability or versatilityVSDevice complexity

Solution Approach 1:

The computational elements are designed as universal units that can perform multiple computational functions through reconfiguration. Each CE can be programmed to execute different operations, eliminating the need for specialized hardware for each algorithm type. This multi-functionality provides algorithm capability while maintaining relatively simple, standardized hardware components.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS20250298773A13D dataflow architecture for a computing device
Publication Date: 2025.09.25 ICAT LLC D B A TURING MICRO
  • US20250298773A1 patent drawing
  • US20250298773A1 patent drawing
  • US20250298773A1 patent drawing

AI summary

A plurality of simplified CPUs (RAPCs) are provided with data from an input or from memory. The first RAPC completes its simplified task, then turns and hands the data downstream to the next RAPC. Data is routed in a programmable, 3D routing scheme through the array, allowing many simultaneous operations to complete an algorithm as in an assembly line. Completed results go downstream for use as needed.