FPGA Orchestrator for Low Latency Disaggregated Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
CPU-based PCIe architectures suffer from significant latency due to the requirement that data must flow through a central CPU for processing, leading to bottlenecks and inefficiencies in data throughput, especially in high-speed data processing tasks.
Innovation Solution
A field programmable gate array (FPGA) is configured to operate as a data communications orchestrator, enabling direct communication between component devices and processing operations without the need for CPU intervention, allowing for parallel processing and reduced latency through the use of an FPGA-based architecture.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a CPU-based PCIe architecture is used to interconnect component devices, then the system is expandable and straightforward to manage, but significant latency occurs due to the requirement that data must flow through the central CPU for every transaction
Solution Approach 1:
The patent segments the centralized CPU processing function into distributed processing across multiple component devices. Each device performs local processing and filtering operations independently, eliminating the need for all data to traverse the central CPU. This segmentation resolves the contradiction by maintaining manageability through standardized interfaces while dramatically reducing latency through distributed decision-making.
Solution Approach 2:
The patent introduces intelligent intermediaries in the form of PCIe switches and component devices with embedded processing capabilities. These intermediaries filter, route, and preprocess data locally before transmission, preventing unnecessary data from reaching the CPU. This intermediary layer maintains system manageability while reducing the CPU bottleneck that causes latency.
2Adaptability or versatility
If data flows through the CPU for every transaction in a PCIe architecture, then all subcomponents can be interconnected, but bottlenecks occur that limit the ability of the system to respond to events rapidly
Solution Approach 1:
The patent implements preliminary action through preprocessing and filtering operations performed at the source devices and intermediate PCIe switches before data reaches the CPU. Data is prepared, validated, and routed in advance, eliminating the need for the CPU to perform these time-consuming operations during critical transactions. This preliminary action maintains versatile interconnection while dramatically improving response speed.
Solution Approach 2:
The patent transitions from a single-dimensional centralized processing model to a multi-dimensional distributed processing architecture. Data can flow through multiple parallel paths and processing levels simultaneously, adding dimensional complexity to the interconnection topology. This dimensional expansion enables both versatile device interconnection and high-speed response through parallel processing pathways.
3Reliability
If the CPU is involved in every single transaction on the bus, then data flow to and from respective subcomponents is ensured, but the cost in terms of latency and data throughput is resource intensive and highly inefficient
Solution Approach 1:
The patent enables self-service by equipping component devices and PCIe switches with autonomous processing capabilities. These elements perform data filtering, validation, routing, and error checking independently without requiring CPU intervention for each transaction. This self-service approach ensures reliable data flow through built-in error handling while dramatically improving throughput efficiency by eliminating the CPU bottleneck.
Data Source
AI summary
An asynchronous computer-implemented disaggregated processing system including an ethernet transceiver configured to receive data, from a computing device over a data communications network. A core processor is configured to execute to process information associated with at least some of the received data and a memory is configured to store, via a memory controller, processed information from the core processor. A field programmable gate array is configured via an execution implementation directive to parse at least some of the received data; preprocess at least some of the parsed data for use by the core processor; route, to the core processor, the preprocessed data; receive, from the core processor, information associated with the preprocessed data; route, to the computing device, a response associated with the information associated with the preprocessed data received from the core processor; and store, via the memory controller, the information associated with the preprocessed data in the memory.


