In-Line Accelerator for Buffer-Less Stream Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional systems using general purpose instruction-based processors for data processing are slower and more power-consuming compared to accelerators, due to high overhead in data transfer between processors and accelerators, which involves copying data between interfaces and processors.
Innovation Solution
The implementation of an auto-generated in-line accelerator that performs buffer-less computations and automates data transfer with a general purpose instruction-based processor, allowing for seamless integration and efficient processing by slicing computations into fast and slow paths, with bailout mechanisms to handle premature termination.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If data is transferred between general purpose instruction-based processor and accelerator using conventional methods, then the accelerator can process computations, but high overhead is incurred due to copying data between interfaces and processors
Solution Approach 1:
The patent merges the accelerator directly into the I/O processing unit, eliminating the separate accelerator component and its associated data transfer interfaces. This integration allows the accelerator to process data in-place within the I/O unit without requiring copying to and from a separate processor, thereby resolving the technical contradiction by eliminating data transfer overhead while maintaining accelerated processing capability.
2Productivity
If accelerator is used for specialized computations, then processing performance increases, but device complexity increases due to integration with general purpose processor
Solution Approach 1:
The I/O processing unit is designed with multi-functionality, incorporating both general I/O processing capabilities and specialized acceleration functions within a single unified device. This allows the system to achieve high processing performance for specialized computations while avoiding the complexity of integrating separate accelerator components, as the acceleration capability is natively part of the I/O unit's universal functionality.
3Productivity
If conventional accelerator architecture is used, then specialized computations are performed, but power consumption increases due to data copying operations
Solution Approach 1:
By merging the accelerator functionality directly into the I/O processing unit, the patent eliminates the need for data copying between separate components. This integration maintains the specialized computation capability while significantly reducing power consumption by removing the energy-intensive data transfer operations between independent accelerator and processor units.
Data Source
AI summary
A data processing system is disclosed that includes machines having an in-line accelerator and a general purpose instruction-based general purpose instruction-based processor. In one example, a machine comprises storage to store data and an Input/output (I/O) processing unit coupled to the storage. The I/O processing unit includes an in-line accelerator that is configured for in-line stream processing of distributed multi stage dataflow based computations. For a first stage of operations, the in-line accelerator is configured to read data from the storage, to perform computations on the data, and to shuffle a result of the computations to generate a first set of shuffled data. The in-line accelerator performs the first stage of operations with buffer less computations.


