Stream Processor and Fixed Dataflow Accelerator for Computer Architecture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer accelerator technologies are not practical for applications with limited demand due to high design and manufacturing costs, and they require complex changes to application programs and tool chains, making them inefficient for tasks that do not justify the development of specialized computer architecture accelerators.
Innovation Solution
A high-speed fixed-function accelerator with a dataflow architecture and a special-purpose stream processor that handles memory accesses independently, allowing long data runs to be processed without general-purpose processor involvement, providing speed advantages and broader applicability through the selection of individual functions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Speed
If a general-purpose processor is used to execute computational tasks, then programming flexibility and ease of operation are maintained, but computational speed and processing efficiency for specialized tasks deteriorate
Solution Approach 1:
The system divides computational work into two segments: general-purpose processing handled by the host CPU and specialized fixed-function processing handled by dedicated accelerator circuits. This segmentation allows each component to be optimized for its specific function, improving overall computational speed without requiring the entire system to be complex.
Solution Approach 2:
The patent introduces an intermediary interface mechanism that allows the general-purpose processor to efficiently transfer control and data to fixed-function accelerators and receive results back. This intermediary layer enables speed optimization for specialized tasks while maintaining the simplicity of the general-purpose programming interface.
2Productivity
If fixed-function accelerators are designed with specialized circuitry, then computational speed for specific functions is improved, but device complexity and manufacturing cost increase
Solution Approach 1:
The patent implements fixed-function accelerators that can be configured to perform multiple specialized functions through a unified architecture. By designing the accelerator with universal interfaces and configurable functional blocks, the system achieves high productivity for various computational tasks without requiring separate specialized circuits for each function, thereby reducing overall design complexity.
3Adaptability or versatility
If data is transferred between general-purpose processor and accelerator, then functional flexibility is maintained, but time loss and computational burden increase
Solution Approach 1:
The patent implements preliminary action by pre-configuring the fixed-function accelerators with predetermined functions and interfaces before execution. The accelerator is prepared in advance with the necessary circuitry and control logic, allowing immediate execution without time-consuming configuration or setup during runtime. This reduces the time loss associated with initiating data transfers and function execution.
4Speed
If specialized computer architecture accelerators are developed, then processing speed for specific applications is improved, but ease of manufacture and development cost increase
Solution Approach 1:
The patent employs copying by implementing standardized interfaces and modular functional blocks that can be replicated across different accelerator designs. Instead of creating entirely new specialized circuits for each application, the system uses copies of proven functional units configured for different tasks, significantly reducing development cost and manufacturing complexity while maintaining high processing speeds.
Data Source
AI summary
A hardware accelerator for computers combines a stand-alone, high-speed, fixed program dataflow functional element with a stream processor, the latter of which may autonomously access memory in predefined access patterns after receiving simple stream instructions and provide them to the dataflow functional element. The result is a compact, high-speed processor that may exploit fixed program dataflow functional elements.


