Coalescing Accelerators via Dynamic Image Generation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current computer systems face inefficiencies in performance due to traditional system architectures, where compute power is not optimally aligned with data processing, leading to bottlenecks and suboptimal performance.
Innovation Solution
The Open Coherent Accelerator Processor Interface (OpenCAPI) enables the development of hardware accelerators that can share virtual addresses with processors, allowing for the creation of a new accelerator image by coalescing multiple accelerators based on coalescence criteria, which is then deployed to programmable devices to enhance performance.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple accelerators are deployed to handle different compute tasks, then processing capability and versatility are improved, but system complexity and resource overhead increase
Solution Approach 1:
The patent merges multiple accelerators into a single coalesced accelerator that can dynamically handle different compute tasks. The accelerator manager monitors performance metrics and combines accelerators that process related data streams, reducing system complexity while maintaining versatility through dynamic task allocation to the coalesced unit.
Solution Approach 2:
The coalesced accelerator is designed to perform multiple functions by dynamically allocating its resources to different compute tasks based on real-time performance monitoring. A single coalesced accelerator unit can handle various data processing functions that previously required multiple separate accelerators, achieving multi-functionality while reducing overall system complexity.
2Adaptability or versatility
If multiple separate accelerators are used for different functions, then functional versatility is improved, but resource utilization efficiency deteriorates
Solution Approach 1:
The patent implements dynamic accelerator coalescing where the system continuously monitors performance metrics and dynamically combines or separates accelerators based on real-time workload characteristics. This dynamic approach allows the same hardware resources to be flexibly allocated to different functions at different times, improving both functional versatility and resource utilization efficiency simultaneously.
Solution Approach 2:
The accelerator manager automatically monitors performance metrics and makes decisions about coalescing or separating accelerators without external intervention. The system self-adjusts the accelerator configuration to optimize resource utilization while maintaining the required functional versatility, eliminating the need for manual resource management.
3Device complexity
If accelerators are coalesced into fewer units, then system complexity is reduced, but processing capacity may be limited
Solution Approach 1:
The coalesced accelerator maintains continuous processing capability by dynamically switching between different compute tasks without idle periods. The accelerator manager ensures that the coalesced unit is always assigned productive work based on monitored performance metrics, maintaining high processing capacity while using fewer accelerator units, thus reducing system complexity without sacrificing processing power.
Data Source
AI summary
An accelerator manager monitors run-time performance of multiple accelerators in one or more programmable devices, and when first and second accelerators satisfy at least one coalescence criterion, the accelerator manager generates a new accelerator image corresponding to a third accelerator that coalesces the first accelerator and the second accelerator. In a first embodiment, the accelerator manager generates the new accelerator image from the accelerator images for the first and second accelerators. In a second embodiment, the accelerator manager generates the new accelerator image from hardware description language representations of the first and second accelerators. The new accelerator image is deployed to a programmable device to provide a new accelerator, and calls to the first and second accelerators are replaced with calls to the new accelerator.


