Edge AI Accelerator Virtualization for Latency and Cost
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Telecommunications networks with edge computing face challenges in providing high-speed, low latency access to hardware accelerators for user equipment, as existing solutions often lock users into proprietary hardware and software.
Innovation Solution
Deployment of a Bitfusion server in edge servers that allows for virtualized access to hardware accelerators, enabling users to share accelerators across the network and providing flexible AI computing power without the need for proprietary hardware or software, through a method that intercepts application calls and executes them using a pool of accelerators.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If hardware accelerators are included on edge nodes, then computing performance is improved, but device complexity increases and users are locked into proprietary hardware or software
Solution Approach 1:
The patent introduces a virtualization layer as an intermediary between hardware accelerators and user applications. This layer abstracts the physical accelerator hardware, allowing multiple users and applications to access the accelerators through standardized interfaces without needing to know or be locked into proprietary hardware details. The virtualization layer handles the complexity of hardware management while providing simple access points for users.
Solution Approach 2:
The patent creates a universal access mechanism where different types of hardware accelerators can be accessed through a common interface. The system allows various applications and users to share the same accelerator hardware without requiring proprietary software or hardware configurations, enabling one hardware resource to serve multiple functions and users simultaneously.
2Productivity
If hardware accelerators are deployed on edge nodes, then AI computing capability is improved, but cost increases due to need for proprietary hardware or software
Solution Approach 1:
The patent enables a single hardware accelerator to serve multiple applications and users through virtualization. Instead of requiring dedicated accelerators for each application or user, the system allows shared access, reducing the total number of accelerators needed and lowering overall hardware costs while maintaining high AI computing capability.
Solution Approach 2:
The patent creates virtual copies or instances of accelerator access through software layers. Rather than physically replicating expensive hardware accelerators for each user, the system creates virtual interfaces that allow multiple users to access the same physical accelerator, reducing hardware costs while maintaining computing capability.
3Device complexity
If edge computing hosts have low compute power, then device complexity is reduced, but access speed to accelerators becomes slow and latency increases
Solution Approach 1:
The patent introduces a virtualization layer as an intermediary that manages the connection between edge computing hosts and hardware accelerators. This layer optimizes the access path, enabling fast communication between clients and accelerators while maintaining the simplicity of the edge host architecture. The virtualization layer handles the complexity of accelerator management and communication optimization.
4Reliability
If hardware accelerators are locked to specific hosts, then reliability of acceleration is improved, but adaptability decreases as users cannot access accelerators flexibly
Solution Approach 1:
The patent creates a universal access mechanism where hardware accelerators can be dynamically assigned to different applications and users. The virtualization layer maintains reliable acceleration by ensuring proper resource allocation and access control, while simultaneously enabling flexible adaptability as users can access accelerators based on their needs without being locked to specific hosts.
Data Source
AI summary
Disclosed herein is the integration into edge nodes of a telecommunications network system of client computer system and server computer system where the server computer system includes a pool of shareable accelerators and the client computer runs an application program that is assisted by the pool of accelerators. The edge nodes connect to user equipment, and some of the user equipment can themselves act as one of the client computer systems. In some embodiments, the accelerators are GPUs, and in other embodiments, the accelerators are artificial intelligence accelerators.


