Dynamic quality of service management for virtualized hardware accelerators
The system addresses inefficiencies in traditional virtualization by translating application demands into hardware configurations and enforcing QoS, ensuring efficient and isolated resource allocation for programmable hardware accelerators, optimizing utilization and maintaining performance in heterogeneous environments.
Patent Information
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- TEJAS NETWORKS LTD
- Filing Date
- 2025-09-12
- Publication Date
- 2026-05-28
AI Technical Summary
Traditional resource virtualization techniques struggle to provide well-defined metrics for programmable hardware accelerators, leading to inefficiencies and performance bottlenecks, especially in heterogeneous computing environments, as they fail to accurately translate application demands into hardware configurations and lack QoS guarantees.
A system and method for virtualizing hardware acceleration resources with guaranteed quality of service, involving a host processing unit, hardware accelerator, and components like a housekeeping unit and embedded hypervisor, which translate requests into resource requirements, allocate resources, isolate address spaces, and enforce QoS parameters through dynamic monitoring and modification.
Enables efficient utilization of hardware resources, maintains performance isolation, and guarantees quality of service by dynamically allocating and managing resources, optimizing utilization and supporting diverse application requirements in heterogeneous environments.
Smart Images

Figure IN2025051489_28052026_PF_FP_ABST
Abstract
Description
DYNAMIC QUALITY OF SERVICE MANAGEMENT FOR VIRTUALIZED HARDWARE ACCELERATORSFIELD
[0001] Various embodiments of the disclosure relate to the field of computer hardware virtualization, specifically to systems and methods for virtualizing hardware acceleration resources with guaranteed quality of service in heterogeneous computing environments. More particularly, the disclosure describes a hardware accelerator architecture that enables the efficient sharing of programmable and nonprogrammable acceleration between multiple virtual machines or applications while maintaining performance isolation and ensuring quality of service guarantees.BACKGROUND
[0002] Traditional resource virtualization techniques have been successfully employed to share computing, networking, and memory resources. However, such traditional techniques have proven inadequate when extended to programmable hardware accelerators. The traditional techniques lack well-defined metrics for requesting a quantum of resources from custom hardware, making it challenging to quantify and guarantee Quality of Service (QoS) in virtualized hardware environments.
[0003] For instance, traditional techniques for virtualization may typically involve creating a "slice" of a resource and sharing it with virtualized or containerized applications. The primary challenges associated with such traditional techniques may include the difficulty in mapping application requirements to actual hardware parameters and the inability to guarantee that requested resources are consistently available. For example, in the context of AI / ML workloads running on specialized hardware accelerators, the resource requirements for such workloads may be difficult to predict, leading to inefficient resource allocation and potential performance bottlenecks. In many instances, the virtualization mechanisms or techniques may fail to accurately translate high-level application demands into appropriate low-level hardware configurations, resulting in suboptimal utilization of the accelerator's capabilities.
[0004] Another significant limitation of traditional techniques is the lack of well-defined ways to partition and provide QoS guarantees for heterogeneous computing systems. Such deficiency becomes challenging in scenarios when multiple applications compete for the same hardware resources, often leading to unpredictable performance and potential service disruptions. Such challenges may lead to, for example, performance degradation, inefficient resource utilization, and increased operational complexity. Moreover, the inability to effectively manage and allocate resources in large-scale heterogeneous computing environments may impede scalability and the potential benefits of hardware acceleration in critical technologies such as RAN and AI / ML.
[0005] Overcoming the above-described limitations by providing an optimal mechanism or technique for the virtualization of resources for programmable and non-programmable hardware accelerators, including providing a robust framework for quantifying, guaranteeing, and managing resources in advanced computing environments, ensuring optimal QoS, may be challenging.
[0006] The limitations and disadvantages of conventional and traditional techniques or mechanisms will become apparent to one of skill in the art through a comparison of described systems with some aspects of the present disclosure, as set forth in the remainder of the present application and with reference to the drawings.SUMMARY
[0007] The system described in the subject specification may implement the steps of the method described. In an embodiment, the system and method are provided for virtualizing hardware acceleration resources with a guaranteed quality of service.
[0008] In an embodiment, the system and method provide for creating virtual hardware accelerators on a hardware accelerator unit coupled to a host server. This virtualization allows for efficient utilization of hardware resources and enables multiple applications to share the same physical hardware. In an embodiment, the system and method provide for assigning address registers, including device configuration address registers, to each virtual hardware accelerator. This assignment enables proper configuration and management of the virtual accelerators. In an embodiment, the system and method provide for receiving requests for executingapplications from the virtual hardware accelerators. These requests are processed to determine the required hardware resources for each application.
[0009] In an embodiment, the system and method provide for translating the received requests into hardware accelerator resource requirements. This translation process maps application needs to specific hardware capabilities. In an embodiment, the system and method provide for partitioning and allocating hardware accelerator resources to the virtual hardware accelerators based on the translated resource requirements. This allocation ensures that each virtual accelerator receives the necessary resources to perform its tasks. In an embodiment, the system and method provide for isolating address spaces between the virtual hardware accelerators. This isolation enhances security and prevents interference between different applications running on the virtual accelerators. In an embodiment, the system and method provide for monitoring the hardware accelerator resources for usage of each virtual hardware accelerator. This monitoring allows for real-time tracking of resource utilization.
[0010] In an embodiment, the system and method provide for comparing the monitored hardware accelerator resource usage with the allocated resources for each virtual hardware accelerator. This comparison helps identify any discrepancies between allocated and actual resource usage. In an embodiment, the system and method provide for enforcing the quality-of-service parameters for each virtual hardware accelerator. This enforcement ensures that each virtual accelerator meets its performance requirements. In an embodiment, the system and method provide for dynamically monitoring and modifying the allocation of hardware accelerator resources based on the comparison of monitored usage with allocated resources. This dynamic monitoring and modification may optimize resource utilization and maintain performance levels.
[0011] In an embodiment, the system and method provide for maintaining the isolation of address spaces between the virtual hardware accelerators during resource modification. This maintenance preserves security and prevents conflicts between virtual accelerators. In an embodiment, the system and method provide for assigning device configuration address registers to each virtual hardware accelerator from the hardware accelerators. This assignment enables the proper configuration of virtual accelerators. In an embodiment, the system and method provide for translating requests into hardware accelerator resource requirements using resource templates.These templates streamline the translation process and ensure consistent resource allocation.
[0012] In an embodiment, the system and method provide for notifying the host processing unit of potential violations of quality-of-service limits. This notification allows for timely intervention and modification of resource allocation. In an embodiment, the system and method provide for utilizing heterogeneous hardware computing resources in the hardware accelerator. This heterogeneity allows for specialized processing of diverse workloads. In an embodiment, the system and method provide for representing analysis of applications as directed acyclic graphs capturing execution and data flow. This representation facilitates efficient resource allocation and scheduling. In an embodiment, the system and method provide for implementing a mixed integer linear programming mechanism to determine resource requirements. This mechanism optimizes resource allocation across multiple virtual accelerators. In an embodiment, the system and method provide for determining the internal interconnect bandwidth required for transfers between the host processing element and accelerators, between different accelerators, and between memory, host processing elements, and accelerators. This determination ensures sufficient data transfer capabilities within the system. The described system and method enable efficient virtualization of hardware acceleration resources while maintaining a guaranteed quality of service. By dynamically allocating resources, isolating address spaces, and enforcing performance parameters, the invention optimizes hardware utilization and supports diverse application requirements.
[0013] These and other features and advantages of the present disclosure may be appreciated by a review of the following detailed description of the present disclosure, along with the accompanying figures in which reference numerals refer to like parts throughout.OBJECTS OF THE INVENTION
[0014] An object of the present invention is to provide a system and method for virtualizing hardware acceleration resources with a guaranteed quality of service. Another object of the invention is to provide a memory and processor configured to execute instructions stored in the memory to execute operations, enabling efficient resource management and task execution.
[0015] Another object of the invention is to provide a host processing unit that interfaces with hardware accelerator components to facilitate virtualization. Another object of the invention is to provide a hardware accelerator that creates virtual hardware accelerators on a hardware accelerator unit coupled to a host server or hosting processing unit, allowing for flexible resource allocation.
[0016] Another object of the invention is to enable assignment of address registers, including device configuration address registers, to each virtual hardware accelerator for proper resource management. Another object of the invention is to provide a housekeeping unit that receives and processes requests for executing applications on the hardware accelerator through virtual hardware accelerators.
[0017] Another object of the invention is to translate received requests into hardware accelerator resource requirements for efficient resource allocation. Another object of the invention is to partition and allocate hardware accelerator resources to virtual hardware accelerators based on translated resource requirements.
[0018] Another object of the invention is to provide a memory management unit that ensures isolation of address spaces between virtual hardware accelerators. Another object of the invention is to provide an embedded hypervisor running on the housekeeping unit that monitors hardware accelerator resource usage for each virtual hardware accelerator.
[0019] Another object of the invention is to compare monitored hardware accelerator resource usage with allocated resources for each virtual hardware accelerator. Another object of the invention is to enforce quality of service parameters through dynamic monitoring and modification of resource allocation based on usage comparison.
[0020] Another object of the invention is to maintain isolation of address spaces between virtual hardware accelerators while enforcing quality of service parameters. Another object of the invention is to assign device configuration address registers to each virtual hardware accelerator from the plurality of virtual hardware accelerators.
[0021] Another object of the invention is to translate requests into hardware accelerator resource requirements using resource templates. Another object of the invention is to notify the host processing unit of potential violations of quality-of- service limits.
[0022] Another object of the invention is to provide heterogeneous hardware computing resources within the hardware accelerator. Another object of the invention is to represent analysis of applications as directed acyclic graphs capturing execution and data flow.
[0023] Another object of the invention is to implement a mixed integer linear programming mechanism to determine resource requirements. Another object of the invention is to provide internal interconnected bandwidth for transfers between host processing elements and accelerators, between accelerators, and between memory and host processing elements.BRIEF DESCRIPTION OF THE DRAWINGS
[0024] FIG. 1 is a block diagram showing an environment 100, including a system for virtualizing hardware acceleration resources with guaranteed quality of service, according to an exemplary embodiment.
[0025] FIG. 2 is a block diagram showing an architecture of a host processing unit, according to an exemplary embodiment.
[0026] FIG. 3 is a block diagram illustrating an architecture of a hardware accelerator, according to an exemplary embodiment.
[0027] FIG. 4 is a block diagram that illustrates an environment for deploying a host processing unit and a hardware accelerator, according to an exemplary embodiment.
[0028] FIG. 5 is an illustration of a flow diagram of process 500 for virtualizing hardware acceleration resources with guaranteed quality of service, according to an exemplary embodiment.
[0029] FIG. 6 is an illustration showing an exemplary hardware configuration of a special-purpose computer 600 that may be used to implement components for virtualizing hardware acceleration resources with guaranteed quality of service, according to exemplary embodiments.DETAILED DESCRIPTION
[0030] Various embodiments of the disclosure relate to a system and method for virtualizing hardware acceleration resources with a guaranteed quality of service.The present disclosure describes various embodiments of systems and methods. It should be understood that the description herein may be presented in terms of a system, a method, or various combinations of both. The embodiments described are not limited to any particular aspect, feature or combination thereof, nor are the embodiments limited to any specific implementation or technique for carrying out or implementing the invention.
[0031] The terms "system" and "method" may be used interchangeably herein unless explicitly stated otherwise. It is to be understood that where a system is described, the description encompasses a corresponding method, and vice versa. The embodiments described herein may be implemented as a system, a method, or both and should not be interpreted as being limited to either a system or a method unless specifically stated.
[0032] In an embodiment, the system described may include a host processing unit that may be communicatively coupled with a hardware accelerator. The hardware accelerator may be configured to create multiple virtual hardware accelerators on a hardware accelerator unit, which may be coupled to a host server or the host processing unit. The hardware accelerator may further be configured to assign multiple address registers, including device configuration address registers, to each virtual hardware accelerator.
[0033] In an embodiment, the system may implement a component, subsystem, module, engine, etc., such as a housekeeping unit that may be configured to receive requests for executing applications from the virtual hardware accelerators. The housekeeping unit may translate the requests into hardware accelerator resource requirements. Based on the translated hardware accelerator resource requirements, the housekeeping unit may partition and allocate hardware accelerator resources to the virtual hardware accelerators. In an embodiment, the component, subsystem, module, engine, etc., such as a memory management unit, may provide isolation of the address spaces between the virtual hardware accelerators. Such isolation may enable each virtual hardware accelerator to operate within its designated address space without interfering with others.
[0034] In an embodiment, the component, subsystem, module, engine, etc., such as an embedded hypervisor (e-hypervisor), may execute operations or run on or in cooperation with the housekeeping unit and may be configured to monitor the hardware accelerator resources for the usage of each of the virtual hardwareaccelerators. The e-hypervisor may compare the monitored hardware accelerator resource usage with the allocated hardware accelerator resources for each of the virtual hardware accelerators. Based on this comparison, the e-hypervisor may enforce QoS parameters for each virtual hardware accelerator.
[0035] In an embodiment, the enforcement of the QoS parameters may be achieved by dynamically monitoring and modifying the allocation of the hardware accelerator resources. Such modification may be based on the comparison of the monitored hardware resource usage with the hardware accelerator resources allocated for each of the virtual hardware accelerators. Throughout this process, the isolation of the address spaces between the virtual hardware accelerators may be maintained.
[0036] In an embodiment, the hardware accelerator may be further configured to assign multiple device configuration address registers to each virtual hardware accelerator. This assignment may enable precise control and configuration of each virtual hardware accelerator. In the mechanism of translating requests into hardware accelerator resource requirements, the housekeeping unit may use or utilize resource templates. The resource templates may provide standardized techniques for mapping application requirements to specific hardware accelerator resource needs.
[0037] In an embodiment, the e-hypervisor may also be configured to notify the host processing unit of potential violations of QoS parameters or QoS limits. Such a function or operation may enable proactive management and intervention if resource allocation or usage approaches predefined thresholds. In an embodiment, the hardware accelerator unit may include multiple heterogeneous computing elements that may be shared and referred to as shared resources. Such diversity in computing elements or resources may enable a wide range of acceleration capabilities to be virtualized and allocated as needed.
[0038] In an embodiment, the housekeeping unit may further be configured to represent analysis of applications as directed acyclic graphs (DAGs) that capture execution and data flow. Such representation enables a detailed understanding of application requirements and resource needs. To determine resource requirements, the e-hypervisor may implement mixed integer linear programming mechanisms or techniques that may enable optimal allocation of resources based on complex constraints and objectives. In an embodiment, the housekeeping unit may be configured to provide an internal interconnect bandwidth required for transfers between the host processing element and accelerators, between different accelerators,and between memory and host processing elements, enabling data flow efficiently within the system and supporting the overall performance and quality of service guarantees.
[0039] In an embodiment, the system described in the subject specification may comprise various software components, subsystems, modules, engines, frameworks, and layers that collectively enable the implementation of the methods, mechanisms, techniques, operations, functions, etc., described in the independent and dependent claims. The terms "software components" or "components," "software routines" or "routines," "software models" or "models," "software engines" or "engines," "software scripts" or "scripts," and "layers" are used interchangeably throughout this specification unless the context requires specific distinctions based on implementation details. The core of the system is a processing engine configured to execute computer-readable code stored in a memory module. This code consists of sequences of instructions that, when executed by a processor of a computing device, such as a special-purpose computer, a general-purpose computer, or a mobile device, perform designated operations within an integrated environment. The computing device may be configured to operate as a special-purpose computer by executing these instructions, thereby enhancing its technical capabilities and performance.
[0040] In an embodiment, the system's architecture may include multiple subsystems, each executing or implementing specific functions or operations. An input / output subsystem may manage data communication between the system and external entities, including components like data receivers, transmitters, and interface modules that handle various data formats and communication protocols. A data processing module may include routines and engines that process incoming data, including data validation routines, transformation engines, and aggregation models that prepare data for further analysis or storage. A storage engine may provide data persistence, using models and scripts to interact with databases or other storage mechanisms, ensuring data integrity and efficient retrieval. An analysis framework may include layers and components that perform computational analysis, which may include statistical models, machine learning engines, or simulation modules that derive insights from the processed data.
[0041] In an embodiment, a user interface layer provides components and scripts that render data and system status to users, including graphical interface modules, interaction routines, and visualization engines that enhance user experience.A resource management subsystem may optimize the allocation of computing resources, including components that monitor system performance, manage memory allocation, and modify processing loads dynamically. The execution of specific operations by these components, individually or in cooperation, collectively provides a robust platform that optimizes computing resource allocation. For example, the resource management subsystem works in tandem with the data processing module to allocate processing power where it's most needed, thereby improving overall system efficiency.
[0042] In an embodiment, the system's modular design allows for the reusability of software models, components, and routines. Modules can be updated or replaced without affecting the entire system, facilitating maintenance and scalability. For instance, the analysis framework can implement new algorithms by updating its models and engines without requiring changes to the input / output subsystem or storage engine. Furthermore, the system may include a security engine that implements authentication routines, encryption models, and monitoring scripts to protect data and system integrity. A networking module manages network communications using protocols and components designed for efficient and secure data transfer, and an API layer provides interfaces for external systems to interact with the system's functionalities through predefined scripts and routines.
[0043] In an embodiment, by executing the instructions stored in memory, the processor enables the computing device to perform as the special-purpose computer designed for the claimed methods, mechanisms, techniques, operations, functions, etc. This specialization improves the technical operation by streamlining processes, reducing computational overhead, and enhancing responsiveness. The system leverages a combination of interchangeable software components, subsystems, modules, engines, frameworks, and layers to implement the claimed methods. The architecture described in the subject specification not only improves the technical operation of the computing device but also provides a flexible and efficient platform for optimizing the allocation of computing resources.
[0044] FIG. 1 is a block diagram showing an environment 100, including a system for virtualizing hardware acceleration resources with guaranteed quality of service, according to an exemplary embodiment. In an embodiment, FIG. 1 shows an environment 100 that includes a system for virtualizing hardware acceleration resources with guaranteed quality of service for the execution of applications. Theenvironment 100 may include a host processing unit 102 (that may interchangeably be referred to a host server or host computing environment or host), including a general- purpose processor (GPP), which may be communicatively coupled with a hardware accelerator 104. The hardware accelerator 104 may include shared resources 104A. The host processing unit 102 may include multiple virtual machines (VMs) (not shown), each with a corresponding virtual hardware accelerator (VHA) driver. The hardware accelerator 102 may include multiple components (not shown) or subsystems, engines, modules, etc., that may execute operations independently or in cooperation to achieve a specific task. For instance, such components, subsystems, modules, engines, etc., may include a hardware accelerator manager (HAM), multiple virtual hardware accelerators (VHAs), a housekeeping unit, an e-hypervisor, and a shareable resource pool (e.g., shared resources 104 A).
[0045] In an embodiment, the shared resources 104 A may be communicatively coupled with both the host processing unit 102 and the hardware accelerator 104 via a shareable interconnect. For example, the shareable interconnect may be a high-speed communication interface, such as PCIe or SERDES, that allows data transfer between the host processing unit, the hardware accelerator, and the shared resources. In an embodiment, the shared resources 104 A may include or be represented by a diverse collection of hardware resources (e.g., also referred to as hardware computing resources, or computing resources) in the hardware accelerator 104 that may be dynamically allocated between the VHAs. For example, the shared resources 104 A may include memory, processing units, memory bandwidth, cache, interconnect bandwidth, etc. In an embodiment, the hardware accelerator 104 may be a specialized processing unit that may execute operations to offload specific computationally intensive tasks from the host processor.
[0046] In an embodiment, the following components may execute operations independently or in cooperation and are described herein. For instance, a Virtual Machine (VM) may correspond to a software-based emulation of a computer system that may provide the functionality of a physical computer. The VM may execute operations based on the computer architecture and functions of a real or hypothetical computer and may execute a full operating system. In an embodiment, an embedded hypervisor (e-hypervisor) may correspond to a specialized virtualization layer implemented directly within the hardware accelerator, either in software or hardware. The e-hypervisor may execute operations to manage and coordinate the creation,execution, and resource allocation of Virtual Hardware Accelerators (VHAs). In an embodiment, a virtual hardware accelerator (VHA) may correspond to a logical partition of a physical hardware accelerator, presenting itself as a dedicated accelerator to a specific VM or application. A hardware accelerator (HA) may correspond to a specialized hardware component that may be designed to perform certain types of computational tasks or operations that may be more efficient than general -purpose CPUs. A hardware accelerator manager (HAM) may correspond to a component that may execute operations for creating and managing VHAs, assigning address registers, and facilitating communication between the host system and VHAs.
[0047] In an embodiment, a housekeeping unit (HKU) may correspond to a component that may execute operations to translate application requests into hardware resource requirements, manage resource allocation, and monitor resource usage to maintain quality of service (QoS). The QoS may correspond to a measure of the overall performance of a service, particularly the performance seen by the users of the network. For instance, the QoS may correspond to the guaranteed level of performance provided to each VHA. The resource mapping tables may correspond to or include data structures that may keep track of the allocation and utilization of hardware resources among various VHAs. In an embodiment, a hardware resource manager may correspond to a component that may execute operations for low-level management and allocation of physical hardware resources based on instructions from the e-hypervisor and the HKU. The device configuration address registers may correspond to registers that may be used to point to memory spaces reserved for configuring underlying hardware partitions in VHAs.
[0048] In an embodiment, a memory management unit (MMU) may correspond to a component that may execute operations for handling memory accesses requested by the VHAs. The MMU may translate virtual memory addresses to physical addresses and enforce memory protection between different VHAs. A direct memory access (DMA) controller may correspond to the execution of operations that may enable hardware subsystems within the accelerator to access system memory independently of the central processing unit (CPU). A single root I / O virtualization (SR-IOV) may correspond to a specification that may enable a PCIe device to appear to be multiple separate physical PCIe devices, facilitating the efficient sharing of PCIe devices in a virtual environment.
[0049] In an embodiment, a register may correspond to a fast storage location within the hardware accelerator or its management components that may be used to hold specific data or control information essential for the operation and virtualization of the hardware acceleration resources. In an embodiment, address registers may store memory addresses that point to specific locations in the system's memory space. The address registers may be used to define the boundaries of memory areas allocated to each VHA, ensuring proper isolation and access control. The device configuration address registers may correspond to registers that may point to memory spaces reserved for configuring the underlying hardware partitions of each VHA. The device configuration address registers may enable the separation of control and data paths in the virtualized environment.
[0050] In an embodiment, base address registers (BARs) may be used in PCIe devices that may update or provide information to the system about the amount of address space that the device may use. For instance, BARs may be assigned to each VHA to define its memory and I / O spaces. The control registers may store configuration data that may define the operational parameters of each VHA, such as resource allocation limits, QoS parameters, or operational modes. The status registers may store real-time information about the state of each VHA, including current resource utilization, error conditions, or performance metrics.
[0051] In an embodiment, the translation registers may be used by the MMU, which may store the rules for translating virtual addresses used by VHAs to physical addresses in the hardware accelerator's memory. The HAM and the e-Hypervisor may write to or read from these registers to configure VHAs, monitor their status, and manage resource allocation. The ability to assign and manage these registers independently for each VHA is a key aspect of the virtualization process described in this invention, allowing for fine-grained control and isolation of virtualized hardware acceleration resources.
[0052] In operation, the specific operations or functions may facilitate virtualizing hardware acceleration resources with guaranteed QoS for the execution of the applications. In an embodiment, when a virtual machine requests the execution of an application, the housekeeping unit may translate such requests into specific hardware resource requirements. The e-hypervisor may execute operations to allocate the necessary resources from the shareable resource pool to the corresponding VHA. The memory management unit may execute operations to provide isolation of addressspaces between different VHAs. The e-hypervisor may automatically and continuously monitor resource usage and enforce quality of service parameters by dynamically monitoring and modifying resource allocations as needed. The environment 100, including the components or subsystems, engines, modules, etc., described in the subject specification may enable efficient sharing of hardware resources among multiple virtual machines while maintaining performance isolation and guaranteeing quality of service for each virtual machine.
[0053] FIG. 2 is a block diagram showing an architecture of a host processing unit, according to an exemplary embodiment. FIG. 2 is described in conjunction with FIG. 1. FIG. 2 is a block diagram showing the architecture 200 of a host processing unit 202. In an embodiment, the host processing unit 202 may be implemented as the GPP (e.g., represented as the host processing unit (GPP) 202). The host processing unit (GPP) 202 may execute operations to run the virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n) and the corresponding VHA drivers (e.g., 206 A, 206B, 206n). Each virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n) may represent an isolated computing environment, including its operating system and applications. The corresponding VHA drivers (e.g., 206 A, 206B, 206n) may facilitate communication between the virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n) and their corresponding virtual hardware accelerators.
[0054] In an embodiment, the host processing unit (GPP) 202 may embody an implementation of a central processing unit (CPU) that may be designed to handle a wide variety of computing tasks. The host processing unit (GPP) 202 may be referenced as "general purpose" because it may execute different types of applications and perform various functions, unlike specialized processors that are optimized for specific tasks. For instance, the host processing unit (GPP) 202 may execute operations for: running multiple virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n), executing the VHA drivers (e.g., 206A, 206B, 206n) for each VM (e.g., VM1 204A, VM2 204B, VMn 204n), managing the overall operations, and coordinating with the hardware accelerator. In an embodiment, the host processing unit (GPP) 202 may work or execute operations in conjunction with the hardware accelerator to offload specific computationally intensive tasks, improving overall system performance and efficiency. The host processing unit (GPP) 202 may execute operations to manage the higher-level operations and general computing tasks, while the hardware accelerator handles specialized, resource-intensive operations.
[0055] Referring to FIG. 2, the architecture 200 for the host processing unit (GPP) 202 may be designed to manage virtualized hardware acceleration resources. FIG. 2 shows that the host processing unit (GPP) 202 may include several components, subsystems, modules, engines, etc., that are embodied in the hypervisor 208. For instance, the hypervisor 208 may be communicatively coupled with the multiple virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n), each with the corresponding VHA driver (e.g., 206A, 206B, 206n), a hardware accelerator (HA) driver 210, a direct memory access (DMA) controller 212, a memory management unit (MMU) 214, a resource allocation manager 216, a quality of service (QoS) monitoring module 218, an interrupt controller 220, and a configuration management module 222.
[0056] In an embodiment, the host processing unit (GPP) 202 may represent and execute operations as the primary computational platform. The host processing unit (GPP) 202 may include the hypervisor 208, which may execute operations to manage and coordinate the multiple virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n) and their associated resources. The hypervisor 208 may provide a layer of abstraction between the physical hardware (e.g., 104A) and the virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n), enabling efficient resource sharing and isolation. In an embodiment, multiple virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n) may be implemented in the host processing unit (GPP) 202. Each virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n) may represent an isolated computing environment, including its operating system and applications. In an embodiment, each virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n) may be associated with a corresponding VHA driver (e.g., 206A, 205B, 206n). In an embodiment, the VHA drivers (e.g., 206 A, 205B, 206n) may be implemented to establish a direct correspondence between the VMs and the corresponding virtual hardware accelerator. Once such correspondence is established, the implementation may facilitate direct communication subsequently between the virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n) and its corresponding virtual hardware accelerator without using the VHA driver (e.g., 206A, 205B, 206n).
[0057] In an embodiment, the HA driver 210 may be a component that interfaces directly with the physical hardware accelerator. The HA driver 210 may execute operations to manage low-level communication and control between the host processing unit (GPP) 202 and the hardware accelerator 104. In an embodiment, theDMA controller 212 may execute operations to manage data transfers efficiently between the host memory and the hardware accelerator 104 without directly involving the CPU. This may enable high-speed data movement and reduce CPU overhead. In an embodiment, the MMU 214 may execute operations for managing and translating virtual memory addresses to physical memory addresses. The MMU 214 may further execute operations to maintain memory isolation between the virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n) and ensure efficient memory utilization.
[0058] In an embodiment, the resource allocation manager 216 may execute operations and be implemented to manage the distribution and management of system resources between the virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n) and their associated virtual hardware accelerators. The resource allocation manager 216 component may work closely with the hypervisor 208 to optimize resource utilization and maintain fair allocation. In an embodiment, the QoS monitoring module 218 may execute operations to continuously monitor the performance and resource usage of each virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n) and its corresponding virtual hardware accelerator. The QoS monitoring module 218 may execute operations to collect metrics and provide feedback to the resource allocation manager 216 and the hypervisor 208 to maintain the desired level of service for each virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n).
[0059] In an embodiment, the interrupt controller 220 may execute operations to manage and route hardware and software interrupts within the system. The interrupt controller 220 may ensure that interrupts are properly handled and directed to the appropriate virtual machines or system components. In an embodiment, the configuration management module 222 may execute operations for managing the configuration settings of the virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n), virtual hardware accelerators, and other system components. The configuration management module 222 may execute operations corresponding to tasks related to initializing and updating configuration parameters, ensuring consistency across the system.
[0060] In operation, the host processing unit (GPP) 202, in cooperation with the hypervisor 208, the HA driver 210, and the other subsystems, modules, components, engines, etc. (e.g., 212, 214, 216, 218, 220, and 222), and when deployed with the hardware accelerator may cooperatively enable efficient virtualization of hardware acceleration resources while maintaining performanceisolation and quality of service guarantees. In an embodiment, the hypervisor 208 may execute operations to create and manage multiple virtual machines (e.g., VM1 204A, VM2 204B, VMn 204n), each with its own VHA driver (e.g., 206A, 206B, 206n). When a virtual machine (e.g., VM1 204A, VM2 204B, VMn 204n) requires hardware acceleration, its VHA driver (e.g., 206A, 206B, 206n) may communicate the request directly with the established corresponding virtual hardware accelerator without knowing the underlying subsystems facilitating such a seamless execution. The resource allocation manager 216, in conjunction with the MMU 214, may allocate the necessary resources and set up appropriate memory mappings. The DMA controller 212 may be utilized for efficient data transfer between the host and the hardware accelerator.
[0061] In an embodiment, throughout the execution of the above-described operation, the QoS monitoring module 218 may continuously monitor performance metrics, providing feedback to the resource allocation manager 216 and hypervisor 208 for dynamic resource modification if needed. The interrupt controller 220 may manage any interrupts generated during the process, ensuring they are properly routed to the correct virtual machine. The configuration management module 222 may maintain and update system configurations as needed. The above-described architecture may enable the efficient sharing of hardware acceleration resources among multiple virtual machines while maintaining performance isolation and guaranteeing quality of service for each virtual machine.
[0062] FIG. 3 is a block diagram illustrating an architecture of a hardware accelerator, according to an exemplary embodiment. FIG. 3 is described in conjunction with FIGs. 1, and 2. FIG. 3 is an illustration showing an architecture 300 of the hardware accelerator 328. In an embodiment, the hardware accelerator 328, as shown in FIG. 3 may be designed for virtualizing hardware acceleration resources with guaranteed quality of service. The architecture of the hardware accelerator 328 is shown in FIG. 3 may include components, subsystems, modules, engines, etc. For instance, the components, subsystems, modules, engines, etc., may include an e- hypervisor 302, a housekeeping unit 304, a hardware accelerator manager (HAM) 306, multiple virtual hardware accelerators (VHAs) (e.g., VHA1 308A, VHA2 308B, VHAn 308n), resource mapping tables 310, a hardware resource manager 312, access / usage barriers (e.g., 322, 324, 326) between resources allocated (e.g., 314, 316, 318) to different VHAs (e.g., 308A, 308B, 308n), and shared resources represented by320. In an embodiment, the access / usage barriers (e.g., 322, 324, 326) may be implemented between resources allocated to different VHAs (e.g., 308A, 308B, 308n) to maintain isolation and security.
[0063] In an embodiment, the e-hypervisor 302 may execute operations to provide functions or operations as the primary orchestrator of the virtualization process within the hardware accelerator 328. The e-hypervisor 302 may further execute operations to manage and coordinate between the multiple virtual hardware accelerators (e.g., 308A, 308B, 308n) and the associated resources (314, 316, 318, etc.). The e-hypervisor 302 may provide a layer of abstraction between the physical hardware resources (e.g., 314, 316, 318) and the virtual instances, thereby enabling efficient resource sharing and isolation. In an embodiment, the housekeeping unit 304 may work in conjunction with the e-hypervisor 302 and execute multiple functions or operations. For instance, the housekeeping unit 304 may execute operations to: translate the application requests into specific hardware resource requirements, determine and calculate resource availability, and execute resource allocation decisions based on such determinations and calculations. The housekeeping unit 304, independently or in cooperation with the QoS monitoring module 218, may execute operations for monitoring resource utilization and enforcing quality of service parameters for each virtual hardware accelerator (e.g., 308 A, 308B, 308n).
[0064] In an embodiment, the HAM 306 may execute operations to function as an intermediary between the host system (e.g., host processing unit (GPP) 202) and the virtual hardware accelerators (e.g., 308A, 308B, 308n). The HAM 306 may execute operations for creating and managing the VHAs (e.g., 308A, 308B, 308n), assigning address registers to each VHA (e.g., 308A, 308B, 308n), and handling the configuration of device-specific parameters. The HAM 306 may further execute operations to notify the VHAs (e.g., 308 A, 308B, 308n) of successful resource allocations and manage the communication between the host system and the VHAs (e.g., 308A, 308B, 308n). In an embodiment, FIG. 3 shows VHAs (e.g., 308A, 308B, 308n), which may represent logical partitions of the physical hardware accelerator. Each VHA (e.g., 308 A, 308B, 308n) may represent a dedicated accelerator to its corresponding virtual machine or application, enabling efficient sharing of the hardware accelerator's resources among multiple users or processes.
[0065] In an embodiment, the resource mapping tables 310 may store data or information related to the allocation and utilization of hardware resources. Theresource mapping tables 310 may be utilized to keep track of the allocation and utilization of hardware resources among the various VHAs (e.g., 308A, 308B, 308n). The resource mapping tables 310 may further include data or information related to memory allocations, processing unit assignments, bandwidth allocations, and other relevant resource metrics. The housekeeping unit 304 may continuously update the resource mapping tables 310, and the e-hypervisor 302 may cooperatively work with the housekeeping unit 304 to determine or execute resource management decisions. In an embodiment, the hardware resource manager 312 may execute operations for the low-level management and allocation of physical hardware resources. The hardware resource manager 312 may execute operations in cooperation with the e-hypervisor 302 and housekeeping unit 304 to guarantee that resources are efficiently distributed among the VHAs (e.g., 308A, 308B, 308n) while maintaining isolation and adhering to the QoS requirements. The elements or components 322, 324, and 326 may represent access / usage barriers between the resource allocations to the different VHAs (e.g., 308A, 308B, 308n). The reference numeral 320 may represent a shareable resource pool that may include memory, processing units, cache interconnect, etc.
[0066] In operation, the hardware accelerator 328, when deployed to work in conjunction with the host processing unit (GPP) 202, may enable efficient virtualization of hardware acceleration resources while maintaining performance isolation and quality of service guarantees. In an embodiment, when an application or virtual machine (e.g., 204A, 204B, 204n) requests hardware acceleration services, the request may be processed by the HAM 306 and forwarded to the housekeeping unit 304. The housekeeping unit 304 may execute operations to translate the application requirements into specific hardware resource requirements. The housekeeping unit 304 may reference and cooperatively work with the resource mapping tables 310 to determine the resource availability. The e-hypervisor 302 may execute operations to make allocation decisions based on the above-described information or data. The e- hypervisor 302 may execute operations in cooperation with the hardware resource manager 312 to partition and allocate the necessary resources to the appropriate VHA (e.g., 308A, 308B, 308n).
[0067] In an embodiment, throughout the above-described operation, housekeeping unit 304 may continuously monitor resource usage and update the resource mapping tables 310. Further, the housekeeping unit 304 may notify the e- hypervisor 302 of any deviations from allocated resources. Such notification mayenable dynamic modification of resource allocations to maintain quality of service guarantees. In an embodiment, the hardware accelerator 328 may implement template-based resource mapping, where pre-defined templates corresponding to specific application configurations may be used to streamline the resource allocation process. The hardware accelerator 328 may enable the implementation of mixed integer linear programming mechanisms or machine learning techniques to optimize resource partitioning in complex scenarios. In an embodiment, the above-described virtualization mechanism may enable the efficient sharing of heterogeneous hardware acceleration resources among multiple virtual machines or applications, providing guaranteed quality of service in diverse computing environments such as 5thgeneration (5G) radio access networks, AI / ML workloads, and other computationally intensive applications.
[0068] In an embodiment, template-based resource mapping may enable the efficient allocation of hardware resources to VHAs based on predefined configurations. For example, consider a scenario where the hardware accelerator is being used to support 5G radio access network (RAN) functions. The system may define a set of templates that correspond to different RAN configurations. Each template may specify a combination of application requirements and the corresponding hardware resource allocations.
[0069] For example, consider below a template for a mid-band 5G RAN configuration:
[0070] {
[0071] Resource template" : {
[0072] "id": "5G_Mid_Band_40MHz",
[0073] "applications": [
[0074] {
[0075] "app name": "5G_Mid",
[0076] "app_properties": {
[0077] "bandwidth": "40MHz",
[0078] "numerology": " 1",
[0079] num tx antennas" : "4",
[0080] num_rx_antennas": "4",
[0081] num component carriers" : " 1 "
[0082] },
[0083] "app_cardinality": "2"
[0084] }
[0085] ],
[0086] hwa_config": {
[0087] "cfg_name": "Mid_Band_Config",
[0088] "cfg_properties": {
[0089] num dsp cores" : "8",
[0090] num_fft_engines" : "4",
[0091] "memory": "256MB",
[0092] dma channels": " 16",
[0093] interconnect bandwidth" : " 1 OOGbps"
[0094] },
[0095] power_envelope": "25W"
[0096] }
[0097] }
[0098] }
[0099] In the above example, the template may include:
[0100] The application properties define the RAN configuration parameters.
[0101] The "app cardinality" indicates that this template supports two instances of this configuration.
[0102] The "hwa config" section specifies the hardware resources required to support this configuration.
[0103] When a VM requests a VHA for a 40MHz mid-band 5G RAN application, the following steps may be executed:
[0104] The VHA driver may communicate this request to the HAM.
[0105] The HAM may forward this request to the HKU.
[0106] The HKU may search its database of templates and identify the"5G_Mid_Band_40MHz" template as the best match for the requested configuration.
[0107] The HKU may then instruct the e-hypervisor to allocate resources according to the "hwa config" specified in the template.
[0108] The e-hypervisor may update the resource mapping tables to reflect this allocation, ensuring that 8 DSP cores, 4 FFT engines, 256MB of memory, 16 DMA channels, and 1 OOGbps of interconnect bandwidth are reserved for this VHA.
[0109] The MMU may be programmed to enforce the memory allocation and isolation for this VHA.
[0110] The HAM may then confirm with the VHA driver that the resources have been allocated and that the VHA is ready for use.
[0111] In an embodiment, the above-described template-based mechanism may enable quick and efficient resource allocation, as the system may not need to calculate resource requirements for each request. The template-based mechanism may also ensure consistency in resource allocation for similar application configurations. The system may store multiple such templates for various configurations (e.g., low- band, high-band, different bandwidths, different MIMO configurations) and for different types of applications beyond 5G RAN (e.g., AI / ML workloads, video processing). The HKU may select the most appropriate template based on the specific request or even combine multiple templates for more complex configurations. Template-based resource mapping may be useful in scenarios where the hardware accelerator is frequently reconfigured to support different types of workloads, allowing for rapid and efficient adaptation to changing application requirements while maintaining QoS guarantees.
[0112] FIG. 4 is a block diagram that illustrates an environment 400 for deploying a host processing unit and a hardware accelerator, according to an exemplary embodiment. FIG. 4 is described in conjunction with FIGs 1, 2, and 3. In an embodiment, FIG. 4 shows an environment 400 that may represent a system architecture for virtualized hardware acceleration within a, for example, cloud-based Radio Access Network (RAN). The host processing unit (GPP) 402 may include multiple subsystems, components, modules, and engines, including multiple VMs (e.g., VM1 404A, VM2 404B, VMn 404n), corresponding VHA drivers (e.g., 406A, 406B, 406n), a hypervisor 408, and the HA driver 410. The hardware accelerator 436 may include, for example, a HAM 414, multiple VHAs (e.g., VHA1 416A, VHA2 416B, VHAn 416n), an e-hypervisor 438, a housekeeping unit 412, resource mapping tables 418, a hardware resource manager 420, access / usage barriers (e.g., 430, 432, 434) between resources allocated (e.g., 422, 424, 426) to different VHAs (e.g., 416A, 416B, 416n), and shared resources represented by 428.
[0113] In an embodiment, the host processing unit (GPP) 402 may execute operations as the primary computational platform, hosting multiple virtual machines (e.g., 404A, 404B, 404n). Each VM (e.g., 404A, 404B, 404n) includes acorresponding VHA driver (e.g., 406 A, 406B, 406n) that may enable communication and control with the virtualized hardware acceleration resources. The HA driver 410 may provide a direct interface between the host processing unit and the hardware acceleration components.
[0114] In an embodiment, the HAM 414 may provide operations or functions as an intermediary between the VMs (e.g., 404A, 404B, 404n) and the virtual hardware accelerators (e.g., VHA1 416A, VHA2 416B, ..., VHAn 416n). The HAM 414 may execute operations to manage the allocation and deallocation of hardware acceleration resources to the VMs (e.g., 404A, 404B, 404n), ensuring efficient utilization and isolation among different virtual environments. The HAM 414 may continuously monitor overall system load and resource usage to reallocate resources based on real-time conditions. In an embodiment, the e-hypervisor 438 may manage the overall virtualization environment, coordinating interactions between the host processing unit (GPP) 402, VMs (e.g., 404A, 404B, 404n), and hardware acceleration resources. Working in conjunction with housekeeping unit 412, the e-hypervisor 438 may execute operations to maintain system stability and enforce resource allocation policies. The e-hypervisor 438 may execute operations to create virtualized hardware accelerators (e.g., 416A, 416B, 416n), enabling each VM (e.g., 404A, 404B, 404n) to access necessary computational resources via its corresponding VHA driver (e.g., 406A, 406B, 406n).
[0115] In an embodiment, the housekeeping unit 412 may monitor and manage the system's resources. The housekeeping unit 412 may utilize the resource mapping tables 418 to track the allocation of hardware resources (e.g., 422, 424, 426) to various VHAs (e.g., 416A, 416B, 416n) and VMs (e.g., 404A, 404B, 404n). The hardware resource manager 420 may work in conjunction with housekeeping unit 412 to dynamically allocate and optimize resource usage based on current system demands and predefined policies. The housekeeping unit 412 may execute operations to monitor the interconnect bandwidth and manage provisioning according to workload requirements to ensure sufficient DMA channels and other interconnect resources support high-throughput data transfer. In an embodiment, the VHAs (e.g., 416A, 416B, 416n) may represent virtualized instances of hardware acceleration resources. These VHAs (e.g., 416A, 416B, 416n) may be dynamically created, modified, or destroyed based on the requirements of the virtual machines and overall system load. The system's architecture allows seamless integration of new VHAs(e.g., 416A, 416B, 416n) as the infrastructure scales, enabling futureproofing and adaptability to technological advancements.
[0116] In an embodiment, advanced scheduling algorithms may be implemented to optimize the allocation of hardware acceleration resources. For example, the Lowest Immediate Follower Exploration (LIFE) algorithm and its variants, such as connected-LIFE and wrap-LIFE, may be implemented to schedule tasks across available processing cores efficiently. The above-described algorithms consider factors like task dependencies, execution times, and timing constraints to generate optimal schedules. In an embodiment, the resource mapping tables 418 may include or store detailed information about available hardware resources (e.g., 422, 424, 426) and the current allocation status. The housekeeping unit 412 and the hardware resource manager 420 may continuously update the resource mapping tables 418. In an embodiment, the resource mapping tables 418 may reflect the dynamic nature of resource utilization in the virtualized environment. The resource mapping tables 418 may store information or data related to the hardware components that may be allocated to each VHA (e.g., 416A, 416B, 416n), ensuring that requests from VMs (e.g., 404A, 404B, 404n) are translated into concrete resource assignments.
[0117] In an embodiment, the MMU (e.g., 214) may be controlled by the housekeeping unit 412 and may provide robust memory isolation for each VHA (e.g., 416A, 416B, 416n). Such memory isolation may prevent cross-contamination between memory spaces allocated to different VMs (e.g., 404 A, 404B, 404n), ensuring that each VM (e.g., 404A, 404B, 404n) operates within its designated boundaries. In an embodiment, the system may also include artificial intelligence (Al) or machine learning components that leverage predictive models to forecast resource usage trends based on historical data. These predictions enable the HAM 414 to preemptively modify resource allocations, optimizing for performance, power consumption, or other metrics depending on the system's objectives.
[0118] In operation, when a VM (e.g., 404A, 404B, 404n) demands hardware acceleration resources, the corresponding VHA driver (e.g., 406A, 406B, 406n) may send a request to the HAM 414. The HAM 414 may communicate with the e- hypervisor 412 and the housekeeping unit 412 to determine resource availability. The housekeeping unit 412 may cooperatively work or execute operations with the resource mapping tables 418 and the hardware resource manager 420 to identify suitable hardware acceleration resources. Once resources are allocated, the e-hypervisor 438 creates or modifies a VHA instance (e.g., 416A, 416B, 416n) to serve the VM's (e.g., 404A, 404B, 404n) request. The HAM 414 may subsequently communicate with the VM (e.g., 404 A, 404B, 404n) and the corresponding assigned VHA (e.g., 416A, 416B, 416n), enabling hardware acceleration tasks to be executed efficiently and isolated from other VMs (e.g., 406A, 406B, 406n).
[0119] In an embodiment, throughout the execution of the above-described operations or functions, the system may implement advanced scheduling algorithms to optimize task execution and resource utilization across all VHAs (e.g., 416A, 416B, 416n) and processing cores. The housekeeping unit 412 may continuously monitor resource usage and performance metrics, modifying allocations as needed to maintain system efficiency and meet the QoS requirements for each VM (e.g., 404 A, 404B, 404n). If a VM (e.g., 404A, 404B, 404n) exceeds the corresponding allocated resources or fails to meet its QoS targets, the housekeeping unit 412 and the HAM 414 may work in cooperation to send notifications to the e-hypervisor 438 and the VMs (e.g., 404A, 404B, 404n), thereby enabling or implementing corrective actions. In an embodiment, the dynamic and flexible architecture described in FIG. 4 enables efficient sharing of hardware acceleration resources in cloud-based RAN environments. The above-described mechanism, when potentially implemented, improves overall system performance and resource utilization while maintaining isolation between different virtual machines and their workloads. The abovedescribed mechanism may be designed with a focus on modularity, enabling seamless scalability and adaptability in various high-performance computing environments.
[0120] In an embodiment, the above-described mechanism and the corresponding deployment in FIG. 4 may be particularly well -suited for telecommunications applications, especially in supporting the deployment of 5G Radio Access Networks (RAN). The above-described mechanism may provide the ability to handle high-throughput, low-latency workloads while maintaining strict QoS guarantees, making it ideal for processing tasks such as multiple input multiple output (MIMO) applications, beamforming, and other advanced 5G operations. Moreover, the system's modularity and scalability allow it to handle increasing demands for network bandwidth and processing power as 5G networks continue to evolve.
[0121] Additionally, the above-described mechanism and system implementing the mechanism may be useful in industries such as artificial intelligenceand machine learning (AI / ML), cloud computing, and high-performance computing (HPC), where multiple virtualized applications need to share specialized hardware resources efficiently. The combination of dynamic resource allocation, robust isolation mechanisms, and proactive monitoring ensures that the system can meet the needs of demanding workloads while remaining adaptable to future technological advancements.
[0122] FIG. 5 is an illustration of a flow diagram of process 500 for virtualizing hardware acceleration resources with guaranteed quality of service, according to an exemplary embodiment. FIG. 5 is described in conjunction with FIGs. 1, 2, 3 and 4. In an embodiment, the process 500 for virtualizing hardware acceleration resources with guaranteed quality of service may be implemented using a system of a host server, including a central processing unit (CPU), memory, and storage that may be communicatively coupled to a hardware accelerator. This hardware accelerator may include specialized processors such as Graphics Processing Units (GPUs), Tensor Processing Units (TPUs), Al processors, or Neural Processing Units (NPUs). The system may be designed with specific components, subsystems, modules, and engines as described in FIGs. 1, 2, 3, and 4 to manage the virtualization and resource allocation processes effectively.
[0123] In an embodiment, process 500 or a mechanism for virtualizing hardware accelerator resources with guaranteed quality of service may implement the following steps: At step 505, multiple virtual hardware accelerators on a hardware accelerator unit may be created. For instance, the hardware accelerator manager may be configured to execute operations for initializing and configuring the virtual environment on the hardware accelerator unit. The hardware accelerator manager creates multiple virtual hardware accelerators on the physical hardware accelerator unit. This process involves partitioning the physical resources of the hardware accelerator, such as compute cores, memory, and bandwidth, into logical units that can be independently allocated and managed. Each virtual hardware accelerator is assigned a unique identifier within the system. At step 510, multiple address registers, including device configuration address registers, are assigned to each virtual hardware accelerator. In an embodiment, the hardware accelerator manager may execute operations to assign a set of address registers, including device configuration address registers, to each virtual hardware accelerator. The address registers, including device configuration address registers, may be mapped to specific memory locations andstore data or information such as the virtual accelerator's resource allocation, priority settings, and quality of service parameters. This step ensures that each virtual accelerator has its own isolated configuration space, laying the groundwork for secure and efficient operation.
[0124] At step 515, one or more requests for executing applications are received. In an embodiment, requests are received for the multiple virtual hardware accelerators. The housekeeping unit may execute operations to coordinate resource management and request processing. The housekeeping unit may receive requests to execute applications on the virtual hardware accelerators. Such requests may typically originate from user applications running on the host server and may be communicated to the virtual accelerators through APIs provided by the virtualization layer. At step 520, the received requests are translated into hardware accelerator resource requirements. In an embodiment, the housekeeping unit translates the received requests into specific hardware accelerator resource requirements. This translation process involves analyzing the computational demands of the requested applications and mapping them to the capabilities of the hardware accelerator. For instance, a deep learning training mechanism may be translated into requirements for a certain number of Al processor cores, a specific amount of high-bandwidth memory, and a minimum data transfer rate.
[0125] At step 525, multiple hardware accelerator resources are partitioned and allocated to the virtual hardware accelerators. In an embodiment, based on these translated resource requirements, the housekeeping unit may partition and allocate the hardware accelerator resources to the virtual accelerators. The allocation process may be dynamic and may factor in information or data such as the current state of resource utilization, the priority of each virtual accelerator, and the overall system load, ensuring efficient use of resources across multiple virtual accelerators. For example, in a cloud computing environment, multiple users may submit machine learning workloads that require acceleration. The housekeeping unit collects these requests and determines the computational power, memory bandwidth, and latency needs of each application based on its characteristics, such as model complexity or dataset size.
[0126] At step 530, an isolation of multiple address spaces between the multiple hardware accelerators is provided. In an embodiment, the memory management unit may execute operations to maintain the security and isolation of the virtual accelerators. The memory management unit may implement memorymanagement techniques such as address translation and access control lists to isolate the address spaces between the virtual hardware accelerators. The memory management unit may execute operations such that each virtual accelerator operates within its isolated memory space, preventing unauthorized access to data or resources belonging to other virtual accelerators. Such isolation may maintain integrity and provide security for the execution of operations performed by different virtual accelerators, particularly in multi-tenant environments where different users or applications may be sharing the same physical hardware. In an embodiment, the memory management unit maps virtual addresses used by applications to physical addresses in the hardware accelerator's memory. By isolating the address spaces, the memory management unit ensures that one virtual accelerator cannot access or modify the memory allocated to another. For example, if virtual accelerator A is processing sensitive data, virtual accelerator B cannot access that data due to the enforced memory isolation, thus maintaining data privacy and security.
[0127] At step 535, the hardware accelerator resources for usage of each of the virtual hardware accelerators are monitored. In an embodiment, an embedded hypervisor (also referred to as e-hypervisor) executing or running on the housekeeping unit manages the ongoing operation and performance of the virtual accelerators. The e-hypervisor may continuously monitor the hardware accelerator resources for usage by each of the virtual accelerators. Such monitoring may include collecting real-time data on resource utilization, including metrics such as compute core usage, memory consumption, and I / O bandwidth. At step 540, the monitored hardware accelerator resource usage with the hardware accelerator resources allocated for each of the virtual hardware accelerators is compared. In an embodiment, the e- hypervisor may execute operations to compare such monitored resource usage with the allocated resources for each virtual accelerator. Such comparison may be performed regularly or triggered by specific events, such as the completion of a task or a significant change in resource utilization.
[0128] At step 545, multiple quality of service parameters for each virtual hardware accelerator are enforced. In an embodiment, based on these comparisons, the e-hypervisor may enforce the QoS parameters for each virtual accelerator. The e- hypervisor may execute operations to modify the allocation of hardware accelerator resources dynamically. For example, if a virtual accelerator is consistently underutilizing its allocated resources, the e-hypervisor may reassign some of thoseresources to other virtual accelerators that require additional capacity. Conversely, if a virtual accelerator is experiencing performance bottlenecks due to resource constraints, it may be allocated additional resources from a pool of available capacity. In an embodiment, the operational efficacies of the process 500 steps, namely 505, 510, 515, 520, 525, 530, 535, 540, and 545, are implemented by the corresponding components, subsystems, engines, and modules as described in the detailed description of FIGs. 1, 2, 3 and 4.
[0129] In an embodiment, throughout the above-described dynamic resource modification process, the e-hypervisor may work in conjunction with the memory management unit to maintain the isolation of address spaces between virtual accelerators. Such cooperative working or execution of operations ensures that the reallocation of resources does not compromise the security or integrity of data belonging to different virtual accelerators. Even as resources are reallocated, the memory management unit updates memory mappings to ensure each virtual accelerator continues to operate within its designated memory space.
[0130] In an embodiment, process 500 may be implemented by the abovedescribed system for virtualizing hardware acceleration resources with guaranteed quality of service, which may be implemented using a combination of hardware and software components. The system may comprise a memory, a processor configured to execute instructions stored in the memory, a host processing unit, and a hardware accelerator coupled to a host server.
[0131] In a practical example, the system may be used to virtualize a hardware acceleration unit for 5G baseband processing connected to the host server via PCIe. The hardware accelerator may contain various processing elements (e.g., also referred to as hardware resources or computing elements or hardware computing devices) such as DSPs, FPGAs, and dedicated ASICs for specific baseband functions. Multiple virtual hardware accelerators may be created on this unit, each assigned to handle baseband processing for a different network slice or mobile network operator.
[0132] When a request is received to process a 5G Mid-band signal with 40MHz bandwidth, mu=2, 4x2 MIMO configuration, and one component carrier, the housekeeping unit may translate the request into specific resource requirements. The housekeeping unit may use a resource template. For instance, it may be determined that this request requires 8 signal processing cores, 2 FEC accelerators, 16MB of memory, and 5GBps of interconnect bandwidth. The housekeeping unit may thenallocate these resources to a virtual hardware accelerator, assigning specific PCIe base address registers for memory access and configuring the necessary interconnects.
[0133] The e-hypervisor may continuously monitor the resource usage of this virtual accelerator. Suppose the e-hypervisor detects that the allocated resources are being underutilized. In that case, it may dynamically modify the allocation, thereby reducing the memory allocation to 12MB or decreasing the interconnect bandwidth to 4GBps. Conversely, suppose the resources are being overtaxed, leading to potential quality of service violations. In that case, the e-hypervisor may allocate additional resources, such as extra signal processing cores or increased memory, to maintain the required performance levels.
[0134] Throughout this process, the memory management unit ensures that each virtual accelerator can only access its assigned memory spaces, preventing any unauthorized access to memory allocated to other virtual accelerators. This isolation may enable security to be maintained and prevent interference between different network slices or operators sharing the same physical hardware. In an embodiment, the system may also implement more advanced scheduling techniques to optimize resource utilization. For instance, the Lowest Immediate Follower Exploration (LIFE) algorithm or its variants like connected-LIFE or wrap-LIFE may be used to schedule tasks across the available processing elements. These algorithms may consider factors such as task dependencies, execution times, and resource availability to create efficient schedules that maximize hardware utilization while meeting timing constraints.
[0135] For example, where the hardware accelerator includes heterogeneous computing resources, such as a mix of general -purpose processors, DSPs, and FPGAs, the housekeeping unit may represent the analysis of applications as DAGs. These DAGs may capture the execution and data flow of the applications, enabling resource allocation and scheduling decisions. For example, a DAG for a 5G baseband processing application might include nodes for channel estimation, MIMO detection, and forward error correction, with edges representing the data dependencies between these tasks. The e-hypervisor may implement a mixed integer linear programming mechanism to determine resource requirements and optimal allocation strategies. This approach may allow for more precise and efficient resource allocation, particularly in complex scenarios with multiple competing demands on the hardware resources.
[0136] To handle the communication needs between accelerators and between memory and accelerators, the housekeeping unit may provide internal interconnect bandwidth allocation. This may involve managing DMA channels or configuring on- chip network resources to ensure adequate data transfer capabilities for each virtual accelerator. For instance, it may ensure that the required number of DMA channels are reserved upfront to enable successful data transfer amongst the various virtual hardware accelerators.
[0137] The system may also include mechanisms for notifying the host processing unit of potential violations of quality of service limits. This notification may be implemented through the HAM using internal communication mechanisms such as mailboxes. This may allow for higher-level management decisions, such as migrating workloads between different hardware units or modifying application priorities.
[0138] By implementing these features and techniques, the system may provide a flexible and efficient means of virtualizing hardware acceleration resources while guaranteeing quality of service for diverse applications. This approach may enable more effective utilization of hardware resources in scenarios such as 5G network deployments, where multiple network slices or operators may need to share common hardware infrastructure while maintaining isolation and performance guarantees.
[0139] In an embodiment, a Field-Programmable Gate Array (FPGA) may, for example, be implemented as a programmable hardware accelerator. FPGAs consist of an array of programmable logic blocks and reconfigurable interconnects that can be reprogrammed to perform various complex functions. In the context of this invention, an FPGA-based accelerator may be virtualized to support multiple applications or virtual machines simultaneously. For instance, a single physical FPGA may be partitioned into multiple VHAs, each configured to accelerate specific tasks, such as signal processing for 5G radio access networks (RAN), neural network inference for AI / ML workloads, video encoding / decoding for streaming applications, cryptographic operations for security-intensive tasks. The system and method described may dynamically allocate FPGA resources (e.g., logic elements, DSP blocks, memory, etc.) to different VHAs based on application requirements and QoS parameters. The programmability of FPGAs enables real-time reconfiguration, enabling the system to adapt to changing workload demands efficiently.
[0140] In an embodiment, an Application-Specific Integrated Circuit (ASIC) designed for a particular function or operation may, for example, be implemented as a non-programmable hardware accelerator. For instance, a Tensor Processing Unit (TPU) ASIC may be designed specifically for neural network computations in machine learning applications. While the TPU itself is not programmable in the same way as an FPGA, its resources may still be virtualized and shared among multiple applications or virtual machines. For instance, the system and method described above may partition the TPU's computational resources, such as matrix multiplication units (MXUs), unified buffer memory, accumulator units, and VO bandwidth. In an embodiment, the system and method described above may create multiple VHAs, each with a guaranteed portion of these resources. For example, one VHA might be allocated 30% of the MXUs and memory for a real-time image recognition application, while another VHA may be given 50% for a large-scale natural language processing task. Although the underlying hardware is not reprogrammable, the abovedescribed system and method may provide QoS guarantees by managing resource allocation, scheduling compute operations, and monitoring usage to ensure each VHA receives the allocated or assigned share of the TPU's capabilities or functions.
[0141] In an embodiment, a system for virtualizing hardware acceleration resources with guaranteed quality of service utilizes a host processing unit that functions as the primary control interface for the virtualization system. For instance, in a data center environment with an NVIDIA Al 00 GPU serving as the hardware accelerator, the system partitions this physical GPU into three virtual GPUs, each allocated to different machine learning workloads. The hardware accelerator executes fundamental operations by implementing dedicated circuitry and processing elements, where the Al 00 GPU's CUDA cores and tensor cores are partitioned among the virtual instances. The system creates these virtual hardware accelerators through a virtualization layer that segments the physical 40GB GPU memory and 19.5 TFLOPS compute capacity into isolated portions.
[0142] Each virtual GPU receives specific address registers and device configuration registers, allowing one virtual GPU to run a natural language processing model using 15GB memory, another to perform image recognition using 15GB memory, and a third to handle real-time video processing using 10GB memory. The housekeeping unit processes incoming requests, such as when a client submits a batchof 1000 images for classification. These requests are translated into specific hardware requirements, for example, the image classification task is analyzed to require 5000 CUDA cores and 8GB memory for optimal performance. The system implements resource partitioning by allocating compute resources based on workload priorities, ensuring the natural language processing model receives 40% of compute resources, image recognition 40%, and video processing 20%.
[0143] The memory management unit creates separate memory spaces for each virtual GPU, preventing the image classification workload from accessing memory allocated to the language model. The embedded hypervisor monitors resource usage, detecting when the image classification task exceeds its allocated memory threshold of 15GB. It compares actual usage against allocations - for instance, if the language model only uses 30% of its allocated computing resources while the image classification requires more, the hypervisor dynamically adjusts the allocation. Quality of service parameters are enforced by ensuring the video processing workload maintains its required 30 frames per second processing rate. The system provides an internal bandwidth of 600 GB / s for data transfer between components. Resource templates define standard configurations - for example, a "high-compute" template allocating 8000 CUDA cores and 12GB memory for complex neural network training. The system uses directed acyclic graphs to represent workflows, such as mapping the sequence of convolution, pooling, and fully connected layers in a neural network. Mixed integer linear programming determines optimal resource allocation - for instance, calculating that splitting compute resources 40-40-20 maximizes overall throughput while meeting quality of service requirements. This implementation enables all three workloads to run concurrently with guaranteed performance levels, maintaining isolation while allowing dynamic resource reallocation based on realtime demands.
[0144] FIG. 6 is an illustration showing an exemplary hardware configuration of a special-purpose computer 600 that may be used to implement components for virtualizing hardware acceleration resources with guaranteed quality of service, according to exemplary embodiments. The special-purpose computer 600, shown in FIG. 6, includes CPU 605, including multi core processors, Al processors 610 including Graphics Processing Units (GPU) 610A, Field-Programmable Gate Arrays (FPGA) 610B, Application-Specific Integrated Circuits (ASIC) 610C, NeuralProcessing Units (NPU) 610D, Tensor Processing Units (TPU) 610E, system memory 615, network interface 620, hard disk drive (HDD) interface 625, external disk drive interface 630, and input / output (I / O) interfaces 635A, 635B, 635C. These abovedescribed components or hardware elements of the special-purpose computer 600 may be communicatively coupled to each other via a system bus 640. In an embodiment, the CPU 605 may execute operations related to arithmetic, logic, and / or control operations by accessing the system memory 615. The CPU 605 and the Al processors 610 may implement the interfaces, the processors, etc., of the exemplary devices and / or systems described above. The Al processors 610 may execute or perform operations for processing Al tasks. The Al processors 610 may include multiple types of processors that may be tailored or customized for implementing specific Al workloads.
[0145] In an embodiment, the GPU 610A may be designed for rendering graphics and may be highly effective for executing operations related to Al tasks due to their parallel processing capabilities. The GPU 610A may execute multiple calculations simultaneously and may be used for training Al models. The GPU 610A with Al-specific hardware (e.g., the ASIC 610C, the NPU 610D, the TPU 610E, etc.) may be integrated to accelerate the Al tasks. For example, tensor cores may be designed to speed up the training of neural networks and machine learning models by enhancing matrix multiplications that may be core functions of many Al algorithms. In an embodiment, the FPGA 610B may be reconfigurable processors that may be implemented for executing specific Al tasks, thereby offering flexibility and efficiency in real-time applications. The FPGA 610B may have a unique design that may include a series of interconnected and configurable logic blocks. The reprogrammability of the FPGA 610B enables a high level of customization, enabling it to implement a wide range of Al applications.
[0146] In an embodiment, the ASIC 610C may be custom-designed processors optimized for specific Al applications, providing high performance and energy efficiency. The ASIC 610C may be built or designed with the singular purpose of accelerating the Al workloads. The ASIC 610C is not reprogrammable like the FPGA 610B, but its specialized design enables significant improvements in speed and efficiency for tasks such as deep learning inference and training. In an embodiment, the NPU 610D may be designed to execute operations or functions related to specialized Al tasks, particularly those involving neural networks. The NPU 610Dmay be designed to accelerate neural network computations, often integrated into CPUs or Systems on Chips (SoCs). The NPU 610D may execute operations to process large volumes of data faster than other general -purpose processors and perform various Al tasks such as image recognition and natural language processing (NLP). The NPU 610D may be used in mobile devices and personal computers to implement the execution of Al tasks without significantly increasing power consumption.
[0147] In an embodiment, the TPU 610E may be designed specifically for accelerating machine learning workloads, particularly those involving tensor computations. The TPU 610E may be used in data centers to power various Al services such as search algorithms and language translation. The TPU 610E processor architecture may be optimized for high throughput and low latency that may facilitate the implementation of large-scale Al applications. In an embodiment, the Al processors 610, including the GPU 610A, the FPGA 61 OB, the ASIC 610C, the NPU 610D, and the TPU 610E, may execute operations or functions to perform multiple complex operations, computations, and calculations simultaneously, thereby enabling faster execution of the Al tasks compared to the sequential processing of general purpose CPUs. The Al processors 610 may be designed to execute the operations in a more energy-efficient way. For example, low-precision arithmetic can be used to reduce power consumption, thereby improving the performance of the implemented Al applications. In an embodiment, the Al processors 610 may be customized to implement specific Al models, machine learning models, machine learning engines, and Al applications, enabling optimized execution of operations or functions.
[0148] In an embodiment, the implementation of Al workloads or Al tasks using the Al processors 610 may provide technical advantages or technical benefits of parallel processing, reduced precision, hardware optimization, energy efficiency, and neural network support. The Al processors 610 may be designed with high parallelism, allowing them to execute multiple Al-related calculations simultaneously. The Al tasks or Al workloads, including matrix multiplication and vector operations, may be used in neural networks. The Al processors 610 may use reduced-precision arithmetic (e.g., 8-bit or 16-bit) to improve computational and power efficiency while maintaining acceptable accuracy levels for Al tasks. The Al processors610 may incorporate specialized hardware components such as multiply-accumulate (MAC) units and on-chip memory for executing specific Al workloads. The Al processors610 may be optimized for neural network tasks, including forward and backward passes during training and inference, and support various neural network architectures and frameworks
[0149] In an embodiment, the special-purpose computer 600 does not necessarily include Al processors 610, for example, in case the special -purpose computer 600 is used for implementing a device other than a central processing device. The system memory 615 may store information and / or instructions for use in combination with the CPU 605. The system memory 615 may include volatile and non-volatile memory, such as random-access memory (RAM) 645 and read-only memory (ROM) 650. A basic input / output system (BIOS) containing the basic routines that help to transfer information between elements within the special-purpose computer 600, such as during start-up, may be stored in the ROM 650. The system bus 640 may be any of several types of bus structures, including a memory bus or memory controller, a peripheral bus, and a local bus using any of a variety of bus architectures.
[0150] The special-purpose computer 600 may include the network interface 620 for communicating with other computers and / or devices via a network.
[0151] Further, the special-purpose computer 600 may include a hard disk drive (HDD) 655 for reading from and writing to a hard disk (not shown) and an external disk drive 660 for reading from or writing to a removable disk (not shown). The removable disk may be a magnetic disk for a magnetic disk drive or an optical disk such as a CD ROM for an optical disk drive. The HDD 655 and the external disk drive 660 are connected to the system bus 640 by the HDD interface 625 and the external disk drive interface 630, respectively. The drives and their associated non- transitory computer-readable media provide non-volatile storage of computer- readable instructions, data structures, program modules, and other data when the special-purpose computer operates as a general -purpose computer. The relevant data may be organized in a database, such as a relational or object database.
[0152] Although the exemplary environment described herein employs a hard disk (not shown) and an external disk (not shown), it should be appreciated by those skilled in the art that other types of computer-readable media which may store data that is accessible by a computer, such as magnetic cassettes, flash memory cards, digital video disks, random access memories, read-only memories, and the like, may also be used in the exemplary operating environment.
[0153] Several program modules may be stored on the hard disk, external disk, the ROM 650, or the RAM 645, including an operating system (not shown), one or more application programs 645A, other program modules (not shown), and program data 645B. The application programs may include at least some of the functionality described above.
[0154] The special-purpose computer 600 may be connected to input device 665, such as mouse and / or keyboard, and display device 670, such as liquid crystal display, via corresponding I / O interfaces 635 A to 635C and the system bus 640. In addition to an implementation using a special-purpose computer 600, as shown in FIG. 6, part or all of the functionality of the exemplary embodiments described herein may be implemented as one or more hardware circuits. Examples of such hardware circuits may include but are not limited to Large Scale Integration (LSI), Reduced Instruction Set Circuits (RISC), etc.
[0155] One or more embodiments are now described with reference to the drawings, wherein reference numerals refer to elements throughout. Numerous specific details are set forth in the following description to provide a thorough understanding of the various embodiments. It is evident, however, that the various embodiments may be practiced without these specific details (and without applying them to any networked environment or standard).
[0156] As used in this application, in some embodiments, the terms "component," "system," and the like are intended to refer to, or comprise, a computer- related entity or an entity related to an operational apparatus with one or more specific functionalities, wherein the entity may be either hardware, a combination of hardware and software, software, or software in execution. As an example, a component may be but is not limited to, a process running on a processor, a processor, an object, an executable, a thread of execution, computer-executable instructions, a program, and / or a computer. By way of illustration and not limitation, both an application running on a server and the server may be a component.
[0157] The above descriptions and illustrations of embodiments, including what is described in the Abstract, are not intended to be exhaustive or to limit one or more embodiments to the precise forms disclosed. While specific embodiments of, and examples for, one or more embodiments are described herein for illustrative purposes, various equivalent modifications are possible within the scope, as those skilled in the relevant art will recognize. These modifications may be madeconsidering the above-detailed description. Rather, the scope is to be determined by the following claims, which are to be interpreted in accordance with established doctrines of claim construction.
Claims
CLAIMSWe Claim:
1. A system (600) for virtualizing hardware acceleration resources with a guaranteed quality of service, comprising: a memory (615); a processor (605, 610) configured to execute instructions stored in the memory (615) to execute one or more operations; a host processing unit (102, 202); a hardware accelerator (104) executing operations to: create a plurality of virtual hardware accelerators (308 A, 308B, 308C) on a hardware accelerator unit, wherein the hardware accelerator unit (328) is coupled to a host server (102) or a hosting processing unit (102); assign a plurality of address registers, including device configuration address registers, to each virtual hardware accelerator (308A, 308B, 308C); a housekeeping unit (304) executing operations to: receive one or more requests for executing one or more applications on the hardware accelerator, wherein the one or more requests are received for the plurality of virtual hardware accelerators (308A, 308B, 308C); translate the received one or more requests into one or more hardware accelerator resource requirements; based on the translated one or more hardware accelerator resource requirements, partitioning and allocating one or more hardware accelerator resources (314, 316, 318) to the plurality of virtual hardware accelerators (308 A, 308B, 308C); a memory management unit (214) executing operations to: provide an isolation of a plurality of address spaces between the plurality of virtual hardware accelerators (308 A, 308B, 308C);an embedded hypervisor (302) running on the housekeeping unit (304) executing operations to: monitor by the one or more hardware accelerator resources (314, 316, 318) for the usage of each of the plurality of virtual hardware accelerators (308A, 308B, 308C); compare the monitored at least one hardware accelerator resource (314 or 316 or 318) usage with the allocated one or more hardware accelerator resources for each of the plurality of virtual hardware accelerators (308A, 308B, 308C); and enforce a plurality of quality of service parameters for each virtual hardware accelerator (308A, 308B, 308C) by: dynamically monitoring and modifying the allocation of the one or more hardware accelerator resources, based on the comparison of the monitored at least one hardware resource usage with the one or more hardware accelerator resources (314, 316, 318) allocated for each of the plurality of virtual hardware accelerators (308A, 308B, 308C); and maintaining the isolation of the plurality of address spaces between the plurality of virtual hardware accelerators (308A, 308B, 308C).
2. The system (600) as claimed in claim 1, wherein the hardware accelerator (328) further executes operations to: assign a plurality of device configuration address registers to each virtual hardware accelerator (308A or 308B or 308C) from the plurality of virtual hardware accelerators (308A, 308B, 308C).
3. The system (600) as claimed in claim 1, wherein the housekeeping unit (304) executes operations to: translate the one or more requests into hardware accelerator resource requirements using one or more resource templates.
4. The system (600) as claimed in claim 1, wherein the embedded hypervisor (304) executes operations to: notify the host processing unit of potential violations of quality of service limits.
5. The system (600) as claimed in claim 1, wherein the hardware accelerator (328) comprises: a plurality of heterogeneous hardware computing resources (314, 316, and 318).
6. The system (600) as claimed in claim 1, wherein the housekeeping unit (304) executes operations to: represent analysis of applications as directed acyclic graphs (DAGs) capturing execution and data flow.
7. The system (600) as claimed in claim 1, wherein the embedded hypervisor (302) executes operations to: implement a mixed integer linear programming mechanism to determine resource requirements.
8. The system (600) as claimed in claim 1, wherein the housekeeping unit (304) executes operations to: provide an internal interconnected bandwidth required for transfers between a plurality of host processing elements and a plurality of accelerators, between a plurality of accelerators, and between memory and the plurality of host processing elements.
9. A method (500) for virtualizing hardware acceleration resources with a guaranteed quality of service, comprising: creating (505) a plurality of virtual hardware accelerators (308A, 308B, 308C) on a hardware accelerator, wherein the hardware accelerator (104) is coupled to a host server (102) or a host processing unit (102); assigning (510) a plurality of address registers, including device configuration address registers, to each virtual hardware accelerator; receiving (515) one or more requests for executing one or more applications, wherein the one or more requests are received for the plurality of virtual hardware accelerators (308 A, 308B, 308C); translating (520) the received one or more requests into one or more hardware accelerator resource requirements; based on the translated one or more hardware accelerator resource requirements, partitioning and allocating (525) one or more hardware accelerator resources (320) to the plurality of virtual hardware accelerators (308A, 308B, 308C);providing (530) an isolation of a plurality of address spaces between the plurality of virtual hardware accelerators (308A, 308B, 308C); monitoring (535) the one or more hardware accelerator resources (320) for usage of each of the plurality of virtual hardware accelerators (308 A, 308B, 308C); comparing (545) the monitored at least one hardware accelerator resource usage with the one or more hardware accelerator resources allocated for each of the plurality of virtual hardware accelerators (308A, 308B, 308C); and enforcing (550) a plurality of quality of service parameters for each virtual hardware accelerator by: dynamically monitoring and modifying the allocation of the one or more hardware accelerator resources (314, 316, 318) based on the comparison of the monitored at least one hardware accelerator resource usage with the one or more hardware accelerator resources allocated for each of the plurality of virtual hardware accelerators (308 A, 308B, 308C); and maintaining the isolation of the plurality of address spaces between the plurality of virtual hardware accelerators (308A, 308B, 308C).
10. The method (500) as claimed in claim 9, comprises: assigning a plurality of device configuration address registers to each virtual hardware accelerator (308A or 308B or 308C).
Citation Information
Patent Citations
Compute resource estimation for function implementation on computing platform
US20210096922A1
Network and edge acceleration tile (NEXT) architecture
US20210117360A1
Function as a service (FAAS) system enhancements
US20210263779A1
Using hypervisor to provide virtual hardware accelerators in an o-ran system
US20220283841A1
Workflow allocation in edge infrastructures
WO2024194240A1