Bandwidth counting apparatus, graphics processor, method, device, medium and product
By introducing a bandwidth counting device into the GPU, using operating system identifiers and memory access request types to determine counting events, and combining event groups and register grouping, the problems of universality and scalability of bandwidth monitoring across multiple virtual machines are solved, achieving efficient and low-latency bandwidth statistics.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MOORE THREADS TECH CO LTD
- Filing Date
- 2026-07-01
- Publication Date
- 2026-07-28
AI Technical Summary
When multiple virtual machines run concurrently, insufficient GPU memory resources lead to bandwidth bottlenecks. Existing technologies lack universal and scalable bandwidth monitoring solutions, making it difficult to accurately measure the bandwidth usage of each virtual machine.
By introducing a bandwidth counting device into the GPU, the counting events are determined by operating system identifiers and memory access request types. Combined with event groups and register grouping, accurate statistics and temporary storage of bandwidth performance data for each virtual machine can be achieved, avoiding excessive consumption of hardware resources.
It enables independent recording of bandwidth usage for multiple virtual machines, reduces hardware resource consumption, improves monitoring accuracy and flexibility, supports the statistics of a large number of counting events, and is suitable for large-scale virtualization environments.
Smart Images

Figure CN122472973A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of high-performance computing with graphics processor chips, and more particularly to a bandwidth counting device, a graphics processor, a bandwidth counting method, a computer device, a computer-readable storage medium, and a computer program product. Background Technology
[0002] In high-performance computing systems, especially in the fields of graphics processing units (GPUs) and artificial intelligence (AI) chips, virtualization technology is widely used to improve resource utilization and flexibility. Through hardware virtualization, a single physical GPU can support the simultaneous operation of multiple virtual machines, thereby achieving efficient scheduling and allocation of computing resources. However, with multiple virtual machines running concurrently, video memory resources can easily become insufficient, forcing the GPU to access system memory via the PCIe interface, which can lead to bandwidth bottlenecks and impact overall performance.
[0003] Current technologies that utilize performance counters integrated within GPUs to monitor bandwidth usage lack versatility. Therefore, there is an urgent need for a bandwidth usage monitoring solution that does not rely on vendor-specific tools. Summary of the Invention
[0004] In view of the above, embodiments of this application provide at least one bandwidth counting device, graphics processor, bandwidth counting method, computer device, computer-readable storage medium, and computer program product.
[0005] The technical solution of this application embodiment is implemented as follows: On one hand, embodiments of this application provide a bandwidth counting device, which includes: a control unit and a counting unit; the control unit is used to determine a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine that issued the memory access request; the counting unit is used to count and temporarily store the bandwidth performance data of the target virtual machine based on the first counting event.
[0006] On the other hand, embodiments of this application provide a bandwidth counting device. The graphics processor includes: multiple computing units, a bandwidth counting device, and multiple memory spaces; each computing unit includes multiple virtual machines; a target virtual machine is used to issue a memory access request; the target virtual machine is any one of the multiple virtual machines; the bandwidth counting device is used to determine a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine; the bandwidth performance data of the target virtual machine is counted and temporarily stored based on the first counting event; and the memory space is used to process memory access requests.
[0007] In another aspect, embodiments of this application provide a bandwidth counting method, which includes: determining a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine that issued the memory access request; and counting and temporarily storing the bandwidth performance data of the target virtual machine based on the first counting event.
[0008] In another aspect, embodiments of this application provide a computer device, including: a memory for storing computer-executable instructions or computer programs; and a graphics processor including a bandwidth counting device for executing some or all of the steps in the above method when executing the computer-executable instructions or computer programs stored in the memory.
[0009] In another aspect, embodiments of this application provide a computer-readable storage medium having a computer program stored thereon, which, when executed by a graphics processor, implements some or all of the steps in the above-described method.
[0010] In another aspect, embodiments of this application provide a computer program including computer-readable code. When the computer-readable code is run in a computer device, the graphics processor in the computer device performs some or all of the steps for implementing the above-described method.
[0011] In another aspect, embodiments of this application provide a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer device, it implements some or all of the steps in the above-described method.
[0012] In this embodiment, the control unit determines the first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type. This first counting event accurately reflects the bandwidth usage behavior of the target virtual machine. The counting unit counts and temporarily stores the bandwidth performance data of the target virtual machine based on the first counting event, enabling precise statistics and temporary storage of the virtual machine's bandwidth performance data. The entire process is implemented through a hardware pipeline, ensuring low latency and high throughput. Simultaneously, the bandwidth counting device can support the statistics of a large number of counting events, improving the utilization of hardware resources. Thus, the bandwidth usage of multiple virtual machines can be independently recorded using the bandwidth counting device. Compared to a bandwidth statistics scheme that sets a counter for each virtual machine, this reduces hardware resource consumption while maintaining bandwidth statistics accuracy, significantly lowering hardware resource consumption. Compared to bandwidth statistics schemes based on GPU built-in performance counters, this does not require penetration of the virtualization layer, offering greater flexibility and scalability.
[0013] It should be understood that the above general description and the following detailed description are merely exemplary and explanatory, and are not intended to limit the technical solutions of this application. Attached Figure Description
[0014] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with this application and, together with the specification, serve to explain the technical solutions of this application.
[0015] Figure 1 A schematic diagram of the composition structure of a bandwidth counting device provided in this application embodiment. Figure 1 ; Figure 2 This application provides an illustration of bandwidth performance data statistics in a bandwidth counting device. Figure 1 ; Figure 3 A schematic diagram illustrating the mapping relationship between virtual machine numbers and operating system identifiers in a bandwidth counting device provided in this application embodiment; Figure 4 A schematic diagram of the composition structure of a bandwidth counting device provided in this application embodiment. Figure 2 ; Figure 5 A schematic diagram of the composition structure of a bandwidth counting device provided in this application embodiment. Figure 3 ; Figure 6 A schematic diagram illustrating the implementation process of a bandwidth counting method provided in this application embodiment; Figure 7 This is a schematic diagram illustrating the implementation process of evaluating bandwidth usage in a bandwidth counting method provided in an embodiment of this application. Figure 8This application provides an illustration of bandwidth performance data statistics in a bandwidth counting device. Figure 2 ; Figure 9 A schematic diagram of the composition structure of a graphics processor provided in an embodiment of this application; Figure 10 This is a schematic diagram of the hardware entity of a computer device provided in an embodiment of this application. Detailed Implementation
[0016] To make the objectives, technical solutions, and advantages of this application clearer, the technical solutions of this application are further described in detail below with reference to the accompanying drawings and embodiments. The described embodiments should not be regarded as limitations on this application. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0017] In the following description, references are made to “some embodiments,” which describe a subset of all possible embodiments. However, it is understood that “some embodiments” may be the same subset or different subsets of all possible embodiments and may be combined with each other without conflict.
[0018] If the application documents contain similar descriptions such as "first / second", the following explanation shall be added: the terms "first / second / third" are used only to distinguish similar objects and do not represent a specific order of objects. It is understood that "first / second / third" may be interchanged in a specific order or sequence where permitted, so that the embodiments of this application described herein can be implemented in an order other than that illustrated or described herein.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application pertains. The terminology used herein is for descriptive purposes only and is not intended to limit the scope of this application.
[0020] In current high-performance computing chips such as GPUs and AI, Virtual Desktop Infrastructure (VDI) is a crucial requirement. In a VDI scenario, a single physical GPU can support multiple virtual machines through hardware virtualization. However, when many virtual machines run on a single physical GPU, the problem of insufficient Video Random Access Memory (VRAM, also known as graphics memory) becomes apparent. When this happens, the GPU will access system memory through the Peripheral Component Interconnect Express (PCIe) interface, which significantly increases the bandwidth pressure on PCIe, causing potential communication congestion, queuing of tasks between virtual machines, and degrading the user experience. Therefore, implementing the function of monitoring the bandwidth of each virtual machine running on the GPU is particularly important. Currently, the mainstream methods for monitoring virtual machine bandwidth include the following: 1. Utilize the performance counters integrated within the GPU to directly measure the data bandwidth of key interfaces such as video memory and the PCIe bus. For example, NVIDIA DCGM can provide fine-grained monitoring of metrics such as video memory bandwidth; Intel XPU Manager can support video memory bandwidth monitoring for Intel GPUs. This method provides accurate data, close to the hardware level, but it is usually a vendor-specific solution, lacks universality, and cannot intuitively obtain bandwidth performance data for each virtual machine, requiring penetration through the virtualization layer.
[0021] 2. SR-IOV and Hardware Virtualization. This method generates multiple independent Virtual Functions (VFs) through physical GPU hardware and directly assigns them to virtual machines. Each VF has its own independent hardware counter. For example, with hardware assistance, NVIDIA vGPUs can query the bandwidth usage of each vGPU instance through management tools; AMD MxGPUs are similar. This method achieves isolation between virtual machines at the hardware level, providing a natural foundation for per-VM (Virtual Machine) bandwidth monitoring. However, it requires allocating one or more counters for each virtual machine, resulting in poor scalability and a less than ideal footprint.
[0022] Existing solutions all have certain drawbacks. Using performance counters integrated within the GPU for virtual machine bandwidth monitoring is typically a field-specific approach, lacking versatility. Furthermore, it often acquires PCIe interface bandwidth data, not per-VM bandwidth usage data, requiring penetration through the virtualization layer to obtain per-VM bandwidth data. While SR-IOV and hardware virtualization technologies can achieve per-VM bandwidth data statistics and monitoring, they require allocating one or more counters to each virtual machine, resulting in poor scalability and a large footprint.
[0023] To address the aforementioned issues, this application proposes an innovative bandwidth counting device that primarily solves two major problems: first, by using hardware virtualization technology to accurately count the bandwidth usage of each virtual machine; and second, by implementing a bandwidth performance counting unit that is highly versatile, highly scalable, and has a low footprint.
[0024] like Figure 1 As shown, the bandwidth counting device 10 provided in this application embodiment includes: a control unit 11 and a counting unit 12; the control unit 11 is used to determine a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine that issued the memory access request; the counting unit 12 is used to count and temporarily store the bandwidth performance data of the target virtual machine based on the first counting event.
[0025] A bandwidth counting device is a hardware structure used to monitor and record the bandwidth usage of various virtual machines or processes within a GPU. It can be installed within the GPU. A bandwidth counting device supports logical grouping of multiple perf events and allows selection of the current event to be counted via an external control interface. Core functions of a bandwidth counting device include event triggering, updating the counting register, and periodic writing to memory, making it a key component for achieving efficient, low-power bandwidth monitoring. A bandwidth counting device can also be referred to as a performance monitor (PFM) or a bandwidth performance counting unit.
[0026] The control unit parses the operating system identifier, address information, and request type in the task payload packet to determine whether the request belongs to a predefined count event category. If the request belongs to a predefined count event category, it indicates a successful match, and the corresponding count event is triggered and passed to the counting unit for further processing. This ensures that the bandwidth usage of each virtual machine is recorded independently, avoiding the data confusion problem caused by multiple virtual machines sharing a counter in traditional solutions. The counting unit is used to count and temporarily store the bandwidth performance data of memory accesses initiated by each virtual machine.
[0027] The Operating System Identifier (OSID) of any virtual machine is used to uniquely identify that virtual machine. The target virtual machine is the one currently issuing the memory access request. The OSID is not only used to distinguish different virtual machines, but also to collect statistics on the resource usage of each virtual machine, such as bandwidth consumption and the number of tasks. Through the OSID, hardware can independently perform virtual machine-level resource management without software intervention, and the hardware can improve response speed and operating efficiency.
[0028] In a virtualization environment, each virtual machine (VM) is assigned a unique Operating System Identifier (OSID) to ensure isolation between different VMs. For example, in a GPU supporting 16 VMs, the OSID ranges from 0 to 15. Each VM carries its OSID when issuing a workload packet, allowing the hardware layer to perform bandwidth performance analysis based on the OSID. Furthermore, the software layer embeds the OSID into the workload packet after assigning it to each VM, enabling the hardware layer to identify which VM the current request belongs to based on the received OSID.
[0029] Memory space refers to the storage area accessed in a request. Memory space can include, but is not limited to, Video Random Access Memory (VRAM) and System Memory. VRAM is the GPU's internal video memory, used to store graphics processing-related data. System Memory is host memory connected via the Peripheral Component Interconnect Express (PCIe) interface. In some implementations, when VRAM is insufficient, the GPU may access System Memory, thereby increasing the bandwidth pressure on the PCIe bus. Therefore, distinguishing whether a request is directed to VRAM or System Memory is crucial for bandwidth monitoring.
[0030] Request type refers to the nature of a memory access request, including read or write types. Each type of access request corresponds to different bandwidth performance data. For example, a System Memory read request triggers a System Memory read bandwidth count event, while a VRAM write request triggers a VRAM write bandwidth count event.
[0031] A perf event refers to a statistical action performed on the bandwidth performance data of a virtual machine's memory access requests. For example, once a perf event is activated, it is triggered every time a System Memory read operation occurs, and the count value of that perf event is updated. Through the above perf event mechanism and statistical method, the monitoring module can perform fine-grained monitoring of the bandwidth usage of each virtual machine. In this application, perf events can include System Memory read, System Memory write, VRAM read, VRAM write, and other events. The first perf event refers to the perf event that matches the current memory access request.
[0032] Bandwidth performance data refers to the bandwidth usage of the target virtual machine, which can be statistically analyzed in bytes. In some implementations, where the memory space includes VRAM and System Memory, bandwidth performance data includes VRAM read bandwidth, VRAM write bandwidth, System Memory read bandwidth, and System Memory write bandwidth.
[0033] Assuming a physical GPU can support 16 virtual machines, the OSID range is 0-15. In monitoring virtual machine bandwidth, System Memory read bandwidth, System Memory write bandwidth, VRAM read bandwidth, and VRAM write bandwidth are key performance metrics to monitor. Therefore, in this case, it's necessary to statistically analyze 16 (VMs) * 4 (Events / VMs) = 64 bandwidth performance data points. Figure 2 As shown, the bandwidth performance data statistics include: OSID0 System Memory read bandwidth, OSID0 System Memory write bandwidth, OSID0 VRAM read bandwidth, OSID0 VRAM write bandwidth, OSID1 System Memory read bandwidth, OSID1 System Memory write bandwidth, OSID1 VRAM read bandwidth, OSID1 VRAM write bandwidth, ..., OSID15 System Memory read bandwidth, OSID15 System Memory write bandwidth, OSID15 VRAM read bandwidth, OSID15 VRAM write bandwidth.
[0034] Temporary storage refers to temporarily saving the counting results instead of immediately writing them to external storage, which reduces frequent access to external storage. Furthermore, by using temporary storage, the software can read and process the counting results uniformly at appropriate times (such as when a threshold is reached or a monitoring cycle is completed), thus enabling a more flexible bandwidth management strategy.
[0035] In some implementations, the bandwidth counting device predefines multiple event groups. In response to a received memory access request, a first counting event matching the memory access request can be arbitrated from the multiple event groups based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type.
[0036] In some implementations, the triggering conditions for each counting event can be set based on the operating system identifier (OSID), memory space, and request type. In response to a received memory access request, a matching triggering condition is determined based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the counting event corresponding to the matching triggering condition is taken as the first counting event.
[0037] In some implementations, the read and write bandwidths of each virtual machine can be the same, such as 64 bytes. In this case, upon receiving a memory access request, the value of the corresponding counter can be incremented by one based on the first counting event to count the bandwidth performance data of the target virtual machine. The incrementing process indicates an increase of 64 bytes in bandwidth.
[0038] In some implementations, after determining the first counting event, the counting unit determines the register corresponding to the first counting event and increments the value of the register to count and temporarily store the bandwidth performance data of the target virtual machine. For example, when the control unit detects a System Memory write request with OSID 5, it determines the counting event corresponding to the System Memory write request, identifies the register corresponding to the counting event, and increments the value of the register by the corresponding bandwidth granularity (e.g., 64 bytes). The register value can then be periodically written to memory for software analysis and decision-making. In this way, real-time monitoring of the bandwidth usage of each virtual machine can be achieved, and rate limiting measures can be taken when necessary, thereby improving overall stability and efficiency.
[0039] In this embodiment, the control unit and the counting unit in the bandwidth counting device have a close collaborative relationship. The control unit determines the first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type. The first counting event can accurately reflect the bandwidth usage behavior of the target virtual machine. The counting unit counts and temporarily stores the bandwidth performance data of the target virtual machine based on the first counting event, enabling accurate statistics and temporary storage of the virtual machine's bandwidth performance data. The entire process is implemented through a hardware pipeline, ensuring low latency and high throughput. At the same time, the bandwidth counting device can support the statistics of a large number of counting events, improving the utilization of hardware resources. Thus, the bandwidth usage of multiple virtual machines can be independently recorded by the bandwidth counting device. Compared with the bandwidth statistics scheme that sets a counter for each virtual machine, it reduces hardware resource consumption while ensuring the accuracy of bandwidth statistics, significantly reducing hardware resource consumption. Compared with the bandwidth statistics scheme based on the GPU's built-in performance counter, it does not need to penetrate the virtualization layer, and has strong flexibility and scalability.
[0040] In some embodiments, the bandwidth counting device further includes: a control unit, further configured to determine the storage space to be accessed by the memory access request based on the address of the task load packet of the memory access request; determine the target virtual machine based on the operating system identifier carried in the memory access request; the operating system identifier carried in the memory access request is determined based on the number of the target virtual machine and the mapping relationship between the number of the virtual machine and the operating system identifier of a preset virtual machine; and verify the storage space to be accessed by the memory access request based on the access relationship between the target virtual machine and the preset virtual machine and the memory space.
[0041] The memory access request payload is a specific data packet sent by the virtual machine to the GPU for execution. The memory access request payload carries metadata related to the current task. For example, the payload may include, but is not limited to, multiple fields such as address, request type (read / write), operating system identifier (OSID), and bit width. By parsing these fields, the control unit can identify whether the current request is to access system memory or video memory. For example, if the address of a memory access request's payload falls within the address range of system memory, the control unit determines that the memory access request is for system memory; otherwise, it determines that it is for video memory.
[0042] It should be noted that the metadata carried in the task payload of memory access requests enables the hardware layer to directly identify and process requests from different virtual machines without relying on software intervention, thereby achieving efficient resource scheduling and performance monitoring.
[0043] The Operating System Identifier (OSID) is determined based on the target virtual machine's ID and the predefined mapping between virtual machine IDs and operating system identifiers. For example... Figure 3 As shown, assuming the GPU supports 16 virtual machines (VMs), such as VM0 to VM15, the mapping relationship between virtual machines (VMs) and operating system identifiers (OSIDs) can be as follows: the OSID of the VM in VM0 is 0, the OSID of the VM in VM1 is 1, ..., and the OSID of the VM in VM15 is 15.
[0044] In some implementations, the control unit can determine the memory space to be accessed by the memory access request based on the address of the task payload packet of the memory access request and a preset address mapping table. This method of quickly determining the target location of the memory access request using a preset address mapping table not only improves processing efficiency but also reduces the need for additional caches or registers, thereby reducing hardware complexity and power consumption.
[0045] Specifically, if the address of the task payload packet of the memory access request is determined to be within the address range of the system memory based on the preset address mapping table, the storage space to be accessed by the memory access request is determined to be the system memory; if the address of the task payload packet of the memory access request is determined to be within the address range of the video random access memory based on the preset address mapping table, the storage space to be accessed by the memory access request is determined to be the video random access memory.
[0046] In some implementations, the target virtual machine corresponding to the operating system identifier of a memory access request can be determined based on a preset mapping relationship between virtual machine numbers and operating system identifiers. In this way, the virtual machine issuing the current request can be quickly identified based on the operating system identifier of the memory access request, ensuring that subsequent access verification and performance statistics accurately match the correct virtual machine. This also helps reduce resource conflicts between different virtual machines, improving overall stability and security.
[0047] In some implementations, each virtual machine has its own memory access permission configuration, which defines the memory regions that each virtual machine can access and the access methods (read / write). The control unit can look up the memory access permission configuration corresponding to the target virtual machine in a preset virtual machine-memory space access relationship table based on the target virtual machine's ID or operating system identifier, and perform a validity check on the current memory access request based on the found memory access permission configuration. For example, if a virtual machine is restricted to reading only a specific area of system memory, the control unit will refuse to execute a write request initiated by that virtual machine.
[0048] In some implementations, the access relationship between the virtual machine and memory space can be configured and maintained by software, and this relationship can be stored in hardware-level registers or memory. This allows for flexible adjustment of access strategies for different virtual machines via software.
[0049] In some implementations, the control unit can comprehensively consider the target virtual machine's permission configuration, access address, and operation type during the verification process to determine whether to allow the memory access request to execute. Thus, through fine-grained access control, device crashes or data corruption caused by erroneous access can be effectively avoided, further improving overall reliability.
[0050] In this embodiment, by introducing an OSID-based virtual machine identification mechanism and a memory access request verification mechanism, precise monitoring and control of the memory access behavior of each virtual machine is achieved. This not only improves the rationality of resource allocation but also significantly enhances overall security and stability, making it suitable for application scenarios with high requirements for resource management and performance optimization in large-scale virtualization environments.
[0051] In some embodiments, the bandwidth counting device defines multiple event groups; the multiple event groups are obtained by grouping multiple counting events according to the type of memory space; the counting unit includes multiple registers, the number of registers matching the number of counting events included in the event group; the control unit is further configured to determine a target event group from the multiple event groups based on the storage space to be accessed by the memory access request; and to determine a first counting event from the multiple counting events of the target event group based on the operating system identifier and request type carried by the memory access request.
[0052] An event group refers to a logical set of events used to count specific events. The counted events in each event group are associated with a particular type of memory access behavior, such as System Memory read, System Memory write, VRAM read, and VRAM write. By grouping counted events according to the type of memory space accessed, the number of registers required can be effectively reduced, improving resource utilization.
[0053] like Figure 4 As shown, the bandwidth counting device can include N+1 event groups, from event group 0 to event group N. Each event group includes four counting events: pfm_event0, pfm_event1, pfm_event2, and pfm_event3.
[0054] In some implementations, multiple counting events in the bandwidth counting device can be divided into multiple event groups according to the type of memory space. For example, when a graphics processor can be virtualized into 16 virtual machines, each virtual machine needs to monitor four types of bandwidth performance data, requiring a total of 64 counting events. If the counting events are divided into two groups according to the type of memory space, one group consists of 32 counting events related to System Memory, and the other group consists of 32 counting events related to VRAM. In this case, 32 registers need to be configured according to the number of counting events contained in each event group. In this way, the System Memory read / write bandwidth of multiple virtual machines can be counted simultaneously, as can the VRAM read / write bandwidth of multiple virtual machines.
[0055] In some implementations, multiple counting events in the bandwidth counting device can be divided into multiple event groups according to the category of bandwidth performance data. For example, when a graphics processor can be virtualized into 16 virtual machines, each virtual machine needs to monitor four types of bandwidth performance data, requiring a total of 64 counting events. If the counting events are divided into two groups according to the category of bandwidth performance data, each group contains four counting events related to one virtual machine, requiring only four registers to be configured. This allows for the statistical analysis of System Memory read / write bandwidth and VRAM read / write bandwidth for the same virtual machine.
[0056] In some implementations, multiple counting events in the bandwidth counting device can be divided into multiple event groups according to the operating system identifier and the type of memory space. For example, when the graphics processor can be virtualized into 16 virtual machines, each virtual machine needs to monitor four types of bandwidth performance data, requiring a total of 64 counting events. If the counting events are divided into four groups according to the operating system identifier and the type of memory space, the first group consists of 16 counting events related to System Memory read bandwidth, the second group consists of 16 counting events related to System Memory write bandwidth, the third group consists of 16 counting events related to VRAM read bandwidth, and the fourth group consists of 16 counting events related to VRAM write bandwidth. In this case, 16 registers need to be configured. In this way, the System Memory read bandwidth of multiple virtual machines can be counted simultaneously, as can the System Memory write bandwidth of multiple virtual machines, the VRAM read bandwidth of multiple virtual machines, and the VRAM write bandwidth of multiple virtual machines.
[0057] In some implementations, multiple counting events in the bandwidth counting device can be divided into multiple event groups according to the category of bandwidth performance data. For example, when a graphics processor can be virtualized into 16 virtual machines, each virtual machine needs to monitor four types of bandwidth performance data, requiring a total of 64 counting events. If the counting events are divided into two groups according to the category of bandwidth performance data, each group consists of four counting events related to one virtual machine, requiring only four registers to be configured.
[0058] It should be noted that the above grouping method not only improves the system's scalability but also facilitates independent analysis and scheduling of different types of bandwidth usage by the software. By reasonably setting the number and size of event groups, the utilization rate of hardware resources can be maximized while ensuring monitoring accuracy. This is particularly suitable for scenarios supporting a large number of virtual machines, such as large data centers or cloud computing platforms, and can effectively improve the overall energy efficiency ratio and reduce deployment costs.
[0059] It should be noted that in existing solutions, each counting event requires an independent counting register, resulting in a large hardware area and difficulty in expansion; while this application, through an event grouping mechanism, only requires setting the register that matches the event group to realize bandwidth statistics for all virtual machines, greatly reducing hardware overhead.
[0060] In some implementations, when a task payload arrives at the hardware layer, it carries an OSID, request type, and access address. The hardware determines whether the access is for System Memory or VRAM based on the access address and selects a target event group from predefined event groups that matches the access type. For example, if the access is for System Memory, the control unit selects the event group containing System Memory as the target event group.
[0061] In some implementations, the control unit can determine the specific first counting event in a target event group based on the requested operating system identifier (OSID) and the request type (e.g., read, write). For example, assuming the current request has an OSID of 1 and is a System Memory read operation, the control unit will find the counting event corresponding to the current request's OSID and operation type in the target event group and activate that counting event to begin counting. This process ensures that the bandwidth usage of each virtual machine is accurately recorded, thereby achieving event-level bandwidth monitoring.
[0062] It's important to note that by combining OSID and request type to select the corresponding counting event, we can ensure that the bandwidth usage of each virtual machine is counted independently, preventing data from being mixed up with that of other virtual machines. This not only improves the accuracy of the statistics but also provides reliable data support for subsequent rate limiting decisions. Furthermore, since each counting event is only activated when needed, it effectively reduces unnecessary computational and storage overhead, thus improving overall efficiency.
[0063] In this embodiment, an intelligent selection mechanism for event groups and control units enables efficient monitoring of bandwidth usage across multiple virtual machines without adding extra hardware resources. This effectively solves the problem of excessive hardware footprint caused by the need for an independent counter for each virtual machine in traditional solutions, and provides a solid foundation for future expansion with more virtual machines. Through reasonable event group partitioning and control logic design, high-precision, low-overhead bandwidth monitoring is achieved, demonstrating broad application prospects. By defining event groups based on memory space type and selecting the corresponding first counting event based on OSID and request type, bandwidth statistics at the virtual machine level are realized. This not only improves the accuracy and flexibility of bandwidth monitoring but also significantly reduces hardware resource consumption, exhibiting good scalability and practicality.
[0064] In some embodiments, the control unit is further configured to raise the level of the physical wire of the first counting event to a first level; the counting unit is further configured to increment the value of the register corresponding to the first counting event by one when the physical wire of the first counting event is raised to the first level, so as to count the bandwidth performance data of the virtual machine; the incrementing process represents that the bandwidth performance of the target virtual machine is increased by a target bit; the target bit is the bandwidth granularity of the computing unit to which the target virtual machine belongs.
[0065] The first level is high. The counting unit can contain multiple registers. Registers are hardware components used for temporary data storage. Registers are used to accumulate the number of times an event occurs. For example... Figure 4 As shown, the register size can be 32 bytes. If the bandwidth value is large, a 64-bit register can be used to support a larger counting range.
[0066] In some implementations, the number of registers can be determined according to the type of counting event, and the specific number of registers can be the same as the number of types of counting events. In this case, the registers occupy less hardware space, which can maximize the reuse rate of hardware resources while ensuring monitoring accuracy.
[0067] In some implementations, the number of registers can be determined based on the number of counting events. Specifically, the number of registers can be the same as the number of all counting events, so that the bandwidth usage corresponding to each counting event can be independently counted.
[0068] Each register stores the cumulative count of a count event. When the control unit pulls the physical wire corresponding to a count event high, the counting unit detects the change and increments the register corresponding to that count event by one. For example, suppose a virtual machine initiates a 32-byte system memory write operation; the counting unit will increment the register corresponding to the system memory write count event to indicate that the virtual machine consumed 32 bytes of bandwidth during this operation.
[0069] Each counting event has a physical wire used to transmit the event's occurrence signal. When the signal generated by pulling the physical wire high is recognized, the counting unit performs a corresponding increment operation. The first counting event's physical wire level rises from low to high, indicating that the first counting event has occurred and needs to be recorded; thus, the bandwidth usage of each virtual machine can be accurately identified and responded to, enabling monitoring of the bandwidth usage of each virtual machine.
[0070] For example, when a virtual machine with OSID 1 makes a 64-byte read request to the system memory, a system memory read count event will occur. At this time, the physical wire of the system memory read count event will be pulled high to indicate that the system memory read count event has occurred.
[0071] Incrementing by one indicates that the virtual machine's bandwidth performance has increased by a target bit, which corresponds to the bandwidth granularity of the computing unit. Bandwidth granularity refers to the amount of data represented by each statistical operation. For example, if the bandwidth granularity is 32 bytes, then each register increment operation represents the completion of 32 bytes of data transmission. The bandwidth granularity is determined by the computing unit and serves as the basic unit for statistically analyzing bandwidth performance. Different computing units can be configured with different bandwidth granularities to adapt to different performance requirements.
[0072] In this embodiment, when the first counting event is activated, the control unit converts the first counting event into a level change on the physical wire and further drives the counting unit to perform an increment operation, thereby recording the bandwidth usage of the target virtual machine. This register-based counting method not only improves counting accuracy but also ensures the real-time performance and accuracy of bandwidth statistics. It can efficiently manage the statistical needs of a large number of counting events while reducing hardware resource consumption. It should be noted that the counting event and the increment operation performed by the counting unit have a direct causal relationship in function; the counting event is the hardware-level manifestation of the increment operation performed by the counting unit.
[0073] In some embodiments, the bandwidth counting device further includes: a selector, configured to receive a first selection signal for a first counting event and to issue a first counting signal for the first counting event; the first selection signal is triggered when the level of the physical wire of the first counting event rises to a first level; and a counting unit, further configured to increment the value of the register corresponding to the first counting event by one upon receiving the first counting signal; the incrementing process characterizes the bandwidth performance of the target virtual machine to increase by a target bit; the target bit is the bandwidth granularity of the computing unit to which the target virtual machine belongs.
[0074] The first level is high. A selector is a hardware logic unit used for signal routing and control. For example... Figure 4 As shown, the selector determines which first counting event will be activated and included in the statistics based on the input selection signal. For example, each event group contains multiple counting events, which represent different bandwidth monitoring items, such as System Memory read bandwidth, VRAM write bandwidth, etc. The first selection signal refers to the currently selected first counting event to record the bandwidth usage of the target virtual machine. The first counting signal refers to recording the bandwidth usage of the target virtual machine through the register corresponding to the first counting event.
[0075] For example, when a virtual machine with OSID 1 initiates a System Memory read request, and the OSID carried in the task payload is also 1, the physical wire controlling the System Memory read count event will be pulled high by the hardware circuit, thereby generating a selection signal for the System Memory read count event. The selector responds to the selection signal and issues a count signal for the System Memory read count event. The counting unit responds to the count signal and increments the value of the register corresponding to the System Memory read count event.
[0076] The counting unit achieves precise counting of each first counting event by incrementing the register. This precise counting not only ensures the accuracy of bandwidth statistics but also provides a reliable data foundation for subsequent software analysis. Furthermore, because the performance monitoring module employs a grouped statistical approach, the counting unit can serve multiple first counting events within different time periods, thereby improving hardware resource utilization and reducing redundant design.
[0077] In this embodiment, the selector can dynamically respond to changes in the first selection signal and flexibly switch the register currently being counted; the counting unit can store the counting results in the corresponding register; thus, without increasing the number of additional registers, efficient reuse and management of a large number of counting events can be achieved, significantly reducing hardware area occupation and improving overall scalability and resource utilization.
[0078] In some embodiments, the selector is further configured to receive a second selection signal from an external input regarding the second counting event and to issue a second counting signal regarding the second counting event; the counting unit is further configured to increment the value of the register corresponding to the second counting event by one upon receiving the second counting signal.
[0079] The second selection signal refers to the selection signal input by the user or external software. The second counting event is the counting event selected by the second selection signal. The second counting signal is used to indicate that the second counting event has been selected to record the bandwidth usage of the target virtual machine. The second counting signal refers to the recording of the virtual machine's bandwidth usage through the register corresponding to the second counting event.
[0080] For example, in a scenario supporting eight virtual machines, if the software wants to monitor the VRAM write bandwidth of the eighth virtual machine, it can send a selection signal to the eighth counting event. Upon receiving this selection signal, the selector will activate the circuit path corresponding to the counting event, temporarily storing the bandwidth performance data of the eighth virtual machine in the register corresponding to the eighth counting event.
[0081] In this embodiment, the second counting event to be selected can be determined by an externally input selection signal, which has good scalability and low hardware cost, and is suitable for situations where multiple virtual machines share GPU resources.
[0082] In some embodiments, the bandwidth counting device further includes: a configuration interface for setting configuration information of the bandwidth counting device; the configuration information is used to characterize the counting mode of the bandwidth counting device; a control interface for receiving externally input control information; the control information is used to control the counting process of the bandwidth counting device; and a write interface for writing bandwidth performance data in the register corresponding to the first counting event into memory every first time period.
[0083] The configuration interface is an input port of the bandwidth counting device. Software can write configuration parameters to the bandwidth counting device through this interface. These parameters define the device's operating mode. For example, configuration information may include the statistical granularity (e.g., how many bytes of bandwidth are counted each time), the statistical interval (i.e., how often data is written to memory), and whether to enable specific groups of performance events. The configuration information determines how the bandwidth counting device handles data acquisition and storage for various performance events during operation. Through the configuration interface, software can flexibly adjust the behavior of the bandwidth counting device to adapt to different application scenarios.
[0084] The use of a configuration interface greatly enhances the flexibility and versatility of bandwidth counting devices. The software can dynamically adjust the configuration of the bandwidth counting device based on the current load, such as increasing the counting frequency under high load or reducing resource consumption under low load. Furthermore, the configuration interface supports future expansion; as the number of virtual machines increases, only the configuration information of the configuration interface needs to be modified through software, without changing the hardware design, thus significantly reducing maintenance costs.
[0085] The control interface is another input port in the bandwidth counting device. It allows external devices (such as processors or software) to send control commands to start, stop, or pause the counting operation. Control information can include start counting commands, stop counting commands, and trigger write commands. For example, when a user wants to start monitoring the bandwidth usage of a virtual machine, the software sends a start counting command through the control interface. Upon receiving the start counting command, the bandwidth counting device begins recording bandwidth performance data for the relevant counting events. Similarly, when the monitoring task is complete, the software can send a stop counting command, which stops the bandwidth counting device from recording and prepares for subsequent data writing operations.
[0086] The control interface can also be used to switch the currently monitored group in the bandwidth counter. Since the bandwidth counter supports grouped statistics, the software can select the group to be monitored through the control interface, thereby enabling the monitoring of different virtual machines or different bandwidth event types at different times. For example, system memory read / write bandwidth can be monitored in the first period, and then switched to VRAM read / write bandwidth in the second period. By switching the monitored object at different time periods, resource utilization can be improved and data overflow can be avoided.
[0087] The design of the control interface enables the bandwidth counting device to flexibly adjust its working state according to external requirements, thereby achieving efficient and accurate bandwidth monitoring. This not only improves the system's response speed to user requests but also enhances the real-time performance and controllability of the bandwidth monitoring function.
[0088] The write interface is the port in the bandwidth counting device used to output data. It periodically writes the bandwidth performance data recorded in the bandwidth counting device into memory for subsequent software reading and analysis. Whenever a set time interval is reached, the bandwidth counting device calls the write interface to write the data from these registers into memory, forming a complete snapshot of the bandwidth performance.
[0089] The frequency of write operations can be determined by parameters set in the configuration interface. For example, the software can set the write interval to 1 second in the configuration interface, and the bandwidth counting device will write the bandwidth data in the current register to memory at the end of each second. By configuring the write interval to 1 second, the software can obtain continuous bandwidth performance data, thereby calculating the average bandwidth usage of each virtual machine per unit time, and determining whether rate limiting measures are needed.
[0090] The periodic write mechanism of the write interface ensures the continuity and availability of bandwidth performance data. The software can accurately calculate the bandwidth consumption of each virtual machine within a specific time period by comparing the data difference between two writes. This periodic write mechanism not only improves the accuracy of bandwidth monitoring but also provides reliable data support for subsequent resource scheduling and optimization.
[0091] In this embodiment, the coordinated operation of the configuration interface, control interface, and write interface enables accurate bandwidth statistics and flexible control based on OSID, significantly improving hardware resource reuse, reducing area overhead, and providing excellent scalability. Through grouped statistics, the bandwidth counting device can support bandwidth monitoring of a large number of virtual machines without adding extra hardware, thus meeting the needs for refined management of high-performance computing resources in VDI scenarios.
[0092] This application provides a graphics processor, which includes: multiple computing units, a bandwidth counting device, and multiple memory spaces; each computing unit includes multiple virtual machines; a target virtual machine, used to issue a memory access request, wherein the target virtual machine is any one of the multiple virtual machines; a bandwidth counting device, used to determine a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine; the bandwidth performance data of the target virtual machine is counted and temporarily stored based on the first counting event; and the memory spaces are used to process memory access requests.
[0093] In some embodiments, the graphics processor further includes an on-chip network; multiple computing units are connected to multiple memory spaces via the on-chip network.
[0094] like Figure 5As shown, a graphics processor may include a graphics processor core (GPU Core) and a video processing unit (VPU). Both the GPU Core and the VPU can support hardware virtualization of 16 virtual machines. The GPU Core may contain eight computing units, namely computing units 0 to computing units 7. Each of these eight computing units has an output connected to a network on chip (NoC). The video processing unit has an interface connected to the network on chip, through which it can access video random access memory (VRAM) and system memory.
[0095] The computing unit's output can also be connected to a bandwidth counting device. This bandwidth counting device contains multiple event groups. These event groups can be grouped based on the type of memory space. For example... Figure 4 As shown, assuming the GPU has only one graphics processing unit (GPU) core, and this GPU supports two virtual machines (VMs), each VM needs to have its System Memory read bandwidth, System Memory write bandwidth, VRAM read bandwidth, and VRAM write bandwidth measured—that is, each VM requires four bandwidth performance data points. The event group within the counting unit can contain four counting events. The first event group measures the System Memory read bandwidth and System Memory write bandwidth for both VMs, and the second event group measures the VRAM read bandwidth and VRAM write bandwidth for both VMs. Furthermore, if the software selects event group 0, and there is a memory access request with OSID 1, accessing System Memory, and a read request type, then the physical wire (perf event wire) corresponding to the third counting event in the third box of event group 0 in the diagram will be pulled high, and the corresponding register value will be incremented by 1. The software starts data statistics through the control interface of the bandwidth counter, records the number of times each counter event in event group 0 is pulled up, and then the software controls the bandwidth counter to finish data statistics. At this time, the registers in the bandwidth counter record the bandwidth performance data within this time period. After that, the data in these registers is written to memory, and the software can read the virtual machine's bandwidth performance data from memory.
[0096] This application provides a bandwidth counting method applied to a graphics processor, such as... Figure 6 As shown, the method includes the following steps 601 to 602: Step 601: Determine the first counting event based on the operating system identifier, the storage space to be accessed, and the request type carried in the memory access request; the operating system identifier is used to identify the target virtual machine that issued the memory access request.
[0097] In some implementations, step 601 can be achieved by the following steps 6011 to 6012: Step 6011: Based on the storage space to be accessed by the memory access request, determine the target event group from multiple event groups.
[0098] Step 6012: Based on the operating system identifier and request type of the memory access request, determine the first counting event from multiple counting events of the target event group.
[0099] Step 602: Count and temporarily store the bandwidth performance data of the target virtual machine based on the first counting event.
[0100] In some implementations, the bandwidth performance data in the register corresponding to the first counting event is written to memory every first time period.
[0101] In some implementations, step 602 can be achieved by the following steps 6021 to 6022: Step 6021: Increase the level of the physical wire of the first counting event to the first level.
[0102] Step 6022: When the physical wire of the first counting event is raised to the first level, the value of the register corresponding to the first counting event is incremented by one to count the bandwidth performance data of the virtual machine; the incrementing process represents that the bandwidth performance of the target virtual machine is increased by a target bit; the target bit is the bandwidth granularity of the computing unit to which the target virtual machine belongs.
[0103] In some implementations, the first level is a high level. A first selection signal for the first counting event is triggered when the level of the physical wire of the first counting event rises to a high level; a first counting signal for the first counting event is issued based on the first selection signal to increment the value of the register corresponding to the first counting event by one.
[0104] In some implementations, a second selection signal for the second counting event may be received from an external input, and a second counting signal for the second counting event may be issued to increment the value of the register corresponding to the second counting event.
[0105] In some implementations, configuration information for the bandwidth counting device can also be set; the configuration information is used to characterize the counting method of the bandwidth counting device.
[0106] In some implementations, externally input control information can also be received, and the counting process of the bandwidth counting device can be controlled based on the control information.
[0107] In some implementations, the bandwidth counting method provided in this application further includes the following steps 603 to 604: Step 603: Determine the storage space to be accessed by the memory access request based on the address of the task load packet of the memory access request.
[0108] Step 604: Determine the target virtual machine based on the operating system identifier carried in the memory access request; the operating system identifier carried in the memory access request is determined based on the target virtual machine number and the preset mapping relationship between the virtual machine number and the operating system identifier.
[0109] Step 605: Verify the storage space to be accessed by the memory access request based on the target virtual machine and the preset virtual machine and memory space access relationship.
[0110] In some embodiments, such as Figure 7 As shown, the bandwidth counting method provided in this application embodiment further includes the following steps 701 to 704: Step 701: Determine the starting and ending values of the bandwidth performance of the target virtual machine in the second time period from the bandwidth performance data.
[0111] Step 702: Evaluate the bandwidth usage of the target virtual machine based on the initial and final values of bandwidth performance within the second time period, the target time period, and the bandwidth threshold, and obtain the evaluation results.
[0112] In some implementations, the bandwidth to be evaluated is determined based on the difference between the starting and ending values of bandwidth performance within the second time period and the target time period; the bandwidth to be evaluated is compared with the bandwidth threshold; if the bandwidth to be evaluated is greater than or equal to the bandwidth threshold, it is determined that the target virtual machine is experiencing bandwidth congestion; if the bandwidth to be evaluated is less than the bandwidth threshold, it is determined that the target virtual machine is not experiencing bandwidth congestion.
[0113] In some implementations, the difference between the starting and ending values of the bandwidth performance within the second time period and the target time period can be used to trigger calculations to obtain the bandwidth to be evaluated.
[0114] In some implementations, the mean or median value of bandwidth performance during the second time period can also be used as the bandwidth to be evaluated.
[0115] Step 703: If the evaluation results indicate that the target virtual machine is experiencing bandwidth congestion, reduce the frequency of memory access requests to the target virtual machine.
[0116] Step 704: If the evaluation results indicate that the target virtual machine does not experience bandwidth congestion, remove the restriction on the frequency of memory access requests issued by the target virtual machine.
[0117] The descriptions of the above method embodiments are similar to those of the above device embodiments, and have similar beneficial effects. In some embodiments, the functions or modules included in the device embodiments of this application can be used to perform the methods described in the above method embodiments. For technical details not disclosed in the method embodiments of this application, please refer to the descriptions of the device embodiments of this application for understanding.
[0118] The following describes the application of the bandwidth counting method provided in the embodiments of this application in a real-world scenario.
[0119] To achieve bandwidth performance monitoring at the virtual machine level, a unique identifier for each virtual machine is necessary. The Operating System Identifier (OSID) in this application serves this purpose. When multiple operating systems (OS) or virtual machines run on a GPU, the software assigns an OSID to each virtual machine. This OSID is used to isolate the memory space of each virtual machine, and the task flows issued by this virtual machine also carry this OSID, thus achieving isolation between virtual machines in GPU virtualization. The mapping relationship between virtual machine numbers and OSIDs is as follows: Figure 3 As shown, the mapping relationship between virtual machine number and OSID can be implemented in software.
[0120] When a virtual machine on the GPU distributes tasks, the task payload packet carries its corresponding OSID. The hardware layer then uses this OSID to calculate the bandwidth performance data for each VM at the output, including System Memory read bandwidth, System Memory write bandwidth, VRAM read bandwidth, and VRAM write bandwidth. The specific bandwidth calculation scheme is as follows: like Figure 5 As shown, assuming the GPU includes a graphics processing core (GPU Core) and a video processing unit (Video Processing Unit), both of which can support hardware virtualization of 16 virtual machines, the GPU Core contains 8 computing units, each of which has an output connected to the on-chip network. The video processing unit has an interface connected to the on-chip network, through which it can access video random access memory and system memory.
[0121] Each outgoing workload packet from the interface carries data such as OSID and access width. The hardware uses the OSID and access width to calculate bandwidth performance data and store it in the registers of the corresponding virtual machine. Simultaneously, the GPU uses the address of the workload packet to determine whether the current access is to VRAM or System Memory. Assuming a physical GPU can support 16 virtual machines, the OSID range is 0-15. In monitoring virtual machine bandwidth, System Memory read bandwidth, System Memory write bandwidth, VRAM read bandwidth, and VRAM write bandwidth are key performance data points to focus on. Therefore, in this case, it is necessary to collect bandwidth performance data for 16 VMs * 4 Events / VMs = 64 points, such as... Figure 2 As shown.
[0122] If there is a read request with OSID 1 that needs to read 64 bytes of data, and the address of the read request indicates that the read request is going to System Memory, then the System Memory read bandwidth counter for OSID 1 will be incremented by 64 bytes.
[0123] To collect a total of 64 bandwidth data points from the 16 virtual machines mentioned above, referred to as perfevents in the following text, a performance data monitoring module is needed to trigger and count these performance statistics points. This leads to the design of the highly scalable and reusable Performance Monitor (PFM, corresponding to the bandwidth counting device mentioned above). The core of this PFM design lies in its ability to group multiple perfevents, allowing different groups of perfevents to be counted at different times. The specific group counted is controlled by the software through the PFM's group selection interface. The software can also control the PFM to write bandwidth performance data into memory at regular intervals via a control interface, calculating the per-VM bandwidth data by comparing the differences in memory data over a given time interval. The design and control flow of the PFM will be described in detail below. First... Figure 4 For the PFM architecture, a group within the PFM can hold 4 counting events. Therefore, 4 registers need to be set up within the PFM to temporarily store bandwidth data. The bit width of each register is set to 32 bytes here. If the bandwidth value is large, it can also be set to a larger bit width, such as bytes.
[0124] like Figure 4As shown, PFM defines a total of N+1 event groups, each containing 4 counting events. Each counting event corresponds to a one-bit wire (physical conductor). When the trigger condition for a counting event is met, the level of the physical conductor for that event is pulled high, and the corresponding register value is incremented by 1. The configuration interface allows setting some configuration information for PFM, such as enabling PFM to automatically dump counter data to memory at fixed time intervals. The control interface allows external users to control the start and stop times of PFM counting, as well as the time for dumping data to memory. The counting events of the N event groups are connected to a selector, and the selection signal is used to select the currently counting event group. This selection signal is also controlled by the external user / software. The PFM counting unit mainly contains four 32-byte registers used to temporarily store the bandwidth performance data of the four counting events in the event group.
[0125] Based on the assumptions above, the GPU can support 16 virtual machines (VMs). Each VM requires four types of bandwidth performance data (System Memory read bandwidth, System Memory write bandwidth, VRAM read bandwidth, and VRAM write bandwidth), resulting in a total of 64 bandwidth performance data points. If the PFM event group size is set to 32, with each event group containing 32 count events, then these 64 count events can be divided into two groups: the first group contains the System Memory read / write bandwidth data for the 16 VMs, and the second group contains the VRAM read / write bandwidth data for the 16 VMs. Figure 8 As shown, event group 0 records the read bandwidth of OSID0 System Memory, the write bandwidth of OSID0 System Memory, the read bandwidth of OSID1 System Memory, the write bandwidth of OSID1 System Memory, ..., the read bandwidth of OSID15 System Memory, and the write bandwidth of OSID15 System Memory; event group 1 records the read bandwidth of OSID0 VRAM, the write bandwidth of OSID0 VRAM, the read bandwidth of OSID1 VRAM, the write bandwidth of OSID1 VRAM, ..., the read bandwidth of OSID15 VRAM, and the write bandwidth of OSID15 VRAM.
[0126] If the bandwidth granularity of the GPU's graphics processing unit (GPU) cores and video processing units is 32 bytes, then each increment of the statistics counter represents an increase of 32 bytes in the bandwidth statistics. The specific bandwidth monitoring logic design is described below: like Figure 9As shown, assuming the GPU has only one graphics processing unit (GPU) core, and this GPU supports two virtual machines (VMs), each VM needs to have its System Memory read / write bandwidth and VRAM read / write bandwidth measured, meaning each VM requires four types of bandwidth performance data. The PFM event group size is 4. The first event group measures the System Memory read / write bandwidth of the two VMs, and the second event group measures the VRAM read / write bandwidth of the two VMs. Further assuming the software selects event group 0, and there is a memory access request with OSID 1, accessing System Memory, and a request type of read, then the physical wire corresponding to the counter event in the third box of event group 0 will be pulled high, and the corresponding register value will be incremented by 1. The software will start PFM data collection through the PFM control interface. PFM will record the number of times each counter event in event group 0 is pulled high, and then the software will control PFM data collection to end. At this point, the registers in PFM will record the bandwidth data for this time period, and PFM will write this register data into memory. The software can then read the VM's bandwidth data from memory.
[0127] If the software detects that the bandwidth usage of one or more virtual machines exceeds a threshold through monitoring, it will restrict the task distribution to these virtual machines, such as inserting a 50ms blocking period between virtual machine tasks to reduce the frequency of task distribution. Once the bandwidth usage of these virtual machines falls below the threshold, the software will remove these restrictions.
[0128] The embodiments provided in this application include at least the following points to be protected: 1. GPU hardware virtualization maps each virtual machine to a unique OSID, thereby achieving isolation between virtual machines. Every task load issued by a virtual machine carries this OSID, and bandwidth performance statistics points record bandwidth data into the corresponding performance counters based on the OSID. This enables precise per-VM bandwidth performance monitoring. 2. A highly reusable and scalable bandwidth performance counting module (PFM) is introduced, supporting the grouping of per-VM bandwidth performance data to be statistically analyzed. Software-configured group selection signals enable the currently selected group. The statistical group can be switched during the monitoring period to achieve statistical counting of all per-VM events. This design can significantly reduce hardware footprint even when the number of virtual machines is large.
[0129] It's important to note that based on the virtual machine's built-in OSID, the hardware can achieve precise traffic tracking and corresponding bandwidth monitoring for the virtual machine. This design significantly increases the reuse rate of counting register resources, resulting in substantial optimization of the hardware design area. Assuming the GPU can now support 128 virtual machines, and each virtual machine requires 4 bandwidth performance data points, the traditional method might require 128 * 4 = 512 counting registers to temporarily store bandwidth data. However, using PFM, by grouping the 512 counting events, the number of counting registers can be greatly reduced. If the event group size is set to 128, it's equivalent to dividing these 512 counting events into 4 groups. PFM only requires 128 counting registers, reducing the area by 25% compared to the traditional approach. If the number of supported virtual machines continues to increase, the PFM event group size and grouping method can be increased to adapt, eliminating the need to add counting registers individually, thus greatly improving scalability.
[0130] It should be noted that, in the embodiments of this application, if the bandwidth counting method described above is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the embodiments of this application, or the part that contributes to the related technology, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, mobile hard drives, read-only memory (ROM), magnetic disks, or optical disks. Thus, the embodiments of this application are not limited to any specific hardware, software, or firmware, or any combination of hardware, software, and firmware.
[0131] This application provides a computer device including: a memory for storing computer-executable instructions or computer programs; and a graphics processor including a bandwidth counting device for executing some or all of the steps in the above method when executing the computer-executable instructions or computer programs stored in the memory.
[0132] This application provides a computer-readable storage medium storing a computer program thereon, which, when executed by a graphics processor, implements some or all of the steps in the above-described method. The computer-readable storage medium can be transient or non-transient.
[0133] This application provides a computer program including computer-readable code, wherein when the computer-readable code is run in a computer device, the graphics processor in the computer device performs some or all of the steps in the above-described method.
[0134] This application provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program. When the computer program is read and executed by a computer, it implements some or all of the steps in the above-described method. This computer program product can be implemented specifically through hardware, software, or a combination thereof. In some embodiments, the computer program product is specifically embodied as a computer storage medium; in other embodiments, the computer program product is specifically embodied as a software product, such as a software development kit (SDK), etc.
[0135] It should be noted that the descriptions of the various embodiments above tend to emphasize the differences between them, while their similarities or commonalities can be referred to interchangeably. The descriptions of the above embodiments of the device, storage medium, computer program, and computer program product are similar to the descriptions of the above method embodiments and have similar beneficial effects. For technical details not disclosed in the embodiments of the device, storage medium, computer program, and computer program product of this application, please refer to the descriptions of the method embodiments of this application for understanding.
[0136] It should be noted that, Figure 10 This is a schematic diagram of a hardware entity of a computer device in an embodiment of this application, such as... Figure 10 As shown, the hardware entity of the computer device 1000 includes: a graphics processor 1001, a communication interface 1002, and a memory 1003, wherein: The graphics processor 1001 typically controls the overall operation of the computer device 1000.
[0137] The communication interface 1002 enables computer devices to communicate with other terminals or servers via a network.
[0138] The memory 1003 is configured to store instructions and applications executable by the graphics processor 1001, and can also cache data to be processed or already processed (e.g., image data, audio data, voice communication data, and video communication data) in the graphics processor 1001 and various modules in the computer device 1000. It can be implemented using flash memory or random access memory (RAM). Data transfer between the graphics processor 1001, the communication interface 1002, and the memory 1003 can be performed via bus 1004.
[0139] It should be understood that the phrase "one embodiment" or "an embodiment" throughout the specification means that a specific feature, structure, or characteristic related to the embodiment is included in at least one embodiment of this application. Therefore, "in one embodiment" or "in an embodiment" appearing throughout the specification does not necessarily refer to the same embodiment. Furthermore, these specific features, structures, or characteristics can be combined in any suitable manner in one or more embodiments. It should be understood that in the various embodiments of this application, the sequence numbers of the above steps / processes do not imply a sequential order of execution; the execution order of each step / process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of this application. The sequence numbers of the above embodiments of this application are merely descriptive and do not represent the superiority or inferiority of the embodiments.
[0140] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.
[0141] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods, such as: multiple units or components can be combined, or integrated into another system, or some features can be ignored or not executed. In addition, the coupling, direct coupling, or communication connection between the various components shown or discussed can be through some interfaces, and the indirect coupling or communication connection between devices or units can be electrical, mechanical, or other forms.
[0142] The units described above as separate components may or may not be physically separate. The components shown as units may or may not be physical units. They may be located in one place or distributed across multiple network units. Some or all of the units may be selected to achieve the purpose of this embodiment according to actual needs.
[0143] In addition, each functional unit in the various embodiments of this application can be integrated into one processing unit, or each unit can be a separate unit, or two or more units can be integrated into one unit; the integrated unit can be implemented in hardware or in the form of hardware plus software functional units.
[0144] Those skilled in the art will understand that all or part of the steps of the above method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, it performs the steps of the above method embodiments. The aforementioned storage medium includes various media that can store program code, such as mobile storage devices, read-only memory (ROM), magnetic disks, or optical disks.
[0145] Alternatively, if the integrated units described above are implemented as software functional modules and sold or used as independent products, they can also be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, or the part that contributes to related technologies, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as mobile storage devices, ROM, magnetic disks, or optical disks.
[0146] The above description is merely an embodiment of this application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application.
Claims
1. A bandwidth counting device, characterized in that, The bandwidth counting device includes: a control unit and a counting unit; The control unit is used to determine a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine that issued the memory access request. The counting unit is used to count the bandwidth performance data of the target virtual machine based on the first counting event.
2. The bandwidth counting device according to claim 1, characterized in that, The bandwidth counting device further includes: The control unit is further configured to: determine the storage space to be accessed by the memory access request based on the address of the task load packet of the memory access request; determine the target virtual machine based on the operating system identifier carried by the memory access request; the operating system identifier carried by the memory access request is determined based on the number of the target virtual machine and the preset mapping relationship between the number of the virtual machine and the operating system identifier; and verify the storage space to be accessed by the memory access request based on the access relationship between the target virtual machine and the preset virtual machine and the memory space.
3. The bandwidth counting device according to claim 1, characterized in that, The bandwidth counting device defines multiple event groups; the multiple event groups are obtained by grouping multiple counting events according to the type of memory space; the counting unit includes multiple registers, and the number of registers matches the number of counting events included in the event group; The control unit is further configured to determine a target event group from the plurality of event groups based on the storage space to be accessed by the memory access request; Based on the operating system identifier and request type carried by the memory access request, the first count event is determined from multiple count events of the target event group.
4. The bandwidth counting device according to claim 1, characterized in that, The control unit is further configured to raise the level of the physical wire of the first counting event to a first level; The counting unit is further configured to increment the value of the register corresponding to the first counting event by one when the physical wire of the first counting event is raised to the first level, so as to count the bandwidth performance data of the virtual machine. The "add one" process represents an increase of a target bit in the bandwidth performance of the target virtual machine; the target bit is the bandwidth granularity of the computing unit to which the target virtual machine belongs.
5. The bandwidth counting device according to any one of claims 1 to 4, characterized in that, The bandwidth counting device further includes: A selector is configured to receive a first selection signal for the first counting event and to issue a first counting signal for the first counting event; the first selection signal is triggered when the level of the physical wire of the first counting event is raised to a first level; The counting unit is further configured to increment the value of the register corresponding to the first counting event by one upon receiving the first counting signal; the incrementing process represents an increase of a target bit in the bandwidth performance of the target virtual machine; the target bit is the bandwidth granularity of the computing unit to which the target virtual machine belongs.
6. The bandwidth counting device according to claim 5, characterized in that, The selector is also configured to receive a second selection signal from an external input regarding the second counting event, and to issue a second counting signal regarding the second counting event; The counting unit is further configured to increment the value of the register corresponding to the second counting event by one upon receiving the second counting signal.
7. The bandwidth counting device according to any one of claims 1 to 4, characterized in that, The bandwidth counting device further includes: A configuration interface is provided for setting the configuration information of the bandwidth counting device; the configuration information is used to characterize the counting method of the bandwidth counting device. A control interface is provided for receiving externally input control information; the control information is used to control the counting process of the bandwidth counting device. Write out an interface to write the bandwidth performance data in the register corresponding to the first counting event into memory every first time interval.
8. A graphics processor, characterized in that, The graphics processor includes: multiple computing units, a bandwidth counting device, and multiple memory spaces; each computing unit includes multiple virtual machines; The target virtual machine is used to issue a memory access request, and the target virtual machine is any one of the plurality of virtual machines; The bandwidth counting device is used to determine a first counting event based on the operating system identifier carried in the memory access request, the storage space to be accessed, and the request type; the operating system identifier is used to identify the target virtual machine; and the bandwidth performance data of the target virtual machine is counted based on the first counting event. The memory space is used to process the memory access request.
9. The graphics processor according to claim 8, characterized in that, The graphics processor also includes an on-chip network; The plurality of computing units are connected to the plurality of memory spaces through the on-chip network.
10. A bandwidth counting method, characterized in that, The bandwidth counting method includes: The first counting event is determined based on the operating system identifier, the storage space to be accessed, and the request type carried in the memory access request; the operating system identifier is used to identify the target virtual machine that issued the memory access request. The bandwidth performance data of the target virtual machine is counted based on the first counting event.
11. The bandwidth counting method according to claim 10, characterized in that, The bandwidth counting method further includes: From the bandwidth performance data, determine the starting and ending values of the bandwidth performance of the target virtual machine in the second time period; The bandwidth usage of the target virtual machine is evaluated based on the starting and ending values of bandwidth performance within the second time period, the target time period, and the bandwidth threshold, to obtain the evaluation result.
12. The bandwidth counting method according to claim 11, characterized in that, The bandwidth counting method further includes: If the evaluation results indicate that the target virtual machine is experiencing bandwidth congestion, the frequency of memory access requests to the target virtual machine shall be reduced. If the evaluation results indicate that the target virtual machine is not experiencing bandwidth congestion, the restriction on the frequency of memory access requests to the target virtual machine is removed.
13. A computer device, characterized in that, The computer device includes: Memory is used to store executable instructions or computer programs. A graphics processor including a bandwidth counting device, when executing computer-executable instructions or computer programs stored in the memory, implements the steps of the method according to any one of claims 10 to 12.
14. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a graphics processor, it implements the steps of the method according to any one of claims 10 to 12.
15. A computer program product comprising a non-transitory computer-readable storage medium storing a computer program, wherein the computer program, when read and executed by a computer device, implements the steps of the method of any one of claims 10 to 12.