Hardware virtualization server system and intelligent scheduling method

Through the modularly designed hardware virtualization server system, flexible scheduling and allocation of heterogeneous resources is realized, the problem of low resource utilization in data center servers is solved, and the resource utilization efficiency is improved.

CN120407126AInactive Publication Date: 2025-08-01INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510888672.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-27
Publication Date
2025-08-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Currently, data center servers adopt a fixed hardware configuration architecture, which makes it difficult for heterogeneous computing units to achieve flexible coordination with storage resources, and low resource utilization.

Method used

A hardware virtualization server system with a modular design is adopted, including a computing module, a switching module and a heterogeneous resource module. The switching module responds to the heterogeneous resource request of the computing module, realizes the scheduling and allocation of heterogeneous resources, and supports the module's hot plugging and dynamic reorganization.

Benefits of technology

It improves resource utilization, can quickly respond to the task's immediate resource requirements, and optimizes the global efficiency of resource allocation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120407126A_ABST
    Figure CN120407126A_ABST
Patent Text Reader

Abstract

The invention discloses a hardware virtualization server system and an intelligent scheduling method, and relates to the technical field of computers, the hardware virtualization server system comprises a calculation module, a switching module and at least one heterogeneous resource module, and the calculation module and the at least one heterogeneous resource module are connected through the switching module. The hot plug replacement and dynamic recombination of each module are supported, the heterogeneous resource request of the calculation module is responded through the exchange module, the heterogeneous resources are scheduled, the target heterogeneous resources are allocated to the calculation module to process the target task, and the utilization rate of the resources can be effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technology, and particularly to a hardware virtualization server system and an intelligent scheduling method. Background Art

[0002] Currently, the servers in data centers generally adopt a fixed hardware configuration architecture, where computing, storage, and network resources are integrated through physical binding, resulting in difficulty in achieving flexible coordination between heterogeneous computing units and storage.

[0003] In related technologies, heterogeneous computing units and high-speed storage media are deployed to improve local performance, and static resource partitioning strategies are used to achieve resource allocation. However, heterogeneous resources are difficult to coordinate across domains, resulting in low resource utilization. Summary of the Invention

[0004] This application provides a hardware virtualization server system and an intelligent scheduling method to at least solve the problem of low resource utilization in the server system in related technologies.

[0005] This application provides a hardware virtualization server system, including: a computing module, a switching module, and at least one heterogeneous resource module. Among them, the switching module communicates with the computing module and the heterogeneous resource module respectively.

[0006] The computing module includes multiple computing units. The computing module is configured to send a heterogeneous resource request to the switching module based on the resource requirements of a target processing task to obtain a target heterogeneous resource, and execute the target processing task based on the target heterogeneous resource through the multiple computing units.

[0007] The switching module is configured to obtain the target heterogeneous resource from at least one heterogeneous resource module in response to the heterogeneous resource request sent by the computing unit, and allocate the target heterogeneous resource to the computing unit in the computing module.

[0008] The heterogeneous resource module is configured to provide heterogeneous resources, and the heterogeneous resources include any one of the following:

[0009] Computing acceleration resources, memory resources, storage resources, network resources, channel resources.

[0010] This application also provides an intelligent scheduling method, which is applied to a hardware virtualization server system. The hardware virtualization server system includes: a computing module, a switching module, and at least one heterogeneous resource module. Among them, the switching module communicates with the computing module and the heterogeneous resource module respectively. The method includes:

[0011] The computing module generates a heterogeneous resource request based on the resource requirements of a target processing task, and sends the heterogeneous resource request to the switching module.

[0012] In response to a heterogeneous resource request, the switching module obtains target heterogeneous resources from at least one heterogeneous resource module and allocates the target heterogeneous resources to computing units in the computing module;

[0013] The computing module executes a target processing task based on the target heterogeneous resources through the computing units.

[0014] The hardware virtualization server system proposed in this application includes a computing module, a switching module, and at least one heterogeneous resource module. The computing module and the at least one heterogeneous resource module are connected through the switching module. By adopting a modular design, it supports hot-pluggable replacement and dynamic reorganization of each module. By responding to the heterogeneous resource request of the computing module through the switching module, scheduling the heterogeneous resources and allocating the target heterogeneous resources to the computing module to process the target task, it can effectively improve the resource utilization rate. Description of the Drawings

[0015] In order to more clearly illustrate the embodiments of the present application, the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 Structural schematic of a hardware virtualization server system provided by an embodiment of the present application Figure 1 ;

[0017] Figure 2 Structural schematic diagram of a computing module provided by an embodiment of the present application;

[0018] Figure 3 Structural schematic of a hardware virtualization server system provided by an embodiment of the present application Figure 2 ;

[0019] Figure 4 Schematic diagram of a switch interconnection topology provided by an embodiment of the present application;

[0020] Figure 5 Schematic flowchart of an intelligent scheduling method provided by an embodiment of the present application;

[0021] Figure 6 Schematic flowchart of an intelligent resource scheduling process implemented by an AI-based load prediction model provided by an embodiment of the present application. Detailed Embodiments

[0022] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.

[0023] It should be noted that in the description of the present application, the terms "including", "comprising" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. The terms "first", "second", etc. in the present application are used to distinguish similar objects, rather than to describe a specific order or sequence.

[0024] Currently, the servers in data centers generally adopt a fixed hardware configuration architecture, where computing, storage, and network resources are integrated through physical binding, resulting in rigid resource allocation, difficulty in achieving flexible coordination of heterogeneous computing and storage resources, insufficient matching between hardware resources and upper-layer service requirements, and difficulty in coping with dynamic load changes.

[0025] In related technologies, local performance is improved by deploying heterogeneous computing units and high-speed storage media, and static resource partitioning strategies are used to achieve resource allocation. However, heterogeneous resources are difficult to collaborate across domains, resulting in low resource utilization.

[0026] The present application proposes a hardware virtualization server system, which includes a computing module, a switching module, and at least one heterogeneous resource module. The computing module and at least one heterogeneous resource module are connected through the switching module. By adopting a modular design, it supports hot-swap replacement and dynamic reorganization of each module. The switching module responds to the heterogeneous resource requests of the computing module, schedules the heterogeneous resources, and allocates the target heterogeneous resources to the computing module to process the target tasks, which can effectively improve the resource utilization rate.

[0027] In order to enable those skilled in the art of the present technology to better understand the solution of the present application, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments.

[0028] Figure 1 Structural schematic of a hardware virtualization server system provided for an embodiment of the present application Figure 1 , as Figure 1 shown, the hardware virtualization server system provided for the embodiment of the present application includes a computing module 101, a switching module 102, and at least one heterogeneous resource module 103. Among them, the switching module 102 communicates with the computing module 101 and the heterogeneous resource module 103 respectively.

[0029] The computing module is the core component responsible for executing target tasks in the hardware virtualization service system. The computing module includes multiple computing units, which are the basic execution units of the computing module. For example, the computing unit can be a processor (Central Processing Unit, CPU), etc. Each computing unit can be connected through a high-speed interconnection network, enabling the computing module to be clustered.

[0030] The computing module is used to send heterogeneous resource requests to the switching module based on the resource requirements of the target processing task, so as to obtain the target heterogeneous resources, and execute the target processing task based on the target heterogeneous resources through multiple computing units. The switching module is used to respond to the heterogeneous resource requests sent by the computing units, obtain the target heterogeneous resources from at least one heterogeneous resource module, and allocate the target heterogeneous resources to the computing units in the computing module. The heterogeneous resource module is used to provide heterogeneous resources, and the heterogeneous resources include any one of the following: computing acceleration resources, memory resources, storage resources, network resources, and channel resources.

[0031] The target processing task is a task that needs to be processed by the computing module. The target processing task has clear resource requirements, such as the amount of computation (the number of required CPU cores, etc.), storage capacity, network bandwidth, and whether acceleration resources are required (such as a Graphics Processing Unit, GPU).

[0032] The switching module is the communication hub in the hardware virtualization service system, used to connect the computing module and the heterogeneous resource module, receive and respond to the resource requests sent by the computing module to coordinate the allocation of heterogeneous resources. The switching module can include various types of switching devices, such as a Peripheral Component Interconnect Express (PCIe) switch, a Compute Express Link (CXL) switch, etc. The heterogeneous resource request is a resource requirement instruction sent by the computing module to the switching module, used to indicate the resources required when processing the target processing task.

[0033] Heterogeneous resources are the resources available in a processor system, such as computing acceleration resources, memory resources, storage resources, network resources, channel resources, etc. Among them, computing acceleration resources refer to hardware units designed specifically for specific computing tasks to improve computing efficiency; memory resources refer to remote memory modules connected to computing units to expand the local memory capacity of computing units and provide temporary storage for high-speed and volatile data; storage resources refer to hardware media for data persistent storage or temporary mixed storage, supporting different read / write speeds, capacities, and access modes; network resources refer to hardware components that support data transmission inside and outside the server, used to determine network bandwidth, latency, and the number of concurrent connections, etc.; channel resources refer to logical or physical communication paths connecting different hardware components. The target heterogeneous resource is the specific resource matched and allocated by the switching module from the heterogeneous resource module to the computing unit, and the processing module can execute the target processing task based on the target heterogeneous resource.

[0034] In the related art, the server system connects the computing module and the downstream device through a PCIe switch on the backplane and connects the computing unit and the downstream device through a high-speed copper cable. In this application, the backplane of the PCIe switch is taken out and made into an independent switch chassis, which can improve the flexibility of the architecture. Further, with the update and replacement of PCIe, the rate of PCIe is getting higher and higher, and the requirements for signal attenuation, latency, crosstalk, etc. are getting higher and higher. In the scenario of large data volume and high bandwidth, the copper cable connection has approached the physical limit. In this application, the switching module communicates with the computing module and the heterogeneous resource module through silicon photonics interconnection, and realizes the communication between the computing module and the heterogeneous resource block through silicon photonics interconnection. Silicon photonics interconnection has advantages in realizing long-distance and high-rate transmission. Currently, chip-level optoelectronic integration is gradually mature, and some CPUs and switches, etc. integrate the optical engine and the chip through "co-packaged optics" (CPO) to reduce the cost of optical communication devices and promote the popularization of optical interconnection. The decoupled design of the server system in this application makes the interconnection between modules relatively long. Therefore, the performance of the server system can be further improved through silicon photonics interconnection. Specifically, in the communication process, the electrical signal sent by the sending device is converted into an optical signal through an electro-optical modulator, and after the optical signal is transmitted to the receiving device, it is converted back into an electrical signal by the electro-optical modulator to realize data transmission between different node devices.

[0035] Furthermore, the computing module includes an array of processor units of multiple instruction set types, and the array of processor units is used to deploy different-architecture processor types in a collaborative manner. Specifically, multiple computing units in the computing module can adopt a processor matrix design method of multiple instruction set types to form an array of processor units. Herein, the processor matrix of multiple instruction set types refers to an array structure composed of processors that support different Instruction Set Architectures (ISA). The array of processor units is used to deploy different-architecture processor types in a collaborative manner, that is, processors of multiple instruction set architectures can be deployed in the computing module at the same time, and processors of different architectures can collaborate to complete the same task or independently process different tasks.

[0036] In the embodiments of the present application, the array of processor units can be a two-dimensional array structure or a three-dimensional array structure, and the present application does not make specific limitations thereon. Instruction set architectures include, for example, x86 (such as Intel, AMD, etc.), ARM (such as Ampere), RISC-V, etc. Taking the two-dimensional array structure as an example, Figure 2 is a schematic diagram of a computing module structure provided by an embodiment of the present application, as Figure 2 shown, including processors of multiple instruction set architectures, such as Intel CPUs, AMD CPUs, Ampere CPUs, etc.

[0037] Optionally, the computing unit is configured with a high-speed interconnect controller, and the computing unit is used to achieve memory consistency access across instruction set architectures through the high-speed interconnect controller. Herein, the high-speed interconnect controller is used to break through the memory access limitations of different instruction set architectures, such as uniformly converting the memory access requests of different-architecture processors into standardized protocol instructions to achieve cross-architecture memory consistency, etc.

[0038] Optionally, the computing module can support a resource scheduling strategy for the Non-Uniform Memory Access Architecture (NUMA-aware). In the Non-Uniform Memory Access Architecture, the access times of different processors to different memory regions are different (local memory access is fast, and remote memory access is slow). The resource scheduling strategy for the Non-Uniform Memory Access Architecture can allocate tasks to the cores close to the memory they need according to the NUMA topology characteristics to reduce access latency and improve overall performance.

[0039] Optionally, the computing module can construct a logically unified computing plane through hardware-level virtualization technology and abstract the computing units into a computing resource pool with dynamically adjustable scale. The hardware-level virtualization technology can abstract physical computing resources into a logically unified resource pool, so that upper-layer applications do not need to pay attention to the heterogeneity of the underlying hardware and only need to call a unified virtual resource interface.

[0040] Further, the switching module may include a management engine and a scheduling engine. The management engine is used to monitor the running status of the devices connected to the switching module, and the scheduling engine is used to dynamically configure the uninstallation and installation of the devices connected to the switching module to achieve elastic allocation of resources.

[0041] The hardware virtualization server system provided by the embodiments of the present application includes a computing module, a switching module, and a heterogeneous resource module. The switching module communicates with the computing module and the heterogeneous resource module respectively. The computing module includes a plurality of computing units. The computing module is used to send a heterogeneous resource request to the switching module based on the resource requirements of the target processing task to obtain the target heterogeneous resource, and execute the target processing task based on the target heterogeneous resource through the plurality of computing units; the switching module is used to respond to the heterogeneous resource request sent by the computing unit, obtain the target heterogeneous resource from at least one heterogeneous resource module, and allocate the target heterogeneous resource to the computing units in the computing module; the heterogeneous resource module is used to provide heterogeneous resources. By adopting a modular design, hot pluggable replacement and dynamic reorganization of each module are supported. By the switching module responding to the heterogeneous resource request of the computing module, scheduling the heterogeneous resources and allocating the target heterogeneous resources to the computing module to process the target task, the resource utilization rate can be effectively improved.

[0042] Figure 3 is a schematic structure diagram of a hardware virtualization server system provided by the embodiments of the present application Figure 2 , on the Figure 1 basis of the shown hardware virtualization server system, the heterogeneous resource module may include any at least one of a remote memory module, a storage module, an acceleration module, and a network module. The switching module may include a device interconnection module and a high-speed switching module. As Figure 3 shown, in the hardware virtualization server system provided by the embodiments of the present application, it includes a computing module 101, a device interconnection module 302, a high-speed switching module 303, a remote memory module 304, a storage module 305, an acceleration module 306, and a network module 307. Among them, the device interconnection module 302 communicates with the computing module 101, the remote memory module 304, the storage module 305, the acceleration module 306, and the network module 307 respectively; the high-speed switching module 303 communicates with the computing module 101, the remote memory module 304, the storage module 305, the acceleration module 306, and the network module 307 respectively.

[0043] The device interconnection module 302 and the high-speed switching module 303 are used to respond to the heterogeneous request sent by the computing unit, obtain the target heterogeneous resource from the remote memory module 304, the storage module 305, the acceleration module 306, and the network module 307, and allocate the target heterogeneous resource to the computing units in the computing module.

[0044] The device interconnection module 302 is a high-speed interconnection switching device based on the CXL protocol, which is used to connect the computing module with heterogeneous resource devices supporting the CXL protocol, such as CXL memory expansion cards, CXL accelerators, etc., and provide low-latency and high-bandwidth communication services with memory consistency support.

[0045] The high-speed switching module 303 is a high-speed I / O switching device based on the PCIe protocol, which is used to connect the computing module and traditional PCIe heterogeneous resource devices, such as GPUs, NVMe SSDs, etc., and provide general high-speed data transfer services.

[0046] The device interconnection module 302 and the high-speed switching module 303 may respectively include multiple switches of corresponding types. For example, the device interconnection module 302 may include multiple CXL Switches, and the high-speed switching module 303 may include multiple PCIe Switches. Each Switch is fully interconnected internally, as Figure 4 shown. Figure 4 This is a schematic diagram of the switch interconnection topology provided by an embodiment of this application. As Figure 4 shown in (a), it includes hosts Host 0 – Host 3, GPU arrays GPU 0 – GPU 15, switches SW 0 – SW 3. The host is the initiating node of the target processing task (such as the computing module), the GPU is used to execute parallel computing tasks (such as the acceleration module), and the switch is used to implement the access layer switching between the host and the GPU, and provide local GPU interconnection. Each host is directly connected to each GPU through a x16 interface, and can achieve nanosecond-level response between the host and the GPU. As Figure 4 shown in (b), it includes switches SW 0 – SW 7. The switches are fully cross-connected, and there are independent physical links between any two switches.

[0047] In a possible implementation, the device interconnection module 302 and the high-speed switching module 303 may respectively include different management engines and scheduling engines. For example, the device interconnection module 302 may include a first management engine and a first scheduling engine. The first management engine is used to monitor the operating status of the devices connected to the device interconnection module 302, and the first scheduling engine is used to implement the elastic allocation of the resources of the devices connected to the device interconnection module 302. The high-speed switching module 303 may include a second management engine and a second scheduling engine. The second management engine is used to monitor the operating status of the devices connected to the high-speed switching module 303, and the second scheduling engine is used to implement the elastic allocation of the resources of the devices connected to the high-speed switching module 303.

[0048] In a possible implementation, the switching module includes a unique management engine and a scheduling engine that can be deployed on top of the device interconnection module 302 and the high-speed switching module 303. That is, the management engine can simultaneously monitor the operating states of the devices connected to the device interconnection module 302 and the high-speed switching module 303, and the scheduling engine can implement elastic allocation of resources of the devices connected to the device interconnection module 302 and the high-speed switching module 303.

[0049] The remote memory module 304 is used to provide memory resources. The remote memory module 304 is also used to implement memory expansion through a memory expansion protocol, such as the CXL 3.0 Type3 device protocol. The memory expansion protocol supports high-speed interconnection between the computing module and the external memory device and memory-consistent access. The external memory device is an external hardware device connected to the storage module through a CXL interface and is used to expand the local memory capacity, such as a CXL memory expansion card.

[0050] The remote memory module 304 can implement memory access through a high-speed interconnection protocol (CXL). The high-speed interconnection protocol can include a device interconnection protocol, a memory semantic access protocol, and a cache synchronization protocol. The device interconnection protocol is the CXL.io protocol. The CXL.io protocol is a sub-protocol of CXL. The CXL.io protocol is used to implement memory mapping of PCIe devices. That is, the device interconnection module can map the resources of external devices (such as PCIe devices) to the memory address space of the computing unit through the CXL.io protocol.

[0051] The memory semantic access protocol is the CXL.mem protocol. The CXL.mem protocol is a sub-protocol of CXL. The CXL.mem protocol supports the computing unit to access the memory space of the CXL device through direct memory access (DMA) mode, and can achieve memory-level access speed with low latency and low bandwidth.

[0052] The cache synchronization protocol is the CXL.cache protocol. The CXL.cache protocol is a sub-protocol of CXL. The CXL.cache protocol is used to maintain cache consistency across nodes (the computing module and the CXL device) to ensure that the access results of multiple computing units to shared data are consistent.

[0053] The storage module 305 is used to provide storage resources. The storage module includes at least one storage device, such as a hard disk, an SSD, etc.

[0054] The storage module integrates a hardware offloading engine, which is used to provide a low-latency block device access interface and integrate at least one storage device in the storage module into a logically unified storage resource pool. Specifically, the hardware offloading engine can be a hardware offloading engine based on the storage network protocol (NVMe-oF). The hardware offloading engine is used to provide a low-latency block device access interface and integrate at least one storage device in the storage module into a logically unified storage resource pool, presenting a single and continuous storage space externally, supporting on-demand allocation and elastic expansion.

[0055] The acceleration module 306 includes at least one acceleration device. The acceleration module is used to provide dedicated computing acceleration services, that is, computing acceleration resources, through the acceleration device. The acceleration device can include various types, such as GPUs, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), etc. The acceleration module is also used to build a heterogeneous acceleration resource pool based on a high-speed interconnection protocol and a switch, and achieve the compatibility of each acceleration device through a unified accelerator abstraction layer. Among them, the unified accelerator abstraction layer (Unified Accelerator Abstraction Layer, AAL) is a software middleware between the acceleration module and the switch module. The unified accelerator abstraction layer shields the differences between each acceleration device by defining a standard interface, achieving the compatibility of each acceleration device. Specifically, the high-speed interconnection protocol can be the CXL.io protocol, and the switch can be a PCIe switch. The heterogeneous acceleration resource pool is a computing cluster including multiple acceleration devices. Specifically, multiple and various types of acceleration devices can be connected through a PCIe switch to expand the I / O capabilities of the server system; the communication between the computing module and the acceleration device can be achieved through the CXL.io protocol, such as configuring the video memory allocation of the GPU and triggering the acceleration task of the FPGA through control instructions. For example, when an AI training task needs to start the GPU, the CPU sends a start instruction to the GPU through the CXL.io protocol, and the GPU returns a status confirmation.

[0056] The network module 307 may include multiple network devices, such as switches, routers, etc. The network module is used to provide network resources and channel resources. The network module is also used to dynamically regulate network traffic based on a traffic awareness engine. Among them, the traffic awareness engine is used to statistically analyze network traffic in real time. Specifically, the traffic awareness engine is a dedicated hardware module integrated in the network device, and obtains statistical results by statistically analyzing network traffic in real time, such as the number of data packets, types, delays, etc. The network module can dynamically regulate the allocation strategy of network resources according to the statistical results of the traffic engine, such as adjusting bandwidth, delay, etc., to ensure the quality of service of critical services. For example, when the video conference traffic surges, the queue priority of voice traffic can be set to the highest to ensure that its delay is less than 50ms, and at the same time, the bandwidth occupancy ratio of file transfer traffic is restricted to avoid network congestion.

[0057] The network module 307 is also used to build a virtualized channel based on the device interconnect protocol (CXL.io). The virtualized channel is used to carry network traffic. Specifically, the low-latency characteristics of the CXL.io protocol can be utilized, combined with the high-bandwidth transmission advantage of PCIe 6.0, to establish multiple virtual channels on the physical PCIe link to achieve the transmission of data between the acceleration module and the network module.

[0058] Optionally, the hardware virtualization server system provided by the embodiments of the present application may further include a power consumption management system, which is used to dynamically adjust the power supply status of each module according to the load to achieve energy efficiency optimization.

[0059] The hardware virtualization server system provided by the embodiments of the present application can perform resource scheduling through virtualization software. The virtualization software virtualizes the computing resources in the computing module, the memory resources in the remote memory module, the storage resources in the storage module, the computing acceleration resources in the acceleration module, and the network resources and channel resources in the network module into a unified resource pool, monitors the resources in the unified resource pool and obtains monitoring information, and performs resource scheduling according to the monitoring information.

[0060] Among them, the monitoring resources can be obtained by the monitoring system. The monitoring system may include in-band telemetry and out-of-band chips. The in-band telemetry is used to monitor the resource usage status in real time. The resource usage status includes the computing resource utilization rate, memory resource utilization rate, storage resource occupancy rate, network bandwidth utilization rate, and acceleration device load conditions; the out-of-band chips are used to monitor the operating environment status of the server. The operating environment status includes node temperature, voltage, fan speed, and power supply status, etc. The hardware virtualization server system provided by the embodiments of the present application can also ensure the reliability and bandwidth expansion of the network connection through link aggregation or dual network card binding technology, and improve the reliability of the system.

[0061] The hardware virtualization server system provided by the embodiments of this application, through modular design, supports the pooling of hardware resources. Through the CXL switch, it realizes the logical decoupling and physical sharing of computing resources, memory resources, storage resources, network resources, and acceleration resources, supports dynamic reorganization capabilities and millisecond-level module hot plugging and topology reconstruction, can meet the elastic expansion requirements of services, realizes heterogeneous computing integration through unified addressing of memory space, and achieves the large-scale and scalability of the system.

[0062] Figure 5 It is a schematic flowchart of an intelligent scheduling method provided by the embodiments of this application, which is applied to a hardware virtualization server system. The hardware virtualization server system includes: a computing module, a switching module, and at least one heterogeneous resource module. Among them, the switching module communicates with the computing module and the heterogeneous resource module respectively, as Figure 5 shown, the method includes:

[0063] S501. The computing module generates a heterogeneous resource request based on the resource requirements of the target processing task, and sends the heterogeneous resource request to the switching module.

[0064] [[ID=१२]]S502. In response to the heterogeneous resource request, the switching module obtains the target heterogeneous resource from at least one heterogeneous resource module, and allocates the target heterogeneous resource to the computing unit in the computing module.

[0065] S503. The computing module executes the target processing task based on the target heterogeneous resource through multiple computing units.

[0066] In a possible implementation manner, intelligent resource scheduling can also be realized through an AI-based load prediction model. For example, the long short-term memory network (LSTM) is used to predict the service load, and a resource topology feature map is constructed based on the graph neural network to construct a multi-dimensional dynamic resource decision model. As Figure 6 shown, Figure 6 It is a schematic flowchart of realizing intelligent resource scheduling through an AI-based load prediction model provided by the embodiments of this application. Specifically, with in-band telemetry and out-of-band chip monitoring resources as the model input, LSTM analyzes the changes in service requirements through the monitoring resources. When the service requirements change, it predicts the resource requirements in a future period of time, and then abstracts the resources into graph nodes through the neural network to construct a resource topology feature map. The model integrates the load prediction results and the topology feature map, combines service requirements (such as cost, real-time performance, etc.) and constraint conditions (such as resource availability, reliability, etc.), generates a multi-dimensional resource allocation strategy, and continuously obtains the monitoring resources to perform resource allocation in real time.

[0067] In the embodiments of the present application, by responding to the received resource scheduling request and allocating heterogeneous resources to the computing unit, it is possible to quickly respond to the immediate resource requirements of tasks, improve the utilization rate of resources. Further, it is also possible to analyze the business requirements in real time based on the monitored resources and generate an automatic resource allocation strategy, which can optimize the global efficiency of resource allocation.

[0068] Through the description of the above embodiments, those skilled in the art can clearly understand that the method according to the above embodiments can be implemented by means of software plus a necessary general hardware platform. Of course, it can also be implemented by hardware, but in many cases the former is a better implementation.

[0069] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present application.

[0070] The above has introduced in detail a hardware virtualization server system and an intelligent scheduling method provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application. It should be noted that for those of ordinary skill in the art in the technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the protection scope of the claims of the present application.

Claims

1. A hardware virtualization server system, characterized in that, Including: A computing module, a switching module, and at least one heterogeneous resource module, wherein the switching module communicates with the computing module and the heterogeneous resource module respectively. The computing module includes a plurality of computing units. The computing module is configured to send a heterogeneous resource request to the switching module based on the resource requirements of a target processing task to obtain a target heterogeneous resource, and execute the target processing task based on the target heterogeneous resource through the plurality of computing units. The switching module is configured to obtain a target heterogeneous resource from the at least one heterogeneous resource module in response to the heterogeneous resource request sent by the computing unit, and allocate the target heterogeneous resource to the computing units in the computing module. The heterogeneous resource module is configured to provide heterogeneous resources, and the heterogeneous resources include any one of the following: Computing acceleration resources, memory resources, storage resources, network resources, channel resources.

2. The hardware virtualization server system according to claim 1, wherein The computing module includes a processor unit array of multiple instruction set types, and the processor unit array is used to deploy different architecture processor types in a compatible manner. The computing unit is configured with a high-speed interconnect controller, and the computing unit is used to achieve memory consistency access across instruction set architectures through the high-speed interconnect controller. The computing module supports a resource scheduling strategy for a non-uniform memory access architecture. The computing module constructs a logically unified operation plane through hardware-level virtualization technology, and abstracts the computing units into a computable resource pool with dynamically adjustable scale.

3. The hardware virtualization server system according to claim 1, wherein, The heterogeneous resource module includes a remote memory module, and the remote memory module is used to implement memory expansion through a memory expansion protocol and access the expanded memory through a high-speed interconnect protocol.

4. The hardware virtualization server system according to claim 1, characterized in that The heterogeneous resource module includes a storage module, and the storage module integrates a hardware offloading engine. The hardware offloading engine is used to provide a low-latency block device access interface and integrate at least one storage device in the storage module into a logically unified storage resource pool.

5. The hardware virtualization server system according to claim 1, wherein The heterogeneous resource module includes an acceleration module, and the acceleration module includes at least one acceleration device. The acceleration module is used to provide dedicated computing acceleration services through the acceleration device. The acceleration module is further configured to construct a heterogeneous acceleration resource pool based on a high-speed interconnect protocol and a switch, and achieve compatibility of each acceleration device through a unified accelerator abstraction layer. The unified accelerator abstraction layer is a software middleware between the acceleration module and the switching module. The unified accelerator abstraction layer shields the differences between each acceleration device by defining a standard interface to achieve compatibility of each acceleration device.

6. The hardware virtualization server system according to claim 1, wherein The heterogeneous resource module includes a network module, and the network module is used to dynamically regulate network traffic based on a traffic awareness engine. The traffic awareness engine is used to statistically analyze the network traffic in real time. The network module is further configured to construct a virtualized channel based on a device interconnect protocol, and the virtualized channel is used to carry the network traffic.

7. The hardware virtualization server system according to claim 1, wherein The switching module includes a device interconnection module and a high-speed switching module. The device interconnection module is used to connect the computing module with devices supporting high-speed interconnection protocols, and the high-speed switching module is used to connect the computing module with devices supporting high-speed switching protocols. The device interconnection module and the high-speed switching module are fully interconnected internally; The switching module includes a management engine and a scheduling engine. The management engine is used to monitor the operating status of devices connected to the switching module, and the scheduling engine is used to dynamically configure the unloading and installation of devices connected to the switching module to achieve elastic allocation of resources.

8. The hardware virtualization server system according to claim 1, wherein The hardware virtualization server system performs resource scheduling through virtualization software; The virtualization software virtualizes the computing resources in the computing module, the memory resources in the remote memory module, the storage resources in the storage module, the computing acceleration resources in the acceleration module, and the network resources and channel resources in the network module into a unified resource pool; The virtualization software is used to monitor the resources in the unified resource pool and obtain monitoring information, and the virtualization software performs resource scheduling according to the monitoring information.

9. The hardware virtualization server system according to claim 1, wherein The switching module communicates with the computing module and the heterogeneous resource module respectively through silicon optical interconnection.

10. An intelligent scheduling method, characterized in that, Applied to a hardware virtualization server system, the hardware virtualization server system includes: a computing module, a switching module, and at least one heterogeneous resource module, wherein the switching module communicates with the computing module and the heterogeneous resource module respectively. The method includes: Based on the resource requirements of the target processing task, the computing module generates a heterogeneous resource request and sends the heterogeneous resource request to the switching module; In response to the heterogeneous resource request, the switching module obtains the target heterogeneous resource from the at least one heterogeneous resource module and allocates the target heterogeneous resource to the computing units in the computing module; The computing module executes the target processing task based on the target heterogeneous resource through the computing unit.

Citation Information

Patent Citations

  • Edge-end-side heterogeneous computing power resource scheduling system and method

    CN119537018A

  • Cloud computing technology-based heterogeneous computing power providing system and method

    WO2024230603A1