Resource Management Method, Electronic Device, and Storage Medium
By scanning PCI devices and registering virtual information to create precise device bindings, the method addresses the inefficiencies of coarsely managed resource allocation in K8S clusters, improving resource management in heterogeneous environments.
Patent Information
- Application Number
- CN202510201019.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2045-02-24
AI Technical Summary
In Kubernetes clusters, existing resource management methods are coarse in granularity and cannot carefully manage the performance differences of heterogeneous computing devices, affecting the performance and efficiency of resource applications.
By deploying device agent components at nodes, scanning PCI devices to obtain device information of computing devices, registering virtual device information, reporting to the cluster, and fine-grained management is carried out, including dynamic management of device information, performance information and status information.
The fine-grained resource management of heterogeneous clusters is realized, the performance and efficiency of resource management are improved, and the resource utilization rate of Kubernetes clusters is improved.
Smart Images

Figure CN119690594B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of artificial intelligence, and in particular, to a resource management method, an electronic device, and a storage medium. Background Art
[0002] In a large-scale cluster environment, thousands or even tens of thousands of computing devices are often deployed. These thousands of computing devices may come from different manufacturers or belong to different product models.
[0003] To achieve dynamic allocation and management of computing device resources, currently in a K8S (Kubernetes) cluster, the manufacturers and quantities of computing devices can be reported through a device plugin, so as to provide a basic resource view for the cluster. Accordingly, the cluster can specify the manufacturers and quantities of the computing devices to be used.
[0004] However, there may also be differences in the performance of computing devices of different product models under the same manufacturer, and these differences will directly affect the performance and efficiency of resource application. How to achieve more fine-grained resource management has become an urgent problem to be solved. Summary of the Invention
[0005] The present invention provides a resource management method, an electronic device, and a storage medium to solve the defects in the related art of coarse resource management granularity and affecting the performance and efficiency of resource application.
[0006] The present invention provides a resource management method, including:
[0007] Obtain device information of computing devices deployed at a node, and report the device information to a cluster side to trigger the cluster side to create device resources of the computing devices based on the device information;
[0008] Register virtual device information of the node and report it to the cluster side;
[0009] When a resource management request is monitored, query a resource deployment unit corresponding to the resource management request, determine target device information from device-bound resources of the resource deployment unit, and mount a computing device corresponding to the target device information to the resource deployment unit;
[0010] The device-bound resources are created by the cluster side based on the resource management request and the virtual device information, and the target device information is device information corresponding to target device resources selected by the cluster side from all unbound device resources based on the device-bound resources;
[0011] The registering the virtual device information of the node and reporting it to the cluster side includes:
[0012] The node agent component based on the node registers virtual device information to trigger the node agent component to report the virtual device information to the cluster side, and the number of virtual devices included in the virtual device information is greater than or equal to the number of computing devices deployed at the node.
[0013] According to a resource management method provided by the present invention, the obtaining of the device information of the computing devices deployed at the node includes:
[0014] Scan the PCI devices at the node to obtain the device models of the PCI devices;
[0015] Based on the device models of the PCI devices, determine the device information of the computing devices corresponding to the PCI devices.
[0016] According to a resource management method provided by the present invention, the determining of the device information of the computing devices corresponding to the PCI devices based on the device models of the PCI devices includes:
[0017] Based on the device models of the PCI devices, determine the device models of the computing devices corresponding to the PCI devices;
[0018] Based on the correspondence between the general device model and the detailed device model, determine the general device model corresponding to the device model of the computing device, and determine the device information based on the general device model.
[0019] According to a resource management method provided by the present invention, the device resources include the device information, performance information, and status information of the computing devices.
[0020] According to a resource management method provided by the present invention, the querying of the resource deployment unit corresponding to the resource management request includes:
[0021] Receive a virtual device allocation request sent by the node agent component of the node;
[0022] Based on the virtual device identifier corresponding to the virtual device allocation request, query the resource deployment unit corresponding to the resource management request.
[0023] The present invention also provides a resource management method, including:
[0024] Receive the device information of the computing devices deployed at the node and the virtual device information of the node, and create the device resources of the computing devices based on the device information. The device information is obtained by the device agent component at the node, the virtual device information is registered and reported based on the node agent component of the node, and the number of virtual devices included in the virtual device information is greater than or equal to the number of computing devices deployed at the node;
[0025] Upon receiving a resource management request, based on the resource management request and the virtual device information, create device-bound resources for the resource deployment unit corresponding to the resource management request.
[0026] Based on the device-bound resources, select target device resources from all unbound device resources, and update the target device information in the device-bound resources based on the device information corresponding to the target device resources, so that when the device proxy component monitors the resource management request, it determines the target device information from the device-bound resources of the resource deployment unit corresponding to the resource management request, and mounts the computing device corresponding to the target device information to the resource deployment unit.
[0027] According to a resource management method provided by the present invention, the device resources include device information, performance information, and status information of the computing device.
[0028] According to a resource management method provided by the present invention, the step of selecting target device resources from all unbound device resources based on the device-bound resources includes:
[0029] Based on the status information of all device resources, determine unbound device resources;
[0030] Based on the device-bound resources, as well as the device information and performance information of the unbound device resources, select target device resources from all unbound device resources.
[0031] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, it implements the resource management method as described in any one of the above.
[0032] The present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements the resource management method as described in any one of the above.
[0033] The present invention also provides a computer program product, including a computer program. When the computer program is executed by a processor, it implements the resource management method as described in any one of the above.
[0034] The resource management method, electronic device, and storage medium provided by the present invention obtain the device information of the computing devices deployed at the nodes and report it to the cluster side to create device resources. Additionally, at the cluster side, device-bound resources are created based on the resource management request, thereby combining the device resources and the device-bound resources to achieve fine-grained resource management and allocation, and further improving the performance and efficiency of resource management in a heterogeneous cluster. Brief Description of the Drawings
[0035] In order to more clearly illustrate the technical solutions in the present invention or related technologies, the following will briefly introduce the drawings required for use in the embodiments or related technology descriptions. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0036] Figure 1 It is a schematic flowchart of resource management based on a K8S cluster in related technologies;
[0037] Figure 2 It is one of the schematic flowcharts of the resource management method provided by the present invention;
[0038] Figure 3 It is another schematic flowchart of the resource management method provided by the present invention;
[0039] Figure 4 It is yet another schematic flowchart of the resource management method provided by the present invention;
[0040] Figure 5 It is one of the schematic structural diagrams of the resource management device provided by the present invention;
[0041] Figure 6 It is another schematic structural diagram of the resource management device provided by the present invention;
[0042] Figure 7 It is the schematic structural diagram of the electronic device provided by the present invention. Detailed Embodiments
[0043] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will clearly and completely describe the technical solutions in the present invention in conjunction with the drawings in the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art without creative efforts based on the embodiments in the present invention fall within the scope of protection of the present invention.
[0044] In a large-scale cluster environment, thousands or even tens of thousands of computing devices are often deployed. These thousands of computing devices may come from different manufacturers or belong to different product models under the same manufacturer.
[0045] Here, the computing device may specifically be an artificial intelligence chip. For example, it may be a GPU (Graphics Processing Unit), or it may also be any one or more of a TPU (Tensor Processing Unit), an NPU (Neural network Processing Unit), a DPU (Deep learning Processing Unit), an APU (Accelerated Processing Unit), and a GPGPU (General-Purpose computing on Graphics Processing Unit). The embodiments of the present invention do not make specific limitations thereto.
[0046] The K8S cluster is an open-source cloud server management platform. Taking the computing device as a GPU as an example, to achieve dynamic allocation and management of GPU resources, currently in the K8S cluster, the manufacturer and quantity of the GPU can be reported through a device plugin, so as to provide a basic resource view for the cluster.
[0047] For example, Figure 1 is a schematic flow chart of resource management based on the K8S cluster in the related art. As Figure 1 shown, at node 1 and node 2, GPUs of two manufacturers, manufacturer A and manufacturer B, are respectively deployed. Among them, the GPUs belonging to manufacturer A include GPU1 - GPU4, and the GPUs belonging to manufacturer B include GPU5 - GPU8. For the GPUs belonging to manufacturer A, manufacturer A provides a device plugin for the K8S cluster; in addition, for the GPUs belonging to manufacturer B, manufacturer B also provides a device plugin for the K8S cluster. Therefore, the device plugin of manufacturer A is installed at node 1, and the device plugin of manufacturer B is installed at node 2, so that the K8S cluster can identify the GPUs belonging to manufacturer A and manufacturer B respectively.
[0048] In the process of resource management, it can be divided into the following 5 steps:
[0049] Step ①: The device plugins of each manufacturer respectively report the quantity of the GPUs belonging to their own manufacturers to the node proxy component (also known as Kubelet) at the corresponding node. For example, at node 1, the information reported by the device plugin of manufacturer A may be A.com / GPU:4.
[0050] Step ②: The node proxy components of each node report the number of GPUs of each manufacturer to the API (Application Programming Interface) server (also known as the API Server).
[0051] Step ③: The API server receives the Pod creation request sent by the client. The Pod creation request here can carry the required resources. For example, if two GPUs of manufacturer B are needed, it can be identified as B.com / GPU:2. Here, the Pod is the smallest deployment unit in the K8S cluster and can represent a container or a combination of multiple containers.
[0052] Step ④: When the K8S scheduler (also known as the K8S Scheduler) detects the Pod creation, it can schedule the Pod to Node 2 according to the resources required for the Pod creation.
[0053] Step ⑤: The node proxy component of Node 2 interacts with the device plugin of manufacturer B required for the Pod creation, and allocates two GPUs from the GPUs of manufacturer B to the Pod.
[0054] It can be seen that in the above solution, for the situation where there are GPUs of multiple manufacturers in the K8S cluster, it is necessary to install the device plugins of each manufacturer to facilitate the recognition of the corresponding GPUs.
[0055] Moreover, since the GPUs of different product models under the same manufacturer may also have differences in performance, simply describing the number of GPUs of the manufacturer without reporting the specific GPU models cannot truly reflect the performance of each GPU. For example, manufacturer A has two GPUs with different models, and the parameters and performance of the two GPUs are completely different, but the device plugin of manufacturer A can only report that there are two GPUs of manufacturer A deployed, in a form similar to A.com / GPU:2. Correspondingly, since the cluster cannot obtain the specific GPU models, when the cluster performs resource management, it can only specify the manufacturer and quantity of the GPUs to be used, and cannot specify the specific GPU models, which results in a relatively coarse granularity of the current resource management method and directly affects the performance and efficiency of resource application.
[0056] In view of the above problems, an embodiment of the present invention provides a resource management method. Figure 2 is one of the flow diagrams of the resource management method provided by the present invention. As Figure 2 shown, this method is applied to the nodes of a heterogeneous cluster, specifically applied to the device agent (also known as the Device Agent) component at the node. Here, the heterogeneous cluster refers to a computing cluster composed of computing devices from different manufacturers or different product models under the same manufacturer. This method includes:
[0057] Step 210: Obtain the device information of the computing devices deployed at the node, and report the device information to the cluster side to trigger the cluster side to create the device resources of the computing devices based on the device information.
[0058] Specifically, in a heterogeneous cluster, a node is a physical server in the heterogeneous cluster, and the node bears the computing resources required for the operation of the heterogeneous cluster. One or more computing devices can be deployed on each node.
[0059] Moreover, in the embodiments of the present invention, for any node, the computing devices deployed at the node can come from different manufacturers or belong to different product models under the same manufacturer. Or, for any node, the computing devices deployed at the node can come from the same manufacturer or belong to the same product model under the same manufacturer, while for different nodes, the computing devices deployed at different nodes can come from different manufacturers or belong to different product models under the same manufacturer.
[0060] The device proxy component can be deployed at the node of the heterogeneous cluster, and the device proxy component can detect all the computing devices deployed at the node, thereby obtaining the device information of all the computing devices deployed at the node. It should be noted that the device proxy component is a component set at the node, different from the node proxy component (also known as Kubelet). The device proxy component can communicate with the node proxy component. In the embodiments of the present invention, the device proxy component is a newly added component at each node in the heterogeneous cluster, used to obtain the device information of the computing devices deployed at the node and cooperate with the node proxy component to perform information interaction with the cluster side.
[0061] To obtain the device information of the computing devices deployed at the node, it can be specifically implemented by scanning the PCI (Peripheral Component Interconnect) devices at the node; or it can be implemented by integrating the communication interfaces of the computing devices under each manufacturer. The embodiments of the present invention do not make specific limitations in this regard.
[0062] The device information obtained thereby can include information identifying the computing devices, specifically, it can be the device ID (Identity document) of the computing devices. In addition, the device information can also include the manufacturer information of the computing devices, for example, it can be the device manufacturer ID of the computing devices. Or, the device information can also include information such as the node name where the computing device is located and the PCI address.
[0063] After obtaining the device information of each computing device deployed at the node, the device proxy component can report the device information of each computing device to the cluster side. The cluster side referred to here is the resource management side of the heterogeneous cluster. For example, it can be a K8S cluster, which specifically includes an API server, a scheduler, etc. in the K8S cluster. The embodiments of the present invention do not make specific limitations on this.
[0064] Correspondingly, after the cluster side receives the device information of the computing devices at the node reported by the device proxy component at the node, it can create the device resources of the computing devices deployed at the node based on this.
[0065] The device resources here are a kind of custom resources. For example, in the case where the computing device is a GPU, the device resources can specifically be GPU CRD (Custom Resource Definitions). It can be understood that the device resources maintained by the cluster side can persistently and granularly reflect the situation of each computing device deployed at the node. Compared with the related technology that only reports the number of GPUs of each manufacturer, it provides a richer resource view for the cluster side.
[0066] Step 220, register the virtual device information of the node and report it to the cluster side.
[0067] Specifically, in order to continue to use the original resource management logic of the K8S cluster and avoid the large workload caused by modifying the resource management logic of the K8S cluster, the device proxy component can also register the virtual device information and report it to the cluster side. The virtual device information here is for the node, and it is the information obtained by registering virtual computing devices. The virtual device information can indicate that a large number of computing devices of any manufacturer are deployed at the node. For example, the virtual device information can be recorded as FakeDevice, and the specific form is example.com / FakeDevice:1000, where 1000 is the number of virtual computing devices in the example, and the actual value can be any number greater than or equal to the number of actually deployed computing devices at the node. By setting a number much larger than the number of actually deployed computing devices, oversubscription can be achieved to improve the resource utilization rate of the K8S cluster.
[0068] Correspondingly, the cluster side can receive the virtual device information of the node.
[0069] Step 230, in the case of monitoring a resource management request, query the resource deployment unit corresponding to the resource management request, determine the target device information from the device binding resources of the resource deployment unit, and mount the computing device corresponding to the target device information to the resource deployment unit;
[0070] The device-bound resource is created by the cluster side based on the resource management request and the virtual device information, and the target device information is the device information corresponding to the target device resource selected by the cluster side from all unbound device resources based on the device-bound resource.
[0071] Specifically, in the resource management scenario, the user side can send a resource management request to the cluster side. Here, the user side is used to implement the interaction between the user and the heterogeneous cluster. The user side can interact with the cluster side. For example, the user side can interact with the API server in the cluster side to implement functions such as querying the cluster status and managing the cluster resources, and then utilize the computing resources provided by the computing devices at each node through the cluster side.
[0072] The resource management request here can be a Pod creation request, that is, a request for creating a Pod and configuring computing resources for the created Pod. Here, the Pod is the smallest deployment unit in the K8S cluster and can represent a container or a combination of multiple containers. In the embodiments of the present invention, the Pod created based on the resource management request is denoted as the resource deployment unit.
[0073] The resource management request often carries the requirements of the user for the computing resources expected to be configured for the created Pod, such as the number and manufacturer of the computing devices to be configured. Thus, after receiving the resource management request, the cluster side can create a corresponding resource deployment unit based on the resource management request, and configure corresponding computing resources for the resource deployment unit based on the requirements for computing resources in the resource management request and the virtual device information reported by the nodes. Here, the computing resources configured for the resource deployment unit are denoted as the device-bound resource.
[0074] Here, the device-bound resource is a custom resource. For example, in the case where the computing device is a GPU, the device-bound resource can specifically be a GPUBinding CRD. It can be understood that the device-bound resources maintained by the cluster side can persistently describe the binding relationship between the resource deployment unit and the computing device.
[0075] Since the virtual device information is the virtual device information registered by the device proxy component, configuring the device binding resource for the resource deployment unit can only achieve the binding between the resource deployment unit and the node, but does not achieve the real binding with the computing devices deployed in the node. Therefore, the cluster side also needs to select target device resources that meet the requirements of the computing resources expected to be configured by the resource management request from all the unbound device resources deployed at the node, that is, select target device resources corresponding to the requirements proposed by the device binding resource from all the unbound device resources, and write the device information corresponding to the target device resources into the device binding resource. At this time, the device information of the actual computing device that is actually bound to the resource deployment unit, that is, the target device information, is written into the device binding resource.
[0076] Thus, the device binding resource can include the name of the resource deployment unit that needs to use the computing device, the actual container name, and also the name of the device resource that needs to be bound. And in the case where the real binding with the computing device has not been achieved, the name of the device resource that needs to be bound included in the device binding resource is empty; while in the case where the computing device to be bound is determined, the name of the device resource that needs to be bound included in the device binding resource can be expressed as the node identifier and the serial number of the computing device deployed on the node, for example, node1-0, indicating the 0th card on node1 is bound.
[0077] In addition, the device proxy component at the node can also listen to the above-mentioned resource management request. And in the case where the resource management request is determined to be listened to, the device proxy component can query the resource deployment unit corresponding to the resource management request, thereby querying the device binding resource of the resource deployment unit and reading the target device information from the device binding resource. That is, the device proxy component can determine the computing device deployed at the node that needs to be bound to the resource deployment unit through the target device information in the device binding resource, that is, the computing device corresponding to the target device information.
[0078] Thus, after determining the computing device corresponding to the target device information, the computing device can be mounted to the resource deployment unit to support the operation of the resource deployment unit.
[0079] In the method provided by the embodiment of the present invention, the device information of the computing devices deployed at the node is obtained and reported to the cluster side to create device resources, and in addition, device binding resources are created at the cluster side based on the resource management request, so as to realize fine-grained resource management and allocation by combining the device resources and the device binding resources, and further improve the performance and efficiency of resource management in a heterogeneous cluster.
[0080] Based on the above embodiment, in step 210, the obtaining the device information of the computing devices deployed at the node includes:
[0081] Scan the PCI devices at the node to obtain the device models of the PCI devices;
[0082] Based on the device models of the PCI devices, determine the device information of the computing devices corresponding to the PCI devices.
[0083] Specifically, compared with the related art where device plugins corresponding to different manufacturers' GPUs need to be configured separately, in the embodiments of the present invention, only a device proxy component needs to be installed at the node to replace the device plugins of each manufacturer. That is, no matter how many manufacturers' computing devices are deployed at the node, only one device proxy component needs to be installed. That is, in the embodiments of the present invention, the device proxy component can obtain the device information of all computing devices deployed at the node.
[0084] In order to obtain the device information of computing devices of different manufacturers, the device proxy component can scan the PCI devices at the node. It can be understood that computing devices of various manufacturers need to be connected to the node through the PCI bus, and the PCI devices here are the computing devices connected through the PCI bus. That is, the device proxy component can obtain all the computing devices deployed at the node by scanning the PCI devices at the node.
[0085] By scanning the PCI devices, the device models of the PCI devices can be obtained. For example, it may include the Vendor ID (supplier identifier) and Device ID (device identifier) of the PCI devices. It can be understood that in the device model of the PCI device, the Vendor ID can reflect the manufacturer of the computing device, and the Device ID can reflect the model of the computing device. For example, the device model is [10de:2503], where 10de is the Vendor ID, representing a certain manufacturer, and 2503 is the device number under this manufacturer; 1ee0 in [1ee0: 0003] represents another manufacturer, and 0003 represents the device number under this manufacturer. Thus, after obtaining the device model of the PCI device, the device information such as the manufacturer and model of the computing device corresponding to the PCI device can be determined based on this.
[0086] In the method provided by the embodiments of the present invention, by scanning the PCI devices at the node, the unified acquisition of the device information of computing devices of each manufacturer deployed at the node is realized, and there is no need to install the device plugins of each manufacturer separately, greatly improving the convenience of resource management.
[0087] Moreover, since the device information carries the device model, the resource requirements for using specific models can be met in resource management.
[0088] Based on any of the above embodiments, in step 210, determining the device information of the computing device corresponding to the PCI device based on the device model of the PCI device includes:
[0089] Determining the device model of the computing device corresponding to the PCI device based on the device model of the PCI device;
[0090] Determining the general device model corresponding to the device model of the computing device based on the correspondence between the general device model and the detailed device model, and determining the device information based on the general device model.
[0091] Specifically, in the device model of the PCI device, the Device ID corresponds to the device model of the computing device. For example, for the case where the device model of the PCI device is [10de:2503], 2503 therein is the Device ID, and the device model of the corresponding computing device can be determined based on 2503; for the case where the device model of the PCI device is [1ee0:0003], 0003 therein is the Device ID, and the device model of the corresponding computing device can be determined based on 0003. Thus, after obtaining the device model of the PCI device, the device model of the corresponding computing device can be obtained.
[0092] In practical applications, there may be multiple models of computing devices under the same series of a manufacturer. These computing devices have different models and slightly different performance parameters, but they are all regarded as a type of computing device for resource allocation and management when used by users.
[0093] To facilitate the unified management of similar computing devices, the correspondence between the general device model and the detailed device model can be pre-constructed. Among them, the general device model is the device model commonly used for displaying to users or the device model that users are accustomed to using, such as the series number of a type of computing device or the abbreviated model; the detailed device model is the detailed and unabbreviated device model of the computing device.
[0094] It can be understood that the device model of the computing device obtained based on the device model of the PCI device is usually the detailed device model. Thus, based on the correspondence between the general device model and the detailed device model, the device model of the computing device can be mapped to the corresponding general device model, and the device information of the computing device can be determined based on this general device model, thereby realizing the unified management of a type of computing device.
[0095] Based on any of the above embodiments, the device resources include the device information, performance information, and status information of the computing device.
[0096] Specifically, in addition to describing the device information of a computing device, the device resources can also describe the performance information and status information of the computing device.
[0097] Among them, the device information of the computing device may include the device ID, and may also include the manufacturer ID, node name, PCI address, etc. The performance information of the computing device may include the video memory size, power, CPU affinity, etc. The status information of the computing device may indicate whether the computing device is allocated for use, and in addition, may also include the health status of the computing device, etc. For example, in the case where the computing device is allocated for use, the status information of the computing device may include the information of the allocated resource deployment unit, for example, may include the name and namespace of the allocated resource deployment unit, and the ID of the user corresponding to the resource deployment unit, etc.
[0098] Correspondingly, when the cluster side selects a target device resource from all unbound device resources based on the device-bound resources, specifically, it can determine the unbound device resources from all device resources based on the status information of all device resources, and then select a target device resource that matches the requirements of the device-bound resources from among all unbound device resources based on the device information and performance information of all unbound device resources.
[0099] Based on any of the above embodiments, in step 220, registering the virtual device information of the node and reporting it to the cluster side includes:
[0100] Registering the virtual device information based on the node proxy component of the node to trigger the node proxy component to report the virtual device information to the cluster side, and the number of virtual devices included in the virtual device information is greater than or equal to the number of computing devices deployed at the node.
[0101] Specifically, at the node, a node proxy component (also known as Kubelet) is also deployed. The device proxy component can communicate with the node proxy component. Specifically, it can register the virtual device information of the node through the node proxy component, and then the node proxy component reports the virtual device information to the cluster side, thereby realizing the registration and reporting of the virtual device information.
[0102] Here, in the case where the number of virtual devices included in the virtual device information is greater than the number of computing devices actually deployed at the node, oversubscription can be achieved to improve the resource utilization rate of the K8S cluster.
[0103] Based on any of the above embodiments, in step 230, querying the resource deployment unit corresponding to the resource management request includes:
[0104] Receiving a virtual device allocation request sent by the node proxy component of the node;
[0105] Query the resource deployment unit corresponding to the resource management request based on the virtual device identifier corresponding to the virtual device allocation request.
[0106] Specifically, at the node, the node proxy component can listen for resource management requests. And when the node proxy component listens to a resource management request, the node proxy component can send a virtual device allocation request to the device proxy component. Here, the virtual device allocation request is used to request the device proxy component to allocate a virtual device to meet the requirements of the resource management request.
[0107] Correspondingly, the device proxy component can receive the virtual device allocation request sent by the node proxy component. And after receiving the virtual device allocation request, it can obtain the device identifier of the virtual device to be allocated therefrom, that is, the virtual device identifier. Since the device-bound resource is created at the cluster side based on the resource management request and the virtual device information, a binding relationship is established at the cluster side between the resource management request and the allocated virtual device identifier. Based on this, the device proxy component can query the resource deployment unit corresponding to the resource management request bound to the virtual device identifier according to the virtual device identifier in the virtual device allocation request, thereby realizing the query of the resource deployment unit, and further providing conditions for mounting the computing device for the resource deployment unit.
[0108] Based on any of the above embodiments, Figure 3 is the second flow diagram of the resource management method provided by the present invention. As Figure 3 shown, this method is applied to the cluster side of a heterogeneous cluster. Here, the cluster side referred to is the resource management side of the heterogeneous cluster. For example, it can be a K8S cluster, specifically including the API server, scheduler, etc. in the K8S cluster. The embodiments of the present invention do not make specific limitations on this.
[0109] This method includes:
[0110] Step 310, receive the device information of the computing device deployed at the node and the virtual device information of the node, and create the device resource of the computing device based on the device information, where the device information is obtained by the device proxy component at the node.
[0111] Specifically, at the nodes of the heterogeneous cluster, a device agent (also known as Device Agent) component can be deployed. The device agent component can detect all the computing devices deployed at the node, and thus obtain the device information of each of the computing devices deployed at the node. It should be noted that the device agent component is a component set at the node, which is different from the node agent component (also known as Kubelet). The device agent component can communicate with the node agent component. In the embodiments of the present invention, the device agent component is a newly added component at each node in the heterogeneous cluster, which is used to obtain the device information of the computing devices deployed at the node, and cooperate with the node agent component to interact with the cluster side for information.
[0112] The device agent component can obtain the device information of the computing devices deployed at the node. Specifically, it can be achieved by scanning the PCI devices at the node; or, it can be achieved by integrating the communication interfaces of the computing devices under each manufacturer. The embodiments of the present invention do not make specific limitations on this.
[0113] The obtained device information may include information identifying the computing device, specifically, it may be the device ID of the computing device. In addition, the device information may also include the manufacturer information of the computing device, for example, it may be the device vendor ID of the computing device. Or, the device information may also include information such as the node name and PCI address where the computing device is located.
[0114] After obtaining the device information of each computing device deployed at the node, the device agent component can report the device information of each computing device to the cluster side.
[0115] Correspondingly, after the cluster side receives the device information of the computing devices at the node reported by the device agent component at the node, it can create the device resources of the computing devices deployed at the node based on this.
[0116] The device resources here are a type of custom resources. For example, in the case where the computing device is a GPU, the device resources can specifically be GPU CRD. It can be understood that the device resources maintained by the cluster side can persistently and granularly reflect the situation of each computing device deployed at the node, providing a richer resource view for the cluster side compared with only reporting the number of GPUs of each manufacturer in the related art.
[0117] In addition, in order to continue to use the original resource management logic of the K8S cluster and avoid the large amount of work caused by modifying the resource management logic of the K8S cluster, the device proxy component can also register virtual device information and report it to the cluster side. The virtual device information here is the information obtained by registering virtual computing devices for a node. The virtual device information can indicate that a large number of computing devices of any manufacturer are deployed at this node. For example, the virtual device information can be recorded as FakeDevice, and the specific form is example.com / FakeDevice:1000, where 1000 is the number of virtual computing devices in the example, and the actual value can be any number greater than or equal to the number of actually deployed computing devices at this node. By setting a number much larger than the number of actually deployed computing devices, oversubscription can be achieved to improve the resource utilization rate of the K8S cluster.
[0118] Correspondingly, the cluster side can receive the virtual device information of the node.
[0119] Step 320, in the case of receiving a resource management request, create a device-bound resource for the resource deployment unit corresponding to the resource management request based on the resource management request and the virtual device information.
[0120] Specifically, in a resource management scenario, the user side can send a resource management request to the cluster side. Here, the user side is used to implement the interaction between the user and the heterogeneous cluster. The user side can interact with the cluster side. For example, the user side can interact with the API server in the cluster side to implement functions such as querying the cluster status and managing the cluster resources, and then use the cluster side to utilize the computing resources provided by the computing devices at each node.
[0121] The resource management request here can be a Pod creation request, that is, a request for creating a Pod and configuring computing resources for the created Pod. Here, the Pod is the smallest deployment unit in the K8S cluster and can represent a container or a combination of multiple containers. In the embodiments of the present invention, the Pod created based on the resource management request is denoted as a resource deployment unit.
[0122] The resource management request often carries the requirements of the user for the computing resources expected to be configured for the created Pod, such as the number and manufacturer of the computing devices required to be configured. Therefore, after receiving the resource management request, the cluster side can create a corresponding resource deployment unit based on the resource management request and configure the corresponding computing resources for the resource deployment unit based on the requirements for computing resources in the resource management request and the virtual device information reported by the node. Here, the computing resources configured for the resource deployment unit are denoted as device-bound resources.
[0123] Here, the device binding resource is a custom resource. For example, in the case where the computing device is a GPU, the device binding resource can specifically be the GPUBinding CRD. It can be understood that the device binding resource maintained on the cluster side can persistently describe the binding relationship between the resource deployment unit and the computing device.
[0124] Step 330: Based on the device binding resource, select a target device resource from all unbound device resources, and update the target device information in the device binding resource based on the device information corresponding to the target device resource, so that when the device proxy component monitors the resource management request, it can determine the target device information from the device binding resource of the resource deployment unit corresponding to the resource management request, and mount the computing device corresponding to the target device information to the resource deployment unit.
[0125] Specifically, since the virtual device information is the virtual device information registered by the device proxy component, configuring the device binding resource for the resource deployment unit can only achieve the binding between the resource deployment unit and the node, but does not achieve the actual binding with the computing devices deployed in the node. Therefore, the cluster side also needs to select, from all the unbound device resources deployed at the node, the target device resources that meet the requirements of the computing resources expected to be configured by the resource management request, that is, select the target device resources corresponding to the requirements proposed by the device binding resource from all the unbound device resources, and write the device information corresponding to the target device resources into the device binding resource. At this time, the device information of the actual computing device actually bound to the resource deployment unit, that is, the target device information, is written into the device binding resource.
[0126] The device proxy component at the node can also monitor the above-mentioned resource management request. And when it determines that it has monitored the resource management request, the device proxy component can query the resource deployment unit corresponding to the resource management request, thereby querying the device binding resource of this resource deployment unit and reading the target device information from the device binding resource. That is, the device proxy component can determine the computing device deployed at the node that needs to be bound to the resource deployment unit through the target device information in the device binding resource, that is, the computing device corresponding to the target device information.
[0127] Thus, after determining the computing device corresponding to the target device information, the computing device can be mounted to the resource deployment unit to support the operation of the resource deployment unit.
[0128] In the method provided by the embodiments of the present invention, the device information of the computing devices deployed at the nodes is obtained and reported to the cluster side to create device resources. Additionally, device-bound resources are created at the cluster side based on resource management requests, so as to achieve fine-grained resource management and allocation by combining the device resources and the device-bound resources, thereby improving the performance and efficiency of resource management.
[0129] Based on any of the above embodiments, the device resources include the device information, performance information, and status information of the computing devices.
[0130] Specifically, in addition to being able to describe the device information of the computing devices, the device resources can also describe the performance information and status information of the computing devices.
[0131] Among them, the device information of the computing devices can include the device ID, and can also include the manufacturer ID, node name, PCI address, etc. The performance information of the computing devices can include the video memory size, power, CPU affinity, etc. The status information of the computing devices can indicate whether the computing devices are allocated for use. In addition, it can also include the health status of the computing devices. For example, in the case where the computing devices are allocated for use, the status information of the computing devices can include the information of the allocated resource deployment units, such as the name and namespace of the allocated resource deployment units, and the ID of the user corresponding to the resource deployment units.
[0132] Correspondingly, in step 330, the selecting of the target device resources from all unbound device resources based on the device-bound resources includes:
[0133] Determining the unbound device resources based on the status information of all device resources;
[0134] Selecting the target device resources from all unbound device resources based on the device-bound resources, as well as the device information and performance information of the unbound device resources.
[0135] Specifically, when the cluster side selects the target device resources from all unbound device resources based on the device-bound resources, it can specifically determine the unbound device resources from all device resources based on the status information of all device resources, and then select the target device resources that match the requirements of the device-bound resources from them based on the device information and performance information of all unbound device resources.
[0136] The target device resources selected here are the unbound device resources whose device information and performance information match the requirements of the device-bound resources. For example, it can be the device resources whose models match the requirements of the device-bound resources, or it can be the device resources whose video memory sizes meet the requirements of the device-bound resources. The embodiments of the present invention do not make specific limitations in this regard.
[0137] Based on any of the above embodiments, Figure 4 is the third schematic flowchart of the resource management method provided by the present invention. As Figure 4 shown, GPUs 1 - 4 are deployed at node 1, and GPUs 5 - 8 are deployed at node 2. GPUs 1 - 4 can belong to the same manufacturer or different manufacturers, and GPUs 5 - 8 can belong to the same manufacturer or different manufacturers. For the convenience of explaining the steps in the resource management process, node 1 is taken as an example below:
[0138] Step ①: The device proxy component at node 1 can detect the GPU devices deployed on node 1, and in cooperation with the configurations in the pre - set ConfigMap, obtain the device information of the detected GPU devices, and report the device information to the API server at the cluster end.
[0139] Here, detecting the GPU devices deployed on node 1 can be achieved by scanning the PCI devices on node 1;
[0140] In addition, the ConfigMap can include the mapping relationships between the general device models and the detailed device models of computing devices of various manufacturers. Based on the ConfigMap, the model of the GPU detected by the device proxy component can be mapped to the general device model. Thus, the device proxy component can generate device information based on the general device model. Further, the ConfigMap can include the display_name of the general device model of the manufacturer and the Device ID of the detailed device model.
[0141] Step ②: The device proxy component at node 1 can create corresponding device resource GPU CRDs for each device information through the API server at the cluster end. The device resources here can represent the specific information of the GPUs deployed on node 1.
[0142] Step ③: The device proxy component at node 1 can register virtual device information through the node proxy component at node 1, thereby using the original K8S control mechanism to achieve resource management.
[0143] Step ④: The node proxy component at node 1 can report the virtual device information to the API server at the cluster end.
[0144] Step ⑤: The API server at the cluster end receives the Pod creation request sent by the user end. The Pod creation request here can declare the required device models and quantities. For example, two computing devices of a certain model are required.
[0145] Step ⑥: The device controller GPU Controller on the cluster side can declare the number of devices required according to the Pod creation request. Create the corresponding number of device binding resources GPUBinding CRD, and at the same time indicate the number of virtual devices required by the Pod corresponding to the Pod creation request. For example, in the case of requiring 2 GPUs, it can be specified as example.com / fakeDevice:2.
[0146] Step ⑦: The scheduler GPU scheduler on the cluster side can, according to the device binding resource GPUBinding CRD of the Pod, find a suitable device resource GPU CRD from all unbound device resources GPU CRD according to the scheduling policy for binding, and return the corresponding node name as the scheduling result. For example, the scheduling result can be Node 1. Since Node 1 declares in the virtual device information that it deploys a large number of virtual devices, such as 1000 virtual devices, it can be ensured that Node 1 necessarily meets the condition of the number of virtual devices required by the Pod corresponding to the Pod creation request. The scheduling policy here can be a pre-set policy for matching requirements and device resources.
[0147] Step ⑧: The node proxy component on Node 1 can listen for Pod creation, and after listening to the Pod creation, request the device proxy component to allocate virtual devices with the number of virtual devices required by the Pod. Correspondingly, the device proxy component can query the Pod of the requested device according to the ID of the allocated virtual device, then the device proxy component can obtain the device binding resource GPUBinding CRD of the Pod through the API server on the cluster side, and find the actual GPU device to be mounted according to the device binding resource GPUBinding CRD of the Pod, and return the device information of the GPU device to the node proxy component, thereby mounting the GPU device to the Pod.
[0148] Next, the resource management device provided by the present invention will be described. The resource management device described below can be correspondingly referred to the resource management method described above.
[0149] Figure 5 is one of the structural schematic diagrams of the resource management device provided by the present invention, as Figure 5 shown. This device can be applied to the device proxy component at the node. This device includes:
[0150] The device information reporting unit 510 is used to obtain the device information of the computing device deployed at the node, and report the device information to the cluster side to trigger the cluster side to create the device resources of the computing device based on the device information.
[0151] The virtual information reporting unit 520 is configured to register the virtual device information of the node and report it to the cluster side;
[0152] The binding unit 530 is configured to, when a resource management request is monitored, query the resource deployment unit corresponding to the resource management request, determine target device information from the device binding resources of the resource deployment unit, and mount the computing device corresponding to the target device information to the resource deployment unit;
[0153] The device binding resources are created by the cluster side based on the resource management request and the virtual device information, and the target device information is the device information corresponding to the target device resources selected by the cluster side from all unbound device resources based on the device binding resources.
[0154] In the device provided in the embodiment of the present invention, the device information of the computing device deployed at the node is obtained and reported to the cluster side to create device resources. Additionally, device binding resources are created by the cluster side based on the resource management request, so as to realize fine-grained resource management and allocation by combining the device resources and the device binding resources, thereby improving the performance and efficiency of resource management in a heterogeneous cluster.
[0155] Based on any of the above embodiments, the device information reporting unit is specifically configured to:
[0156] Scan the PCI devices at the node to obtain the device models of the PCI devices;
[0157] Based on the device models of the PCI devices, determine the device information of the computing devices corresponding to the PCI devices.
[0158] Based on any of the above embodiments, the device information reporting unit is specifically configured to:
[0159] Based on the device models of the PCI devices, determine the device models of the computing devices corresponding to the PCI devices;
[0160] Based on the correspondence between the general device model and the detailed device model, determine the general device model corresponding to the device model of the computing device, and determine the device information based on the general device model.
[0161] Based on any of the above embodiments, the device resources include the device information, performance information, and status information of the computing device.
[0162] Based on any of the above embodiments, the virtual information reporting unit is specifically configured to:
[0163] The node proxy component based on the node registers virtual device information, so as to trigger the node proxy component to report the virtual device information to the cluster side, and the number of virtual devices included in the virtual device information is greater than or equal to the number of computing devices deployed at the node.
[0164] Based on any of the above embodiments, the binding unit is specifically configured to:
[0165] Receive a virtual device allocation request sent by the node proxy component of the node;
[0166] Based on the virtual device identifier corresponding to the virtual device allocation request, query the resource deployment unit corresponding to the resource management request.
[0167] Figure 6 It is the second structural schematic diagram of the resource management device provided by the present invention. As Figure 6 shown, this device can be applied to the cluster side, and this device includes:
[0168] A device resource creation unit 610, configured to receive the device information of the computing devices deployed at the node and the virtual device information of the node, and create device resources for the computing devices based on the device information, where the device information is obtained by the device proxy component at the node;
[0169] A bound resource creation unit 620, configured to create device bound resources for the resource deployment unit corresponding to the resource management request based on the resource management request and the virtual device information when receiving the resource management request;
[0170] A resource management unit 630, configured to select target device resources from all unbound device resources based on the device bound resources, and update the target device information in the device bound resources based on the device information corresponding to the target device resources, so that when the device proxy component monitors the resource management request, it determines the target device information from the device bound resources of the resource deployment unit corresponding to the resource management request, and mounts the computing device corresponding to the target device information to the resource deployment unit.
[0171] In the device provided by the embodiments of the present invention, the device information of the computing devices deployed at the node is obtained and reported to the cluster side to create device resources, and in addition, device bound resources are created at the cluster side based on the resource management request, so as to realize fine-grained resource management and allocation by combining the device resources and the device bound resources, thereby improving the performance and efficiency of resource management in a heterogeneous cluster.
[0172] Based on any of the above embodiments, the device resources include the device information, performance information, and status information of the computing devices.
[0173] Based on any of the above embodiments, selecting a target device resource from all unbound device resources based on the device-bound resources includes:
[0174] Determining unbound device resources based on the status information of all device resources;
[0175] Selecting a target device resource from all unbound device resources based on the device-bound resources, the device information, and the performance information of the unbound device resources.
[0176] Figure 7 An example of a schematic physical structure diagram of an electronic device is shown as Figure 7 shown. The electronic device may include: a processor 710, a communications interface 720, a memory 730, and a communication bus 740. Among them, the processor 710, the communications interface 720, and the memory 730 complete mutual communication through the communication bus 740. The processor 710 may call logical instructions in the memory 730 to execute a resource management method, and the method includes:
[0177] Obtaining the device information of the computing devices deployed at the nodes and reporting the device information to the cluster side to trigger the cluster side to create the device resources of the computing devices based on the device information;
[0178] Registering the virtual device information of the nodes and reporting it to the cluster side;
[0179] When a resource management request is monitored, querying the resource deployment unit corresponding to the resource management request, determining the target device information from the device-bound resources of the resource deployment unit, and mounting the computing device corresponding to the target device information to the resource deployment unit;
[0180] The device-bound resources are created by the cluster side based on the resource management request and the virtual device information, and the target device information is the device information corresponding to the target device resource selected by the cluster side from all unbound device resources based on the device-bound resources.
[0181] Alternatively, the method includes:
[0182] Receiving the device information of the computing devices deployed at the nodes and the virtual device information of the nodes, and creating the device resources of the computing devices based on the device information, where the device information is obtained by a device proxy component at the nodes;
[0183] In the case of receiving a resource management request, based on the resource management request and the virtual device information, create a device-bound resource for the resource deployment unit corresponding to the resource management request;
[0184] Based on the device-bound resource, select a target device resource from all unbound device resources, and update the target device information in the device-bound resource based on the device information corresponding to the target device resource, so that when the device proxy component monitors the resource management request, it determines the target device information from the device-bound resource of the resource deployment unit corresponding to the resource management request, and mounts the computing device corresponding to the target device information to the resource deployment unit.
[0185] In addition, when the logical instructions in the above-mentioned memory 730 are implemented in the form of a software functional unit and sold or used as an independent product, they can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that makes a contribution to the related technology, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROMs, Read-Only Memories), random access memories (RAMs, Random Access Memories), magnetic disks, or optical discs that can store program codes.
[0186] On the other hand, the present invention also provides a computer program product, the computer program product includes a computer program, the computer program can be stored on a non-transitory computer-readable storage medium, and when the computer program is executed by a processor, the computer can execute the resource management method provided by the above-mentioned various methods. The method includes:
[0187] Obtain the device information of the computing device deployed at the node, and report the device information to the cluster side to trigger the cluster side to create the device resource of the computing device based on the device information;
[0188] Register the virtual device information of the node and report it to the cluster side;
[0189] In the case of monitoring a resource management request, query the resource deployment unit corresponding to the resource management request, and determine the target device information from the device-bound resource of the resource deployment unit, and mount the computing device corresponding to the target device information to the resource deployment unit;
[0190] The device-bound resource is created by the cluster side based on the resource management request and the virtual device information, and the target device information is the device information corresponding to the target device resource selected by the cluster side from all unbound device resources based on the device-bound resource.
[0191] Alternatively, the method includes:
[0192] Receiving the device information of the computing devices deployed at the receiving node and the virtual device information of the node, and creating the device resources of the computing devices based on the device information, where the device information is obtained by the device proxy component at the node;
[0193] When receiving a resource management request, creating a device-bound resource for the resource deployment unit corresponding to the resource management request based on the resource management request and the virtual device information;
[0194] Based on the device-bound resource, selecting a target device resource from all unbound device resources, and updating the target device information in the device-bound resource based on the device information corresponding to the target device resource, so that when the device proxy component monitors the resource management request, it determines the target device information from the device-bound resource of the resource deployment unit corresponding to the resource management request, and mounts the computing device corresponding to the target device information to the resource deployment unit.
[0195] In another aspect, the present invention also provides a non-transitory computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, it implements a resource management method provided by the above methods. The method includes:
[0196] Obtaining the device information of the computing devices deployed at the node and reporting the device information to the cluster side to trigger the cluster side to create the device resources of the computing devices based on the device information;
[0197] Registering the virtual device information of the node and reporting it to the cluster side;
[0198] When monitoring a resource management request, querying the resource deployment unit corresponding to the resource management request, and determining the target device information from the device-bound resource of the resource deployment unit, and mounting the computing device corresponding to the target device information to the resource deployment unit;
[0199] The device-bound resource is created by the cluster side based on the resource management request and the virtual device information, and the target device information is the device information corresponding to the target device resource selected by the cluster side from all unbound device resources based on the device-bound resource.
[0200] Alternatively, the method includes:
[0201] Receiving device information of a computing device deployed at a receiving node and virtual device information of the node, and creating device resources of the computing device based on the device information, where the device information is obtained by a device proxy component at the node;
[0202] In the case of receiving a resource management request, creating device binding resources of a resource deployment unit corresponding to the resource management request based on the resource management request and the virtual device information;
[0203] Based on the device binding resources, selecting target device resources from all unbound device resources, and updating target device information in the device binding resources based on device information corresponding to the target device resources, so that when the device proxy component monitors the resource management request, determining target device information from the device binding resources of the resource deployment unit corresponding to the resource management request, and mounting a computing device corresponding to the target device information to the resource deployment unit.
[0204] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative labor.
[0205] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, and of course also by hardware. Based on such an understanding, the above technical solution, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0206] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A resource management method, characterized in that, A device proxy component applied to a node in a heterogeneous cluster. The device proxy component is used to detect all computing devices deployed at the node, thereby obtaining the device information of each of all computing devices deployed at the node. The device information carries the device model. The method includes: Obtain the device information of the computing devices deployed at the node and report the device information to the cluster side to trigger the cluster side to create the device resources of the computing devices based on the device information; Register the virtual device information of the node and report it to the cluster side; When a resource management request is monitored, query the resource deployment unit corresponding to the resource management request, and determine the target device information from the device-bound resources of the resource deployment unit, and mount the computing device corresponding to the target device information to the resource deployment unit; The device-bound resources are created by the cluster side based on the resource management request and the virtual device information. The target device information is the device information corresponding to the target device resource selected by the cluster side from all unbound device resources based on the device-bound resources; The registering the virtual device information of the node and reporting it to the cluster side includes: Register the virtual device information based on the node proxy component of the node to trigger the node proxy component to report the virtual device information to the cluster side. The number of virtual devices included in the virtual device information is greater than or equal to the number of computing devices deployed at the node; The obtaining the device information of the computing devices deployed at the node includes: Scan the PCI devices at the node to obtain the device models of the PCI devices; Based on the device models of the PCI devices, determine the device information of the computing devices corresponding to the PCI devices; The determining the device information of the computing devices corresponding to the PCI devices based on the device models of the PCI devices includes: Based on the device models of the PCI devices, determine the device models of the computing devices corresponding to the PCI devices; Based on the correspondence between the general device model and the detailed device model, determine the general device model corresponding to the device model of the computing device, and determine the device information based on the general device model. The general device model is the serial number of the computing device or the abbreviated model, and the detailed device model is the unabbreviated device model of the computing device.
2. The resource management method according to claim 1, characterized in that The device resources include the device information, performance information, and status information of the computing devices.
3. The resource management method according to claim 1, characterized in that The querying the resource deployment unit corresponding to the resource management request includes: Receive a virtual device allocation request sent by the node proxy component of the node; Based on the virtual device identifier corresponding to the virtual device allocation request, query the resource deployment unit corresponding to the resource management request.
4. A resource management method, characterized in that, Applied to the cluster side of a heterogeneous cluster, the method includes: Receive the device information of the computing devices deployed at the receiving node and the virtual device information of the node, and create device resources for the computing devices based on the device information. The device information is obtained by the device proxy component at the node, and the virtual device information is registered and reported based on the node proxy component of the node. The number of virtual devices included in the virtual device information is greater than or equal to the number of computing devices deployed at the node. The device proxy component is used to detect all the computing devices deployed at the node, thereby obtaining the device information of each of the computing devices deployed at the node. The device information carries the device model. The device proxy component is used to scan the PCI devices at the node to obtain the device models of the PCI devices; based on the device models of the PCI devices, determine the device models of the computing devices corresponding to the PCI devices; based on the correspondence between the general device model and the detailed device model, determine the general device model corresponding to the device model of the computing device, and determine the device information based on the general device model. The general device model is the serial number of the computing device or the abbreviated model, and the detailed device model is the unabbreviated device model of the computing device. In the case of receiving a resource management request, create device binding resources for the resource deployment unit corresponding to the resource management request based on the resource management request and the virtual device information. Based on the device binding resources, select target device resources from all unbound device resources, and update the target device information in the device binding resources based on the device information corresponding to the target device resources, so that when the device proxy component monitors the resource management request, it determines the target device information from the device binding resources of the resource deployment unit corresponding to the resource management request, and mounts the computing device corresponding to the target device information to the resource deployment unit.
5. The resource management method according to claim 4, characterized in that, The device resources include the device information, performance information, and status information of the computing devices.
6. The resource management method according to claim 5, characterized in that, The selecting target device resources from all unbound device resources based on the device binding resources includes: Determine unbound device resources based on the status information of all device resources. Based on the device binding resources, and the device information and performance information of the unbound device resources, select target device resources from all unbound device resources.
7. An electronic device, comprising a memory, a processor, and a computer program stored on the memory and running on the processor, characterized in that, When the processor executes the computer program, it implements the resource management method according to any one of claims 1 to 6.
8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the resource management method according to any one of claims 1 to 6.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the resource management method according to any one of claims 1 to 6.
Citation Information
Patent Citations
GPU resource management method, system and device and storage medium
CN116089009A