Method for allocating graphics processor resources
By generating configuration instructions in the computing cluster and dynamically generating configuration templates based on hardware data of various GPU types, the problem of poor compatibility in MIG configuration management is solved, realizing automated and intelligent management in multi-type GPU environments, and improving GPU resource utilization and the efficiency of artificial intelligence platforms.
Patent Information
- Application Number
- CN202511426249.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-30
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-09-30
AI Technical Summary
Existing MIG configuration and management methods lack compatibility and cannot adapt to various GPU environments, resulting in low management efficiency and low compatibility, which affects the GPU resource management efficiency and user experience of artificial intelligence platforms.
By generating configuration instructions at the management node in the computing cluster, configuration templates are dynamically generated based on hardware data of various graphics processor types, enabling adaptive allocation of different types of GPUs. By utilizing the collaborative work of the MIG-CORE and MIG-AGENT components, automated and intelligent management of MIG configuration is achieved.
It enables automated and intelligent management of MIG configuration in various GPU environments, improving GPU resource utilization and the management efficiency of the artificial intelligence platform, as well as enhancing user experience and productivity.
Smart Images

Figure CN120892216B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of computers, and particularly relates to a graphics processor resource allocation method. BACKGROUND
[0002] In order to improve the utilization rate of graphics processor (Graphics Processing Unit, GPU for short) resources, a multi-instance GPU (Multi-Instance GPU, MIG for short) technology is proposed. The technology can divide a physical GPU into multiple independent MIG instances according to different requirements, each instance can be independently scheduled and used, and effectively meets the differentiated requirements of different tasks for GPU resources. Different models of GPUs have great differences in MIG configuration, including the number of divisible slices, supported profile types, maximum instance number, etc.
[0003] At present, the configuration and management of MIG mainly adopts a method of developing a special script or tool, integrating MIG configuration steps of a specific GPU model, and realizing part of automatic operation through node labels of a container orchestration platform component (Kubernetes, K8S) for configuration and information synchronization.
[0004] However, the related MIG configuration and management method has poor compatibility, and can only integrate the automatic configuration process for a specific GPU model, and cannot meet the adaptive configuration requirements of different types of GPUs. SUMMARY
[0005] The present application provides a graphics processor resource allocation method to at least solve the problem of low MIG configuration and management efficiency and low compatibility in a multi-type GPU environment caused by the fact that only the automatic configuration process for a specific GPU model can be integrated in the related art.
[0006] The application provides a graphics processor resource allocation method, a management node running in a computing cluster, comprising: generating a configuration instruction according to a configuration request sent by a target object; the configuration request comprises a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates are generated according to hardware data of a plurality of graphics processor types; the plurality of configuration templates correspond to the plurality of graphics processor types one by one; a configuration template in the plurality of configuration templates comprises a plurality of configuration instance combinations of a graphics processor of a corresponding graphics processor type; a configuration instance combination in the plurality of configuration instance combinations refers to a split configuration of splitting resources of a graphics processor of a corresponding graphics processor type into a plurality of independent units; and sending the configuration instruction to a target computing node in the target computing node list to instruct the target computing node to allocate resources of a graphics processor on the target computing node according to the configuration instruction.
[0007] The application also provides a graphics processor resource allocation method, a target computing node running in a computing cluster, comprising: allocating resources of a graphics processor on the target computing node according to a configuration instruction sent by a management node in the computing cluster; the configuration instruction is generated based on a configuration request sent by a target object; the configuration request comprises a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates are generated by the management node according to hardware data of a plurality of graphics processor types; the plurality of configuration templates correspond to the plurality of graphics processor types one by one; a configuration template in the plurality of configuration templates comprises a plurality of configuration instance combinations of a graphics processor of a corresponding graphics processor type; a configuration instance combination in the plurality of configuration instance combinations refers to a split configuration of splitting resources of a graphics processor of a corresponding graphics processor type into a plurality of independent units.
[0008] The application also provides an electronic device, comprising: a memory for storing a computer program; and a processor for executing the computer program to implement the steps of any of the above methods.
[0009] The application also provides a computer readable storage medium, wherein the computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of any of the above methods.
[0010] The application also provides a computer program product, comprising a computer program, and the computer program is executed by a processor to implement the steps of any of the above methods.
[0011] By the present application, by running on the management node in the computing cluster, the corresponding configuration instruction is generated according to the configuration request content, and the configuration instruction is sent to the computing node in the specified target computing node list, so as to guide the adaptive allocation of the graphics processor resource on the computing node. A plurality of configuration templates are dynamically generated according to the hardware data of a plurality of graphics processor types of graphics processors in the computing cluster, and the plurality of configuration templates correspond to the plurality of graphics processor types one by one, ensuring the accuracy of the configuration. Each configuration template contains a plurality of configuration instance combinations, i.e. segmentation configurations, of the corresponding graphics processor type, which plans a resource allocation scheme for different types of nodes, lays a foundation for automatic resource allocation, avoids the disadvantages of manual configuration and specific model integrated configuration, realizes the automation and intelligent management of MIG configuration in a multi-type GPU environment, improves the GPU resource management efficiency of the artificial intelligence platform, improves the user experience and output efficiency, and solves the problem of low MIG configuration and management efficiency and low compatibility in the related art due to the integration of automatic configuration process for specific GPU models. BRIEF DESCRIPTION OF DRAWINGS
[0012] In order to more clearly illustrate the embodiments of the present application, the drawings needed in the embodiments will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.
[0013] Figure 1 is an application scenario diagram of a graphics processor resource allocation method according to an embodiment of the present application.
[0014] Figure 2 is a flowchart of an optional graphics processor resource allocation method according to an embodiment of the present application.
[0015] Figure 3 is an interaction diagram of information verification between a management node and a computing node according to an embodiment of the present application.
[0016] Figure 4 is a timing diagram of a management node performing graphics processor resource allocation according to an embodiment of the present application.
[0017] Figure 5 is a timing diagram of a computing node responding to a configuration instruction according to an embodiment of the present application.
[0018] Figure 6 is a timing diagram of a management node and a computing node responding to a restart event according to an embodiment of the present application.
[0019] Figure 7is a structural block diagram of an optional package option default value configuration device according to an embodiment of the present application. DETAILED DESCRIPTION
[0020] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by a person of ordinary skill in the art without creative work fall within the protection scope of the present application.
[0021] It should be noted that, in the description of the present application, the terms “comprise”, “contain” or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device comprising a series of elements not only includes those elements, but also includes other elements not explicitly listed or inherent to such a process, method, article or device. The terms “first”, “second” and the like in the present application are used to distinguish similar objects, and are not used to describe a specific order or sequence.
[0022] In order for those skilled in the art to better understand the present application, the present application will be further described in detail below with reference to the drawings and specific embodiments.
[0023] According to an aspect of an embodiment of the present application, a graphics processor resource allocation method is provided. Optionally, in the present embodiment, the above-mentioned graphics processor resource allocation method can be applied to, but is not limited to, a hardware environment including a computing cluster as shown in Figure 1 . The computing cluster can include one management node 102 and a plurality of computing nodes 104. Here, the management node 102 is the control center of the computing cluster, used to monitor the cluster state, schedule tasks, allocate resources and coordinate communication between the computing nodes 104. The computing node 104 is the core component of the computing cluster, used to execute computing tasks and data processing. One or more GPUs can be installed on each computing node 104. The types of graphics processors of the plurality of GPUs installed on each computing node 104 can be the same or different. Optionally, the computing node 104 can be composed of multiple servers.
[0024] The graphics processor resource allocation method of the present embodiment can be executed by the computing cluster, running on the management node in the computing cluster. Taking the graphics processor resource allocation method in the present embodiment as an example, which is executed by the computing cluster, Figure 2 is a flowchart of an optional graphics processor resource allocation method according to an embodiment of the present application. As shown in Figure 2 , the flow of the method can include the following steps S202 to S204.
[0025] Step S202, according to the configuration request sent by the target object, generating a configuration instruction; the configuration request includes a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates are generated according to the hardware data of a plurality of graphics processor types of graphics processors; the plurality of configuration templates correspond one-to-one to the plurality of graphics processor types; the configuration template in the plurality of configuration templates includes a plurality of configuration instance combinations of the graphics processor of the corresponding graphics processor type; the configuration instance combination in the plurality of configuration instance combinations refers to a split configuration of splitting the resources of the graphics processor of the corresponding graphics processor type into a plurality of independent units;
[0026] Step S204, sending the configuration instruction to the target computing node in the target computing node list to instruct the target computing node to allocate the resources of the graphics processor on the target computing node according to the configuration instruction.
[0027] The graphics processor resource allocation method in the embodiment can be applied to the field of artificial intelligence and applied to the scene of allocating resources to the graphics processor of an artificial intelligence platform. At present, with the rapid development of artificial intelligence, graphics processors (Graphics Processing Unit, GPU for short) have become the core acceleration hardware for deep learning model training and inference due to their powerful parallel computing capabilities. Among them, GPU is a hardware component specially used for processing graphics rendering and large-scale parallel computing, which is commonly used in the field of artificial intelligence to accelerate deep learning model training and inference. Related GPUs occupy a high share in the market due to their excellent performance and perfect software ecosystem. However, with the rapid iteration of technology, the release frequency of GPUs is accelerating, and different models of GPUs have significant differences in core size, memory specifications, and computing capabilities. For example, different models of GPUs have unique hardware parameters and technical characteristics.
[0028] To improve the utilization of GPU resources, a multi-instance GPU (MIG) technology is introduced. MIG is a technology that can divide a physical GPU into multiple independent virtual GPU instances, each instance has independent computing cores, memory and resources, and can be used independently by different applications or tasks. Among them, the MIG technology can divide a physical GPU into multiple independent MIG instances according to different needs, each instance can be scheduled and used independently, effectively meeting the differentiated needs of different tasks for GPU resources. However, different models of GPUs have great differences in MIG configuration, including the number of slices that can be divided, the supported profile type, the maximum number of instances, etc. In the MIG technology, Slice refers to the basic resource unit of the GPU after being divided, and different Slice combinations can form MIG instances of different specifications.
[0029] Currently, the market mainly uses the following two methods for MIG configuration and management: (1) manual maintenance and configuration according to the relevant configuration manual. That is, technical personnel need to manually execute a series of commands to configure MIG instances according to the official documents of different GPU models, including enabling MIG mode, selecting Profile, creating instances, etc. This method relies heavily on manual operation, is extremely inefficient, and is prone to configuration errors due to human error. For large artificial intelligence platforms that include multiple GPU models, the workload of manual configuration is huge, and it is difficult to meet the needs of platform-scale management. (2) For a specific GPU type, integrate the configuration steps to achieve automatic configuration and feedback. That is, by developing a special script or tool, the MIG configuration steps for a specific GPU model are integrated, and the configuration and information synchronization are realized through the node label of a container orchestration platform component (Kubernetes, K8S). K8S is an open-source platform for container orchestration that can manage the deployment, expansion and running state of containers. However, this method has poor compatibility and relies heavily on K8S label information synchronization. Frequent MIG configuration may cause the MIG configuration and platform information to be out of sync due to delayed label information synchronization. At the same time, this method can only support pre-integrated specific GPU models and cannot adapt to the rapid update pace of GPU models. For scenarios where multiple different types of GPUs exist on the same computing node, this method cannot achieve differentiated MIG configuration management.
[0030] Therefore, the above two methods cannot properly solve the problem of automatic and intelligent management of MIG configuration in a multi-type GPU environment, resulting in low efficiency of artificial intelligence platforms in GPU resource management, which adversely affects the user's experience and output efficiency.
[0031] In order to at least partially solve the above technical problems, in the embodiment, the low efficiency and poor compatibility problems of MIG configuration management in the prior art are solved, and an artificial intelligence platform MIG automatic configuration and adaptive management method is innovatively proposed. A MIG configuration management method capable of adapting to multiple GPU models is constructed, automatic generation, execution and management of MIG configuration are realized, and the system realizes automatic and intelligent management of MIG configuration of different types of GPUs by the cooperative operation of the first component deployed on the management node in the computing cluster and the second component running on each computing node, effectively supports the rapid adaptation of new GPUs, improves the utilization rate of GPU resources and the operation and maintenance efficiency of the platform, and significantly improves the management efficiency and utilization rate of GPU resources of the artificial intelligence platform.
[0032] In the embodiment, the first component is a control component running on the management node in the computing cluster. Optionally, the first component can be a MIG configuration core (MIG Configuration Core, MIG-CORE for short). The MIG-CORE is a core component in the embodiment, responsible for MIG configuration generation of the computing node, combined calculation of the MIG instance, and issuing of MIG configuration, change, and destruction actions. The second component refers to a component running on the target computing node in the target computing node list. Optionally, the second component can be a MIG configuration agent component (MIG Configuration Agent, MIG-AGENT for short). The MIG-AGENT is another core component in the embodiment, running on the computing node, responsible for obtaining GPU type information, executing MIG configuration and destruction operations, performing state synchronization and node monitoring.
[0033] The target object is an object that sends a configuration request to the first component according to the computing task demand and the resource utilization strategy. For example, the target object can be a user, a user interface, a system administrator, or an automatic task scheduling system. The target object can indicate the specific demand of resource allocation by selecting a specific configuration template, a configuration instance combination, and a specified target computing node list. The configuration request is a request sent by the target object to the first component, used to express the configuration demand of the graphics processor resource in the computing cluster. The configuration request includes a target configuration template selected by the target object from multiple configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination.
[0034] The plurality of configuration templates are generated by the management node according to hardware data of the plurality of graphics processor types of the graphics processors, each of the configuration templates corresponding to a specific graphics processor type, wherein the configuration template comprises a plurality of configuration instance combinations of the graphics processors of the corresponding graphics processor type, and the graphics processor type refers to a type or model of the graphics processors on the computing nodes in the computing cluster. The configuration instance combination in the plurality of configuration instance combinations refers to a split configuration of splitting the resources of the graphics processors on the computing nodes of the corresponding graphics processor type into a plurality of independent units. For example, the split configuration of the plurality of independent units refers to a plurality of MIG instances that are split by the MIG technology to split the GPU on the computing node into a plurality of independent running MIG instances that have independent computing cores and memory, and each of the MIG instances can be individually scheduled for use to adapt to the requirements of different computing tasks.
[0035] Each of the configuration instance combinations represents a different split configuration of splitting the resources of the graphics processors on the computing nodes into a plurality of independent units to adapt to diversified computing requirements and ensure comprehensive coverage and support for different GPU models. Optionally, the configuration instance combination can comprise a configuration instance with a name of 1g.5gb, a profile_id of 19, a slice of 1, and a memory of 5, and a configuration instance with a name of 2g.10gb, a profile_id of 14, a slice of 2, and a memory of 10. Alternatively, the configuration instance combination can comprise two configuration instances with a name of 3g.20gb, a profile_id of 9, a slice of 3, and a memory of 20.
[0036] The target configuration template is one template selected from the plurality of configuration templates by the target object in the configuration request and corresponding to the graphics processor type of the graphics processors on the target computing node, and is generated by the management node based on hardware data of the graphics processors of the graphics processor type that is the same as the graphics processors on the target computing node. The target configuration instance combination is a split configuration selected from the target configuration template and used to split the graphics processors of the graphics processor type corresponding to the target configuration template on the target computing node into a plurality of independent units. The target computing node list is a set of computing nodes that receive and execute the configuration instruction, and in the embodiments of the present application, the target computing node list comprises at least one graphics processor on at least one computing node.
[0037] In some embodiments, the target computing node is determined by the target object, a target configuration template corresponding to the type of the graphics processor of the target computing node is automatically selected from a plurality of configuration templates, hardware information of the image processor in the target computing node (including the total number of slices, the size of the display memory, and the supported MIG configuration file, etc.) is analyzed and compared with a plurality of configuration instance combinations in the target configuration template, a target configuration instance combination maximizing resource utilization is selected from the plurality of configuration instance combinations in the target configuration template, and a configuration request is generated based on the target computing node, the target configuration template, and the target configuration instance combination.
[0038] In some embodiments, the target computing node is determined by the target object, a target configuration template corresponding to the type of the graphics processor of the target computing node is automatically selected from a plurality of configuration templates, a task list of a plurality of tasks to be executed on the target computing node is obtained, the task list includes task requirements (computing requirements, display memory requirements, task types, etc.) of various tasks, affinity scores between each configuration instance combination in the target configuration template and each task are scored based on the capability matching degree between the task requirements of each task and each configuration instance combination in the target configuration template, an affinity matrix between the configuration instance combination and the task is obtained, each element in the affinity matrix represents the affinity score between the configuration instance combination and a specific task, one or more multi-objective optimization functions are constructed according to a preset optimization target (such as maximizing resource utilization, minimizing cost, and minimizing total task execution time), the multi-objective optimization function is used to find the best matching combination in the affinity matrix, the affinity matrix is solved based on the multi-objective optimization function, and a target configuration instance combination is obtained, wherein the target configuration instance combination is determined based on the task-resource affinity matrix matching strategy, highly matches the task requirements, simultaneously achieves the preset optimization target, and a configuration request is generated based on the target computing node, the target configuration template, and the target configuration instance combination.
[0039] The configuration instruction is a specific instruction generated by the management node according to the received configuration request content, used to instruct the target computing node to allocate resources of the graphics processor on the target computing node according to the configuration instruction. Optionally, the configuration instruction includes detailed information of the target configuration instance combination and an indication of the target computing node, and the configuration instruction is used to guide the corresponding graphics processor on the computing node to perform resource allocation and creation or adjustment of MIG instances. Optionally, the hardware resources of the target computing node include GPU device models and their related computing cores, display memory, bandwidth, and other resources on the computing node.
[0040] Optionally, through the embodiment, through the bidirectional communication between the first component MIG-CORE (running on the management node) and the second component MIG-AGENT (running on the computing node), the MIG configuration full process automation of multiple types of GPUs is realized, the dependence on manual operation or specific GPU models in the related art is broken, the GPU hardware information (such as model, display memory, and total number of slices) is acquired in real time through the MIG-AGENT, the configuration combination is dynamically calculated and configured in combination with the MIG-CORE, all GPUs supporting the MIG technology can be supported without manual adaptation, the system code does not need to be modified when a new type of GPU is put into operation, and the platform adaptation cost is significantly reduced.
[0041] According to the embodiment of the application, the management node in the computing cluster is running, and a corresponding configuration instruction is generated according to a configuration request sent by a target object, wherein the configuration request includes a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the configuration instruction is sent to the target computing node in the target computing node list to guide the corresponding graphics processor resource on the target computing node to perform adaptive allocation; wherein the plurality of configuration templates are dynamically generated according to the hardware data of a plurality of graphics processor types of graphics processors in the computing cluster, the plurality of configuration templates correspond to the plurality of graphics processor types one by one, ensuring the accuracy of the configuration, each configuration template contains a plurality of configuration instance combinations of the corresponding graphics processor type, that is, a split configuration, and a resource allocation scheme is planned for the graphics processors of different graphics processor types, laying a foundation for automatic resource allocation, avoiding the disadvantages of manual configuration and specific model integrated configuration, realizing the automation and intelligent management of MIG configuration in a multi-type GPU environment, improving the GPU resource management efficiency of the artificial intelligence platform, improving the user experience and output efficiency, and solving the problem of low MIG configuration and management efficiency and low compatibility in a multi-type GPU environment caused by the automatic configuration process integrated for specific GPU models in the related art.
[0042] In one example embodiment, before the configuration instruction is generated according to the configuration request sent by the target object, the above method further includes: after starting the plurality of computing nodes and the management node in the computing cluster, sending a hardware information acquisition request with an authentication token and / or a timeout time to the plurality of computing nodes to instruct the plurality of computing nodes to verify the validity of the authentication token in the hardware information acquisition request, and in the case that the authentication token is valid, loading a graphics processor driver within the timeout time in response to the hardware information acquisition request of the management node, acquiring the hardware information of the corresponding graphics processor through the graphics processor driver, and reporting a data packet containing the hardware information of the graphics processor to the management node; and generating a plurality of configuration templates according to the hardware data of a plurality of graphics processor types of graphics processors.
[0043] In this embodiment, the hardware information acquisition request includes an authentication token and / or a timeout time, the authentication token is a key for verifying the identity and authority of the requester, and the timeout time is used to guarantee the timeliness of data transmission and the efficiency of system response. For example, the validity of the authentication token in the hardware information acquisition request is verified by the second component on the plurality of computing nodes. In the case of valid authentication token, the graphics processor driver is loaded in response to the hardware information acquisition request of the management node within the timeout time. The graphics processor driver is a software layer for communication between the GPU device and the operating system, which is used to provide functions such as GPU state monitoring, configuration and resource management. For example, the graphics processor driver can be a management library program. Through the GPU driver, the computing node can directly access and read the hardware information of the GPU, which includes but is not limited to the model of the GPU, the size of the video memory, the total number of slices, the list of supported MIG-Profile, etc.
[0044] Optionally, Figure 3 is an interaction diagram of information verification between a management node and a computing node according to an embodiment of the present application, as Figure 3 shown, the second component (MIG-AGENT) runs on the computing node, and the first component (MIG-CORE) runs on the management node. MIG-CORE is deployed and started in the form of POD on the management node, and MIG-AGENT is deployed and started in the form of DaemonSet on all GPU nodes. MIG-CORE requests to acquire GPU information of the computing node (including carrying an authentication token and setting a timeout time of 5s), and MIG-AGENT requests the GPU driver to acquire GPU information, wherein the GPU information is acquired by calling the specified function DeviceGetCount_v2() and traversing each GPU MIG-Profile, and the GPU information (including model, video memory, slice and profile list) is returned.
[0045] In this embodiment, the first component on the management node in the computing cluster runs in a minimum execution unit. Optionally, the minimum execution unit on the management node is a container group (Pod) in the Kubernetes environment. In Kubernetes, a Pod is a minimum execution unit, and a Pod represents a group of containers in the cluster. These containers share resources within a network namespace and are scheduled and managed as a whole. For example, the first component MIG-CORE is started in the form of POD on the management node.
[0046] Optionally, a plurality of second components are deployed on each computing node of the computing cluster, each second component corresponding to a specific computing node, and each second component running in a minimum execution unit on the corresponding computing node. Optionally, the minimum execution unit on the computing node can be a Pod, and the second component MIG-AGENT is started in a DaemonSet manner on all GPU nodes. Among them, DaemonSet is a workload resource type in Kubernetes (k8s), and its main role is to run a copy of a Pod on every node (or selected set of nodes) in the cluster. This means that no matter how the nodes of the cluster increase or decrease, DaemonSet will ensure that there is a Pod copy running on each node to perform some services that are needed on all nodes, such as log collection, monitoring, storage management, etc. Through the plurality of second components, the hardware information of the plurality of computing nodes in the computing cluster is collected, and the hardware information of the plurality of computing nodes is sent to the first component in the management node. For example, the hardware information of the plurality of computing nodes includes the model of the GPU, the size of the video memory, whether the MIG technology is supported, the total number of divisible slices, and the like.
[0047] For example, the second component (MIG-AGENT) runs on the computing node, and the first component (MIG-CORE) runs on the management node. After the MIG-AGENT is started, the Management Library of the GPU driver is quickly loaded, wherein the Management Library is a set of Application Programming Interfaces (APIs) for monitoring and managing GPU devices, allowing developers to query GPU status, configuration, and control GPU behavior. The communication connection with the GPU is initialized, and the GPU information of the computing node is obtained through the driver, including the number, ID, name, size of the GPU memory, whether the MIG is supported, the total number of slices, the configuration file (MIG-Profile, in the MIG technology, MIG-Profile refers to a pre-defined MIG instance specification, including a specific number of computing cores, video memory, and other resource configurations), whether the MIG has been configured, MIG instance conditions, etc. The computing node GPU information is synchronized, that is, the MIG-AGENT system organizes the collected node information (node host name, IP address) and detailed information of each GPU, encapsulates it in accordance with the preset standard format, and then synchronizes it to the MIG-CORE in a safe and fast manner through Remote Procedure Call (gRPC).
[0048] Optionally, the computing nodes report hardware information to the management node, and periodically report the current load status (such as GPU usage, GPU occupancy, etc.) of the computing nodes to the management node. The management node comprehensively analyzes the hardware information and the current load status of the nodes, analyzes the available resources and the predicted resource consumption of each GPU instance, and predicts possible resource conflicts. When a resource conflict is detected, the management node automatically adjusts the configuration instructions. For example, if it is found that the GPU resources of the computing nodes of type A are tight during the peak period, the management node may prefer to configure the computing nodes of type B, or adjust the configuration instance combination to reduce the allocation of slices until the load state tends to be stable, so that the management node dynamically generates multiple configuration templates according to the resource allocation status of each node in the cluster.
[0049] Optionally, the management node generates multiple configuration templates according to the hardware data of the graphics processors of multiple graphics processor types after receiving the hardware data of the graphics processors of the graphics processor types of all the computing nodes. The generation method of the configuration template includes but is not limited to the following examples: for example, the management node pre-processes the hardware information of multiple computing nodes, including but not limited to standardization processing, missing value filling, and outlier detection, extracts key features such as GPU type, total number of slices, and memory size from the pre-processed hardware information as input of a machine learning model; the management node trains the machine learning model (such as GPU type, total number of slices, and memory size) using historical data, and uses the trained model to predict the optimal configuration instance combination for the hardware information of a new computing node, and generates a preliminary configuration template according to the prediction result.
[0050] For example, for each graphics processor type, a configuration template corresponding to the graphics processor type is generated according to a group of configuration instance combinations in the computing nodes of the corresponding graphics processor type that meet a preset condition. The preset condition means that the total number of slices occupied by the configuration instance combination on the computing nodes of the corresponding graphics processor type is less than or equal to the maximum number of slices of the corresponding computing nodes, and the memory size occupied by the configuration instance combination is less than or equal to the maximum memory size of the corresponding computing nodes.
[0051] Through the embodiment, the management node sends a hardware information acquisition request with an authentication token and / or a timeout time to the computing nodes on the plurality of computing nodes, wherein the authentication token is used to ensure the security of information query and data interaction, effectively preventing the intrusion of illegal requests, and protecting sensitive information and resources within the computing cluster from unauthorized access; the timeout time limits the response time, which helps to improve the efficiency of data transmission, reduce system congestion caused by long waiting time, and maintain good responsiveness and stability of the computing cluster. The computing nodes load the GPU driver to obtain the hardware information of the corresponding graphics processors, which directly ensures the accuracy and integrity of the information from the hardware layer. The management node generates a plurality of configuration templates according to the collected hardware information, without the need to rewrite the configuration rules every time the GPU model is changed, greatly improving the flexibility and scalability of resource management.
[0052] In an exemplary embodiment, the hardware information of the plurality of computing nodes adopts a specified format string, and the specified format string corresponds to a first checksum. Before generating the plurality of configuration templates according to the hardware data of the plurality of graphics processors of the plurality of graphics processor types, the method further includes: taking a computing node in the plurality of computing nodes as a current computing node, and performing the following operations: receiving a data packet of the current computing node, and generating a second checksum corresponding to the specified format string in the data packet of the current computing node; in the case that the second checksum matches the first checksum in the data packet of the current computing node, and the syntax of the specified format string is correct, storing the data packet of the current computing node to a database according to a specified storage path, with the specified information of the current computing node as a key name.
[0053] In the embodiment, the specified format string is the hardware information of the graphics processor after structured processing, which conforms to the text form of the data exchange standard. For example, the specified format string can be in JSON format or XML format. The first checksum is a fixed-length numerical value generated by a hash algorithm, which is used to check the integrity and tamper of the data. For example, the hash algorithm includes MD5, SHA-1, SHA-256, etc., and the generated checksum can be used to detect any changes in the data during transmission.
[0054] The second checksum is a check value recalculated by the management node according to the specified format string of the hardware information in the reported data packet.
[0055] After receiving a data packet from the compute node, the management node first decapsulates the packet and extracts a string in a specified format (i.e., a standardized data representation of hardware information). The management node then performs the same hash operation on this string as the first checksum to generate a second checksum. The management node compares this second checksum with the first checksum corresponding to the specified format string in the data packet. If the second and first checksums match exactly, it indicates that the data has not been tampered with during transmission. Simultaneously, the management node can also check the syntactic correctness of the specified format string to ensure the integrity of the data structure and prevent storage failures or subsequent processing anomalies caused by incorrect data format.
[0056] If the second checksum matches the first checksum and the specified format string syntax is correct, the management node stores the current compute node's data packets in the database using the specified information of the current compute node as the key and according to the specified storage path. The specified information refers to information that uniquely identifies the current compute node and its corresponding hardware resource status, while the specified storage path refers to the specific location where the current compute node's data packets are stored. For example, ` / mig / nodes / {hostname}`, where ` / {hostname}` is the hostname of the compute node.
[0057] Optionally, such as Figure 3 As shown, taking the second component (MIG-AGENT) running on the compute node and the first component (MIG-CORE) running on the management node as an example, MIG-CORE receives data packets from the current compute node, generates a second checksum corresponding to the specified format string in the current compute node's data packets, and performs data integrity verification on the data through MIG-CORE. Data integrity verification includes comparing the first and second checksums and verifying the JSON format. If the second checksum matches the first checksum in the current compute node's data packets, and the syntax of the specified format string is correct, MIG-CORE stores the node information in the ETCD database using the compute node's hostname as a subdirectory or key, storing the current compute node's data packets in the database according to the storage path / mig / nodes / {hostname}. Upon successful storage, the ETCD database sends a storage success response to MIG-CORE. For example, MIG-CORE integrates the hardware information obtained, closely linking information such as the number of GPUs on each computing node, the MIG-Profile of each GPU, and the status of each GPU, storing it in the system's high-performance database (such as etcd), and creating an index to greatly improve query efficiency and facilitate rapid retrieval when users make subsequent requests.
[0058] Through this embodiment, by generating a second checksum for the hardware information in the received data packet and comparing it with the reported first checksum, the preliminary verification of data integrity is realized, ensuring that the hardware information has not been tampered with during transmission. Syntax checking of the specified format string ensures the quality and readability of the data, avoiding storage failures due to data format errors, and ensuring the consistency and effectiveness of internal system data. By using the specified information of the computing node as the key name, storing the data to the database according to the preset path, not only improves the efficiency of the storage process, but also establishes an index system for easy retrieval.
[0059] In an exemplary embodiment, after storing the data packet of the current computing node to the database according to the specified storage path with the specified information of the current computing node as the key name, the method further comprises: setting the life period of the data packet of the current computing node to indicate that the database deletes the data packet of the current computing node if the data packet of the current computing node is not updated within the life period.
[0060] After the management node successfully stores the data packet of the current computing node to the database, the life period of the data packet is set. The life period (Time-To-Live, TTL for short) refers to the time limit for the data packet to exist in the database. If the time limit is exceeded, the data packet will be considered expired.
[0061] The management node associates the life period information with the stored data packet, instructing the database management system to automatically delete the expired data packet after the life period of the data packet expires.
[0062] Optionally, as shown in Figure 3 For example, the second component (MIG-AGENT) runs on the computing node, and the first component (MIG-CORE) runs on the management node. The MIG-CORE stores the node information to the ETCD database with the hostname of the computing node as the subdirectory or key name, and stores the data packet of the current computing node to the database according to the storage path / mig / nodes / {hostname}. The TTL is set to 300s, meaning that if the value of this data entry is not updated within 300 seconds, ETCD will automatically delete this entry.
[0063] Through this embodiment, setting the life period of the data packet can effectively avoid expired, invalid or duplicate data occupying limited storage space for a long time, ensure the rational use of database system resources, and improve the storage efficiency. And by regularly updating and automatically cleaning up expired data packets, the system can ensure that the hardware information processed is the latest, reflecting the real-time state of the computing node, which helps to reduce data redundancy and potential version conflicts in the database, avoid resource configuration errors caused by the existence of old data, and improve the consistency and stability of the computing cluster management.
[0064] In an exemplary embodiment, the hardware information of the same type of graphics processor is the same, and the hardware information of the different types of graphics processors includes the maximum memory size, the maximum slice number and a plurality of configuration instances of the graphics processor of the corresponding graphics processor type; the configuration instance in the plurality of configuration instances corresponding to the different types of graphics processors refers to a slice mode supported by the graphics processor of the corresponding graphics processor type; and the configuration instance combination in the plurality of configuration instance combinations includes at least one configuration instance.
[0065] Here, the maximum memory size is the maximum available memory capacity of the GPU on the computing node, the maximum slice number is the maximum number of independent computing units into which the GPU is divided, and the plurality of configuration instances is a set of MIG instance types of different specifications predefined in advance. Among them, the configuration instance combination is composed of one or more configuration instances, and the configuration instance combination takes the configuration instance as the basic unit and combines a plurality of configuration instances to meet specific resource requirements and constraint conditions. In addition, the configuration instance focuses on the segmentation mode of a single GPU resource, while the configuration instance combination focuses on how to combine a plurality of configuration instances on a single computing node to maximize resource utilization. That is, the configuration instance aims to define the segmentation specification of the GPU resource, and the configuration instance combination is to find the best combination scheme on the basis of these configuration instances to meet the preset conditions, such as the constraints of the slice number and the memory size.
[0066] For example, taking node name node1 and node ip 100.125.60.2 as an example, the corresponding hardware information is as follows:
[0067] {
[0068] "nodename": "node1",
[0069] "ip": "100.125.60.2",
[0070] "GPU_CNT": 2,
[0071] "GPU_LIST": [
[0072] {
[0073] "GPU_ID": "GPU-6fsrx67e...",
[0074] "GPU_NAME": "A",
[0075] "MEM_SUM": 40,
[0076] "SLICE_CNT": 7,
[0077] "MIG_PROFILE": [
[0078] {
[0079] "name": "1g.5gb",
[0080] "profile_id": 19,
[0081] "slice": 1,
[0082] "memory": 5
[0083] },
[0084] {
[0085] "name": "1g.10gb",
[0086] "profile_id": 15,
[0087] "slice": 1,
[0088] "memory": 10
[0089] },
[0090] {
[0091] "name": "2g.10gb",
[0092] "profile_id": 14,
[0093] "slice": 2,
[0094] "memory": 10
[0095] },
[0096] {
[0097] "name": "3g.20gb",
[0098] "profile_id": 9,
[0099] "slice": 3,
[0100] "memory": 20
[0101] },
[0102] {
[0103] "name": "4g.20gb",
[0104] "profile_id": 5,
[0105] "slice": 4,
[0106] "memory": 20
[0107] },
[0108] {
[0109] "name": "7g.40gb",
[0110] "profile_id": 0,
[0111] "slice": 7,
[0112] "memory": 40
[0113] }
[0114] ],
[0115] "FULL_CARD": false,
[0116] "MIG_MSG": [
[0117] {
[0118] "name": "3g.20gb",
[0119] "count": 1
[0120] },
[0121] {
[0122] "name": "4g.20gb",
[0123] "count": 1
[0124] } ]
[0126] },
[0127] {
[0128] "GPU_ID": "GPU-57dsgse..."
[0129] } ]
[0131] }
[0132] In the above content, the node name (nodename) is node1, the node ip is 100.125.60.2; the number of GPU_CNT is 2; the ID of the GPU is GPU-6fsrx67e, the name of the GPU GPU_NAME is A, the maximum memory size MEM_SUM is 40, the memory is the memory, the maximum slice number SLICE_CNT is 7, the count is the number, and the plurality of configuration instances can be as follows: configuration instance one, "name": "1g.5gb", "profile_id": 19, "slice": 1, "memory": 5; configuration instance two, "name": "1g.10gb", "profile_id": 15, "slice": 1, "memory": 10; configuration instance three, "name": "2g.10gb", "profile_id": 14, "slice": 2, "memory": 10; configuration instance four, "name": "3g.20gb", "profile_id": 9, "slice": 3, "memory": 20; configuration instance five, "name": "4g.20gb", "profile_id": 5, "slice": 4, "memory": 20; configuration instance six, "name": "7g.40gb", "profile_id": 0, "slice": 7, "memory": 40.
[0133] In some embodiments, a plurality of configuration templates are generated according to hardware data of a plurality of graphics processor types, including: taking a graphics processor type in the plurality of graphics processor types as a current graphics processor type, taking a graphics processor corresponding to the current graphics processor type as a current graphics processor, and performing the following processing operations to obtain a plurality of configuration templates corresponding to the plurality of graphics processor types: performing permutation and combination on a plurality of configuration instances in hardware information of the current graphics processor, determining a set of configuration instance combinations satisfying a preset condition, and generating a configuration template corresponding to the current graphics processor type according to the set of configuration instance combinations; the preset condition refers to a condition that a total number of slices occupied by the set of configuration instance combinations on the current graphics processor is less than or equal to a maximum number of slices of the current graphics processor, and a memory size occupied by the set of configuration instance combinations is less than or equal to a maximum memory size of the current graphics processor.
[0134] In the present embodiment, a graphics processor type in the plurality of graphics processor types is taken as a current graphics processor type, a graphics processor corresponding to the current graphics processor type is taken as a current graphics processor, permutation and combination are performed on a plurality of configuration instances in hardware information of the current image processor to obtain a corresponding configuration instance combination and a corresponding configuration template.
[0135] Optionally, taking the node name as node1 and the node ip as 100.125.60.2 as an example, since the maximum memory size MEM_SUM is 40 and the maximum slice number SLICE_CNT is 7, a configuration instance with a name of 3g.20gb can be selected, and a configuration instance with a name of 4g.20gb can be selected.
[0136] In an optional embodiment, the second component (MIG-AGENT) runs on the computing node, and the first component (MIG-CORE) is taken as an example. Figure 4 is a timing diagram for managing a node to allocate a graphics processor resource according to an embodiment of the present application, as shown in Figure 4As shown, the MIG-CORE reads GPU information from the Kubermetes cluster, the GPU information including model, memory, total number of slices, and supported Profile list, after successful reading, the Kubermetes cluster returns the stored GPU information data to the MIG-CORE, the MIG-CORE arranges all possible MIG instance combinations according to the GPU information of the computing node through a recursive combination algorithm, generates a MIG configuration template, wherein the MIG configuration template can be sorted according to resource utilization, and the MIG configuration template includes instance type, number, total number of slices, stores the MIG configuration template to the ETCD and displays it to the user. The front-end page calls API to query the available templates from the MIG-CORE, the MIG-CORE reads the template list from the Kubermetes cluster, the Kubermetes cluster returns the configuration template data to the MIG-CORE, and the MIG-CORE returns the template list to the front-end page, wherein the template list includes resource details and applicable GPUs. The front-end page submits the selected template ID and configuration request to the MIG-CORE, the MIG-CORE issues MIG configuration instructions to the MIG-AGENT, the MIG configuration instructions include configuration parameters parsed from the template and a list of target GPU IDs, the MIG-AGENT calls the driver to execute MIG configuration to the GPU driver, and sets the MIG mode and creates an instance creation command, the GPU driver returns the configuration execution result to the MIG-AGENT, the MIG-AGENT returns the configuration result to the MIG-CORE, the configuration result includes success / failure status and a list of instances after configuration, the MIG-CORE returns the configuration result to the front-end page, including instance details and operation log links. In this way, through the embodiment, the optimal MIG configuration template is automatically generated through a recursive algorithm instead of the inefficient manual configuration and specific model integration, one-key configuration is realized through front-end interaction, the whole process is fully automated from information collection, combination calculation to configuration execution, the single-node MIG configuration time is shortened from hours to minutes, and the platform operation and maintenance efficiency is greatly improved.
[0137] Through the embodiment, a plurality of configuration instance combinations in the hardware information of the current graphics processing unit are arranged and combined, and a suitable configuration instance combination is screened out, which not only fully considers the hardware limitation of the GPU, but also flexibly adjusts the resource allocation strategy according to different GPU types, and realizes efficient utilization of resources. And a group of configuration instance combinations are determined according to the preset conditions, and a configuration template is generated, which prevents the generation of configuration instance combinations exceeding the upper limit of hardware resources, avoids waste caused by excessive resource allocation, and at the same time guarantees the stable operation of the system.
[0138] In an exemplary embodiment, before generating a configuration template corresponding to the current graphics processor type based on a set of configuration instance combinations, the method further includes: removing configuration instance combinations from the set of configuration instance combinations whose graphics processor resource utilization is less than a preset utilization threshold, and based on the removed set of configuration instance combinations, performing the step of generating a configuration template corresponding to the current graphics processor type based on the set of configuration instance combinations.
[0139] In this embodiment, during the configuration instance combination process, the graphics processor resource utilization rate (GPU resource utilization) of each configuration instance combination is calculated. Here, GPU resource utilization rate refers to the actual degree to which GPU resources are used under a specific MIG configuration instance combination. Optionally, GPU resource utilization rate includes slice utilization rate and video memory utilization rate. For example, for a configuration instance combination, the total number of slices occupied by the configuration instance combination on the current compute node group is calculated and divided by the maximum number of slices of the nodes in the configuration instance combination to obtain the slice utilization rate; simultaneously, the total video memory size occupied by the configuration instance combination on the current compute node group is calculated and then divided by the maximum video memory size of the nodes in the configuration instance combination to obtain the video memory utilization rate; based on the slice utilization rate and video memory utilization rate, a weighted average or other mathematical formula is used to calculate the GPU resource utilization rate of the configuration instance combination.
[0140] Remove configuration instance combinations whose graphics processor resource utilization is less than a preset utilization threshold. The preset utilization threshold is used to measure whether the resource utilization of the configuration instance combination meets the standard. Optionally, the preset utilization threshold is a predefined value.
[0141] Optionally, such as Figure 4 As shown, taking the first component, MIG-CORE, running on the management node as an example, MIG-CORE uses a recursive combination algorithm to arrange all possible MIG instance combinations based on the GPU information of the compute node, and excludes combinations with low GPU resource utilization (filtering out invalid combinations with low resource utilization), generating a MIG configuration template, storing it, and displaying it to the user. Thus, this embodiment supports differentiated configurations of multiple types of GPUs on the same node (e.g., independently splitting different types of GPU models when they coexist), and allows some GPUs to be configured as MIG instances and some to be used as whole cards, meeting diverse business needs. Through the resource utilization filtering mechanism, only high-efficiency configuration combinations (e.g., Slice and memory utilization ≥ 90%) are retained, avoiding resource waste.
[0142] Optionally, because the combination supported by MIG is very free, a GPU with 40GB of display memory can only configure a 1g.5gb MIG instance, which is equivalent to wasting 6 slices and 35GB of display memory, so MIG-CORE will filter out combinations with relatively high resource utilization to display to users. Table 1 is an example of a configuration template generated by MIG-CORE. As shown in Table 1, for a GPU with 40G display memory, there are 28 combinations, which can meet the configuration needs of different users and different scenarios.
[0143] Table 1
[0144]
[0145] Through the embodiment, by eliminating configuration instance combinations with low graphics processor resource utilization, the quality of the final configuration template is improved, ensuring that the GPU resources on the computing node can be more fully utilized, and avoiding resource waste due to improper resource configuration.
[0146] In one exemplary embodiment, before generating the configuration instruction according to the configuration request sent by the target object, the method further comprises: calling an eviction application interface to initiate a minimum execution unit eviction request to the target computing node; the minimum execution unit eviction request is used to instruct the target computing node to maintain state data and safely shut down the minimum execution unit within a preset time length.
[0147] In this embodiment, the eviction application interface is used to initiate a minimum execution unit eviction request to the minimum execution unit on the target computing node. Optionally, the eviction application interface can be Evict API.
[0148] Optionally, Figure 5 is a timing diagram of a computing node responding to a configuration instruction according to an embodiment of the present application, as Figure 5 shown, taking the management node running the first component MIG-CORE as an example, the user selects or the system automatically selects a configuration template according to the demand on the front-end page, the front-end page submits the selected template ID to MIG-CORE and submits a configuration request, MIG-CORE initiates a target node POD eviction request to the Kubermetes cluster, including calling Evict API and setting a graceful termination time of 30s. The Kubermetes cluster returns the eviction result to MIG-CORE.
[0149] In some embodiments, after sending the configuration instruction to the target computing node in the target computing node list to instruct the target computing node to allocate resources of the graphics processor on the target computing node according to the configuration instruction, the graphics processor resource allocation method further includes: initiating a recovery signal to the target computing node to instruct the target computing node to remove the minimum execution unit eviction restriction and allow a new minimum execution unit to be started on the target computing node.
[0150] Optionally, as shown in Figure 5 Taking the management node running the first component MIG-CORE as an example, the MIG-CORE initiates a recovery node POD scheduling to the Kubermetes cluster, including removing the eviction restriction and allowing a new POD scheduling, after the removal succeeds, the Kubermetes cluster returns a scheduling recovery success to the Kubermetes cluster, the MIG-CORE sends an update GPU configuration state to the ETCD database, after the update succeeds, the ETCD database returns a state update success to the MIG-CORE, and the MIG-CORE sends a configuration result to the front-end page, the configuration result includes instance details and an operation log link.
[0151] Through the embodiment, the eviction application interface is called to initiate a minimum execution unit eviction request to the target computing node, so that the running tasks on the computing node are stopped in an orderly manner in the resource allocation process, resource conflicts and data loss are avoided, and the security of the hardware resource allocation is improved. After the resource allocation is completed, the management node recovers the scheduling of the minimum execution unit, so that the system can quickly recover the business processing capability of the computing node, the influence of the resource allocation on the ongoing business is reduced, and the high availability of the computing cluster is maintained.
[0152] In one exemplary embodiment, the graphics processor resource allocation method further includes: in the case where the management node detects an abnormal recovery event, obtaining a computing node list of the computing cluster, actively and cyclically calling a plurality of computing nodes in the computing cluster to obtain hardware information of graphics processors on the plurality of computing nodes, and merging and storing the hardware information of the graphics processors on the plurality of computing nodes with hardware information of the corresponding plurality of computing nodes stored in the database.
[0153] In the embodiment, the management node can detect an abnormal recovery event, the abnormal recovery event refers to an event that the management node in the computing cluster needs to be restarted or recovered due to an abnormality (such as a network problem), wherein the abnormal recovery event is caused by the abnormality of the management node, and the node restart event is triggered by the restart of the computing node itself. Optionally, the abnormal recovery event includes container drift, container restart, and host restart.
[0154] In the case that the management node detects an abnormal recovery event, a computing node list of the computing cluster is obtained, wherein the computing node list is collection information of all available computing nodes in the computing cluster, and the computing node list optionally includes ID, IP address and other information of the nodes.
[0155] The management node actively and cyclically calls multiple computing nodes in the computing cluster to request the latest hardware information of the graphic processors on the multiple computing nodes according to the obtained computing node list. After the management node receives the hardware information of the graphic processors on the computing nodes, the hardware information of the graphic processors on the multiple computing nodes can be stored in combination with the hardware information of the corresponding multiple computing nodes stored in the database. The hardware information is stored in combination with the hardware information of the multiple computing nodes stored in the database in order to maintain the integrity of the data in the database and the continuity of the historical records. If an overwrite strategy is adopted, the old data will be completely replaced by the new data, which may cause important historical information to be lost, such as the use of GPU resources, configuration history, etc. In the abnormal recovery event, the hardware information of the computing nodes may not be completely updated due to network delay, non-response of the nodes, etc. If the overwrite strategy is used, the incomplete data may cover the previous valid information, causing further loss of data. The combination storage strategy in the embodiment can ensure that even if part of the data update fails, the existing information is still retained, which is beneficial to system fault tolerance and subsequent recovery operations. In addition, in the computing cluster, multiple computing nodes may update their corresponding graphic processor hardware information at the same time, and direct overwrite may cause inconsistency of data, because there may be a time difference in the information update of different nodes, and the overwrite strategy cannot ensure that the information update of all nodes is at the same time point. The combination storage strategy can better handle concurrent updates, update only the changed part by comparing the old and new information, and maintain the consistency and accuracy of the information in the database.
[0156] Optionally, Figure 6 is a timing diagram of a management node and a computing node responding to a restart event according to an embodiment of the present application, as Figure 6 shown, taking a first component MIG-CORE running on the management node as an example, in the case that the MIG-CORE detects an abnormal recovery event, the MIG-CORE obtains a computing node list through the K8S API. For example, if the MIG-CORE container drifts, the container restarts, and the host restarts, the MIG-CORE will obtain all computing nodes through K8S, actively and cyclically call the MIG-AGENT of each computing node, and obtain the GPU information of the host.
[0157] Through the embodiment, in the case that the management node detects an abnormal recovery event, the latest hardware information of the graphics processor of all the computing nodes can be acquired, ensuring instant update of the information and avoiding information delay or inconsistency caused by abnormal events; through active cyclic calling of the computing nodes and information merging and storage, the management node can obtain complete hardware information, reducing system delay caused by waiting for information synchronization and improving operation efficiency of the computing cluster and user experience.
[0158] In an exemplary embodiment, the embodiment also provides a graphics processor resource allocation method, which is run on a target computing node in a computing cluster, wherein the target computing node refers to any computing node in the computing cluster that needs to perform graphics processor resource splitting, and the method comprises: allocating resources of a graphics processor on the target computing node according to a configuration instruction sent by a management node in the computing cluster; the configuration instruction is generated based on a configuration request sent by a target object; the configuration request comprises a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates are generated by the management node according to hardware data of graphics processors of a plurality of graphics processor types; the plurality of configuration templates correspond one-to-one to the plurality of graphics processor types; a configuration template in the plurality of configuration templates comprises a plurality of configuration instance combinations of the graphics processor of the corresponding graphics processor type; a configuration instance combination in the plurality of configuration instance combinations refers to a splitting configuration of splitting the resources of the graphics processor of the corresponding graphics processor type into a plurality of independent units.
[0159] The embodiment can be understood with reference to the foregoing embodiments of the method run on the management node in the computing cluster, and details are not repeated here.
[0160] In an exemplary embodiment, before allocating resources of a graphics processor on the target computing node according to a configuration instruction sent by a management node in the computing cluster, the method further comprises: receiving a hardware information acquisition request sent by the management node with an authentication token and / or a timeout time, verifying validity of the authentication token in the hardware information acquisition request, and in the case that the authentication token is valid, loading a graphics processor driver within the timeout time in response to the hardware information acquisition request of the management node, acquiring hardware information of the graphics processor on the target computing node through the graphics processor driver, and reporting a data packet containing the hardware information of the graphics processor to the management node, so as to instruct the management node to generate a configuration template corresponding to a graphics processor type of the graphics processor on the target computing node according to the hardware information of the graphics processor on the target computing node in the case that the data packet passes the verification.
[0161] The embodiment can be understood with reference to the foregoing embodiments, and details are not repeated here.
[0162] In an exemplary embodiment, reporting a data packet containing hardware information of a graphics processor to the management node includes: converting the hardware information of the graphics processor into a specified format string, generating a first checksum corresponding to the specified format string, encapsulating the specified format string and the first checksum to obtain a data packet containing the hardware information of the graphics processor, and reporting the data packet to the management node.
[0163] In this embodiment, the specified format string and the first checksum can be understood with reference to the foregoing embodiments, and will not be repeated here.
[0164] Optionally, in order to verify the integrity of the transmitted data, the computing node can perform a hash operation on a specified format string to generate a first checksum corresponding to the original data.
[0165] After generating the first checksum corresponding to the specified format string, the specified format string and the first checksum are encapsulated together to obtain the data packet of the current computing node. The purpose of encapsulation is to facilitate network transmission, while also ensuring the security and integrity of the data.
[0166] Optionally, such as Figure 3 As shown, taking the second component (MIG-AGENT) running on the compute node and the first component (MIG-CORE) running on the management node as an example, after the MIG-AGENT obtains GPU information (including model, memory, slice, and profile list), it performs data encapsulation and verification (including converting hardware information into standard JSON format and the first checksum of the compute data). It encapsulates the specified format string and the first checksum to obtain the data packet for the current compute node, which is then reported to the MIG-CORE. The data packet reported to the MIG-CORE contains returned node information, including the node IP, hostname, first checksum, and timestamp. For example, the encapsulated data packet is reported to the first component via a secure communication protocol (such as gRPC).
[0167] In this embodiment, after obtaining the GPU hardware information of the computing node, it is converted into a specified format string to ensure the standardization and normalization of the information. Furthermore, the first checksum is generated through the specified format string, which effectively guarantees the integrity and tamper-proof status of the data during transmission. By encapsulating the specified format string and the first checksum, the data packet of the current computing node is obtained, which not only improves the security of data transmission but also avoids the tampering and loss of information during transmission.
[0168] In an exemplary embodiment, allocating resources of a graphics processor on a target computing node according to a configuration instruction sent by a management node in a computing cluster comprises: stopping processes running on the target computing node, calling a driver to enable a resource allocation mode of the target computing node after the processes on the target computing node are successfully stopped, and creating a target configuration instance combination on the graphics processor on the target computing node; and resuming the processes on the target computing node that have been stopped after the resource configuration of the graphics processor on the target computing node is successful.
[0169] In the embodiment, to avoid competition of processes for GPU resources in a configuration process and ensure smooth MIG configuration, the processes running on the target computing node are stopped. Optionally, all related processes running on the computing node and occupying GPU resources are stopped.
[0170] Optionally, taking an example of running a first component MIG-CORE on a management node and a second component MIG-AGENT on a computing node, a front-end page calls an API of the MIG-CORE to select a MIG template for configuration; the MIG-CORE converts the configuration template into configuration information, calls the MIG-AGENT after evicting a POD of the computing node to perform configuration; the MIG-AGENT stops processes occupying the GPU, calls a GPU driver to perform MIG configuration; returns a result after the configuration is completed; the MIG-AGENT and the MIG-CORE resume a node state, resume related services, and return information to the front-end.
[0171] Optionally, as shown in FIG. 6, the MIG-CORE and the MIG-AGENT are running on the same computing node. Figure 5As shown, taking the example that the first component MIG-CORE runs on the management node and the second component MIG-AGENT runs on the computing node, the MIG-CORE sends the MIG configuration instruction to the MIG-AGENT, including the configuration parameters of the template parsing and the template GPU ID list, the MIG-AGENT sends the node GPU occupation process to the Kubermetes cluster, and the dcgm monitoring and related driver services, the Kubermetes cluster returns the process stop success to the MIG-AGENT. The MIG-AGENT calls the graphics processor driver to execute the configuration setting mode, and creates an instance creation command; after the graphics processor driver creates the instance, the configuration execution result is returned; in the case that the configuration execution result represents that the execution is successful, the related processes of the graphics processor on the target computing node are restored; after the processes are restored, the computing cluster feeds back the process restoration success message to the MIG-AGENT, and the MIG-AGENT returns the configuration result to the MIG-CORE, wherein the configuration result includes the success / failure state and the instance list after the configuration. In the case that the configuration result represents that the configuration is successful, the MIG-CORE restores the scheduling of the computing cluster, and in the case that the scheduling of the computing cluster is restored, the configuration state of the graphics processor stored in the database is updated; in the case that the database state is updated successfully, the MIG-CORE returns the configuration result to the front-end page, wherein the configuration result returned to the front-end page includes the instance details and the operation log link.
[0172] Through the embodiment, by stopping the processes running on the target computing node and then calling the driver to enable the resource allocation mode of the target computing node, the influence on the running computing tasks in the configuration process can be avoided, and the accuracy and safety of the resource allocation action are ensured. Moreover, the processes are quickly restored after the resource configuration is successful, thereby ensuring the continuity of the computing tasks and services.
[0173] In one exemplary embodiment, when a new GPU is launched, the adaptation and configuration mode of the MIG technology thereof is also often changed, and in the related MIG automatic configuration method, the tool needs to be redeveloped and integrated, the period is long, and the timely use of the new GPU by the user is seriously affected, it is difficult to meet the rapid adaptation demand of the new GPU, and the flexibility and expansibility of the platform are reduced. Therefore, to solve the above problem of being difficult to adapt to the new GPU, the above method further includes:
[0174] The hardware incremental information of the target computing node is acquired in a timely manner, the incremental update package is generated based on the hardware incremental information of the target computing node, the incremental update package is differentially compressed, and the compressed incremental update package is reported to the management node to instruct the management node to update the database according to the incremental update package; the database is used to store the hardware information of a plurality of computing nodes in the computing cluster.
[0175] In this embodiment, incremental hardware information of the corresponding target computing node is acquired periodically. Incremental information refers to the updated hardware information of the graphics processor on the computing node. For example, the target computing node acquires incremental changes in hardware information from GPU device and system monitoring every 30 seconds. These incremental changes include GPU usage status, MIG configuration updates, Slice allocation, the graphics processor type of the newly added GPU, and the hardware information of the newly added GPU.
[0176] Based on the incremental hardware information of the target compute node, an incremental update package is generated. Optionally, the incremental update package may include information that differs from the previously reported data, such as newly created MIG instances, changes in Slice status, etc., while omitting unchanged data to reduce the amount of data.
[0177] After generating the incremental update packet, in order to reduce the size of the data packet and improve network transmission efficiency, the incremental update packet of the target computing node can be differentially compressed, and the corresponding compressed incremental update packet is reported to the management node.
[0178] In an optional embodiment, such as Figure 6 As shown, taking the second component (MIG-AGENT) running on the compute nodes and the first component (MIG-CORE) running on the management node as an example, MIG-AGENT is deployed on each compute node. Every 30 seconds, it obtains incremental changes in hardware information from the GPU device, that is, it obtains GPU status changes from the GPU driver, including detecting MIG mode changes and checking Slice allocation status. The GPU driver returns incremental change data to MIG-AGENT, which generates an incremental update package. The incremental update package includes the changed fields and is compressed using differential compression. MIG-AGENT reports the node status changes to MIG-CORE. Based on the incremental update package corresponding to the compute nodes among multiple compute nodes, MIG-CORE updates the ETCD database and performs partial update operations. The ETCD database returns the update success information to MIG-CORE, and then MIG-CORE reports a confirmation response to MIG-AGENT.
[0179] In some embodiments, where the incremental information includes the graphics processor type and hardware information of the newly added GPU, the incremental update package is further used to instruct the management node to generate a configuration template corresponding to the graphics processor type of the newly added GPU based on the incremental update package corresponding to the target computing node. The process of generating the configuration template has been explained in the above embodiments and will not be repeated here.
[0180] When a new type of GPU is online, the new type of GPU collects its hardware information and feeds back to the management node in an incremental transmission manner, the management node generates a configuration template corresponding to the new type of GPU accordingly, without modifying the system code, breaking the limitation of related methods in the adaptation of new types of GPUs, and quickly responding to the online of new types of GPUs to meet the rapid adaptation demand; and only the incremental changes of the hardware information are transmitted, rather than the full amount of information, which significantly reduces the data transmission amount and reduces the demand for network bandwidth. The generation of the incremental update package and the timely reporting of the incremental update package to the management node can quickly capture and feedback the changes of the hardware state of the target computing node, improve the real-time performance of data synchronization, and reduce the resource allocation errors caused by data delay; and the database is updated using the differential compressed incremental update package, which reduces the write amount of the database and avoids the negative impact of frequent full data update on the performance of the database.
[0181] In an exemplary embodiment, in order to disperse the time points of reporting data by the computing nodes and reduce the network pressure caused by the simultaneous reporting of data by all computing nodes at the same time, in this embodiment, the hardware incremental information of the target computing node is obtained at regular intervals, including:
[0182] The random delay corresponding to the target computing node is determined, and the hardware incremental information of the corresponding target computing node is obtained after the specified time corresponding to the target computing node is reached. The specified time corresponding to the target computing node refers to the sum of the predetermined period and the random delay corresponding to the target computing node.
[0183] Optionally, according to the sum of the predetermined period and the random delay calculated by itself, the time point of actually obtaining and reporting the hardware incremental information is determined. For example, if the predetermined period is 30 seconds and the random delay range is 0 to 5 seconds, the actual reporting time point will be between 30 seconds and 35 seconds.
[0184] Optionally, as shown in Figure 6 , taking the case that the second component (MIG-AGENT) runs on the target computing node, the MIG-AGENT starts a timing task with a period of 30s and a random delay of 0-5s to avoid concurrency.
[0185] Through this embodiment, the random delay corresponding to the target computing node is determined, which avoids the reporting of data by all computing nodes at the same time, effectively balances the network traffic, reduces network congestion and delay, and obtains the hardware incremental information of the corresponding target computing node after the specified time is reached, which can automatically and periodically obtain the hardware incremental information, reducing the need for manual intervention and simplifying the management process of the hardware information.
[0186] In an exemplary embodiment, the above-mentioned GPU resource allocation method further comprises:
[0187] In the case that the target computing node detects a node restart event, the hardware information of the graphics processor is reacquired, and the hardware information carrying the event type field is reported to the management node to instruct the management node to overwrite the hardware information of the target computing node stored in the database with the hardware information carrying the event type field.
[0188] In the embodiment, the node restart event refers to the process that the computing node in the computing cluster experiences shutdown and restart due to system abnormality, hardware failure, software update, etc. Optionally, the node restart event can be planned, for example, during periodic maintenance or software upgrade update; or unplanned, for example, caused by sudden hardware failure or operation abnormality.
[0189] Optionally, when the hardware information of the corresponding computing node is reported to the management node, an event type field can be attached to the hardware information of the corresponding computing node to indicate that the hardware information of the corresponding computing node is full information update triggered by the node restart event. The event type field is used to mark what kind of event triggers the hardware information. For example, the event type field can mark that the hardware information is triggered by the node restart event. In addition, attaching an event type field to the hardware information of the corresponding computing node indicates that the hardware information of the graphics processor of the corresponding computing node is full information update, so as to ensure that the management node can more correctly and comprehensively acquire the latest hardware state and resource availability of the computing node after the node restart event.
[0190] After the management node acquires the hardware information of the graphics processor carrying the event type field, the hardware information of the corresponding graphics processor of the plurality of computing nodes stored in the database is overwritten with the hardware information carrying the event type field, so as to ensure that the database stores the information in the latest state and is not affected by any residual data before the restart.
[0191] Optionally, the second component (MIG-AGENT) runs on the target computing node, and the first component (MIG-CORE) runs on the management node. For example, MIG-AGENT acquires GPU information of the computing node at regular intervals and reports the GPU information to MIG-CORE; wherein MIG-AGENT detects host restart, container restart, and MIG-AGENT restart, which all trigger state reporting to MIG-CORE. In this way, through the embodiment, the dual synchronization mechanism of regular reporting (such as 30 seconds period) and event triggering (host / container restart) is combined to ensure the information consistency of MIG-CORE and MIG-AGENT; through MIG-CORE actively pulling data (when abnormality is recovered) and data verification (signature + hash comparison), the information inconsistency problem caused by the dependence on K8S label synchronization in the prior art is solved, and the system availability is improved to more than 99.9%.
[0192] Optionally, such as Figure 6 As shown, taking the second component (MIG-AGENT) running on the target compute node and the first component (MIG-CORE) running on the management node as an example, when MIG-AGENT on the target compute node detects a node restart event (i.e., MIG-AGENT restarts), MIG-AGENT re-acquires the full GPU information from the GPU driver. The GPU driver returns the complete GPU information to MIG-AGENT, and MIG-AGENT forces the reporting of the full GPU information to MIG-CORE, where the full GPU information is marked with an event type field. Then, MIG-CORE uses the full GPU information carrying the event type field to overwrite the node information stored in the ETCD database. The ETCD database returns a successful storage response to MIG-CORE, and MIG-CORE sends a reporting confirmation response to MIG-AGENT. For example, when node GPUs are plugged in or removed, or when the cluster is scaled up or down, MIG-AGENT can detect changes in real time and trigger information retransmission. MIG-CORE automatically updates its configuration combination, adapting to cluster topology changes without manual intervention, meeting the elastic scaling requirements of large-scale artificial intelligence platforms.
[0193] In this embodiment, when the target computing node detects a node restart event, it re-acquires the hardware information of the graphics processor, ensuring that the received hardware information is up-to-date. This avoids residual or inaccurate hardware information of the graphics processor caused by node restart, improving the reliability of hardware information in the database. Furthermore, by marking the hardware information reported to the management node through the event type field, the target computing node can identify and take corresponding response measures according to different types of events, promptly overwrite and update the database, thereby quickly restoring to the normal state before the restart and reducing the impact of the restart on system state synchronization.
[0194] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory (ROM) / random access memory (RAM), magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this application.
[0195] According to another aspect of the embodiments of the present application, there is also provided a graphic processor resource allocation apparatus which can be used to implement the graphic processor resource allocation method provided in the above-mentioned embodiments, and the description has been made above. As used below, the term "unit" can be a combination of software and / or hardware which implements a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, implementation in hardware, or a combination of software and hardware is also possible and contemplated.
[0196] Figure 7 is a structural block diagram of an optional graphic processor resource allocation apparatus according to the embodiments of the present application, as shown in Figure 7 The graphic processor resource allocation apparatus, running in a management node in a computing cluster, comprises:
[0197] The first execution unit 702 is configured to generate a configuration instruction according to a configuration request sent by a target object; the configuration request comprises a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list of target configuration instance combinations specified by the target object; the plurality of configuration templates are generated according to hardware data of a plurality of graphic processor types; the plurality of configuration templates correspond to the plurality of graphic processor types one by one; a configuration template in the plurality of configuration templates comprises a plurality of configuration instance combinations of graphic processors of the corresponding graphic processor type; a configuration instance combination in the plurality of configuration instance combinations refers to a split configuration of splitting resources of the graphic processors of the corresponding graphic processor type into a plurality of independent units.
[0198] The second execution unit 704 is configured to send the configuration instruction to the target computing nodes in the target computing node list to instruct the target computing nodes to allocate resources of graphic processors on the target computing nodes according to the configuration instruction.
[0199] It should be noted that the first execution unit 702 in this embodiment can be used to execute the above-mentioned step S202, and the second execution unit 704 in this embodiment can be used to execute the above-mentioned step S204.
[0200] According to another aspect of the embodiments of the present application, there is also provided another graphic processor resource allocation apparatus which runs in a target computing node in a computing cluster, comprising:
[0201] The third execution unit is configured to allocate resources of the graphic processor on the target computing node according to a configuration instruction sent by the management node in the computing cluster; the configuration instruction is generated based on a configuration request sent by the target object; the configuration request comprises a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates are generated by the management node according to hardware data of a plurality of graphic processor types; the plurality of configuration templates correspond to the plurality of graphic processor types one by one; and the configuration template in the plurality of configuration templates comprises a plurality of configuration instance combinations of the graphic processor of the corresponding graphic processor type; and the configuration instance combination in the plurality of configuration instance combinations refers to a split configuration of splitting the resources of the graphic processor of the corresponding graphic processor type into a plurality of independent units.
[0202] It should be noted that the above various units can be implemented by software or hardware, and for the latter, the following implementation manners can be used, but are not limited thereto: the above modules are located in the same processor; or the above various modules are located in different processors in any combination.
[0203] According to another aspect of the embodiments of the present application, a computer readable storage medium is provided, which includes a stored program, wherein the program performs the steps in any of the above method embodiments when running.
[0204] In an example embodiment, the above computer readable storage medium can include, but is not limited to, a U disk, a ROM, a RAM, a mobile hard disk, a magnetic or optical disk, and various computer program storage media.
[0205] According to another aspect of the embodiments of the present application, an electronic device is provided, which includes a memory, a processor, and a computer program stored in the memory and executable on the processor, and the processor is configured to perform the steps in any of the above method embodiments by the computer program. In an example embodiment, the electronic device can further include a transmission device and an input and output device, wherein the transmission device is connected to the processor, and the input and output device is connected to the processor.
[0206] The specific examples in the embodiments can refer to the examples described in the above embodiments and example implementations, which will not be described herein again.
[0207] The embodiments of the present application further provide a computer program product, which includes a computer program, and the computer program performs the steps in any of the above graphic processor resource allocation method embodiments when executed by a processor.
[0208] The embodiment of the present application further provides another computer program product, comprising a nonvolatile computer readable storage medium, the nonvolatile computer readable storage medium storing a computer program, the computer program being executed by a processor to implement the steps in any of the above-mentioned embodiments of the method for allocating resources of a graphics processor.
[0209] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present application can be realized by general computing devices, which can be centralized on a single computing device or distributed on a network composed of multiple computing devices, and can be realized by program codes executable by the computing devices, so that they can be stored in storage devices and executed by the computing devices, and in some cases, the steps shown or described can be executed in different orders, or they can be respectively manufactured into individual integrated circuit modules, or multiple modules or steps among them can be manufactured into a single integrated circuit module. Thus, the present application is not limited to any specific combination of hardware and software.
[0210] The above merely illustrates the preferred embodiments of the present application, and is not used to limit the present application. Any modification, equivalent replacement, improvement, etc. within the principles of the present application should be included in the protection scope of the present application.
Claims
1. A method of allocating resources to a graphics processor, the method comprising: A management node running in a computing cluster, comprising: generating a configuration instruction according to a configuration request sent by a target object; the configuration request comprising a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates being generated according to hardware data of a plurality of graphics processor types of graphics processors; the plurality of configuration templates corresponding one-to-one to the plurality of graphics processor types; a configuration template in the plurality of configuration templates comprising a plurality of configuration instance combinations of a graphics processor of a corresponding graphics processor type; a configuration instance combination in the plurality of configuration instance combinations referring to a split configuration of splitting resources of the graphics processor of the corresponding graphics processor type into a plurality of independent units; sending the configuration instruction to a target computing node in the target computing node list to instruct the target computing node to allocate resources of a graphics processor on the target computing node according to the configuration instruction; wherein a graphics processor type in the plurality of graphics processor types is taken as a current graphics processor type, and a graphics processor corresponding to the current graphics processor type is taken as a current graphics processor, and the following processing operations are performed to obtain the plurality of configuration templates corresponding to the plurality of graphics processor types: permutation and combination of a plurality of configuration instances in hardware information of the current graphics processor, determination of a group of configuration instance combinations satisfying a preset condition, and generation of a configuration template corresponding to the current graphics processor type according to the group of configuration instance combinations; the preset condition referring to a condition that a total number of slices occupied by a configuration instance combination on the current graphics processor is less than or equal to a maximum slice number of the current graphics processor, and a size of a video memory occupied by the configuration instance combination is less than or equal to a maximum video memory size of the current graphics processor.
2. The method of claim 1, wherein, Before the generating of the configuration instruction according to the configuration request sent by the target object, the method further comprises: after starting a plurality of computing nodes and the management node in the computing cluster, sending a hardware information acquisition request with an authentication token and / or a timeout time to the plurality of computing nodes to instruct the plurality of computing nodes to verify validity of the authentication token in the hardware information acquisition request, and in a case where the authentication token is valid, load a graphics processor driver within the timeout time in response to a hardware information acquisition request of the management node, acquire hardware information of a corresponding graphics processor through the graphics processor driver, and report a data packet containing the hardware information of the graphics processor to the management node; generating the plurality of configuration templates according to hardware data of a plurality of graphics processor types of graphics processors.
3. The method of claim 2, wherein, The hardware information of the plurality of computing nodes adopts a specified format string, and the specified format string corresponds to a first checksum, and before the generating of the plurality of configuration templates according to the hardware data of the plurality of graphics processor types of graphics processors, the method further comprises: The method comprises the following steps: taking a computing node in the plurality of computing nodes as a current computing node, receiving a data packet of the current computing node, and generating a second check sum corresponding to a specified format string in the data packet of the current computing node; and storing the data packet of the current computing node into a database according to a specified storage path and taking specified information of the current computing node as a key name in the case that the second check sum matches a first check sum in the data packet of the current computing node and the syntax of the specified format string is correct.
4. The method of claim 3, wherein, After the step of storing the data packet of the current computing node into the database according to the specified storage path and taking the specified information of the current computing node as the key name, the method further comprises: Setting a life cycle of the data packet of the current computing node to indicate that the database deletes the data packet of the current computing node in the case that the data packet of the current computing node is not updated within the life cycle.
5. The method of claim 2, wherein, The hardware information of the graphics processors of the same graphics processor type is the same, the hardware information of the graphics processors of different graphics processor types comprises maximum memory sizes, maximum slice numbers and a plurality of configuration instances of graphics processors of corresponding graphics processor types; a configuration instance in the plurality of configuration instances corresponding to different graphics processor types refers to a slice mode supported by a graphics processor of a corresponding graphics processor type; and a configuration instance combination in the plurality of configuration instance combinations comprises at least one configuration instance.
6. The method of claim 5, wherein, Before the step of generating the configuration template corresponding to the current graphics processor type according to the set of configuration instance combinations, the method further comprises: Discarding a configuration instance combination in the set of configuration instance combinations with a graphics processor resource utilization rate less than a preset utilization rate threshold, and performing the step of generating the configuration template corresponding to the current graphics processor type according to the set of configuration instance combinations after the discarding.
7. The method of claim 1, wherein, Before the step of generating the configuration instruction according to the configuration request sent by the target object, the method further comprises: Calling an eviction application interface to initiate a minimum execution unit eviction request to the target computing node; the minimum execution unit eviction request is used to instruct the target computing node to maintain state data and safely shut down a minimum execution unit within a preset time length; After the step of sending the configuration instruction to a target computing node in the list of target computing nodes to instruct the target computing node to allocate resources of a graphics processor on the target computing node according to the configuration instruction, the method further comprises: Initiating a recovery signal to the target computing node to instruct the target computing node to remove the minimum execution unit eviction limitation and allow a new minimum execution unit to be started on the target computing node.
8. The method according to any one of claims 1 to 7, characterized in that, The method further comprises: In a case where the management node detects an abnormal recovery event, a list of computing nodes of the computing cluster is acquired, a plurality of computing nodes in the computing cluster are actively loop-called to acquire hardware information of graphic processors on the plurality of computing nodes, and the hardware information of the graphic processors on the plurality of computing nodes is stored in combination with corresponding hardware information stored in a database.
9. A method of allocating resources to a graphics processor, the method comprising: A target computing node running in a computing cluster comprises: According to a configuration instruction sent by a management node in the computing cluster, resources of graphic processors on the target computing node are allocated; the configuration instruction is generated based on a configuration request sent by a target object; the configuration request comprises a target configuration template selected by the target object from a plurality of configuration templates, a target configuration instance combination selected from the target configuration template, and a target computing node list specified by the target configuration instance combination; the plurality of configuration templates are generated by the management node according to hardware data of graphic processors of a plurality of graphic processor types; the plurality of configuration templates correspond one-to-one to the plurality of graphic processor types; a configuration template in the plurality of configuration templates comprises a plurality of configuration instance combinations of graphic processors of a corresponding graphic processor type; a configuration instance combination in the plurality of configuration instance combinations refers to a splitting configuration that splits resources of graphic processors of the corresponding graphic processor type into a plurality of independent units; The management node is configured to take a graphic processor type in the plurality of graphic processor types as a current graphic processor type, take a graphic processor corresponding to the current graphic processor type as a current graphic processor, and perform the following processing operations to obtain the plurality of configuration templates corresponding to the plurality of graphic processor types: arrange and combine a plurality of configuration instances in the hardware information of the current graphic processor, determine a group of configuration instance combinations that satisfy a preset condition, and generate a configuration template corresponding to the current graphic processor type according to the group of configuration instance combinations; the preset condition refers to a condition that a total number of slices occupied by a configuration instance combination on the current graphic processor is less than or equal to a maximum number of slices of the current graphic processor, and a size of a display memory occupied by the configuration instance combination is less than or equal to a maximum size of a display memory of the current graphic processor.
10. The method of claim 9, wherein, Before the resources of the graphic processors on the target computing node are allocated according to the configuration instruction sent by the management node in the computing cluster, the method further comprises: receive the hardware information acquisition request with the authentication token and / or the timeout time sent by the management node, verify the validity of the authentication token in the hardware information acquisition request, and in the case that the authentication token is valid, load a graphic processor driver within the timeout time in response to the hardware information acquisition request of the management node, acquire hardware information of a graphic processor on the target computing node through the graphic processor driver, and report a data packet containing the hardware information of the graphic processor to the management node, so as to instruct the management node to generate a configuration template corresponding to a graphic processor type of the graphic processor on the target computing node according to the hardware information of the graphic processor on the target computing node in the case that the data packet passes the verification.
11. The method of claim 10, wherein, The reporting of the data packet containing the hardware information of the graphic processor to the management node comprises: converting the hardware information of the graphic processor into a specified format string, generating a first checksum corresponding to the specified format string, encapsulating the specified format string and the first checksum to obtain the data packet containing the hardware information of the graphic processor, and reporting the data packet to the management node.
12. The method of claim 9, wherein, The allocation of the resource of the graphic processor on the target computing node according to the configuration instruction sent by the management node in the computing cluster comprises: stopping a process running on the target computing node, calling a driver to enable a resource allocation mode of the target computing node after the process on the target computing node is successfully stopped, and creating the target configuration instance combination on the graphic processor on the target computing node; restoring the stopped process on the target computing node after the resource configuration of the graphic processor on the target computing node is successful.
13. The method according to any one of claims 9 to 12, characterized in that, The method further comprises: acquiring hardware incremental information of the target computing node in a timely manner, generating an incremental update packet based on the hardware incremental information of the target computing node, differentially compressing the incremental update packet, and reporting the compressed incremental update packet to the management node to instruct the management node to update a database according to the incremental update packet; the database is used to store hardware information of a plurality of computing nodes in the computing cluster.
14. The method of claim 13, wherein, The acquiring of the hardware incremental information of the target computing node in a timely manner comprises: determining a random delay corresponding to the target computing node, and acquiring the hardware incremental information corresponding to the target computing node after a specified time corresponding to the target computing node is reached; the specified time corresponding to the target computing node refers to a sum of a predetermined period and the random delay corresponding to the target computing node.
15. The method of claim 13, wherein, The method further comprises: in the case that the target computing node detects a node restart event, reacquiring hardware information of a graphic processor, and reporting the hardware information carrying an event type field to the management node to instruct the management node to overwrite the hardware information of the target computing node stored in the database by using the hardware information carrying the event type field.
Citation Information
Patent Citations
Cluster node management method and device, equipment and medium
CN119728689A
Application deployment method and device based on cloud computing power and storage medium
CN120256025A