Intelligent dynamic management method for GPU (Graphics Processing Unit) computing power and cloud platform

By collecting task data and GPU status data in multi-GPU systems, and combining this with a scheduling plugin to generate task queues and resource allocation strategies, this method optimizes resource allocation using reinforcement learning and task dependencies. This solves the problems of resource management complexity and fragmentation in multi-GPU systems, achieving intelligent and efficient resource scheduling, and is suitable for large-scale AI training and high-performance computing.

CN120994376APending Publication Date: 2025-11-21ZHEJIANG XIANGONG CLOUD TECH CO LTD
View PDF 0 Cites 7 Cited by

Patent Information

Application Number
CN202511089272.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-05
Publication Date
2025-11-21

AI Technical Summary

Technical Problem

In existing technologies, the hardware heterogeneity and resource management complexity of multi-GPU systems lead to rigid and fragmented resource allocation. Traditional static scheduling strategies are difficult to adapt to dynamic task loads, and older GPU architectures are unable to meet the needs of large-scale AI training and high-performance computing due to insufficient computing power density and lagging energy efficiency.

Method used

By collecting task data and GPU status data, and combining the scheduling plugin to generate task queues and resource allocation strategies, the resource allocation ratio is dynamically adjusted based on reinforcement learning. Task dependencies are analyzed and migration plans are generated to optimize communication paths. Synchronization state marking and hot migration strategies are introduced to release redundant resources in real time.

Benefits of technology

It significantly improves the resource utilization and task execution efficiency of multi-GPU systems, solves the resource fragmentation problem caused by task priority conflicts and communication bandwidth contention, realizes intelligent and efficient resource scheduling, and provides a scalable computing power management solution for large-scale AI training and high-performance computing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120994376A_ABST
    Figure CN120994376A_ABST
Patent Text Reader

Abstract

The invention is suitable for the field of GPU management, and provides an intelligent dynamic management method for GPU computing power and a cloud platform, and the method comprises the following steps: collecting task data and GPU state data, and generating a task queue and a resource allocation strategy in combination with a scheduling plug-in; based on the task queue and the GPU real-time load, dynamically adjusting a resource allocation proportion through reinforcement learning to analyze a task dependency relationship and generate a migration plan; optimizing a communication path and adjusting asynchronous transmission delay according to the task dependency relationship and the GPU communication topology, and outputting a synchronous state mark; and monitoring abnormity in combination with the synchronization state and the GPU hardware state, executing thermal migration according to the migration plan, performing video memory recovery, and updating the resource idle list. According to the invention, through algorithm innovation and hardware collaborative optimization, intelligent, dynamic and efficient resource scheduling in the multi-GPU system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application relates to the field of GPU management, in particular to a GPU computing power intelligent dynamic management method and a cloud platform. BACKGROUND

[0002] With the popularity of artificial intelligence, high-performance computing (HPC) and large-scale deep learning models, the global GPU market has reached a key turning point. The parameter quantity of AI training tasks has increased from hundreds of billions to trillions, and the memory occupation of a single model has increased from dozens of GB to hundreds of GB. However, the old architecture GPU is difficult to meet the needs of new scenarios due to insufficient computing power density and low energy efficiency ratio.

[0003] Taking distributed training as an example, if the gradient synchronization operation is not combined with the GPU communication topology (such as NVLink direct connection priority), high-priority tasks may be blocked due to the occupation of communication bandwidth by low-priority tasks. More seriously, in super-large-scale model training, the superposition effect of memory overflow and communication delay further aggravates resource contention, forcing the system to frequently trigger the memory recycling mechanism, significantly reducing training efficiency.

[0004] The hardware heterogeneity and resource management complexity of multi-GPU systems have become a significant pain point in the current technology ecosystem. On the one hand, the computing power density, memory bandwidth and communication capability of different types of GPUs differ greatly, and traditional static scheduling strategies are difficult to adapt to dynamic task loads. In addition, hardware manufacturers gradually terminate the support for old architecture GPUs, forcing users to accelerate the hardware upgrade cycle, but this also exacerbates the fragmentation of computing power resources and cost pressure. Therefore, a GPU computing power intelligent dynamic management method and a cloud platform are proposed to solve the above problems. SUMMARY

[0005] In view of the deficiencies in the prior art, the purpose of the present application is to provide a GPU computing power intelligent dynamic management method and a cloud platform to solve the problems in the background art.

[0006] The present application is implemented as follows: a GPU computing power intelligent dynamic management method, the method comprising the following steps: Collecting task data and GPU state data and generating a task queue and a resource allocation strategy in combination with a scheduling plug-in, the task data including task type and task priority; Based on the task queue and the GPU real-time load, the resource allocation proportion is dynamically adjusted through reinforcement learning to analyze the task dependency relationship and generate a migration plan, the GPU real-time load being obtained through the GPU state data; According to the task dependency relationship and the GPU communication topology, the communication path is optimized and the asynchronous transmission delay is adjusted, and a synchronization state mark indicating whether the current task has completed data consistency processing with other related tasks is output. According to the migration plan, the hot migration is performed and the video memory is recycled, and the resource idle list is updated, which is the status of all available computing resources.

[0007] As a further scheme of the application, the step of collecting task data and GPU state data and generating a task queue and a resource allocation strategy in combination with a scheduling plug-in specifically includes: Collecting data of the task to be processed, the data including task type and priority label; Obtaining GPU state data of each GPU device, the GPU state data including video memory occupation, computing load, temperature and communication bandwidth; Inputting the collected task data and GPU state data into the scheduling plug-in, so that the scheduling plug-in preliminarily sorts the tasks according to the task priority and GPU availability to generate an ordered task queue; Based on the task type and the GPU performance characteristics, a dynamic matching model is constructed, and a resource allocation strategy is generated, which can dynamically adjust the resource allocation proportion according to the real-time performance of the task.

[0008] As a further scheme of the application, the step of analyzing the task dependency relationship and generating a migration plan by dynamically adjusting the resource allocation proportion based on the task queue and the real-time load of the GPU through reinforcement learning specifically includes: Preliminarily dividing a resource allocation candidate set according to the priority and type of the task in the task queue; Based on the resource allocation candidate set, a resource-task matching matrix is constructed in combination with the data of the real-time load of the GPU, the matching matrix being used to record the adaptability of the available resources of each GPU to the candidate tasks; Based on the task dependency relationship, a migration priority ranking is generated according to the result of predicting the influence of migration on system performance by a graph neural network; In combination with a reinforcement learning model, the resource-task matching matrix and the migration priority ranking are taken as inputs to dynamically adjust the resource allocation proportion and generate a migration plan; Checking whether the GPU communication topology meets the data transmission requirement required by migration, which is used to determine whether the migration plan is executed.

[0009] As a further scheme of the application, the step of constructing a resource-task matching matrix based on the resource allocation candidate set in combination with the data of the real-time load of the GPU specifically includes: Real-time indicators of the GPU are collected, and time sequence features are extracted through an encoder to capture the mutation trend of the GPU load. The GPU and the task matched therewith are modeled as a dynamic relationship graph in the form of a node to quantify the adaptation degree. The matching score is calculated through a dynamic GNN in combination with the dynamic relationship graph to generate a real-time updated matching matrix. A real-time task migration triggering mechanism for dynamically adjusting the migration threshold is established to enable the migration of low-priority tasks and ensure the stable migration of high-priority tasks.

[0010] As a further scheme of the application, the step of optimizing the communication path and adjusting the asynchronous transmission delay according to the task dependency relationship and the GPU communication topology and outputting the synchronization state mark specifically comprises: The dependency relationship between tasks is extracted, and a communication path graph is constructed based on the GPU interconnection structure; The low-delay cross-GPU data transmission path is dynamically selected according to the dependency relationship and the communication path graph; When high-priority task transmission is performed, the communication bandwidth is preferentially allocated and the transmission interval of non-critical tasks is compressed; After the cross-GPU data transmission is completed, the transmitted data is verified using the version number to ensure the consistency of the data at the receiving end and the sending end; After the verification is passed, the task is marked as a synchronization completed state.

[0011] Another object of the application is to provide a GPU computing power intelligent dynamic management cloud platform, which comprises: A collection and scheduling module is configured to collect task data and GPU state data and generate a task queue and a resource allocation strategy in combination with a scheduling plug-in, wherein the task data includes task type and task priority; A dynamic allocation module is configured to analyze task dependency relationship and generate a migration plan by dynamically adjusting resource allocation proportion through reinforcement learning based on the task queue and GPU real-time load, wherein the GPU real-time load is obtained through the GPU state data; An optimization and synchronization module is configured to optimize the communication path and adjust the asynchronous transmission delay according to the task dependency relationship and the GPU communication topology, and output a synchronization state mark, wherein the synchronization state mark indicates whether the current task has completed data consistency processing with other related tasks; A migration execution and recycling update module is configured to combine the synchronization state and the GPU hardware state monitoring exception, execute hot migration according to the migration plan, recycle the video memory, and update a resource idle list, wherein the resource idle list is the status of all available computing resources.

[0012] As a further scheme of the application, the collection and scheduling module comprises: A collection sorting unit is configured to collect data of tasks to be processed, the data including a task type and a priority label; A state monitoring unit is configured to acquire GPU state data of each GPU device, the GPU state data including a memory occupation, a calculation load, a temperature, and a communication bandwidth; A task queue construction unit is configured to input the collected task data and the GPU state data to a scheduling plug-in, so that the scheduling plug-in generates an ordered task queue by preliminarily sorting tasks according to a task priority and GPU availability; A dynamic allocation adjustment unit is configured to construct a dynamic matching model based on a task type and a GPU performance feature, and generate a resource allocation strategy, the resource allocation strategy being capable of dynamically adjusting a resource allocation proportion according to a task real-time performance.

[0013] As a further scheme of the application, the dynamic allocation module comprises: A candidate set division unit is configured to preliminarily divide a resource allocation candidate set according to a priority and a type of a task in the task queue; A matching matrix construction unit is configured to construct a resource-task matching matrix based on the resource allocation candidate set and in combination with real-time load data of the GPU, the matching matrix being used to record an adaptability of available resources of each GPU to the candidate tasks; A prediction sorting unit is configured to generate a migration priority sorting according to a result of predicting an influence of migration on system performance based on a task dependency relationship and a graph neural network; A migration plan generation unit is configured to dynamically adjust a resource allocation proportion and generate a migration plan by taking the resource-task matching matrix and the migration priority sorting as inputs in combination with a reinforcement learning model; An adaptability checking unit is configured to check whether a GPU communication topology meets a data transmission requirement required by migration, and to determine whether the migration plan is executed.

[0014] As a further scheme of the application, the optimization and synchronization module comprises: A communication path construction unit is configured to extract a dependency relationship among tasks, and construct a communication path graph based on a GPU interconnection structure; A path selection unit is configured to dynamically select a low-delay cross-GPU data transmission path according to the dependency relationship and the communication path graph; A preferential allocation unit is configured to preferentially allocate a communication bandwidth and compress a transmission interval of a non-critical task when performing a high-priority task transmission; A data checking unit is configured to check data transmitted after a cross-GPU data transmission is completed by using a version number, to ensure consistency of data at a receiving end and a sending end; A state marking unit is configured to mark the task as a state of synchronization completion after the verification is passed.

[0015] Compared with the prior art, the present application has the following advantages: The present application significantly improves the resource utilization and task execution efficiency of the multi-GPU system through the collaborative optimization of dynamic resource scheduling and intelligent task migration. Specifically, the resource allocation ratio is dynamically adjusted based on reinforcement learning, and a migration plan is generated combining the real-time load of the GPU and the task dependency relationship, effectively alleviating the resource fragmentation problem caused by task priority conflicts and communication bandwidth contention; through task dependency modeling and GPU communication topology optimization, a low-latency transmission path is dynamically selected and the communication bandwidth of high-priority tasks is preferentially allocated, avoiding the problem of low-priority tasks blocking high-priority tasks in traditional static scheduling; a synchronization state marking and data consistency checking mechanism is introduced, combined with the hot migration and video memory recycling strategy, to release redundant resources in real time and update the resource idle list, reducing the abnormal interruption triggered by video memory overflow. In summary, through algorithm innovation and hardware collaborative optimization, the present application realizes the intelligentization, dynamization and high efficiency of resource scheduling in the multi-GPU system, providing a scalable computing power management solution for large-scale AI training and high-performance computing scenarios. BRIEF DESCRIPTION OF DRAWINGS

[0016] Figure 1 A flowchart of a method for intelligent dynamic management of GPU computing power.

[0017] Figure 2 A flowchart of a method for intelligent dynamic management of GPU computing power, in which task data and GPU state data are collected and combined with scheduling plugins to generate a task queue and resource allocation strategy.

[0018] Figure 3 A flowchart of a method for intelligent dynamic management of GPU computing power, in which based on the task queue and the real-time load of the GPU, the resource allocation ratio is dynamically adjusted through reinforcement learning to analyze the task dependency relationship and generate a migration plan.

[0019] Figure 4 A flowchart of a method for intelligent dynamic management of GPU computing power, in which a resource-task matching matrix is constructed based on the resource allocation candidate set and the data of the real-time load of the GPU.

[0020] Figure 5 A flowchart of a method for intelligent dynamic management of GPU computing power, in which the communication path is optimized and the asynchronous transmission delay is adjusted according to the task dependency relationship and the GPU communication topology, and the synchronization state marking is output.

[0021] Figure 6 A structural schematic diagram of a cloud platform for intelligent dynamic management of GPU computing power.

[0022] Figure 7 This is a schematic diagram of the acquisition and scheduling module in an intelligent dynamic management cloud platform for GPU computing power.

[0023] Figure 8 This is a schematic diagram of the structure of a dynamically allocated module in an intelligent dynamic management cloud platform for GPU computing power.

[0024] Figure 9 This is a schematic diagram of the optimization and synchronization module in an intelligent dynamic management cloud platform for GPU computing power. Detailed Implementation

[0025] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0026] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.

[0027] like Figure 1 As shown in the figure, this embodiment of the invention provides an intelligent dynamic management method for GPU computing power, the method comprising the following steps: S100: Collect task data and GPU status data, and combine them with the scheduling plugin to generate a task queue and resource allocation strategy. The task data includes task type and task priority. S200, based on task queues and GPU real-time load, analyzes task dependencies and generates migration plans by dynamically adjusting resource allocation ratios through reinforcement learning. The GPU real-time load is obtained through GPU status data. S300 optimizes the communication path and adjusts the asynchronous transmission delay based on task dependencies and GPU communication topology, and outputs a synchronization status flag, which indicates whether the current task has completed data consistency processing with other related tasks. S400, by combining the synchronization status and GPU hardware status monitoring for anomalies, performs hot migration and memory reclamation according to the migration plan, and updates the resource free list, which represents the current status of all available computing resources.

[0028] It should be noted that the scheduling plug-in is to integrate task priority, GPU performance characteristics and real-time state data, dynamically generate task queue and resource allocation strategy. The scheduling plugin collects task type (such as AI training, inference) and priority label (such as high / medium / low), and preliminarily sorts the tasks based on priority (such as high priority task first in queue); real-time acquisition of GPU device memory occupation, computing load, temperature and communication bandwidth and other state data as the basis for resource allocation; dynamically adjust the resource allocation proportion according to the real-time load change (such as GPU utilization sudden increase) (such as temporary expansion of low priority task resource quota), output task queue and migration candidate set, provide input for subsequent reinforcement learning. The scheduling plug-in solves the problem of resource allocation rigidity and fragmentation in traditional static scheduling through the dynamic scheduling logic driven by task priority.

[0029] In the embodiment of the application, the application significantly improves the resource utilization and task execution efficiency of the multi-GPU system through the cooperative optimization of dynamic resource scheduling and intelligent task migration. Specifically, the resource allocation proportion is dynamically adjusted based on reinforcement learning, and the migration plan is generated combining the real-time load of the GPU and the task dependency relationship, which effectively alleviates the resource fragmentation problem caused by task priority conflict and communication bandwidth contention; through task dependency modeling and GPU communication topology optimization, the low-delay transmission path is dynamically selected and the communication bandwidth of high-priority tasks is preferentially allocated, avoiding the problem of low-priority tasks blocking high-priority tasks in traditional static scheduling; the synchronization state marker and data consistency verification mechanism are introduced, combined with the hot migration and memory recycling strategy, to release redundant resources in real time and update the resource idle list, reducing the abnormal interruption triggered by memory overflow. In summary, through algorithm innovation and hardware cooperative optimization, the application realizes the intelligence, dynamics and efficiency of resource scheduling in the multi-GPU system, and provides a scalable computing power management solution for large-scale AI training and high-performance computing scenarios.

[0030] As shown in Figure 2 As a preferred embodiment of the application, the step of collecting task data and GPU state data and generating task queue and resource allocation strategy in combination with the scheduling plug-in specifically includes: S101, collecting data of the task to be processed, the data including task type and priority label; S102, acquiring GPU state data of each GPU device, the GPU state data including memory occupation, computing load, temperature and communication bandwidth; S103, inputting the collected task data and GPU state data to the scheduling plug-in, so that the scheduling plug-in preliminarily sorts the tasks according to the task priority and GPU availability to generate an ordered task queue; S104, a dynamic matching model is constructed based on the task type and the GPU performance characteristics, and a resource allocation strategy is generated, which can dynamically adjust the resource allocation ratio in real time according to the task performance.

[0031] In the embodiment of the application, the intelligent matching of tasks and GPU resources is realized through a scheduling plug-in. First, the data of the tasks to be processed (such as the task type is "deep learning training" and the priority label is "high") are collected, and the states of each GPU (such as the GPU1 memory occupancy is 70% and the calculation load is 85%, and the GPU2 memory occupancy is 30% and the load is 50%) are obtained in real time. Then, the scheduling plug-in inputs these data, and based on the task priority (high-priority tasks are given priority) and the GPU availability (such as the GPU2 memory is sufficient), the tasks are preliminarily sorted to generate an ordered task queue (such as the high-priority task A is preferentially allocated to the GPU2). The scheduling plug-in combines the task type (such as the task A is a memory-intensive task) and the GPU performance characteristics (such as the GPU2 has a higher memory bandwidth), generates a resource allocation strategy (such as 70% of the memory of the GPU2 is allocated to the task A) through a dynamic matching model. However, if the load of the GPU2 suddenly increases to 95% during the running of the task A, the scheduling plug-in will dynamically adjust the resource allocation ratio (such as temporarily expanding the memory allocation of the task A to 80%), thereby ensuring the efficient execution of the task. This process solves the problem of rigid resource allocation in traditional static scheduling, such as avoiding the low-priority task from preempting the memory resources of the high-priority task, and significantly improves the resource utilization and task response efficiency in the multi-GPU environment.

[0032] As shown in Figure 3 As a preferred embodiment of the application, the step of analyzing the task dependency relationship and generating a migration plan by dynamically adjusting the resource allocation ratio through reinforcement learning based on the task queue and the real-time load of the GPU, specifically includes: S201, preliminarily dividing a resource allocation candidate set according to the priority and type of the tasks in the task queue; S202, constructing a resource-task matching matrix based on the resource allocation candidate set and the data of the real-time load of the GPU, the matching matrix being used to record the adaptability of the available resources of each GPU to the candidate tasks; S203, generating a migration priority order according to the result of predicting the influence of migration on the system performance by the graph neural network based on the task dependency relationship; S204, combining the reinforcement learning model, taking the resource-task matching matrix and the migration priority order as inputs, dynamically adjusting the resource allocation ratio and generating a migration plan; S205, checking whether the GPU communication topology meets the data transmission requirements required for migration, which is used to determine whether the migration plan is executed.

[0033] In the embodiments of the present application, resource allocation and task migration are optimized through reinforcement learning and graph neural network. First, the system preliminarily divides the resource allocation candidate set according to the priority (such as high-priority task A and low-priority task B) and type (such as A being a memory-intensive and B being a computing-intensive) of the tasks in the task queue (such as A being a candidate for GPU1 and GPU2, and B being a candidate for GPU1). Then, based on the real-time load of the GPU (such as GPU1 having a memory occupancy of 80% and GPU2 having a memory occupancy of 30%), a resource-task matching matrix is constructed (for example, the adaptation degree of GPU1 to A is 0.6, and the adaptation degree of GPU1 to B is 0.8), and the resource matching of each GPU and task is quantified. Next, the graph neural network (GNN) is used to model the task dependency relationship (such as A waiting for the data output of B), predict the impact of migration on the system performance (such as if A is migrated to GPU2, the load of GPU1 decreases by 5%, but B needs to transmit data across GPUs, increasing the delay by 20ms), and generate a migration priority ranking (such as A is prioritized for migration). The reinforcement learning model takes the matching matrix and the migration ranking as input, and dynamically adjusts the resource allocation ratio (for example, 70% of the memory of GPU2 is allocated to A, and the low-priority task B on GPU1 is released). Finally, the system checks the GPU communication topology (such as the direct connection bandwidth between GPU1 and GPU2 through NVLink being 50GB / s), confirms whether the data transmission required for migration can be met (for example, B requires 10GB bandwidth for data transmission, which can be supported by NVLink), and executes the migration plan (such as migrating A to GPU2 and recycling the resources of GPU1) if it is met. Through dynamic resource reallocation and dependency-aware migration, this process solves the problem of resource waste caused by uneven load or communication blocking in traditional scheduling. For example, in a distributed training scenario, if GPU1 is overloaded with low-priority tasks and is close to memory overflow, the system can prioritize migrating high-priority tasks to idle GPUs, while ensuring data synchronization efficiency through NVLink to avoid training interruption.

[0034] As shown in Figure 4 As a preferred embodiment of the present application, the step of constructing a resource-task matching matrix based on the resource allocation candidate set in combination with the data of the real-time load of the GPU specifically includes: S212, real-time indicators of the GPU are collected, and time series features are extracted through an encoder to capture the mutation trend of the GPU load; S222, the dynamic relationship graph is modeled in the form of nodes for the GPU and the tasks matched therewith to quantify the adaptation degree; S232, the matching score is calculated through dynamic GNN in combination with the dynamic relationship graph to generate a real-time updated matching matrix; S242, a real-time task migration triggering mechanism for dynamically adjusting the migration threshold is established to enable the migration of low-priority tasks and ensure the stable migration of high-priority tasks.

[0035] In the embodiments of the present application, real-time matching and optimization of resources and tasks are realized through a dynamic graph neural network (GNN) and a migration threshold mechanism. First, the system collects real-time indicators such as GPU memory occupancy, temperature, and computing load (for example, the memory occupancy of GPU1 increases from 60% to 85%), and extracts time series features (for example, identifies the short-term trend of load surge) through an encoder (for example, a Transformer) to predict the resource contention risk that may occur in the next few seconds (for example, low-priority task B may block high-priority task A due to the high load of GPU1). Subsequently, the GPUs (such as GPU1 and GPU2) and tasks (such as tasks A and B) are modeled as nodes of a dynamic relationship graph, and the edge weight quantifies the degree of adaptation (assuming that the degree of adaptation of GPU1 to task A is 0.9, and the degree of adaptation of GPU2 to task B is 0.6), capturing the impact of real-time load changes on adaptability (for example, when the memory occupancy of GPU1 approaches the threshold, its degree of adaptation to task B drops to 0.4). In combination with the dynamic relationship graph, the matching score is calculated through a dynamic GNN to generate a real-time updated matching matrix, ensuring that the resource allocation strategy is dynamically adjusted with load fluctuations. Finally, a real-time task migration triggering mechanism for dynamically adjusting the migration threshold is established, for example, the migration threshold is dynamically set according to the variance of the GPU load, and when the load variance of GPU1 exceeds the set value, low-priority task B is triggered to migrate to GPU2. In this way, the stable execution of high-priority task A on GPU1 is ensured, and redundant resources are released through the migration of low-priority tasks, avoiding global blocking caused by sudden changes in the load of GPU1. For example, in a distributed training scenario, if the memory occupancy of GPU1 is too high due to the low-priority task B, the system can predict that the load of GPU1 will decrease after migration through a dynamic GNN, and trigger task B to migrate to GPU2, while ensuring the data synchronization efficiency through NVLink to ensure that the training process of high-priority task A is not disturbed.

[0036] As shown in Figure 5 As a preferred embodiment of the present application, the step of optimizing the communication path and adjusting the asynchronous transmission delay according to the task dependency relationship and GPU communication topology, and outputting the synchronization state marker, specifically includes: S301, extracting the dependency relationship between tasks, and constructing a communication path graph based on the GPU interconnection structure; S302, dynamically selecting a low-delay cross-GPU data transmission path according to the dependency relationship and the communication path graph; S303, when transmitting a high-priority task, preferentially allocating communication bandwidth and compressing the transmission interval of non-critical tasks; S304, after the cross-GPU data transmission is completed, verifying the transmitted data using a version number to ensure the consistency of the data at the receiving end with that at the sending end; S305, after the verification, mark the task as a synchronization completed state.

[0037] In the embodiment of the application, the efficient cooperation of cross-GPU data transmission is realized through task dependency perception and GPU communication topology optimization. First, the dependency relationship between tasks (such as task A needs to wait for the output result of task B) is extracted, and a communication path graph is constructed based on the GPU interconnection structure (such as NVLink). According to the dependency relationship and the communication path graph, a low-delay path is dynamically selected (such as when task A depends on task B, data is transmitted through NVLink direct connection between GPU1 and GPU2, rather than through GPU3). When executing a high-priority task, the system preferentially allocates communication bandwidth (such as reserving part of the bandwidth for high-priority tasks), and reduces the blocking of high-priority tasks by compressing the transmission interval of non-critical tasks (such as data preprocessing tasks) (such as reducing the transmission frequency of low-priority tasks). After the data transmission is completed, the system will use the version number verification mechanism to ensure data consistency. If the receiving end detects that the version numbers do not match, it will trigger the retransmission mechanism. After the verification, the system marks the task as a "synchronization completed" state, and then allows the subsequent task to continue execution. Through dynamic path planning and data consistency guarantee, this process solves the delay problem caused by improper path selection or bandwidth contention in traditional communication, and significantly improves the cooperation efficiency of multi-GPU systems.

[0038] As shown in Figure 6 The embodiment of the application also provides an intelligent dynamic management cloud platform of GPU computing power, which comprises: A collection and scheduling module 100 is configured to collect task data and GPU state data, generate a task queue and a resource allocation strategy in combination with a scheduling plug-in, and the task data comprises a task type and a task priority. A dynamic allocation module 200 is configured to analyze task dependency relationships and generate a migration plan by dynamically adjusting a resource allocation ratio through reinforcement learning based on a task queue and a GPU real-time load, and the GPU real-time load is obtained through GPU state data. An optimization and synchronization module 300 is configured to optimize a communication path and adjust an asynchronous transmission delay according to a task dependency relationship and a GPU communication topology, and output a synchronization state mark, wherein the synchronization state mark indicates whether a current task has completed data consistency processing with other related tasks. A migration execution and recycling update module 400 is configured to combine a synchronization state and a GPU hardware state monitoring exception, execute hot migration according to a migration plan, recycle a display memory, and update a resource idle list, wherein the resource idle list is a status of all available computing resources.

[0039] In the embodiment of the application, the cloud platform realizes intelligent dynamic management of GPU computing power through four modules. The collection and scheduling module first collects task type, priority and GPU state data (such as memory occupation, load), generates a task queue and an initial resource allocation strategy in combination with the scheduling plug-in; the dynamic allocation module dynamically adjusts the resource allocation ratio based on the task queue and the real-time load of the GPU (such as memory utilization), analyzes the task dependency relationship and generates a migration plan (such as high-priority task migration first); the optimization and synchronization module optimizes the cross-GPU data transmission path in combination with the task dependency relationship and the GPU communication topology (such as NVLink / PCIe bandwidth), compresses the transmission interval of low-priority tasks, and ensures data consistency through version number checking, and outputs a synchronization state marker; the migration execution and recycling update module combines the synchronization state and the GPU hardware monitoring (such as abnormal load), executes hot migration and memory recycling (such as releasing resources occupied by low-priority tasks), and updates the resource idle list in real time. Through the closed-loop process of dynamic perception-intelligent decision-cooperative optimization, the platform solves the problems of resource fragmentation and communication blocking in the multi-GPU scenario, and improves the computing power utilization rate and task stability.

[0040] As shown in Figure 7 , as a preferred embodiment of the application, the collection and scheduling module 100 comprises: a collection and sorting unit 101 for collecting data of tasks to be processed, the data comprising task type and priority label; a state monitoring unit 102 for obtaining GPU state data of each GPU device, the GPU state data comprising memory occupation, calculation load, temperature and communication bandwidth; a task queue construction unit 103 for inputting the collected task data and GPU state data to a scheduling plug-in, so that the scheduling plug-in preliminarily sorts the tasks according to task priority and GPU availability to generate an ordered task queue; a dynamic allocation unit 104 for constructing a dynamic matching model based on task type and GPU performance characteristics, and generating a resource allocation strategy, the resource allocation strategy being capable of dynamically adjusting the resource allocation ratio according to real-time performance of the task.

[0041] As shown in Figure 8 , as a preferred embodiment of the application, the dynamic allocation module 200 comprises: a candidate set division unit 201 for preliminarily dividing a resource allocation candidate set according to the priority and type of the tasks in the task queue; a matching matrix construction unit 202 for constructing a resource-task matching matrix based on the resource allocation candidate set in combination with the real-time load data of the GPU, the matching matrix being used to record the adaptability of the available resources of each GPU to the candidate tasks; The prediction ordering unit 203 is configured to generate a migration priority ranking according to the results of predicting the influence of migration on system performance based on task dependency relationships and according to a graph neural network. The migration plan generation unit 204 is configured to combine a reinforcement learning model, input the resource-task matching matrix and the migration priority ranking, dynamically adjust the resource allocation ratio, and generate a migration plan. The adaptability checking unit 205 is configured to check whether the GPU communication topology meets the data transmission requirements required by migration, and to determine whether the migration plan is executed.

[0042] As shown in Figure 9 As a preferred embodiment of the present application, the optimization and synchronization module 300 includes: The communication path construction unit 301 is configured to extract the dependency relationships between tasks, and construct a communication path graph based on the GPU interconnection structure; The path selection unit 302 is configured to dynamically select a low-latency cross-GPU data transmission path according to the dependency relationships and the communication path graph; The priority allocation unit 303 is configured to preferentially allocate communication bandwidth and compress the transmission interval of non-critical tasks when performing high-priority task transmission; The data verification unit 304 is configured to verify the transmitted data using the version number after the cross-GPU data transmission is completed, to ensure the consistency of the data at the receiving end with that at the sending end; The state marking unit 305 is configured to mark the task as a state of synchronization completion after the verification is passed.

[0043] The above only describes the preferred embodiments of the present application in detail, and does not limit the present application. Any modification, equivalent replacement, and improvement made within the spirit and principles of the present application shall be included in the protection scope of the present application.

[0044] It should be understood that although each step in the flowchart of each embodiment of the present application is displayed in sequence according to the arrow, these steps are not necessarily executed in the order indicated by the arrow. Unless otherwise specified herein, the execution of these steps is not strictly limited in sequence, and these steps can be executed in other orders. Moreover, at least part of the steps in each embodiment can include multiple sub-steps or multiple stages, which are not necessarily executed at the same time, but can be executed at different times, and the execution order of these sub-steps or stages is not necessarily sequential, but can be executed in rotation or alternation with at least part of other steps or sub-steps or stages of other steps.

[0045] Those skilled in the art can understand that all or part of the processes in the above-mentioned embodiment methods can be completed by instructing the relevant hardware through a computer program. The program can be stored in a non-volatile computer readable storage medium, and when the program is executed, the processes of the above-mentioned embodiments of the methods can be included. Any reference to memory, storage, databases, or other media in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. As an illustration but not limitation, RAM is available in many forms such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), Synchlink DRAM (SLDRAM), Rambus DRAM (RDRAM), direct Rambus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0046] Other embodiments of the present disclosure will be apparent to those skilled in the art with the disclosure herein. The present application is intended to cover any variations, uses, or adaptations of the present disclosure, including its general principles and specific embodiments, which are disclosed herein. This application is intended to cover such processes or methodologies falling within the scope of the present disclosure. The specification and examples are to be regarded as illustrative only, and the true scope and spirit of the present disclosure are indicated by the claims.

Claims

1. A method for intelligent dynamic management of GPU computing power, characterized in that, The method includes the following steps: Collect task data and GPU status data, and combine them with a scheduling plugin to generate a task queue and resource allocation strategy. The task data includes task type and task priority. Based on task queues and real-time GPU load, task dependencies are analyzed and migration plans are generated by dynamically adjusting resource allocation ratios through reinforcement learning. The real-time GPU load is obtained through GPU status data. Based on task dependencies and GPU communication topology, the communication path is optimized and the asynchronous transmission delay is adjusted, and a synchronization status flag is output. The synchronization status flag indicates whether the current task has completed data consistency processing with other related tasks. By combining the synchronization status and GPU hardware status monitoring for anomalies, a hot migration is performed according to the migration plan, and the video memory is reclaimed. The resource free list is updated, which represents the current status of all available computing resources.

2. The intelligent dynamic management method for GPU computing power according to claim 1, characterized in that, The steps of collecting task data and GPU status data and combining them with the scheduling plugin to generate task queues and resource allocation strategies specifically include: Collect data on tasks to be processed, including task type and priority label; Acquire GPU status data for each GPU device, including memory usage, computing load, temperature, and communication bandwidth; The collected task data and GPU status data are input into the scheduling plugin, which then performs a preliminary sorting of tasks based on task priority and GPU availability to generate an ordered task queue. A dynamic matching model is constructed based on task type and GPU performance characteristics, and a resource allocation strategy is generated. The resource allocation strategy can dynamically adjust the resource allocation ratio according to the real-time performance of the task.

3. The intelligent dynamic management method for GPU computing power according to claim 1, characterized in that, The steps described above, which analyze task dependencies and generate migration plans by dynamically adjusting resource allocation ratios through reinforcement learning based on task queues and real-time GPU load, specifically include: The resource allocation candidate set is initially divided based on the priority and type of tasks in the task queue; Based on the resource allocation candidate set and the GPU real-time load data, a resource-task matching matrix is ​​constructed. The matching matrix is ​​used to record the suitability of each GPU's available resources with the candidate tasks. Based on task dependencies, a migration priority ranking is generated by predicting the impact of migration on system performance using a graph neural network. By combining a reinforcement learning model with a resource-task matching matrix and migration priority ranking as input, the resource allocation ratio is dynamically adjusted and a migration plan is generated. Check whether the GPU communication topology meets the data transfer requirements for migration to determine whether the migration plan should be executed.

4. The intelligent dynamic management method for GPU computing power according to claim 3, characterized in that, The step of constructing a resource-task matching matrix based on the resource allocation candidate set and GPU real-time load data specifically includes: Real-time metrics of the GPU are collected, and temporal features are extracted through an encoder to capture abrupt trends in GPU load. A dynamic relationship graph is modeled in the form of nodes between the GPU and its matching tasks to quantify the fit. By combining dynamic relationship graphs with dynamic GNN to calculate matching scores, a matching matrix that is updated in real time is generated. Establish a real-time task migration triggering mechanism that dynamically adjusts migration thresholds to enable the migration of low-priority tasks and ensure the stable migration of high-priority tasks.

5. The intelligent dynamic management method for GPU computing power according to claim 1, characterized in that, The steps of optimizing communication paths and adjusting asynchronous transmission latency based on task dependencies and GPU communication topology, and outputting synchronization status flags, specifically include: Extract the dependencies between tasks and construct a communication path graph based on the GPU interconnect structure; Dynamically select low-latency cross-GPU data transfer paths based on dependencies and communication path graphs; When transmitting high-priority tasks, prioritize the allocation of communication bandwidth and compress the transmission intervals of non-critical tasks. After cross-GPU data transfer is completed, the version number is used to verify the transmitted data to ensure the consistency between the data at the receiving end and the data at the sending end. After successful verification, the task is marked as synchronously completed.

6. A cloud platform for intelligent dynamic management of GPU computing power, characterized in that, The cloud platform includes: The acquisition and scheduling module is used to collect task data and GPU status data and combine them with the scheduling plugin to generate task queues and resource allocation strategies. The task data includes task type and task priority. The dynamic allocation module is used to analyze task dependencies and generate migration plans by dynamically adjusting the resource allocation ratio based on task queues and GPU real-time load through reinforcement learning. The GPU real-time load is obtained through GPU status data. The optimization and synchronization module is used to optimize the communication path and adjust the asynchronous transmission delay according to the task dependency and GPU communication topology, and output a synchronization status flag, which indicates whether the current task has completed data consistency processing with other related tasks. The migration execution and recycling update module is used to monitor anomalies by combining synchronization status and GPU hardware status, perform hot migration and memory reclamation according to the migration plan, and update the resource free list, which is the status of all available computing resources at present.

7. The intelligent dynamic management cloud platform for GPU computing power according to claim 6, characterized in that, The data acquisition and scheduling module includes: The data acquisition and sorting unit is used to acquire data of the tasks to be processed, including task type and priority label. The status monitoring unit is used to acquire GPU status data of each GPU device, including memory usage, computing load, temperature and communication bandwidth. The task queue construction unit is used to input the collected task data and GPU status data into the scheduling plugin, so that the scheduling plugin can initially sort the tasks according to task priority and GPU availability to generate an ordered task queue. A dynamic adjustment unit is allocated to build a dynamic matching model based on task type and GPU performance characteristics, and generate a resource allocation strategy. The resource allocation strategy can dynamically adjust the resource allocation ratio according to the real-time performance of the task.

8. The intelligent dynamic management cloud platform for GPU computing power according to claim 6, characterized in that, The dynamic allocation module includes: The candidate set partitioning unit is used to initially partition the resource allocation candidate set based on the priority and type of tasks in the task queue; A matching matrix construction unit is used to construct a resource-task matching matrix based on the resource allocation candidate set and the GPU real-time load data. The matching matrix is ​​used to record the suitability of each GPU's available resources with the candidate tasks. The prediction and ranking unit is used to generate a migration priority ranking based on the results of the graph neural network's prediction of the impact of migration on system performance, according to the task dependencies. The migration plan generation unit is used to combine a reinforcement learning model, taking the resource-task matching matrix and migration priority ranking as input, dynamically adjust the resource allocation ratio and generate a migration plan. The compatibility check unit is used to check whether the GPU communication topology meets the data transmission requirements of the migration and to determine whether the migration plan should be executed.

9. The intelligent dynamic management cloud platform for GPU computing power according to claim 6, characterized in that, The optimization and synchronization module includes: The communication path construction unit is used to extract the dependencies between tasks and to construct a communication path graph based on the GPU interconnect structure. The path selection unit dynamically selects low-latency cross-GPU data transfer paths based on dependencies and the communication path graph; The priority allocation unit is used to prioritize the allocation of communication bandwidth and compress the transmission interval of non-critical tasks when high-priority task transmission is performed. The data verification unit is used to verify the transmitted data using the version number after the cross-GPU data transfer is completed, in order to ensure the consistency between the data at the receiving end and the data at the sending end. The status marking unit is used to mark the task as synchronously completed after the verification is passed.

Citation Information

Cited By

  • Fully-mechanized coal mining three-machine management and control system

    CN121187144A

  • GPU-oriented multi-queue adaptive scheduling method and system

    CN121210077A

  • A Multi-Queue Adaptive Scheduling Method and System for GPUs

    CN121210077B

  • Video memory control method, device, equipment and system and computer storage medium

    CN121255472A

  • Industrial data migration dynamic priority adjustment method for productivity platform

    CN121434188A