Cloud platform-based computing power resource dynamic scheduling and monitoring method

By using cloud platform computing resources dynamic scheduling and monitoring methods, combined with task profiling and resource trend prediction, efficient and stable GPU cluster scheduling was achieved. This solved the static nature and lag issues of existing scheduling platforms, and improved resource utilization and task SLA achievement rate.

CN121070513APending Publication Date: 2025-12-05北京娱广科技有限公司

Patent Information

Application Number
CN202511071317.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-31
Publication Date
2025-12-05

AI Technical Summary

Technical Problem

Existing scheduling platforms such as Kubernetes and YARN have overly static scheduling strategies in high-performance GPU clusters. They cannot dynamically identify task types and latency tolerances, lack fine-grained awareness of underlying GPU resources, fail to drive task SLAs in a closed loop, and experience lag in migration response when resources are hot or abnormal nodes occur, affecting system stability.

Method used

A cloud-based dynamic scheduling and monitoring method for computing resources is adopted. Through node status collection, task profile modeling, resource trend prediction, SLA tracking, and scheduling scoring and decision-making, intelligent scheduling with task awareness, topology awareness, and risk prediction is achieved.

Benefits of technology

Significantly improves GPU resource utilization, reduces task queuing time and overall latency, increases task SLA achievement rate, shortens fault node recovery response time, and ensures system stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121070513A_ABST
    Figure CN121070513A_ABST
Patent Text Reader

Abstract

The invention relates to a computing power resource dynamic scheduling and monitoring method based on a cloud platform. The method is suitable for an intelligent scheduling scene of a high-performance GPU cluster. The method comprises seven steps of task portrait modeling, GPU node state acquisition, resource trend prediction, SLA tracking, scheduling scoring and deployment, operation monitoring and task migration, and SLA feedback optimization. According to the system, task semantics are represented by constructing task vectors, node health states, topology affinity, SLA historical performance conditions and resource prediction risks are fused, a multi-factor adjustable scheduling scoring mechanism is constructed, and second-level perception and task thermal migration of high-temperature nodes are achieved. Compared with a traditional Kubernetes static scheduling scheme, the method has the advantages that the GPU utilization rate, the task SLA achievement rate and the system stability are remarkably improved, the learning ability, the self-adaptive ability and the high availability are achieved, and the method is an intelligent scheduling closed-loop system oriented to AI reasoning and training scenes.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of cloud computing and artificial intelligence infrastructure management, in particular to a dynamic computing power scheduling and real-time performance monitoring method for high-performance GPU clusters, especially suitable for video live inference, recommendation system, model training and other scenarios. BACKGROUND

[0002] With the wide application of big data and deep learning, AI computing power infrastructure grows exponentially. High-performance GPU clusters represented by NVIDIA's related products are widely deployed in data centers for high-load tasks such as video encoding inference and machine learning model training. However, existing scheduling platforms such as Kubernetes and YARN generally have the following problems: the scheduling strategy is too static and cannot dynamically identify task types and delay tolerance; there is a lack of fine perception of GPU underlying resources such as Tensor Core and memory fragmentation; task SLA (such as response time and training completion time) cannot drive scheduling optimization in a closed loop; when the resource temperature is high or an abnormal node fails, the task migration response is lagging, affecting the overall system stability.

[0003] Therefore, a unified solution is needed that can integrate task feature modeling, GPU resource topology analysis, real-time monitoring feedback and intelligent decision scheduling. SUMMARY

[0004] The present application provides a cloud platform-based computing power resource dynamic scheduling and monitoring method to overcome at least one technical problem in the related art.

[0005] According to a first aspect of the embodiments of the present application, a cloud platform-based computing power resource dynamic scheduling and monitoring method is provided, characterized by comprising the following steps: S1, node state acquisition: real-time acquisition of node historical sampling data of each GPU node; the node historical sampling data includes temperature, memory usage, SM utilization and ECC error code, and is stored as time series data; S2, task portrait modeling: the characteristics of the task to be scheduled are encoded as a task vector; the task vector includes model type, delay tolerance, frame rate, historical completion time and error rate, etc.; S3, resource trend prediction: according to the node historical sampling data, a long short-term memory network or an autoregressive integrated moving average model is used to predict future resource load and temperature trend, and a risk label is generated; S4, SLA tracking: real-time recording of SLA indicators during task running; the SLA indicators include response time, completion rate and frame loss rate; S5, scheduling score and decision: call the data in the above steps; for all healthy and available nodes, calculate the comprehensive score of the network topology affinity, node health, historical SLA compliance probability and resource risk trend of the task in turn, and select the node with the highest score; S6, computing power execution and hot migration: deploy the task to the selected node with the highest score, and trigger container hot migration when detecting that the node with the highest score fails or high temperature alarm; S7, SLA feedback and optimization: record the response time delay, frame loss rate and completion time of the task, and feed back the compliance data to the task portrait and scheduling module to form a closed loop optimization.

[0006] Preferably, the score formula in step S5 is as follows: ; Wherein: T(N,T i ) represents the affinity (Topo) between the task vector T i and the node N in network and storage topology; H(N) represents the good degree of the current running state of the node; S(T i ,N) represents the historical SLA compliance rate of the node N in executing similar tasks T i ; R(N) represents the trend judgment of the prediction model on the resource availability of the node N in the future 1-3 minutes; and α, β, γ, δ are adjustable parameters.

[0007] Preferably, the sliding window width used in the resource trend prediction is not less than 60 seconds, and the resource index in the future 1 to 3 minutes is predicted.

[0008] Preferably, the node state acquisition acquires GPU index through DCGM interface with a period not greater than 1 second, and stores through Prometheus time series database.

[0009] Preferably, the scheduling score and decision module adopts a multi-factor weighted scoring mechanism which can be self-adaptively adjusted, and the scoring mechanism dynamically optimizes the scoring weight according to the historical SLA compliance rate.

[0010] Preferably, the historical SLA compliance probability is calculated based on the average response time delay and task completion rate of the node in executing similar tasks.

[0011] Another aspect of the application is to provide a computing power resource dynamic scheduling and monitoring system based on a cloud platform, characterized in that it comprises the following modules: A node state acquisition module is used to acquire the temperature, power consumption, video memory usage and SM utilization rate of each GPU node in real time, and store them in a time series database. The task profiling and modeling module is used to extract the structured semantic features of the tasks to be scheduled and generate task vectors; The GPU topology and prediction module is used to predict the future operating trend of GPU nodes based on historical data from a sliding window, obtain prediction results, and mark risky nodes. SLA tracking module: Records the SLA metrics for each task in real time for dynamic scoring and scheduling reinforcement learning; the SLA metrics include response time, completion rate, and frame drop rate; The scheduling scoring and decision module is used to calculate the scheduling score based on the task vector, the GPU node status and the prediction result, and select the optimal node for task scheduling. The computing power execution module is used to deploy tasks to the selected GPU node through the Kubernetes Operator or NVIDIA Operator interface and supports hot migration of containers; The event monitoring and linkage module is used to receive alarm information from the GPU node and trigger task rescheduling or load reduction operations; at the same time, it records the task performance and feeds it back to the scheduling scoring and decision module to optimize the scheduling strategy.

[0012] Another aspect of the present invention is to provide a computer-readable storage medium storing program instructions for performing the above-described method, wherein the program instructions, when loaded into and executed by a processor, cause the processor to complete each step of the above-described method.

[0013] The core technological advancement of this invention lies in proposing a scheduling and scoring mechanism that integrates task semantic modeling and multi-dimensional resource status assessment. By using task profile vectors and prediction-driven scoring formulas in synergy, the system can proactively select the optimal node based on task characteristics, thereby achieving a resource scheduling mechanism driven by task awareness, topology awareness, risk prediction, and SLA.

[0014] To achieve the above objectives, the technical innovations proposed in this invention are as follows: A multi-factor scoring scheduling function and a scheduling scoring and decision module were used: combining GPU topology affinity, node health status, historical SLA fulfillment capability, and predictive risk trends, an adjustable-weight scheduling scoring function Score[N] was constructed to accurately determine the optimal scheduling node. This scoring mechanism achieves interpretability, foresight, and adaptability in scheduling decisions, solving the problems of traditional scheduling's inability to dynamically avoid hot nodes and guarantee task SLA.

[0015] In addition, for the task portrait modeler: by extracting the structured semantic information of the task, a high-dimensional feature vector Ti is constructed, realizing intelligent scheduling driven by task semantics. This mechanism breaks through the static method of relying only on resource request parameters in traditional scheduling, enabling the scheduler to have task understanding ability and be able to develop personalized scheduling strategies for different task types (such as delay-sensitive and throughput-oriented).

[0016] Compared with existing scheduling methods, the application can significantly improve the utilization rate of GPU resources, reduce the task queuing waiting time and overall delay, effectively improve the SLA achievement rate of the task, and significantly shorten the response time of the fault node recovery. BRIEF DESCRIPTION OF DRAWINGS

[0017] In order to more clearly illustrate the technical solutions in the embodiments of the present specification or the related art, the drawings needed to be used in the embodiments or related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can also be obtained without creative labor on the basis of these drawings.

[0018] Figure 1 : System overall architecture diagram of the embodiment of the specification; Figure 2 : System overall operation flowchart of the embodiment of the specification; Figure 3 : GPU scheduling closed-loop flowchart based on task portrait and resource prediction of the embodiment of the specification; Figure 4 Structure diagram of a storage medium provided by the embodiment of the specification. DETAILED DESCRIPTION

[0019] The technical solutions in the embodiments of the present specification will be described clearly and completely below in combination with the drawings in the embodiments of the present specification. Obviously, the described embodiments are only some of the embodiments of the present application, not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0020] It should be noted that the terms "include" and "have" and any variations thereof in the embodiments of the present specification and the drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device that includes a series of steps or units is not limited to the listed steps or units, but can optionally include steps or units not listed, or can optionally include other steps or units inherent to the process, method, product or device.

[0021] The following embodiments are combined with Figures 1 to 2The technical solutions of the application are described in detail to show the system function modules and their cooperative mechanism.

[0022] Node state collector 01: Call DCGM interface at second level frequency, collect GPU temperature, power consumption, video memory usage, SM utilization, ECC error code and other indicators, and push to Prometheus database (Prometheus is an open source monitoring and alarm toolkit, and its core is a powerful time series database (Time Series Database, TSDB). It was originally developed by SoundCloud and has been widely used in containerized environments (especially Kubernetes) and microservice architecture monitoring). The Prometheus database has: strong time series data storage and compression capability; Support PromQL query language, flexible filtering of node state; Highly integrated with GPU monitoring API such as DCGM, especially suitable for "monitoring-alarm-linkage" links in GPU scheduling system. The collected data is used for subsequent module calculation.

[0023] Task image modeling device 02: Extract the structural features of the task (such as model type, delay level, frame rate, inference depth, etc.) and historical running records to generate task vector T i ∈Rⁿ(n=64 or 128), for the scheduler to call. This module is one of the key innovative modules of the application, which converts abstract task requirements into machine recognizable vector features by constructing task semantic expression, thereby realizing accurate matching between tasks and resources.

[0024] The above-mentioned task image modeling device 02 breaks through the traditional Kubernetes scheduling method of relying only on resource requests (such as video memory size and GPU number) and solves the problem of not being able to identify delay-sensitive tasks and task behavior differences. Task image is not only used for current scheduling, but also as an important input for SLA compliance history feedback and policy reinforcement learning, providing basic support for the system to build a "task-resource-behavior" three-dimensional closed-loop scheduling capability. GPU topology and prediction module 03: Collect the NVLink topology structure between GPUs in the cluster, and use the LSTM model to predict the resource trend in the next 1-3 minutes based on the sliding window historical data. If the temperature rises or the resource is tight, output the high-risk label R (N). The sliding window historical data is the state indicators such as GPU temperature, video memory usage, SM utilization, and ECC error code collected by the node state collector 1, which are stored in time series form.

[0025] SLA tracking module 04: Real-time record the response time, completion rate, frame loss rate and other SLA indicators of each task, for dynamic scoring and scheduling reinforcement learning.

[0026] To achieve accurate tracking, the system automatically injects lightweight logs and monitoring probes when deploying tasks, and collects execution data in real time through the Prometheus Exporter interface built into the container. The collected data is pushed to the Prometheus time series database at a frequency of seconds, and is aggregated and labeled by task ID. This module supports recording and updating the following SLA indicators: Response time (Latency): automatically calculated from task reception time and response return time; Completion rate (Success Rate): calculated by container return status code; Frame loss rate (Frame Drop Rate): frame count for FFmpeg or TensorRT codec modules.

[0027] Scheduling scoring and decision module 05: This module is a core innovation of the invention, which builds a multi-factor weighted scoring mechanism based on task profiling , node state information and future trend prediction results. The scheduling score no longer depends on traditional static resource matching logic, but integrates topology structure, health status, task performance history and predicted risk in multiple dimensions, achieving dynamic, intelligent and interpretable scheduling strategy selection.

[0028] The specific scoring formula is as follows: Equation (1); Where: : represents the affinity (Topo) between the task vector and the node in the network and storage topology, high indicating small data transmission delay and sufficient bandwidth, suitable for deployment; : represents the goodness of the current running state of the node, including temperature, power consumption, GPU fragmentation and other factors; : represents the historical SLA compliance rate of the node in executing similar tasks , which reflects the task performance ability; : the trend judgment of the availability of the node in the next 1-3 minutes by the prediction model, such as continuous temperature rise, power consumption close to the upper limit, etc., representing potential risks.

[0029] This formula (1) is used to comprehensively evaluate whether each GPU node N is suitable for scheduling a specific task The design idea is to weight and sum four key indicators: resource topology affinity, node health status, SLA compliance probability, and resource future risk trend, so as to select the node with the highest comprehensive score. The score function solves the following technical problems of traditional scheduling methods: The inability to identify the topology between GPUs leads to cross-card communication bottlenecks; ignoring node health leads to improper scheduling of high-temperature nodes, reducing stability; lack of SLA awareness leads to misallocation of critical tasks to low-reliability nodes; and lack of forward-looking prediction mechanism, unable to avoid nodes that are about to fail or have resource shortages. Through scoring and sorting by the Score function, the scheduler can optimally select the most suitable GPU node for executing the task in a data-driven manner, thereby achieving an efficient, stable, and self-adaptive intelligent scheduling strategy.

[0030] wherein is an adjustable parameter, reflecting the relative importance of each scoring dimension in the scheduling strategy. The system can use the following recommended default values in the initial stage: The above weights can be manually configured according to specific business needs, or optimized adaptively through historical SLA achievement rate and scheduling success rate statistical data. The value range of each parameter is [0, 1], and the sum of the four is 1, to ensure the normalization and comparability of the scoring results. The system supports dynamic fine-tuning of this set of parameters based on reinforcement learning mechanism, to adapt to scheduling needs in different task types (such as delay-sensitive tasks, throughput-priority tasks) and different load environments, further improving the adaptive ability and robustness of scheduling decisions. The higher the Score, the higher the priority.

[0031] Computing power execution module 06: responsible for completing task deployment, GPU resource mounting, and container hot migration operations by calling Kubernetes API or NVIDIA Operator interface according to the target node instructions of the scheduler; if deployment exceptions or running failures occur during task execution, the module will also feedback the execution status to the event monitoring module; Event monitoring and linkage module 07: receives high-temperature, abnormal power consumption, and high GPU utilization events from the Prometheus warning manager (AlertManager), and receives task running exception feedback from the computing power execution module, and then automatically triggers task migration, priority adjustment, or resource load reduction linkage response actions, forming a closed-loop operation safety guarantee mechanism.

[0032] To more clearly illustrate the scheduling scoring and task allocation process, the following gives the pseudo code of the core algorithm logic of the scheduler. The algorithm takes the task For input, traverse all GPU nodes, score available nodes based on the aforementioned scoring function, and select the highest scorer as the target deployment node. The scheduler ensures that tasks run on nodes that are healthy in resources, topologically affinitive, reliable in compliance, and avoid predicted risk nodes, thereby achieving an efficient, stable, and task-aware scheduling strategy: The scheduling pseudo code is as follows: Input: T i , GPU Nodes {N1,..., N n} Output: BestNode N* For each N in GPU Nodes : If N is Healthy and Available: Else: N*=argmax Score Dispatch T i to N* The computing power execution module 06 receives the target node and task configuration given by the scheduler, calls the Kubernetes native command or NVIDIA Operator interface to complete container deployment, GPU resource allocation and hot migration operation, ensures that the task can fall on the best target node according to the scoring strategy, and has hot migration capability for fault avoidance.

[0033] The event monitoring and linkage module 07 is responsible for receiving events triggered by Prometheus alarm rules (such as high temperature, tight memory, etc.), and notifying the scheduler to start re-scoring decisions or trigger hot migration through the Webhook mechanism, forming a closed-loop linkage capability of scheduling-running-monitoring-rescheduling.

[0034] The system forms a real-time closed loop through data interfaces and API calls between modules: task input→ portrait modeling→ resource evaluation→ scheduling decision→ execution deployment→ running monitoring→ SLA feedback→ optimization and rescheduling, realizing an intelligent GPU scheduling system with learning ability, self-adaptability and high availability. The following is a summary of the key points of the functions of each module and the corresponding closed-loop steps.

[0035] Table 1 Module composition of the GPU scheduling system of the application and its role in the closed-loop process Referring to Figure 3, the functions of each module and the corresponding closed-loop steps are described in combination with a typical scenario to show how the system works in real business, dynamically adapts, and responds in time: An AI platform deploys real-time live image inference services, and the service object is more than 2000 concurrent users of video streams. Each video needs to complete inference, encoding and response within 200ms. The system SLA requires an average response time delay ≤200ms, an allowed frame loss rate <0.5%, and a requirement that the annual service availability is not less than 99.95%.

[0036] Under the traditional Kubernetes scheduling strategy, the scheduler is difficult to identify the load trend of the GPU resources under high concurrency, resulting in that part of the tasks are scheduled to the nodes with high temperature or serious memory fragmentation, causing the response delay fluctuation in a short time, occasional frame loss, and even the problem of GPU downtime and unable to transfer tasks in time.

[0037] In the scheme of the present application, the system first identifies the live inference task as "delay sensitive type" through the task portrait modeler, and constructs its task feature vector including the inference model size, input frame rate, delay tolerance, etc. In the scheduling stage, the scheduler selects Node12 with high topology affinity and good health indicators as the target node. The memory capacity of the node is 10GB, and the initial temperature is 76°C.

[0038] After running for about 10 minutes, the temperature of Node12 continues to rise, and exceeds 89°C at the 642nd second. Prometheus real-time monitoring data triggers a high temperature alarm event, and the system judges that the node has a resource overheating risk. The scheduler confirms that the resource state of Node12 will continue to deteriorate in the next 2 minutes according to the trend judgment of the prediction sub-module LSTM model, and then re-performs the scheduling scoring.

[0039] After integrating the scores of each node, the system selects Node14 with a current temperature of 72°C and a memory free of 12GB as the migration target. The computing power execution module triggers container cold migration through NVIDIA Operator, and at the same time, uses the MIG mechanism to retain the task state, and the migration process takes 1.2 seconds. During this period, the system automatically retains the frame sequence in the edge cache to realize the non-jitter transition. After migration, the system automatically reduces the priority of Node12 and enters the cooling state.

[0040] After the migration is completed, the task runs stably in Node14, and the average inference delay is , which is much lower than the upper limit of SLA. The system records this migration event in the scheduling log and uses it as a reinforcement learning sample for subsequent policy adjustment.

[0041] The scene fully verifies the scheduling capability of the system in identifying hot node, dynamically predicting resource trend, and guaranteeing high concurrency and low delay task.

[0042] In the prior art, the Kubernetes default scheduling scheme (referring to the standard task scheduling strategy adopted by the Kubernetes system without customized modification), the main features include: scheduling based on static resource request (CPU / GPU number, memory); without perception of task type, delay demand or SLA index; without considering the interconnection topology structure of GPU (such as NVLink bandwidth difference), lacking collection and processing of GPU core index (temperature, power consumption, Tensor Core usage rate, etc.); in the case of resource overload or GPU high temperature, the response is not timely, and it is difficult to realize dynamic hot migration of task. The present application has significant improvement in GPU resource utilization, task response delay control, and fault handling efficiency. For example, the traditional scheduling method often relies on static configuration and cannot identify GPU hot nodes in real time, while the present system can realize forward-looking task migration and SLA protection through the linkage of prediction mechanism and monitoring.

[0043] In the same physical cluster composed of 32 A100 GPUs, we built two test environments: one uses the dynamic scheduling system of the present application, and the other uses the Kubernetes native scheduling mechanism for comparison test. The experimental results are as follows: In terms of average GPU utilization, the present system significantly improves the resource matching accuracy through task portrait modeling and prediction scheduling, achieving an average GPU usage rate of 76.5%, while the Kubernetes default strategy cannot dynamically perceive the task structure, only reaching 61.2%; In terms of average queuing waiting delay, the present system realizes rapid allocation through node resource trend prediction and scheduling scoring mechanism, with an average waiting time of only 24 seconds, while the Kubernetes scheme needs to poll and wait due to the lack of prediction ability, with an average of 38 seconds.

[0044] In terms of SLA achievement rate, the present system introduces SLA backtracking and reinforcement feedback mechanism, prioritizes key business tasks, and the achievement rate is improved to 98.4%, while the task compliance rate in the traditional strategy is only 92.1%.

[0045] In terms of fault response, the present system realizes second-level high temperature perception and hot migration with the help of DCGM and AlertManager, and the task can be smoothly migrated within 1.2 seconds; while the Kubernetes lacks fine-grained GPU state collection, and the average fault recovery time is 3 minutes.

[0046] The specific comparison of the above experiments is shown in the following table: Table 2 GPU scheduling of the application and comparison of the experimental scheme of the Kubernetes default scheduling scheme In summary, the scheme of the application is superior to the traditional scheduling platform in core indicators, and fully verifies its practical value and technical advancement in dynamic AI scenarios.

[0047] Figure 4 A structural diagram of a storage medium provided by an embodiment of the present application is shown in the figure. Figure 4 As shown in the figure, a storage medium 100 stores a computer program 110 used in the computing device, which is executed by a processor to complete each step of the cloud platform-based computing power resource dynamic scheduling and monitoring method described above.

[0048] Those skilled in the art can understand that the accompanying drawings are only schematic diagrams of an embodiment, and the modules or flows in the drawings are not necessarily necessary for implementing the application.

[0049] Those skilled in the art can understand that the modules in the device in the embodiment can be distributed in the device of the embodiment according to the embodiment description, or can be changed and located in one or more devices different from the embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.

[0050] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the application, and not to limit them; although the application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that: it can still modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacement for part of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the application.

Claims

1. A cloud platform-based computing power resource dynamic scheduling and monitoring method, characterized in that, Comprising the following steps: S1, node state acquisition: real-time acquisition of node historical sampling data of each GPU node; the node historical sampling data includes temperature, video memory usage, SM utilization and ECC error code, and is stored as time series data; S2, task image modeling: the feature encoding of the to-be-scheduled task is taken as a task vector; the task vector includes model type, delay tolerance, frame rate, historical completion time and error rate, etc.; S3, resource trend prediction: according to the node historical sampling data, a long short-term memory network or an autoregressive integrated moving average model is used to predict future resource load and temperature trend, and a risk label is generated; S4, SLA tracking: real-time recording of SLA indicators during task running; the SLA indicators include response time delay, completion rate, frame loss rate; S5, scheduling score and decision: calling the data in the above steps; for all healthy and available nodes, the network topology affinity, node health degree, historical SLA compliance probability and resource risk trend of each node are calculated in turn, and the node with the highest score is selected; S6, computing power execution and hot migration: deploying the task to the selected node with the highest score, and triggering container hot migration when the node with the highest score fails or high temperature alarm is detected; S7, SLA feedback and optimization: recording the response time delay, frame loss rate and completion time of the task, and feeding back the compliance data to the task image and scheduling module to form a closed-loop optimization.

2. The method of claim 1, wherein, The formula of the comprehensive score in step S5 is as follows: Score[N] = a · T(N, T i ) + β · H(N) + γ · S(N, T i ) - δ · R(N) wherein: T(N, T i ) represents the affinity (Topo) between the task vector T i and the node N on the network and storage topology; H(N) represents the goodness of the current running state of the node; S(T i , N) represents the historical SLA compliance rate of the node N in executing the similar task T i ; R(N) represents the trend judgment of the prediction model on the resource availability of the node N within 1-3 minutes in the future; and a, b, g, d are adjustable parameters.

3. The method of claim 1, wherein, The sliding window width used in the resource trend prediction in step S3 is not less than 60 seconds, and the resource indicators in the future 1 to 3 minutes are predicted.

4. The method of claim 1, wherein, The node state acquisition in step S1 acquires GPU indicators through DCGM interface with a period not greater than 1 second, and stores them through Prometheus time series database.

5. The method of claim 1, wherein, The scheduling score and decision in step S5 adopts a multi-factor weighted scoring mechanism that can be adaptively adjusted, and the scoring weight is dynamically optimized according to the historical SLA compliance rate.

6. The method of claim 5, wherein, The historical SLA compliance probability in step S5 is calculated based on the average response time delay and task completion rate of the node executing the same type of task.

7. A cloud platform-based computing power resource dynamic scheduling and monitoring system, characterized in that, Comprising the following modules: A node state acquisition module (01) for real-time acquisition of temperature, power consumption, video memory usage, SM utilization indicators of each GPU node, and storage in a time series database; A task image modeling module (02) for extracting structured semantic features of the to-be-scheduled task and generating a task vector; A GPU topology and prediction module (03) for predicting future running trend of GPU nodes based on sliding window historical data, obtaining prediction results, and marking risk nodes; An SLA tracking module (04) for real-time recording of SLA indicators of each task for dynamic scoring and scheduling reinforcement learning; the SLA indicators include response time, completion rate, frame loss rate; A scheduling score and decision module (05) for calculating a scheduling score based on the task vector, the GPU node state and the prediction result, and selecting an optimal node for task scheduling; The computing power execution module (06) is configured to deploy tasks to selected GPU nodes and support container live migration through a Kubernetes Operator or an NVIDIA Operator interface; The event monitoring and linkage module (07) is configured to receive the GPU node alarm information, trigger task rescheduling or load reduction operation, record task fulfillment and feedback to the scheduling scoring and decision module (05) to optimize the scheduling strategy. 8.A computer readable storage medium, having stored thereon program instructions for executing the method of any one of claims 1-6, which, when loaded into a processor and executed, cause the processor to perform the steps of any one of claims 1-6.

Citation Information

Patent Citations

  • Computing power data management system and method based on distributed computing

    CN119025283A

  • Computing system and method for GPU (Graphics Processing Unit) computing power scheduling

    CN119645661A

  • Intelligent cluster fault tolerance method based on distributed component dynamic migration

    CN119961039A

  • Deep learning task hybrid deployment method and system

    CN119987974A

  • One-cloud multi-core heterogeneous resource hybrid scheduling method

    CN120263792A

Cited By

  • Multi-scene-oriented computing power dynamic scheduling optimization analysis method

    CN122044886A

  • A multi-scene-oriented computing power dynamic scheduling optimization analysis method

    CN122044886B

  • Task scheduling method and device based on intelligent agent, equipment and storage medium

    CN122195687A

  • An AI computing power dynamic scheduling method and system based on a hybrid cloud architecture

    CN122387580A