GPU computing power scheduling method based on one-cloud multi-core heterogeneous computing power platform

Through heterogeneous resource registration and modeling, virtualized resource reconstruction, topology-aware matching and dynamic load scheduling, the problems of low utilization and insufficient reliability in heterogeneous GPU resource management are solved, and efficient and flexible GPU computing power scheduling is achieved, which is suitable for a variety of computing scenarios.

CN120295785AActive Publication Date: 2025-07-11SAISI TECH (XIAN) CO LTD

Patent Information

Application Number
CN202510373465.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-27
Publication Date
2025-07-11
Estimated Expiration
2045-03-27

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively manage and schedule heterogeneous GPU resources, resulting in low resource utilization, unbalanced task scheduling, poor hardware topology adaptability, dynamic load response lag and insufficient system reliability.

Method used

Through heterogeneous resource registration and modeling, virtualized resource reconstruction, topology-aware matching, dynamic load scheduling and full-link monitoring, improved Hungarian algorithms and deep reinforcement learning technology are used to achieve unified management and efficient scheduling of multiple types of GPU resources.

Benefits of technology

It significantly improves the resource utilization rate, task execution efficiency and system reliability of heterogeneous computing power platforms, supports flexible scheduling of multiple chip architectures, ensures low-latency execution of critical tasks and early failure warning, and is suitable for AI training, scientific computing, and edge reasoning and other scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120295785A_ABST
    Figure CN120295785A_ABST
Patent Text Reader

Abstract

The invention provides a GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform, and the method comprises the following steps: S1, carrying out heterogeneous resource registration and modeling, accessing hardware equipment containing multiple types of GPUs through a resource registration module, collecting the equipment model, the video memory capacity and performance index data, and carrying out heterogeneous resource modeling; constructing a resource feature database containing a topological relation, and supporting hybrid access of chips; s2, virtualized resource reconstruction: pooling a physical GPU into virtual GPU resources by adopting a hardware abstraction layer technology, realizing video memory isolation and calculation unit division through a containerization technology, and configuring each virtual GPU instance with an independent drive stack and a security sandbox; and S3, submitting a multi-modal task, receiving a CUDA / OpenCL calculation task submitted by a user, analyzing task demand parameters including a calculation core number, a video memory occupation amount and a data throughput threshold, and generating a task descriptor containing a priority label.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computing power scheduling, and particularly relates to a GPU computing power scheduling method based on a heterogeneous computing power platform with one cloud and multiple chips. Background Art

[0002] Currently, with the continuous acceleration of enterprise digital transformation and application innovation, the connotation of "one cloud, multiple chips" is also constantly expanding. It not only refers to the management of different GPU architecture resources but also includes the management of computing power such as CPUs, FPGAs, and DPUs. "One cloud, multiple chips" uses a set of cloud platforms to manage computing resources of different chip architectures, realizing unified management and scheduling of heterogeneous resources, and shielding the underlying architecture differences through the cloud platform to provide users with consistent cloud computing services.

[0003] Based on different types of processors (such as CPUs, GPUs, etc.) and their corresponding memory, storage devices, etc., the hardware resources together constitute a heterogeneous computing environment. Through hardware virtual pooling technology, physical resources are abstracted into virtual resources to achieve flexible scheduling and management of resources. Virtualize hardware resources such as GPUs so that they can be uniformly managed and scheduled by the cloud platform. According to the characteristics of tasks and the availability of GPU resources, tasks are reasonably allocated to GPUs for execution. It includes scheduling algorithms and a scheduling engine, which are responsible for task allocation and scheduling according to task requirements and resource status. Users define the required computing power resources through declarative interfaces, and the scheduling system transparently allocates and manages resources for users. Users only need to focus on the requirements at the application layer without caring about the specific implementation of the underlying hardware resources.

[0004] The implementation of GPU computing power scheduling by "one cloud, multiple chips" has profound background significance. It not only meets the diverse computing power needs of users but also optimizes resource configuration and utilization, reduces supply chain risks, promotes technological innovation and industrial upgrading, and promotes the development of the digital economy. Therefore, on this premise, the heterogeneous computing power platform with one cloud and multiple chips needs to shield the hardware differences of different GPU architectures downward and provide a service-oriented computing power model encapsulation upward, providing computing power-related storage, network, operating system, monitoring and operation and maintenance, etc. in a service-oriented manner. Therefore, the heterogeneous computing power platform with one cloud and multiple chips supporting different GPU architectures of the Xinchuang industry to solve the "one cloud, multiple chips" problem is the most basic ability requirement.

[0005] With the explosive growth of business innovation in all walks of life, the differences in computing power requirements for different business scenarios are also increasing. For artificial intelligence, the requirement for computing accuracy is relatively low, but there is a higher requirement for parallel computing performance, and the GPU is the key component to meet this high-performance computing demand. By implementing GPU computing power scheduling through "one cloud, multiple chips", it can meet the diverse computing power needs of users in different application scenarios and improve computing efficiency and performance.

[0006] In the cloud computing industry, the performance of server chips varies widely, resulting in inconsistent user experiences and significantly different application effects. By implementing GPU computing power scheduling through "one cloud, multiple chips", the risk of choosing a technology route can be minimized because users can flexibly switch between different chip architectures without relying on a specific vendor or architecture. This helps improve the stability of the business and the flexibility of business transformation.

[0007] In a cloud computing environment, the utilization rates of various computing resources (including GPUs, etc.) are often uneven. By implementing GPU computing power scheduling through "one cloud, multiple chips", these heterogeneous resources can be incorporated into a unified cloud platform for management and dynamically allocated and scheduled according to actual needs. This can not only improve the utilization rate of resources but also optimize resource allocation and reduce operating costs.

[0008] Implementing GPU computing power scheduling through "one cloud, multiple chips" is a technological innovation in the field of cloud computing, which promotes the development of cloud computing platforms towards a more flexible, efficient, and intelligent direction. At the same time, this also provides a broader space for the application of acceleration processors such as GPUs in cloud computing, helping to promote the upgrading and development of related industries.

[0009] Computing power has become the core supporting force and driving force for promoting the development of the digital economy. By implementing GPU computing power scheduling through "one cloud, multiple chips", more computing power resources can be released, providing strong support for the innovative development of the digital economy. This helps to accelerate the pace of digital transformation and promote the high-quality development of the economic society. Summary of the Invention

[0010] The present invention proposes a GPU computing power scheduling method based on a heterogeneous computing power platform with one cloud and multiple chips. This method solves the problems of low resource utilization rate, unbalanced task scheduling, poor hardware topology adaptability, lag in dynamic load response, and insufficient system reliability in a heterogeneous GPU computing power platform, and realizes the unified management, efficient scheduling, and real-time optimization of multiple types of GPU resources.

[0011] The technical solution of the present invention is realized as follows: A GPU computing power scheduling method based on a heterogeneous computing power platform with one cloud and multiple chips, the method comprising the following steps:

[0012] Step S1: Heterogeneous resource registration and modeling. Connect to hardware devices containing multiple types of GPUs through a resource registration module, collect data on device models, video memory capacities, and performance indicators, and construct a resource feature database containing topological relationships to support the hybrid access of chips;

[0013] Step S2: Virtualized resource reconstruction. The physical GPU pool is virtualized into virtual GPU resources using the hardware abstraction layer technology. Memory isolation and computing unit partitioning are achieved through containerization technology. Each virtual GPU instance is equipped with an independent driver stack and a security sandbox.

[0014] Step S3: Multi-modal task submission. Receive the CUDA / OpenCL computing tasks submitted by the user, parse the task requirement parameters, including the number of computing cores, memory occupancy, and data throughput threshold, and generate a task descriptor containing a priority label.

[0015] Step S4: Topology-aware resource matching. Based on the real-time updated resource status graph, use the improved Hungarian algorithm for multi-dimensional resource matching. When selecting computing nodes, synchronously evaluate the PCIe bandwidth, NVLink connection status, and memory fragmentation degree.

[0016] Step S5: Dynamic load scheduling. Deploy a distributed scheduling engine to collect the temperature, utilization rate, and task queue depth data of the node GPUs in real time. When a local hot spot is detected, trigger the load balancing strategy and use a combination mechanism of task migration and computing power reallocation.

[0017] Step S6: Adaptive task execution. Inject a runtime monitoring agent into the GPU instance, dynamically adjust the CUDA stream priority and memory page locking strategy, and use the MPS multi-process service technology to achieve exclusive resource guarantee for critical tasks.

[0018] Step S7: Full-link monitoring and diagnosis. Build a three-dimensional monitoring system: Monitor the SM unit utilization rate and ECC error rate of the GPU hardware through the resource layer; Track the computing progress and memory leak risk at the task layer; Implement a health score at the node layer and trigger an automatic isolation mechanism.

[0019] As a preferred implementation, after completing the construction of the three-dimensional monitoring system, perform steps S8 and S9.

[0020] Step S8: Parameterized computing power optimization. Use an improved Adam optimizer to dynamically adjust the scheduling strategy parameters; Calculate the first-order moment estimate and second-order moment estimate of the computing cluster load gradient; Perform deviation correction based on historical load fluctuations; Update the task allocation weight coefficient according to the corrected momentum value; Introduce Nesterov accelerated gradient to predict the future load trend.

[0021] Step S9: Closed-loop feedback update. Collect the actual resource consumption data during task execution, compare and analyze it with the preset requirement parameters, iteratively optimize the resource matching algorithm through a reinforcement learning model, and update the specification template library of virtual GPUs.

[0022] Existing technologies usually design scheduling strategies for a single type of GPU, making it difficult to meet the mixed access requirements of multi-vendor and multi-architecture GPUs (such as NVIDIA, AMD, and domestic chips). Through the heterogeneous resource registration and modeling in step S1, this solution constructs a resource feature database containing topological relationships, supports the mixed access of chips, and solves the problems of device model differences, mismatched video memory capacities, and fragmented performance metrics. The technical difficulty lies in how to uniformly abstract the hardware characteristics of different architecture GPUs (such as the number of CUDA cores and Tensor Core configurations) and establish a cross-vendor topological relationship model (such as the scenario of mixed NVLink and PCIe connections). Traditional GPU virtualization technologies (such as GPU passthrough) cannot achieve fine-grained resource partitioning and have risks of video memory conflicts and driver compatibility. Through the hardware abstraction layer and containerization technology in step S2, this solution pools physical GPUs into independent virtual instances, equipped with isolated driver stacks and security sandboxes, and solves the problems of video memory leakage, computing unit preemption, and security isolation among multiple tasks. The technical difficulty lies in how to achieve the dynamic reconstruction of virtual GPU instances through the video memory page locking strategy and driver stack hot plug mechanism while avoiding performance losses caused by containerization.

[0023] Existing scheduling algorithms (such as the greedy algorithm) are difficult to handle the complex task requirements of multi-dimensional resource constraints (such as video memory, bandwidth, and computing cores). Through the task descriptor generation in step S3 and the topology-aware matching in step S4, this solution uses an improved Hungarian algorithm combined with the evaluation of PCIe bandwidth, NVLink connection status, and video memory fragmentation degree to solve the problem of accurate matching between task requirements and hardware topology. The technical difficulty lies in how to design a multi-dimensional weight scoring function in a dynamic resource environment and reduce the algorithm complexity to meet the requirements of real-time scheduling. Traditional scheduling systems rely on periodic resource collection and are difficult to respond promptly to local hotspots (such as sudden increases in GPU temperature and task queue congestion). Through the distributed scheduling engine and load balancing strategy in step S5, this solution collects temperature, utilization, and queue depth data in real time and uses a combination mechanism of task migration and computing power reallocation to solve the problems of local resource overload and low heat dissipation efficiency. The technical difficulty lies in how to design a low-latency task migration protocol to avoid performance jitter caused by data transfer.

[0024] Existing systems lack the ability to guarantee exclusive resources for critical tasks (such as AI inference), and runtime parameter adjustment relies on manual intervention. Through the runtime monitoring agent and MPS multi-process service technology in step S6, this solution dynamically adjusts the CUDA stream priority and video memory page locking strategy, solving the problem of latency fluctuations caused by resource competition for critical tasks. The technical difficulty lies in how to achieve dynamic switching of the video memory allocation strategy without interrupting task execution and balance the resource preemption conflicts among multiple processes. Traditional monitoring systems only focus on hardware-level metrics and lack a full-link health assessment of the task layer and node layer. Through the three-dimensional monitoring system in step S7, this solution synchronously tracks the SM unit utilization rate, video memory leakage risk, and node health score, solving the problems of ECC error accumulation, high concealment of video memory leakage, and fault diffusion. The technical difficulty lies in how to design a cross-layer correlation analysis model to achieve causal reasoning between hardware anomalies (such as overheating of SM units) and task anomalies (such as video memory leakage).

[0025] As a preferred implementation method, the dynamic load scheduling in step S5 includes a scheduling strategy optimization engine based on deep reinforcement learning. The specific implementation includes: constructing a state feature space, collecting the node GPU utilization rate, video memory occupancy rate, PCIe link bandwidth utilization rate, and task queue waiting duration to form a 16-dimensional state vector; designing a dual-channel Dueling DQN network architecture, where the value stream network evaluates the global scheduling benefit, and the advantage stream network calculates the local scheduling action advantage value; defining the reward function R = α * (1 - load variance) + β * task completion rate - γ * migration overhead, where α, β, and γ are adjustable weight coefficients; deploying an experience replay pool to store historical scheduling decision data and using a priority sampling mechanism to screen high-value training samples; implementing an online policy update mechanism that triggers immediate fine-tuning of network parameters when detecting changes in the cluster topology or sudden changes in the load pattern; introducing an action mask technology to constrain invalid scheduling operations and ensure that task migration only occurs between nodes with hardware compatibility.

[0026] As a preferred implementation method, the virtualized resource reconstruction in step S2 also includes: creating a securely isolated vGPU instance, using hardware-assisted SR-IOV technology to divide the physical GPU into multiple VFs, and binding each VF to an independent secure enclave; implementing dynamic video memory quota management, and completing the elastic expansion of vGPU video memory from 1GB to 24GB within 0.1 seconds according to task requirements; deploying a driver-level protection module to intercept abnormal CUDA API requests through system call hijacking technology and block unauthorized memory access operations; designing a hot migration channel to migrate the vGPU instance along with the running tasks to a standby node without loss when detecting overheating of the physical GPU; constructing an energy consumption-aware scheduler to dynamically adjust the core frequency according to the SM unit utilization rate of the vGPU instance and achieve a 30% power reduction with a performance loss of <5%.

[0027] As a preferred embodiment, the resource matching in step S4 further includes: constructing a multi-dimensional topological feature map, integrating the number of NVLink connections between nodes, the PCIe Switch hierarchy, and the RDMA network latency data; developing a graph neural network matching model, abstracting task requirements into feature subgraphs, and searching for connected subgraphs that meet the constraints in the topological map; implementing transmission cost estimation, calculating the bandwidth-delay product of the data transfer path for candidate node combinations, and selecting the topological path with the minimum cost; deploying a prefetch optimization module to synchronously schedule relevant data sets to the HBM video memory of the target node during task allocation; implementing NUMA-aware memory binding to lock the task process to the CPU memory domain directly connected to the target GPU.

[0028] As a preferred embodiment, the full-link monitoring in step S7 further includes: deploying an LSTM fault prediction model to analyze the historical SM unit error rate, temperature gradient changes, and the number of ECC corrections, and predicting hardware failures 30 minutes in advance; establishing a hierarchical alarm mechanism to define yellow warnings, orange warnings, and red warnings; implementing an active self-healing strategy to start driver reloading and video memory remapping for yellow warning nodes, trigger service migration for orange warnings, and isolate the entire fault domain in case of red warnings; designing a fault-tolerant training mode to automatically switch to redundant computing nodes to continue execution and reconstruct the computing context when irreparable errors are detected; generating a fault traceability map to correlate and analyze hardware logs, task characteristics, and environmental variables, locate the root cause, and update the health model.

[0029] After adopting the above technical solutions, the beneficial effects of the present invention are as follows: Through the collaborative design of heterogeneous resource registration and modeling, virtualized resource reconstruction, topology-aware matching, dynamic load scheduling, and full-link monitoring, the GPU computing power scheduling method significantly improves the resource utilization rate, task execution efficiency, and system reliability of heterogeneous computing power platforms. The heterogeneous resource registration module supports the mixed access of multiple types of GPUs, and realizes the unified management of cross-vendor devices through topology relationship modeling, solving the problem of hardware fragmentation; the virtualized resource reconstruction technology ensures the security and stability of multi-task parallel execution through containerized video memory isolation and independent deployment of the driver stack. The improved Hungarian algorithm combined with the multi-dimensional resource scoring mechanism realizes the precise matching of task requirements and hardware topology, reducing scheduling delays caused by PCIe bandwidth bottlenecks or video memory fragmentation; the dynamic load balancing strategy effectively avoids local resource overload and improves the overall throughput of the cluster through real-time hot spot detection and task migration mechanisms. The runtime monitoring agent and MPS technology provide exclusive resource guarantees for critical tasks, and ensure the low-latency execution of high-priority tasks by dynamically adjusting the CUDA stream priority and video memory page locking strategy. The three-dimensional monitoring system covers the resource layer, task layer, and node layer, and realizes early warning and automatic isolation of faults through SM unit utilization analysis, video memory leakage tracking, and health scoring, reducing the risk of system downtime. The full-link diagnosis function quickly locates the causal relationship between hardware anomalies and task anomalies through cross-layer correlation analysis, improving the operation and maintenance efficiency. This method provides a highly elastic, highly reliable, and low-latency scheduling solution for heterogeneous GPU computing power platforms, and is applicable to various scenarios such as AI training, scientific computing, and edge inference.

[0030] Multi-core support improves resource utilization. The multi-core heterogeneous computing power platform with one cloud and multiple cores can adapt to and manage resources of multiple chip architectures (such as GPUs, CPUs, FPGAs, etc.), effectively supporting diversified computing power. This enables the platform to flexibly schedule different types of computing resources according to different business needs, thereby improving resource utilization and efficiency. Heterogeneous computing power works collaboratively. The platform realizes the collaborative work of heterogeneous computing power resources through a unified cloud management platform. Different types of chip architectures can work together in the same cloud computing environment, providing users with more powerful computing capabilities. Users do not need to concern themselves with the differences in underlying hardware architectures and can simply apply for, allocate, and use computing power through a unified interface or platform.

[0031] Computing power pooling: The one-cloud multi-core heterogeneous computing power platform integrates different types of computing power resources into a GPU computing power pool, achieving the standardization and pooling management of computing power. This enables more efficient utilization of computing power resources, avoiding idle and wasted resources. Intelligent scheduling algorithm: The platform adopts an advanced intelligent scheduling algorithm that can dynamically adjust the computing power allocation according to business requirements and resource status. This allows the computing power resources to be intelligently allocated based on factors such as the priority and computing requirements of different tasks, improving resource utilization and computing efficiency.

[0032] Reducing hardware investment costs: The one-cloud multi-core heterogeneous computing power platform can be compatible with computing resources of different architectures, avoiding waste of hardware investment caused by limitations of a single chip architecture. Users can choose the most suitable chip architecture according to actual needs, thus reducing hardware investment costs. Optimizing computing power costs: Through the intelligent scheduling algorithm, the platform can dynamically adjust computing resources according to business requirements, avoiding idle and wasted resources. This enables users to utilize computing power resources more efficiently and reduces computing power costs.

[0033] Improving computing efficiency: The one-cloud multi-core heterogeneous computing power platform can make full use of different types of computing capabilities by optimizing the computing power scheduling strategy. When processing large-scale image data, using a GPU can complete the processing task faster than using a CPU, thus improving computing efficiency. Shortening computing time: Through intelligent scheduling and load balancing technologies, the platform can allocate user tasks to the most suitable computing resources, thereby shortening the computing time. This enables faster obtaining of computing results and improves work efficiency.

[0034] Flexible expansion: The one-cloud multi-core heterogeneous computing power platform supports the flexible expansion of computing resources as the business grows and demands change. Users can dynamically increase or decrease computing resources according to actual needs, thus meeting the ever-changing business requirements. Multi-architecture compatibility: The platform can adapt to multiple chip architectures and operating systems, providing users with a wider range of choices. This allows users to more flexibly select the most suitable computing resources and software environment.

[0035] Intelligent monitoring and management: The one-cloud multi-core heterogeneous computing power platform provides intelligent monitoring and management functions, which can monitor the status and usage of computing resources in real time. This enables users to promptly discover and solve potential problems, ensuring the stable operation of computing resources. Simplifying the operation and maintenance process: Through a unified cloud management platform, the platform simplifies the operation and maintenance process. Users can manage resources, monitor, and troubleshoot through a unified portal, reducing the operation and maintenance difficulty and cost. In addition, the platform also provides rich operation and maintenance tools and documents to facilitate users' daily operation and maintenance and troubleshooting. Description of the Drawings

[0036] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the accompanying drawings required for the description of the embodiments or the prior art. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0037] Figure 1 It is the system structure diagram in the embodiment of the present invention;

[0038] Figure 2 It is the flowchart for the present invention to achieve GPU computing power scheduling;

[0039] Figure 3 It is the connection relationship diagram between the components of the system in the embodiment of the present invention;

[0040] Figure 4 It is the system diagram of the new generation cloud platform for the energy sales Internet application service in the second embodiment of the present invention. Specific implementation manners

[0041] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present invention.

[0042] Embodiment:

[0043] As Figures 1 to 3 shown, the working process of a GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform is a complex and delicate system engineering, which involves multiple links such as the registration and access of hardware resources, the submission and allocation of tasks, the dynamic management of GPU resources, and the task execution and result feedback. Through reasonable steps such as hardware abstraction and integration, task submission and allocation, dynamic management of GPU resources, and task execution and result feedback, the cloud platform can make full use of the performance advantages of heterogeneous computing resources such as GPUs, improve computing efficiency, reduce power consumption, and accelerate task execution.

[0044] 1. Resource registration: First, through the resource registration module, register the resource information of various different types of processors (GPUs), storage devices, memories, etc., and support the access of a variety of domestic chip resources. These hardware resources together form the basis of the heterogeneous computing environment. So that the platform can, through this module, identify and discover available resource information, and obtain the computing power topology information and status information in real time to ensure the effective utilization of GPU resources.

[0045] 2. Resource Modeling: Model the registered GPU resources on the cloud platform, including key information such as the GPU model, performance, and video memory capacity. This information will serve as the basis for scheduling decisions.

[0046] 3. Resource Reconfiguration: Through virtualization and pooling technologies, the cloud platform can achieve dynamic reconfiguration of GPU resources, abstracting physical hardware resources into virtual resources to meet the needs of different tasks, enabling resource pooling and flexible scheduling. At this layer, hardware resources such as GPUs are virtualized into virtual GPUs (vGPUs), dividing GPU resources into multiple independent virtual environments for sharing and scheduling among multiple virtual machines or containers, realizing unified management and scheduling on the cloud platform. When a task requires more GPU resources, the cloud platform can dynamically allocate more GPU resources from the resource pool to this task.

[0047] 4. Resource Isolation: When multiple tasks share GPU resources, containerization technology is used to achieve resource isolation, which can prevent interference and conflicts between different tasks to ensure the stability and security of tasks.

[0048] 5. Security Policy: Adopt container-level sandbox security technology, formulate and execute strict security policy tasks to ensure that GPU resources are not accessed or misused without authorization. The platform can prevent unauthorized access and malicious attacks, protecting user data and privacy.

[0049] 6. Task Submission: Users create new tasks through the interfaces or tools provided by the cloud platform and submit the tasks to be executed (such as computing tasks, data processing tasks, etc.) to the cloud platform. Support for multiple task types and programming models, such as CUDA, OpenCL, etc., so that users can flexibly use GPU resources for accelerated computing.

[0050] 7. Task Analysis: Analyze the requirements of the task to be executed for GPU resources, including the amount of computation, memory requirements, data transfer speed, etc.

[0051] 8. Task Parsing: The cloud platform parses the submitted tasks to understand the specific requirements of the tasks, including the required computing resources, storage resources, etc.

[0052] 9. Resource Evaluation: The cloud platform evaluates the currently available GPU resources, including the number, performance, and utilization rate of GPUs, to determine which GPU resources can meet the needs of the current task.

[0053] 10. Resource Matching: According to the task requirements, filter out the GPU resources that meet the conditions from the resource pool. In the matching process, topology-aware scheduling technology can be considered to optimize data transfer speed and computing efficiency.

[0054] 11. Task Allocation: Based on the result of resource matching, the cloud platform reasonably allocates tasks to available GPUs for execution. The scheduling algorithm takes into account multiple factors such as task priority, GPU resource availability, task execution time, etc. to optimize the overall computing performance.

[0055] 12. Task Execution: Once a task is allocated to a GPU for execution, the GPU starts executing the task and stores the result on the specified storage device. The cloud platform monitors the execution progress and status of the task in real time.

[0056] 13. Computing Power Scheduling: It is achieved through intelligent scheduling of the load scheduling strategy according to the real-time requirements of task execution and the availability of GPU resources, dynamically adjusting the resource computing power allocation. The scheduling engine can perform scheduling based on different scheduling strategies (such as fairness, priority, performance, etc.) to meet the needs of different applications. It can automatically adjust resource allocation according to task requirements and system status, and reasonably allocate tasks to different cloud hosts to achieve efficient utilization and load balancing.

[0057] Its load scheduling strategy monitors the GPU load of each node in the cluster in real time. The scheduler can understand which nodes are currently idle or under low load, and which nodes are under high load or overloaded. Based on this information, the scheduler can make decisions to preferentially allocate new tasks or jobs to nodes with lower load to reduce task waiting time and improve the overall system performance.

[0058] Node Selection: Based on the processed load data, the scheduler evaluates the load of each node. It preferentially selects nodes with lower load and more sufficient resources to allocate new tasks.

[0059] Resource Allocation: According to the requirements of the task and the resource situation of the node, the scheduler determines the specific GPU resources allocated to each task. This may include specifying the number of GPUs, the size of video memory, etc.

[0060] Load Balancing: The load scheduling strategy also considers how to achieve load balancing among different nodes. By dynamically adjusting task allocation, it ensures that the load of each node remains at a relatively balanced level, thereby improving the overall performance and stability of the system.

[0061] 14. Resource Sharing: When multiple executing tasks need to access GPU resources simultaneously, GPU sharing scheduling technology can be adopted. It enables multiple containers or tasks to share the same GPU resources, thereby improving resource utilization and reducing computing costs.

[0062] 15. Task Feedback: When the task execution is completed, the cloud platform returns the task execution result to the user. At the same time, the cloud platform also updates the status information of the current GPU resources for reference in subsequent task scheduling.

[0063] 16. Task Monitoring: The cloud platform supports full-process observability, provides the observation capabilities in terms of resources and tasks, and can monitor the execution status and progress of tasks throughout the process, facilitating users to monitor tasks and manage resources.

[0064] 17. Resource Monitoring: Through real-time monitoring on the cloud platform, continuously track the usage of GPUs, including temperature, utilization rate, fault information, etc. Discover potential problems in a timely manner and make corresponding adjustments.

[0065] 18. Node Diagnosis: The cloud platform provides a one-key diagnosis function, which can diagnose the software and hardware configurations of nodes, etc., and quickly locate and solve faults.

[0066] 19. Computing Power Optimization: The platform can automatically adjust GPU configurations and parameters to adapt to different application requirements and load conditions. By adopting adaptive optimization algorithm technology, continuously optimize the scheduling algorithm and resource management mechanism, and conduct real-time analysis and optimization of GPU computing power to improve computing efficiency and performance, thereby enhancing the efficiency and accuracy of GPU computing power scheduling. The implementation is as follows:

[0067] Status Initialization: The algorithm also needs to initialize some state variables, such as the average gradient of parameters, the average value of momentum, etc. These state variables will be used for subsequent parameter updates and adaptive adjustments.

[0068] Loss Function Calculation: The algorithm needs to calculate the value of the loss function under the current parameters, which is usually completed by passing forward through the network and comparing the prediction results with the actual labels.

[0069] Gradient Calculation: Next, the algorithm uses the backpropagation algorithm to calculate the gradient of the loss function with respect to the parameters. These gradients will be used to guide the update direction of the parameters.

[0070] Adaptive Learning Rate Adjustment: This is the core step of the adaptive optimization algorithm. The algorithm dynamically adjusts the learning rate based on the current gradient, historical gradient information (such as the average value of gradients, cumulative sum, etc.) and preset parameters (such as learning rate decay factor, momentum factor, etc.).

[0071] The adopted Adam algorithm combines the idea of calculating the average value of gradients to dynamically adjust the learning rate and momentum optimization, and also considers the average value of gradients and momentum to dynamically adjust the learning rate.

[0072] The mathematical principle of its Adam algorithm is to adaptively adjust the learning rate of each parameter based on the first-order moment estimation and second-order moment estimation of gradients. The iterative formula of the algorithm is as follows:

[0073] (Calculate the first-order moment estimation)

[0074] (Calculation of second - order moment estimation)

[0075] $\hat{m}_t = m_t - \beta_1 t$ (Bias correction for first - order moment estimation)

[0076] $\hat{v}_t = v_t - \beta_2 t$ (Bias correction for second - order moment estimation)

[0077] $\theta_{t + 1}=\theta_t-\alpha\sqrt{v_t}$ (Update parameters)

[0078] Where $m_t$ and $v_t$ represent the first - order moment estimation and second - order moment estimation of the gradient in the $t$-th iteration respectively. $\hat{m}_t$ is the estimated value after bias correction for $m_t$ and $v_t$. $\alpha$ is the learning rate, $\beta_1$ and $\beta_2$ are momentum decay factors, and $\epsilon$ is a small constant added to avoid the denominator being zero.

[0079] Parameter update: Use the adjusted learning rate and the calculated gradient to update the parameters. This is usually done through simple gradient descent or its variants (such as momentum optimization, Nesterov accelerated gradient, etc.).

[0080] Performance evaluation: After each parameter update, the algorithm needs to evaluate the performance (such as loss function value, accuracy, etc.) under the current parameters. This can be done by evaluating on the validation set.

[0081] Iteration and termination: If the performance meets the preset stop conditions (such as reaching the maximum number of iterations, the loss function value is lower than a certain threshold, etc.), the algorithm terminates; otherwise, the algorithm will continue to iterate and repeat the above steps.

[0082] A multi - core heterogeneous computing power platform based on cloud usually constructs based on cloud computing technology, adopts a hierarchical architecture, including a resource layer, a security layer, a task management layer, a monitoring layer, and an application layer.

[0083] Main modules of the structure:

[0084] The resource layer is responsible for resource registration: providing registration for resource information of various different types of processors, storage devices, memory, etc. These hardware resources together form the basis of the heterogeneous computing environment. Resource modeling: providing modeling for key information of the registered GPU resources on the cloud platform. Resource reconstruction: Through virtualization and pooling technologies, the cloud platform can achieve dynamic reconstruction of GPU resources.

[0085] Resource evaluation: providing determination of which GPU resources can meet the current task requirements and evaluating the currently available GPU resources. Resource matching: screening out the GPU resources that meet the conditions from the resource pool according to the task requirements.

[0086] The security layer is responsible for security policies: adopting sandbox security technology at the container level, formulating and executing strict security policy tasks to ensure that GPU resources are not accessed or misused without authorization. Resource isolation: providing containerization technology to achieve resource isolation when multiple tasks share GPU resources.

[0087] The task layer is responsible for task submission: providing an interface or tool through the cloud platform for users to create new tasks and submit the tasks to be executed to the cloud platform. Task analysis: analyzing the GPU resource requirements of the tasks to be executed. Task parsing: the cloud platform parses the submitted tasks to understand the specific requirements of the tasks, including the required computing resources, storage resources, etc. Task allocation: based on the result of resource matching, the cloud platform reasonably allocates tasks to available GPUs for execution. Task execution: the GPU executes the tasks and stores the results on the specified storage device. Task feedback: when the task execution is completed, the cloud platform returns the task execution result to the user.

[0088] The monitoring layer is responsible for resource monitoring: providing real-time monitoring through the cloud platform to continuously track the usage of GPUs. Task monitoring: providing full-process monitoring of the execution status and progress of tasks. Node diagnosis: the cloud platform provides a one-key diagnosis function to diagnose the software and hardware configuration of nodes, etc. Computing power optimization: the platform can automatically adjust GPU configurations and parameters to adapt to different application requirements and load conditions.

[0089] The application layer is responsible for computing power scheduling: achieving intelligent scheduling through intelligent scheduling algorithms to dynamically adjust resource computing power allocation. Resource sharing: providing the ability for multiple containers or tasks to share the same GPU resource.

[0090] Such as Figure 3 The component composition and connection relationship are as follows:

[0091] (1) Resource layer and task layer: Tasks drive resource allocation. In a cloud computing environment, when a user submits a compute-intensive task, the system allocates corresponding computing resources (such as GPUs, memory, etc.) according to the computing requirements of the task. In the task layer, after a specific task is defined, the system allocates corresponding resources according to the task's requirements. These resources can be computing resources, storage resources, network resources, etc., depending on the type and goal of the task. Resources support task execution. The resources provided by the resource layer are the basis for the task layer to execute tasks. Without the support of resources, the tasks in the task layer cannot be effectively executed. In an information system, resources such as processors (GPUs), storage devices, and memory are crucial for completing work tasks. The performance and availability of these resources will directly affect the execution efficiency and quality of manufacturing tasks. Dynamic adjustment and coordination: During actual operation, the task layer may dynamically adjust the allocation and use of resources according to the execution situation and changing requirements of the tasks. At the same time, the resource layer also needs to coordinate different resources according to the status and usage of the resources to meet the task's requirements. This dynamic adjustment and coordination are the key to ensuring the effective cooperation between the task layer and the resource layer.

[0092] (2) Task layer and application layer: When implementing the dynamic allocation of computing power resources, in a cloud computing environment, the computing power scheduling system can dynamically adjust the number and configuration of virtual machines according to the load situation of the task layer to meet the performance requirements of the task layer. Computing power scheduling can dynamically allocate computing power resource sharing according to the actual needs of the task layer, avoiding the idle and waste of computing power resources and improving the utilization rate of resource sharing. Improvement of task execution efficiency: In a big data processing scenario, the computing power scheduling system can allocate tasks to multiple computing nodes for parallel execution, thus significantly improving the data processing speed. Through computing power scheduling, the task layer can obtain the required computing power resources faster, thereby shortening the task execution time and improving the execution efficiency. Enhancement of task reliability: In a distributed computing environment, the computing power scheduling system can replicate tasks to multiple nodes for execution to ensure that tasks can continue to execute and be completed in case of a single node failure. The computing power scheduling system can also provide fault recovery and fault tolerance mechanisms to ensure that the task layer can continue to execute in case of computing power resource failures.

[0093] (3) Resource Layer and Security Layer: The resource layer is the foundation for resource isolation security. The resources such as servers, computing power, storage, and networks provided by the resource layer are the basis for implementing resource isolation security. Without the support of these resources, resource isolation security cannot be achieved. Through technical means such as virtualization and containerization, the resource layer can divide physical resources into multiple independent logical resources, thus providing a basis for resource isolation security. Different architectures of hardware resources are adapted through drivers, and the hardware resources are virtualized through virtualization technology to form a virtual resource pool. Resource isolation security protects the security of the resource layer. By implementing measures such as access control and authentication, resource isolation security ensures that only authorized users or applications can access the resources of the resource layer. Such isolation measures can prevent resources from being illegally accessed or misused, thus protecting the security of the resource layer. The resource layer and resource isolation security work together to ensure the overall security of the system. The resource layer provides basic resources, while resource isolation security is responsible for ensuring the secure use of these resources. The two cooperate with each other to jointly address various security threats and ensure the overall security of the system. Resource isolation security improves the utilization rate of the resource layer. Through resource isolation, resources can be divided into multiple independent units, and each unit can independently allocate, manage, and use resources. This division can improve the utilization rate of resources and enable resources to better meet the needs of different users or applications.

[0094] (4) Resource Layer and Monitoring Layer: The resource layer is the data source for the monitoring layer. The monitoring layer needs to collect various data in the resource layer in real time, such as GPU usage rate, memory occupancy rate, disk I / O, etc., to evaluate the running status and performance of the system. These data are the basis for the monitoring layer to analyze and make decisions, and are also important references for system operation and maintenance personnel to troubleshoot faults and optimize performance. The monitoring layer ensures the stability and security of the resource layer. By monitoring the status and performance indicators of the resource layer in real time, the monitoring layer can promptly detect potential problems and anomalies, such as resource overload and performance bottlenecks. The monitoring layer can also provide an alarm and notification mechanism. When an anomaly occurs in the resource layer, it promptly notifies the operation and maintenance personnel for handling, thus ensuring the stability and security of the resource layer. The two work together to improve the overall performance of the system. The resource layer and the monitoring layer cooperate with each other to jointly improve the overall performance of the system. The resource layer provides basic resources, while the monitoring layer is responsible for monitoring and optimizing the use of these resources to ensure that the system can operate efficiently and stably. The monitoring layer provides optimization for the resource layer. By analyzing the data of the resource layer, the monitoring layer can discover the patterns and trends of resource usage and provide suggestions for optimizing the resource layer. For example, the monitoring layer can suggest increasing or decreasing resource capacity, adjusting resource allocation strategies, etc., to optimize the usage efficiency and performance of resources.

[0095] (5) Task layer and monitoring layer: Functional collaboration. The task layer is responsible for the specific task execution and management, while the monitoring layer is responsible for the real-time monitoring and early warning of the operation status of the task layer. The two cooperate with each other functionally to jointly ensure the smooth progress of the project. The monitoring layer provides data support for the decision-making of the task layer by collecting the operation data of the task layer in real time. When the monitoring layer discovers that a certain task is executed abnormally, it can issue an early warning to the task layer in time so that the task layer can take corresponding measures to correct it. Mutual influence. The operation status of the task layer directly affects the output of the monitoring layer. When problems occur in the task layer, the monitoring layer will issue an early warning in time, thus reminding the project manager to take corresponding measures to solve them. The accuracy and timeliness of the monitoring layer also have an important impact on the execution efficiency of the task layer. If there are errors or delays in the data collection and processing of the monitoring layer, it may lead to wrong decisions made by the task layer, thus affecting the overall progress and quality of the project. Information circulation. A good information circulation mechanism needs to be established between the task layer and the monitoring layer. On the one hand, the task layer needs to provide the real-time task execution data to the monitoring layer for data processing and analysis; on the other hand, the monitoring layer also needs to timely feedback the processed data and analysis results to the task layer so that the task layer can make corresponding decisions and adjustments based on this information.

[0096] (6) Application Layer and Monitoring Layer: Data Collection and Analysis in the Monitoring Layer. The monitoring layer is responsible for real-time collection of the operating data of computing resources, such as the usage of GPUs, memory, network bandwidth, etc., as well as key performance indicators such as response time and system load during task execution. These data are important inputs for computing power scheduling decisions, helping the scheduling system understand the current resource distribution and task load. During the intelligent decision-making of computing power scheduling in the application layer, the computing power scheduling system makes intelligent resource allocation and task scheduling based on the data provided by the monitoring layer. It dynamically adjusts the allocation of computing resources according to factors such as task priority, resource requirements, and real-time load to ensure the efficient operation of tasks and the maximization of resource utilization. The impact of the accuracy of the monitoring layer on computing power scheduling. If the data collection in the monitoring layer is inaccurate or there is a delay, then the computing power scheduling system may make decisions based on incorrect information, resulting in unreasonable resource allocation or low task execution efficiency. Therefore, the accuracy and real-time nature of the monitoring layer have an important impact on the effect of computing power scheduling. The requirements of the application layer's computing power scheduling optimization for the monitoring layer. With the continuous progress of computing power scheduling technology, its data requirements for the monitoring layer are also increasing. To manage resources more finely, the computing power scheduling system may need to monitor more types of resource metrics or more fine-grained data, which prompts the monitoring layer to continuously upgrade and optimize its data collection and processing capabilities. The monitoring layer provides real-time data to computing power scheduling. The monitoring layer provides accurate information support to the computing power scheduling system by real-time collecting and transmitting the operating data of computing resources. These data are the basis for the computing power scheduling system to make intelligent decisions. The application layer's computing power scheduling feeds back the scheduling results to the monitoring layer. After the computing power scheduling system completes resource allocation and task scheduling, it can feed back the scheduling results to the monitoring layer. The monitoring layer can further monitor the task execution situation and resource usage according to these feedback data, providing a reference for subsequent scheduling decisions.

[0097] Embodiment 2

[0098] As Figure 4 shown, specifically in the new generation cloud platform system for energy sales Internet application services, the present invention relates to a new generation cloud platform system for energy sales Internet application services, aiming to support heterogeneous management capabilities and be able to uniformly manage, monitor, and allocate GPU computing power resources of different architectures. A large energy company needs to process a large number of image recognition tasks, which require high-performance GPU computing power support. At the same time, the company also needs to process some general computing tasks, which can be completed by the CPU.

[0099] This platform can shield the hardware differences of different GPU architectures. By shielding the differences of underlying chips and devices, it decouples the application's dependence on the underlying hardware, forms a unified "cloud-based resource pool", and provides a stable operating environment and rich ecological services. Administrators can use the same set of systems to manage all resources and allocate heterogeneous resources to different tenants. Users only need to log in to one platform to apply for and use different architecture resources shared in the environment. It provides a service-oriented computing power model encapsulation to support the application of the cloud platform in different information technology innovation industries. The system provides general cloud computing services by uniformly managing and scheduling resources of different chip architectures, maximizes the utilization of resources of different chip architectures, and provides a stable operating environment and rich ecological services during the process of user business expansion and application innovation. This embodiment demonstrates the optimal configuration of the system and its working principle.

[0100] The platform architecture of the entire system is as follows:

[0101] ① Resource layer

[0102] It includes resource registration, resource modeling, resource reconstruction, resource evaluation, and resource matching capabilities, and provides access to various types of resources such as processors, storage devices, and memory. These resources are encapsulated into a standard computing power resource pool through virtualization technology for unified management and scheduling by the cloud platform layer. Together, these hardware resources form the basis of the heterogeneous computing environment.

[0103] ② Security layer

[0104] It includes security policies and resource isolation capabilities. When multiple tasks share GPU resources, containerization technology is used to achieve resource isolation. Container-level sandbox security technology is adopted to formulate and execute strict security policy tasks to ensure that GPU resources are not accessed or misused without authorization.

[0105] ③ Task layer

[0106] It includes task submission, task analysis, task parsing, task allocation, task execution, and task feedback capabilities, and is responsible for the creation, submission, execution, and monitoring of GPU tasks. Through this module, users can submit tasks that require GPU acceleration to the platform and monitor the execution status and progress of the tasks. The cloud platform parses the submitted tasks, provides results based on resource matching, and the cloud platform reasonably allocates the tasks to available GPUs for execution. After the task execution is completed, the cloud platform returns the task execution results to the user.

[0107] ④ Monitoring layer

[0108] It includes resource monitoring, task monitoring, node diagnosis, and computing power optimization capabilities, providing real-time monitoring through the cloud platform to continuously track the usage of GPUs. It also monitors the execution status and progress of tasks throughout the process. It provides a one-key diagnosis function that can diagnose the software and hardware configurations of nodes. The platform can automatically adjust GPU configurations and parameters to adapt to different application requirements and load conditions.

[0109] ⑤ Application layer

[0110] It includes computing power scheduling and resource sharing capabilities, as well as an interface for users to interact with the platform. Users can submit computing power requirements through methods such as Web interfaces and API interfaces. Based on application requirements and GPU resource status, it dynamically schedules GPU computing power. The scheduling engine can perform scheduling based on different scheduling policies (such as fairness, priority, performance, etc.) to meet the needs of different applications.

[0111] The platform functions are as follows: Unified management of heterogeneous resources: The platform can achieve unified management of computing resources with different chip architectures, including resource allocation, monitoring, scheduling, etc. Dynamic scheduling of computing power resources: According to the requirements of upper-layer applications, the platform can dynamically adjust the allocation of computing power resources to ensure the maximum utilization of resources. Cross-architecture application deployment: The platform supports cross-architecture application deployment. Users can run the same application on different chip architectures without additional adaptation work. Security isolation and protection: The platform adopts advanced security isolation technologies to ensure the isolation of data and applications between different tenants, and at the same time provides comprehensive security protection measures to ensure the data security of users.

[0112] When implementing specifically, the implementation steps are as follows: Resource preparation: On the cloud platform layer, prepare a computing power resource pool containing two chip architectures, CPU and GPU. Ensure that these resources have been encapsulated and standardized through virtualization technology.

[0113] Application deployment: Deploy image recognition tasks and general computing tasks to the corresponding chip architectures respectively. For image recognition tasks, select high-performance GPUs for computing; for general computing tasks, select CPUs for computing.

[0114] Resource scheduling: According to the requirements of tasks and the status of underlying resources, the resource scheduling layer performs intelligent scheduling. When image recognition tasks require more GPU computing power, the resource scheduling layer will automatically allocate more resources from the GPU resource pool to these tasks; when general computing tasks require more CPU computing power, it will also automatically allocate resources from the CPU resource pool.

[0115] Monitoring and optimization: Through the monitoring function provided by the application layer, view the execution status and results of tasks in real time. Based on the monitoring data, perform necessary optimizations and adjustments on the platform to ensure the maximum utilization of resources and the smooth completion of tasks.

[0116] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. A GPU computing power scheduling method based on a heterogeneous computing power platform with one cloud and multiple cores, characterized in that, The method includes the following steps: Step S1: Heterogeneous resource registration and modeling. Access hardware devices containing multiple types of GPUs through the resource registration module, collect data such as device models, video memory capacities, and performance index data, and construct a resource feature database containing topological relationships to support the hybrid access of chips; Step S2: Virtualized resource reconstruction. Use the hardware abstraction layer technology to pool physical GPUs into virtual GPU resources, and achieve video memory isolation and computing unit division through containerization technology. Each virtual GPU instance is equipped with an independent driver stack and a security sandbox; Step S3: Multi-modal task submission. Receive CUDA / OpenCL computing tasks submitted by users, parse task requirement parameters, including the number of computing cores, video memory occupancy, and data throughput thresholds, and generate a task descriptor containing a priority label; Step S4: Topology-aware resource matching. Based on the resource status graph updated in real time, use an improved Hungarian algorithm for multi-dimensional resource matching. When selecting computing nodes, synchronously evaluate the PCIe bandwidth, NVLink connection status, and video memory fragmentation degree; Step S5: Dynamic load scheduling. Deploy a distributed scheduling engine to collect data on the temperature, utilization rate, and task queue depth of the GPUs on the nodes in real time. When detecting local hotspots, trigger the load balancing strategy and adopt a combination mechanism of task migration and computing power reallocation; Step S6: Adaptive task execution. Inject a runtime monitoring agent into the GPU instance, dynamically adjust the priority of CUDA streams and the video memory page locking strategy, and use the MPS multi-process service technology to achieve exclusive resource guarantee for critical tasks; Step S7: Full-link monitoring and diagnosis. Construct a three-dimensional monitoring system: Monitor the utilization rate of SM units and the ECC error rate of GPU hardware through the resource layer; Track the computing progress and video memory leakage risk at the task layer; Implement a health score at the node layer and trigger an automatic isolation mechanism.

2. The GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform according to claim 1, wherein: After completing the construction of the three-dimensional monitoring system, perform steps S8 and S9; Step S8: Parameterized computing power optimization. Use an improved Adam optimizer to dynamically adjust the scheduling strategy parameters; calculate the first-order moment estimate and the second-order moment estimate of the load gradient of the computing cluster; perform deviation correction based on historical load fluctuations; update the task allocation weight coefficient according to the corrected momentum value; introduce Nesterov accelerated gradient to predict future load trends; Step S9: Closed-loop feedback update. Collect the actual resource consumption data during task execution, compare and analyze it with the preset requirement parameters, iteratively optimize the resource matching algorithm through a reinforcement learning model, and update the specification template library of virtual GPUs.

3. A GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform according to claim 1, characterized in that: The dynamic load scheduling in step S5 includes a scheduling strategy optimization engine based on deep reinforcement learning, and the specific implementation includes: Construct a state feature space, and collect the GPU utilization rate, video memory occupancy rate, PCIe link bandwidth utilization rate, and task queue waiting duration of the nodes to form a 16-dimensional state vector; Design a dual-channel Dueling DQN network architecture, where the value stream network evaluates the global scheduling benefit, and the advantage stream network calculates the local scheduling action advantage value; Define the reward function \(R = \alpha\times(1 - \text{load variance})+\beta\times\text{task completion rate}-\gamma\times\text{migration overhead}\), where \(\alpha\), \(\beta\), and \(\gamma\) are adjustable weight coefficients; Deploy an experience replay pool to store historical scheduling decision data, and adopt a prioritized sampling mechanism to screen high-value training samples; Implement an online policy update mechanism. When detecting changes in the cluster topology or sudden changes in the load pattern, trigger an immediate fine-tuning of the network parameters; Introduce an action mask technique to restrict invalid scheduling operations, ensuring that task migration only occurs between nodes with hardware compatibility.

4. A GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform according to claim 1, characterized in that: The virtualized resource reconstruction in step S2 further includes: Create a securely isolated vGPU instance. Use hardware-assisted SR-IOV technology to divide a physical GPU into multiple VFs, and bind each VF to an independent secure enclave; Implement dynamic video memory quota management, and complete the elastic expansion of vGPU video memory from 1GB to 24GB within 0.1 seconds according to task requirements; Deploy a driver-level protection module. Intercept abnormal CUDA API requests through system call hijacking technology to block unauthorized memory access operations; Design a live migration channel. When detecting overheating of a physical GPU, migrate the vGPU instance along with the running tasks to a standby node without loss; Build an energy consumption-aware scheduler. Dynamically adjust the core frequency according to the SM unit utilization rate of the vGPU instance, and achieve a 30% power reduction when the performance loss is <5%; 5. A GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform according to claim 1, characterized in that: The resource matching in step S4 further includes: Build a multi-dimensional topology feature map, integrating data on the number of NVLink connections between nodes, the PCIe Switch hierarchy, and the RDMA network latency; Develop a graph neural network matching model. Abstract task requirements into feature subgraphs and search for connected subgraphs that meet the constraints in the topology map; Implement transmission cost estimation. Calculate the bandwidth-delay product of the data transfer path for candidate node combinations, and select the topology path with the minimum cost; Deploy a prefetch optimization module. Synchronously schedule relevant data sets to the HBM video memory of the target node during task allocation; Implement NUMA-aware memory binding, and lock the task process to the CPU memory domain directly connected to the target GPU.

6. A GPU computing power scheduling method based on a one-cloud multi-core heterogeneous computing power platform according to claim 1, characterized in that: The full-link monitoring in step S7 further includes: Deploy an LSTM fault prediction model. Analyze historical SM unit error rates, temperature gradient changes, and the number of ECC corrections to predict hardware failures 30 minutes in advance; Establish a hierarchical alarm mechanism, defining yellow alerts, orange alerts, and red alerts; Implement an active self-healing strategy. Start driver reloading and video memory remapping for yellow-alert nodes, trigger service migration for orange alerts, and isolate the entire faulty domain during red alerts; Design a fault-tolerant training mode. When detecting irreparable errors, automatically switch to redundant computing nodes to continue execution and reconstruct the computing context; Generate a fault tracing map, correlatively analyze hardware logs, task characteristics, and environmental variables to locate the root cause and update the health model.

Citation Information

Patent Citations

  • Multi-source computing power data integration and intelligent scheduling system and method

    CN118916147A

Cited By

  • Algorithm deployment method of heterogeneous cluster and storage medium

    CN120596275A

  • Cooperative scheduling and optimization method for heterogeneous resources of algorithm training platform

    CN120872535A

  • Heterogeneous AI computing virtualization architecture and slice enhancement method and system

    CN121118038A

  • Heterogeneous ai compute virtualization architecture and slice enhancement methods and systems

    CN121118038B

  • GPU virtualization system based on CUDA cross-level translation and multi-pooling scheduling

    CN121143952A