AI intelligent management system based on multi-cluster heterogeneous computing power scheduling
Through the AI intelligent management system and adaptive scheduling mechanism, the problems of uneven resource allocation and cross-platform scheduling complexity in multi-cluster heterogeneous computing power scheduling are solved, realizing efficient scheduling and optimized allocation of heterogeneous computing resources, and improving task processing efficiency and resource utilization.
Patent Information
- Application Number
- CN202511319902.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-16
- Publication Date
- 2026-02-03
AI Technical Summary
Existing technologies lack the ability to dynamically and adaptively schedule heterogeneous computing power across multiple clusters, resulting in low resource utilization, increased task latency, and cross-platform collaboration complexity, making it difficult to operate efficiently under large-scale tasks.
An AI-powered intelligent management system based on multi-cluster heterogeneous computing power scheduling is adopted. The system generates a set of computing resource feature vectors through a feature recognition unit. Combined with an adaptive scheduling mechanism and deep learning algorithms, it achieves efficient scheduling and optimized allocation of heterogeneous computing resources. The system includes a computing power scheduling module, a collaborative scheduling module, an edge-cloud scheduling module, and a heterogeneous resource management module. The intelligent optimization module is used to optimize resource strategies.
It significantly improves task processing efficiency and computing resource utilization, dynamically adjusts resource allocation, ensures efficient operation of edge nodes and cloud clusters, and improves overall system performance and response speed.
Smart Images

Figure CN121455656A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application relates to the technical field of artificial intelligence computing and resource management, and particularly relates to an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling. BACKGROUND
[0002] The rapid development of current artificial intelligence promotes the diversification and scaling of computing demand. Traditional computing platforms gradually shift from single servers to multi-cluster collaborative modes, and heterogeneous computing resources such as GPUs, FPGAs, TPUs and edge nodes jointly participate in operation. The existing technology schedules computing tasks through centralized management software and relies on static rules to allocate tasks and resources. The concept of distributed resource pool and edge computing is introduced to improve the real-time performance of cross-platform adaptation and computing. In deep learning and large-scale data processing tasks, the existing computing power scheduling platform can provide basic resource coordination functions to meet the computing demand within a certain range and is widely used in scientific research institutions and enterprise applications.
[0003] However, the existing technology still has obvious deficiencies, and lacks dynamic self-adaptive scheduling capability for multi-cluster heterogeneous computing power. The traditional platform often relies on fixed adaptation rules when processing different types of computing resources, and it is difficult to flexibly schedule according to the real-time load, complexity and priority of the task. When multiple clusters run training and inference tasks simultaneously, the problem of low resource utilization and increased task delay occurs, and the complexity of cross-platform collaboration is further amplified. The static and single scheduling mechanism cannot fully exert the performance potential of heterogeneous computing power, limiting the efficient operation and expansion of the AI system under large-scale tasks. SUMMARY
[0004] In view of the deficiencies of the prior art, the application provides an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling, which solves the problems of uneven distribution of heterogeneous computing resources and complexity of cross-platform scheduling through the AI intelligent management system and the self-adaptive scheduling mechanism.
[0005] To achieve the above purpose, the application is implemented by the following technical scheme: an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling, comprising: The computing power scheduling module comprises a feature recognition unit and a scheme generation unit. The feature recognition unit accesses the heterogeneous computing resources to perform feature recognition and generate a computing resource feature vector set. The feature recognition adopts a deep learning algorithm. The scheme generation unit generates a computing power allocation scheme based on the computing resource feature vector set. The computing process adopts a self-adaptive scheduling mechanism combined with task scheduling parameters. A cooperative scheduling module allocates the heterogeneous computing resources based on the computing power allocation scheme to form a plurality of heterogeneous computing clusters, integrates the plurality of heterogeneous computing clusters by a distributed resource pool to generate a unified resource pool, and performs cross-platform adaptation based on the unified resource pool to generate a cross-domain scheduling result; An edge-cloud scheduling module includes an edge node and a cloud cluster, performs task switching between the edge node and the cloud cluster based on the cross-domain scheduling result, and generates an edge-cloud scheduling result by an intelligent decision mechanism; A heterogeneous resource management module parses the edge-cloud scheduling result into task demand parameters, establishes a mapping relationship between the task demand parameters and performance parameters of the heterogeneous computing resources to generate a hardware performance mapping relationship, and performs heterogeneous resource allocation according to the hardware performance mapping relationship to generate a heterogeneous resource strategy; An intelligent optimization module analyzes and compares the heterogeneous resource strategy by an AI self-adaptive decision algorithm to generate an optimization instruction, corrects the computing power allocation scheme, the cross-domain scheduling result, and the edge-cloud scheduling result based on the optimization instruction, and outputs an optimized intelligent scheduling result.
[0006] Preferably, the heterogeneous computing resources are formed by GPU, FPGA, TPU, and CPU hardware access and abstract representation, and the task scheduling parameters include task complexity, priority, and real-time load.
[0007] Preferably, the model formula of the deep learning algorithm is: .
[0008] wherein, is the scheduling input vector of the i-th task, is a dimensionless index, is the corresponding task target scheduling result, is a dimensionless index, is the predicted output, is a dimensionless index, is the model parameter, is a dimensionless index, is the regularization coefficient, is a dimensionless index, is the overall loss function value, is a dimensionless index.
[0009] Preferably, the model formula of the self-adaptive scheduling mechanism is: .
[0010] wherein, is the computing power allocation weight, is a dimensionless index, is the normalized task complexity, is a dimensionless index, is a normalized task priority, a dimensionless index, is a normalized real-time load, a dimensionless index, is a scheduling coefficient, a dimensionless index, satisfying .
[0011] Preferably, the plurality of heterogeneous computing clusters are composed of the heterogeneous computing resources at a physical deployment layer, the distributed resource pool performs normalization processing and resource indexing on the plurality of heterogeneous computing clusters to generate a unified resource pool with comparability and schedulability, and the cross-platform adaptation includes cross-platform mapping and hardware difference adaptation.
[0012] Preferably, the edge node performs inference tasks during the task switching process, the cloud cluster performs training tasks during the task switching process, and the intelligent decision mechanism dynamically determines the task scheduling direction according to the load condition of the edge node and the high-throughput condition of the cloud cluster when performing the task switching, and configures low-delay inference tasks to the edge node and high-throughput training tasks to the cloud cluster.
[0013] Preferably, the task demand parameters include a computing demand amount, a storage demand amount and a bandwidth demand amount, the performance parameters include a computing core number, a storage bandwidth, a network throughput rate and a delay index, and the mapping relationship is established by using a normalization matrix mapping.
[0014] Preferably, a model formula of the AI adaptive decision algorithm is as follows: .
[0015] wherein, is a scheduling optimization objective function value, a dimensionless index, is a task number, a dimensionless index, is a normalized weight coefficient of an i th task, a dimensionless index, is a normalized complexity parameter of an i th task, a dimensionless index, is a normalized priority parameter of an i th task, a dimensionless index, is a normalized real-time load parameter of an i th task, a dimensionless index, is a scheduling coefficient, satisfying . .
[0016] The application provides an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling. The application has the following beneficial effects: This AI intelligent management system, based on multi-cluster heterogeneous computing power scheduling, achieves efficient scheduling and optimized allocation of heterogeneous computing resources through feature recognition using deep learning algorithms, cross-domain scheduling, and edge-cloud collaborative scheduling technologies. The feature recognition unit generates a set of computing resource feature vectors and optimizes the allocation of computing resources by combining them with task scheduling parameters through an adaptive scheduling mechanism, forming a cross-platform adapted scheduling scheme, thereby significantly improving task processing efficiency and computing resource utilization.
[0017] An adaptive scheduling mechanism and intelligent optimization module are employed, using AI adaptive decision-making algorithms to optimize resource allocation, making the task scheduling process more flexible and efficient. This mechanism can dynamically adjust the allocation of computing resources and switch tasks according to real-time task requirements and load conditions, ensuring that edge nodes and cloud clusters operate efficiently on their respective task types, thereby significantly improving the overall performance and response speed of the system. Attached Figure Description
[0018] Figure 1 This is a schematic diagram of the system structure of the present invention; Figure 2 This is an internal flowchart of the algorithm scheduling module of the present invention; Figure 3 This is a schematic diagram of task allocation for the edge scheduling module of the present invention. Detailed Implementation
[0019] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention. Example 1
[0020] like Figures 1-3 As shown, this embodiment of the invention provides an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling, including a computing power scheduling module, comprising a feature recognition unit and a scheme generation unit. The feature recognition unit accesses heterogeneous computing resources to perform feature recognition and generate a set of computing resource feature vectors. Feature recognition employs a deep learning algorithm. The scheme generation unit performs calculations based on the computing resource feature vector set to generate a computing power allocation scheme. The calculation process employs an adaptive scheduling mechanism combined with task scheduling parameters. Heterogeneous computing resources are formed by the hardware access and abstract representation of GPUs, FPGAs, TPUs, and CPUs. Task scheduling parameters include task complexity, priority, and real-time load. The model formula of the deep learning algorithm is: .
[0021] in, For the first The scheduling input vector for each task is a dimensionless index. The scheduling result for the corresponding task objective is a dimensionless indicator. For predictive output, the indicator is dimensionless. Here are the model parameters, and here are the dimensionless indices. Here, is the regularization coefficient, and is a dimensionless index. is the overall loss function value, and is a dimensionless index.
[0022] The model formula for the adaptive scheduling mechanism is: .
[0023] in, Weights are assigned to computing power; these are dimensionless indicators. Let be the normalized task complexity, and be a dimensionless index. The normalized task priority is represented by a dimensionless index. The normalized real-time load is a dimensionless metric. Let be the scheduling coefficient, and be a dimensionless index, satisfying . .
[0024] The collaborative scheduling module allocates heterogeneous computing resources based on a computing power allocation scheme to form multiple heterogeneous computing clusters. It then integrates these clusters into a unified resource pool via a distributed resource pool. Finally, based on this unified resource pool, the module performs cross-platform adaptation to generate cross-domain scheduling results. At the physical deployment layer, these heterogeneous computing clusters consist of heterogeneous computing resources. The distributed resource pool performs normalization and resource indexing on these clusters to generate a unified resource pool that is comparable and schedulable. Cross-platform adaptation includes cross-platform mapping and hardware difference adaptation.
[0025] The edge-cloud scheduling module, comprising edge nodes and a cloud cluster, performs task switching between edge nodes and the cloud cluster based on cross-domain scheduling results. The module generates edge-cloud scheduling results through an intelligent decision-making mechanism. Edge nodes execute inference tasks during task switching, while the cloud cluster executes training tasks. The intelligent decision-making mechanism dynamically determines the task scheduling direction based on the load of the edge nodes and the high throughput of the cloud cluster, allocating low-latency inference tasks to the edge nodes and high-throughput training tasks to the cloud cluster.
[0026] The heterogeneous resource management module parses the edge-cloud scheduling results into task requirement parameters. It then establishes a mapping relationship between these parameters and the performance parameters of heterogeneous computing resources, generating a hardware performance mapping relationship. Based on this mapping, the module executes heterogeneous resource allocation to generate heterogeneous resource strategies. Task requirement parameters include computational requirements, storage requirements, and bandwidth requirements. Performance parameters include the number of computing cores, storage bandwidth, network throughput, and latency. The mapping relationship is established using a normalized matrix mapping.
[0027] The intelligent optimization module analyzes and compares heterogeneous resource strategies using an AI adaptive decision-making algorithm to generate optimization instructions. Based on these instructions, the module corrects the computing power allocation scheme, cross-domain scheduling results, and edge-cloud scheduling results, and outputs the optimized intelligent scheduling result. The model formula for the AI adaptive decision-making algorithm is as follows: .
[0028] in, The objective function value for scheduling optimization is a dimensionless index. Let be the number of tasks, and be a dimensionless indicator. For the normalized first The weight coefficients for each task are dimensionless indicators. For the normalized first The complexity parameter of each task is a dimensionless index. For the normalized first The priority parameter for each task is a dimensionless indicator. For the normalized first The real-time load parameters for each task are dimensionless metrics. Let be the scheduling coefficient, satisfying .
[0029] The adaptive scheduling mechanism dynamically adjusts the computing power allocation scheme based on the complexity, priority, and real-time load status of the tasks. This mechanism can flexibly optimize resource allocation as tasks change, effectively reducing resource waste and improving computing efficiency.
[0030] The collaborative scheduling module integrates multiple heterogeneous computing clusters through a distributed resource pool to build a unified resource pool, significantly enhancing resource schedulability. Its cross-platform adaptability ensures efficient collaboration between different hardware platforms.
[0031] The edge-cloud scheduling module leverages an intelligent decision-making mechanism to dynamically determine task scheduling direction based on the load status of edge nodes and the high throughput characteristics of the cloud cluster. Low-latency inference tasks are preferentially allocated to edge nodes, while high-throughput training tasks are scheduled to the cloud cluster, thereby significantly optimizing overall system performance.
[0032] The heterogeneous resource management module establishes a mapping relationship between hardware performance indicators and task requirements, ensuring that computing, storage, and bandwidth requirements are effectively matched with hardware resources, thereby achieving precise resource allocation and improving system operating efficiency.
[0033] The intelligent optimization module employs an AI adaptive decision-making algorithm to optimize a given resource allocation scheme. Based on factors such as task weight, complexity, priority, and real-time load, this module continuously optimizes the scheduling strategy to achieve optimal resource utilization.
[0034] Example 2 This embodiment is an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling. By combining deep learning models and adaptive scheduling mechanisms, it achieves efficient task scheduling of heterogeneous computing resources, optimizes the allocation of computing resources, and improves task execution efficiency and resource utilization. The specific implementation method is as follows: 1. Deep learning models for task scheduling The core of task scheduling is using deep learning algorithms to predict the scheduling outcome based on task characteristics. The loss function of the scheduling model uses the following formula: The model formula for deep learning algorithms is: .
[0035] in, For the first The scheduling input vector for each task is a dimensionless index. The scheduling result for the corresponding task objective is a dimensionless indicator. For predictive output, the indicator is dimensionless. Here are the model parameters, and here are the dimensionless indices. Here, is the regularization coefficient, and is a dimensionless index. is the overall loss function value, and is a dimensionless index.
[0036] Data preparation: Assume the task is characterized by computational complexity, storage requirements, and bandwidth requirements, and the target scheduling result is the completion time.
[0037] The input features and actual scheduling results for each task are as follows, assuming there are a total of 3 tasks: Table 1: Task Data Table.
[0038] Task Computational complexity Storage requirements Bandwidth requirements Target scheduling result Task 1 0.7 0.8 0.9 0.9 Task 2 0.6 0.7 0.8 0.8 Task 3 0.9 0.7 0.7 1.0 Model training: Using a deep neural network to process the input feature vector Train to predict task scheduling results And through actual scheduling results The loss function is calculated based on the difference between the predicted and actual values. .
[0039] The calculation process of the loss function is as follows: .
[0040] Assuming that after training, the model's prediction results are: Task 1: .
[0041] Task 2: .
[0042] Task 3: .
[0043] If regularization term Then the loss function is calculated as follows: .
[0044] .
[0045] optimization: Update model parameters using gradient descent. Minimize the loss function to optimize the accuracy of task scheduling.
[0046] 2. Adaptive scheduling mechanism According to the adaptive scheduling mechanism, the computational power allocation weight of a task is calculated using the following formula: The model formula for the adaptive scheduling mechanism is: .
[0047] in, Weights are assigned to computing power; these are dimensionless indicators. Let be the normalized task complexity, and be a dimensionless index. The normalized task priority is represented by a dimensionless index. The normalized real-time load is a dimensionless metric. Let be the scheduling coefficient, and be a dimensionless index, satisfying . .
[0048] Task characteristics: The task complexity, priority, and real-time load are as follows: Table 2: Task Feature Data Table.
[0049]
[0050] Assume the scheduling coefficient is: , , .
[0051] Calculation of computing power allocation weight: For each task, the computing power allocation weight is calculated using a formula. : Task 1: .
[0052] Task 2: .
[0053] Task 3: .
[0054] Resource allocation: Assume the total computing resources are 100 units: Task 1: Highest weight, allocated approximately 37 units of resources.
[0055] Task 3: Second highest weight, allocated approximately 35 units of resources.
[0056] Task 2: Lowest weight, allocated approximately 28 units of resources.
[0057] Through the above steps, the deep learning model provides accurate predictions for task scheduling, ensuring the accuracy of the scheduling results, while the adaptive scheduling mechanism optimizes resource allocation in real time according to task requirements. Experimental results show that the system can achieve optimal task scheduling in a heterogeneous computing resource environment, improving overall computing performance, reducing latency, and optimizing resource utilization efficiency.
[0058] Example 3 This embodiment is an AI intelligent management system based on multi-cluster heterogeneous computing power scheduling. Through an intelligent edge-cloud collaborative scheduling mechanism, it optimizes the allocation of tasks between edge nodes and cloud clusters according to task requirements and system load, thereby improving resource utilization efficiency, reducing latency, and optimizing overall system performance. The specific implementation method is as follows: 1. Task Identification and Requirements Analysis The first step in the edge-cloud scheduling module is to identify tasks and analyze their requirements. Tasks can be inference tasks or training tasks, and the requirements for each type of task are different, so they need to be categorized based on the following factors: Computational requirements: The computational resource requirements of the task.
[0059] Storage requirements: The storage resource requirements of the task.
[0060] Bandwidth requirements: The network bandwidth requirements of the task.
[0061] Latency requirements: The task's response time requirements, which are especially critical for inference tasks.
[0062] Priority: The urgency of the task.
[0063] Task A is a reasoning task. Task B is a training task.
[0064] Task A: Computational requirements: 100 GFLOPS.
[0065] Storage requirement: 5GB.
[0066] Bandwidth requirement: 50GB / s.
[0067] Latency requirement: Low latency, less than or equal to 10ms.
[0068] Priority: High, the task requires a fast response.
[0069] Task B: Computational requirements: 1200 GFLOPS.
[0070] Storage requirement: 200GB.
[0071] Bandwidth requirement: 500GB / s.
[0072] Latency requirement: Not critical, higher latency is acceptable.
[0073] Priority: Low, primarily for calculation.
[0074] 2. Load monitoring of edge nodes and cloud clusters The second step in task scheduling is to monitor the load of edge nodes and cloud clusters in real time to determine the optimal scheduling location for tasks.
[0075] Edge node load status: CPU load: 80%.
[0076] Memory usage: 70%.
[0077] Network latency: 6ms, suitable for low-latency tasks.
[0078] Cloud cluster load status: CPU load: 30%.
[0079] Memory usage: 40%.
[0080] Network latency: 50ms, suitable for tasks requiring a large amount of computation.
[0081] 3. Intelligent decision-making mechanism The intelligent decision-making mechanism makes task scheduling decisions based on task requirements and current resource load.
[0082] Task A: Inference tasks have extremely high latency requirements and require rapid response. Although edge nodes have high load, their low latency (6ms) makes them suitable for executing inference tasks. Therefore, task A will be scheduled to the edge node.
[0083] Task B: The training task has very high computational resource requirements but low sensitivity to latency. The cloud cluster has low load and sufficient computational resources to handle the task. Therefore, task B will be scheduled to the cloud cluster.
[0084] 4. Task switching and execution The task switching process includes the actual allocation and execution of tasks. After the task scheduling decision is made, the task will be switched and executed according to the result of the decision.
[0085] Task A execution flow: Scheduled to edge nodes.
[0086] The edge node has a CPU load of 80% and a memory utilization of 70%.
[0087] Task A has a computational requirement of 100 GFLOPS, which can be met by edge nodes.
[0088] The edge node begins executing task A with a latency of 8ms, meeting the low latency requirement.
[0089] Task B execution flow: Dispatch to cloud cluster.
[0090] The CPU load of the cloud cluster is 30%, and the memory usage is 40%.
[0091] Task B requires 1200 GFLOPS of computing power, 200 GB of storage, and 500 GB / s of bandwidth. The cloud cluster is fully capable of handling this task.
[0092] The cloud cluster begins executing task B with an execution latency of 50ms and throughput that meets the task requirements.
[0093] 5. Task execution monitoring and adjustment The system will continuously monitor resource usage during task execution to ensure the task is completed smoothly as expected. Feedback data during task execution will be used for subsequent optimization decisions.
[0094] Task A monitoring data: CPU utilization: 75%.
[0095] Memory usage: 65%.
[0096] Network latency: 8ms, slightly increased, but still meets requirements.
[0097] Task B monitoring data: CPU utilization: 70%.
[0098] Memory usage: 60%.
[0099] Network bandwidth utilization: 480GB / s, close to maximum demand.
[0100] During task execution, the edge-cloud scheduling module adjusts resource allocation strategies based on load conditions. For example, if the load on edge nodes continues to rise and approaches 100%, some inference tasks may be transferred to the cloud cluster to avoid overloading the edge nodes.
[0101] 6. Task scheduling result output and optimization The scheduling module will output the scheduling results for each task, including: Task A: Scheduled to an edge node, with a latency of 8ms, meeting real-time requirements.
[0102] Task B: Scheduled to the cloud cluster, with an execution latency of 50ms, which meets the throughput requirements of the training task.
[0103] The system also continuously optimizes the scheduling process and collects feedback data to adjust scheduling strategies. For example, if edge nodes are overloaded and cannot meet low-latency requirements, the intelligent decision-making mechanism will prioritize allocating inference tasks to cloud clusters with lower loads.
[0104] Through the above steps, the edge-cloud scheduling module not only ensures low-latency response for inference tasks at edge nodes but also guarantees full utilization of high-throughput computing resources for training tasks in the cloud cluster. The intelligent decision-making mechanism enables the system to respond to load changes in real time, dynamically optimize task scheduling, improve resource utilization efficiency, and reduce latency.
[0105] Although embodiments of the invention have been shown and described, it will be understood by those skilled in the art that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the appended claims and their equivalents.
Claims
1. An AI intelligent management system based on multi-cluster heterogeneous computing power scheduling, characterized in that, include: The computing power scheduling module includes a feature recognition unit and a scheme generation unit. The feature recognition unit accesses heterogeneous computing resources to perform feature recognition and generate a set of computing resource feature vectors. The feature recognition adopts a deep learning algorithm. The scheme generation unit performs calculations based on the set of computing resource feature vectors to generate a computing power allocation scheme. The calculation process adopts an adaptive scheduling mechanism combined with task scheduling parameters. The collaborative scheduling module allocates heterogeneous computing resources based on the computing power allocation scheme to form multiple heterogeneous computing clusters. The collaborative scheduling module integrates the multiple heterogeneous computing clusters into a unified resource pool through a distributed resource pool. The collaborative scheduling module performs cross-platform adaptation based on the unified resource pool to generate cross-domain scheduling results. The edge-cloud scheduling module includes edge nodes and cloud clusters. Based on the cross-domain scheduling results, it performs task switching between the edge nodes and the cloud clusters. The edge-cloud scheduling module generates edge-cloud scheduling results by performing the task switching through an intelligent decision-making mechanism. The heterogeneous resource management module parses the edge-cloud scheduling result into task requirement parameters. The heterogeneous resource management module establishes a mapping relationship between the task requirement parameters and the performance parameters of the heterogeneous computing resources to generate a hardware performance mapping relationship. The heterogeneous resource management module executes heterogeneous resource allocation to generate a heterogeneous resource strategy based on the hardware performance mapping relationship. The intelligent optimization module analyzes and compares the heterogeneous resource strategy using an AI adaptive decision-making algorithm to generate optimization instructions. Based on the optimization instructions, the intelligent optimization module corrects the computing power allocation scheme, the cross-domain scheduling result, and the edge-cloud scheduling result, and outputs the optimized intelligent scheduling result.
2. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The heterogeneous computing resources are formed by hardware access and abstract representation of GPUs, FPGAs, TPUs and CPUs, and the task scheduling parameters include task complexity, priority and real-time load.
3. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The model formula for the deep learning algorithm is: , in, For the first The scheduling input vector for each task. The scheduling result is for the corresponding task objective. To predict the output, For model parameters, The regularization coefficient is . This represents the overall loss function value.
4. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The model formula for the adaptive scheduling mechanism is as follows: , in, Assign weights to computing power. The normalized task complexity, For normalized task priorities, This represents the normalized real-time load. Let be the scheduling coefficient, satisfying .
5. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The multiple heterogeneous computing clusters are composed of heterogeneous computing resources at the physical deployment layer. The distributed resource pool performs normalization processing and resource indexing on the multiple heterogeneous computing clusters. The cross-platform adaptation includes cross-platform mapping and hardware difference adaptation.
6. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The edge nodes perform inference tasks during the task switching process, and the cloud cluster performs training tasks during the task switching process. When performing the task switching, the intelligent decision-making mechanism dynamically determines the task scheduling direction based on the load of the edge nodes and the high throughput of the cloud cluster, and configures low-latency inference tasks to the edge nodes and high-throughput training tasks to the cloud cluster.
7. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The task requirement parameters include computational requirements, storage requirements, and bandwidth requirements. The performance parameters include the number of computing cores, storage bandwidth, network throughput, and latency. The mapping relationship is established using a normalized matrix mapping.
8. The AI intelligent management system based on multi-cluster heterogeneous computing power scheduling according to claim 1, characterized in that: The model formula for the AI adaptive decision-making algorithm is as follows: , in, To optimize the objective function value for scheduling, For the number of tasks, For the normalized first The weighting coefficients of each task. For the normalized first The complexity parameter of each task For the normalized first The priority parameter of each task. For the normalized first Real-time load parameters for each task. Let be the scheduling coefficient, satisfying .