Heterogeneous computing power resource dynamic cooperative scheduling method and system, and computer device

By acquiring multi-dimensional features and state information of heterogeneous computing environments, and using time series models for prediction and constructing multi-objective optimization functions, the resource scheduling problem in heterogeneous computing environments is solved, achieving efficient collaborative scheduling of tasks and reasonable allocation of resources, thereby improving system performance and resource utilization.

CN121300980APending Publication Date: 2026-01-09BEIJING ELECTRONIC DIGITAL INTELLIGENCE TECHNOLOGY CO LTD
View PDF 0 Cites 5 Cited by

Patent Information

Application Number
CN202511375130.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-25
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In heterogeneous computing environments, existing technologies suffer from insufficient multi-dimensional perception and fusion, lack of short-term prediction capabilities, failure of multi-objective dynamic trade-offs, and a lack of cross-level collaboration mechanisms, resulting in low resource utilization, rising task default rates, and uncontrolled operating costs.

Method used

By acquiring multi-dimensional feature information of the tasks to be scheduled, the current state of the heterogeneous computing power resource pool, and scheduling environment information, a multi-objective optimization function is constructed using a time series model for training. Based on a preset strategy, a scheduling instruction set is generated to achieve dynamic collaborative scheduling of tasks.

Benefits of technology

It improves resource utilization, reduces task latency and operating costs, and enhances system stability and reliability, making it suitable for high-demand scenarios such as artificial intelligence computing and edge real-time processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121300980A_ABST
    Figure CN121300980A_ABST
Patent Text Reader

Abstract

The invention discloses a heterogeneous computing power resource dynamic cooperative scheduling method and system and a computer device. The method comprises the steps of obtaining multi-dimensional feature information of a to-be-scheduled task, current state information of each node in a heterogeneous computing power resource pool and scheduling environment state information; training a time sequence model by adopting historical associated data, inputting the multi-dimensional feature information, the current state information and the scheduling environment state information into the trained model to obtain associated trend information in a future preset time window, and then combining the multi-dimensional feature information and the current state information of all the nodes to obtain the scheduling environment state information. A multi-objective optimization function fusing task delay, cost consumption and energy consumption is constructed, the function is optimized based on a preset strategy, a target decision strategy is obtained, a scheduling instruction set is generated after analysis, allocation information and target nodes are determined according to the scheduling instruction set, and tasks are issued. According to the method, dynamic collaborative scheduling of heterogeneous computing power resources can be realized, and the overall utilization efficiency of the resources is accurately, efficiently and effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of distributed computing resource management technology, and in particular to a method and system for dynamic collaborative scheduling of heterogeneous computing resources, and a computer device. Background Technology

[0002] With the rapid development of technologies such as artificial intelligence and the Industrial Internet of Things, global demand for computing power is experiencing explosive growth, leading to profound changes in the form of computing power supply. Traditional centralized cloud computing centers are struggling to meet the demands of low-latency and high-privacy scenarios. Edge computing nodes and smart terminal devices are being integrated into the computing power resource system, forming a three-tiered "cloud-edge-device" computing power network. However, this transformation brings complex challenges: heterogeneous hardware resources exhibit significant performance differences, task requirements are highly differentiated, and the dynamic nature of the environment exacerbates scheduling complexity.

[0003] Among the mainstream resource scheduling schemes publicly available in existing technologies, rule-based static schemes, represented by the Kubernetes default scheduler, mainly rely on basic indicators such as CPU / memory for decision-making. They fail to consider the differences in computing power between accelerators like GPUs and NPUs, nor do they take into account multi-dimensional factors such as network topology and energy costs, resulting in low hardware resource utilization. Furthermore, their preset weighting strategies (such as Binpacking or Spread strategies) lack dynamic adaptability, easily leading to scheduling imbalances during traffic surges. They are only applicable to a single resource domain and cannot achieve cross-layer collaboration between cloud, edge, and endpoint resources. In addition, heuristic methods based on historical data are insufficient for predicting unknown tasks and sudden traffic surges, have a single optimization objective, and traditional optimization techniques are inefficient. Edge computing-specific scheduling frameworks are limited to edge nodes, ignore environmental parameters, and their rescheduling mechanisms lack foresight, increasing system overhead.

[0004] In summary, resource scheduling in heterogeneous computing environments suffers from four core bottlenecks: insufficient multi-dimensional perception and integration, lack of short-term prediction capabilities, failure of dynamic trade-offs between multiple objectives, and absence of cross-level collaboration mechanisms. These bottlenecks lead to low utilization of computing resources, rising task default rates, and uncontrolled operating costs, urgently requiring systematic solutions to address these problems. Summary of the Invention

[0005] In view of this, the present disclosure provides a method and system for dynamic collaborative scheduling of heterogeneous computing resources, as well as a computer device, which can solve the problems existing in the prior art such as insufficient multi-dimensional perception and fusion, lack of short-term prediction capability, failure of multi-objective dynamic trade-off, and lack of cross-level collaborative mechanism in resource scheduling under heterogeneous computing environment.

[0006] In a first aspect, embodiments of this disclosure provide a method for dynamic collaborative scheduling of heterogeneous computing resources, including: Obtain multi-dimensional feature information of the task to be scheduled; The system can acquire the current status information of each node in the heterogeneous computing power resource pool in real time. The heterogeneous computing power resource pool includes cloud computing centers, edge computing nodes, and terminal devices. Obtain scheduling environment status information; Historical correlation data is obtained based on historical task logs, and the time series model is trained using the historical correlation data. The multi-dimensional feature information, the current state information of all nodes, and the corresponding scheduling environment state information are input into the trained time series model to obtain the correlation trend information within the future preset time window. Based on the multi-dimensional feature information, the current state information of all nodes, the scheduling environment state information, and the correlation trend information, a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption is constructed. The multi-objective optimization function is optimized based on a preset strategy to obtain the target decision strategy; The target decision-making strategy is analyzed to generate a scheduling instruction set; The task allocation information and the corresponding target node are obtained according to the scheduling instruction set, and the task is sent to the corresponding target node.

[0007] Secondly, this disclosure also provides a computer device, which adopts the following technical solution: The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform any of the above-described heterogeneous computing resource dynamic collaborative scheduling methods.

[0008] Thirdly, embodiments of this disclosure also provide a computer-readable storage medium storing computer instructions for causing a computer to execute any of the above-described methods for dynamic collaborative scheduling of heterogeneous computing resources.

[0009] Fourthly, embodiments of this disclosure also provide a computer program product, including a computer program / instructions that, when executed by a processor, implement the steps of any of the methods described above.

[0010] The heterogeneous computing resource dynamic collaborative scheduling method disclosed in this application first obtains multi-dimensional feature information of the task to be scheduled, current status information of each node in the heterogeneous computing resource pool, and scheduling environment status information, which enables a comprehensive understanding of the task characteristics, the real-time status of each node in the resource pool, and the real-time network and energy environment. Second, it obtains historical correlation data based on historical task logs and uses this data to train a time series model. The multi-dimensional feature information, the current status information of all nodes, and the corresponding scheduling environment status information are input into the trained time series model to obtain correlation trend information within a future preset time window. By predicting future trends, task scheduling planning can be done in advance, avoiding scheduling only when resources are scarce, thus improving the foresight and effectiveness of scheduling. Then, based on the multi-dimensional feature information and the current status information of all nodes... Based on the current state information, scheduling environment state information, and related trend information, a multi-objective optimization function is constructed that integrates task latency, cost consumption, and energy consumption. This comprehensively considers multiple important factors to achieve a balance between task latency, cost consumption, and energy consumption in scheduling decisions, thereby optimizing overall performance. The multi-objective optimization function is then optimized based on a preset strategy to obtain the target decision strategy. Finally, the target decision strategy is parsed to generate a scheduling instruction set. Based on the scheduling instruction set, task allocation information and corresponding target nodes are obtained, and tasks are distributed to the corresponding target nodes. Through optimization algorithms, the optimal solution can be found among many possible task-resource allocation schemes, improving the accuracy and efficiency of scheduling. Simultaneously, the optimized decision is transformed into actually executable scheduling instructions, ensuring that tasks are accurately allocated to suitable nodes for execution, achieving efficient task scheduling. Attached Figure Description

[0011] To more clearly illustrate the technical solutions of the embodiments of this disclosure, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 A flowchart illustrating the dynamic collaborative scheduling method for heterogeneous computing resources provided in this embodiment of the disclosure.

[0013] Figure 2 This is a flowchart illustrating a method for training a time series model according to an embodiment of the present disclosure.

[0014] Figure 3 This is a flowchart illustrating the method for constructing a multi-objective optimization function provided in an embodiment of this disclosure.

[0015] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of the present disclosure. Detailed Implementation

[0016] The embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0017] Reference Figure 1 This application discloses a dynamic collaborative scheduling method for heterogeneous computing resources, which is a dynamic collaborative scheduling method for heterogeneous computing resources based on multi-dimensional perception and prediction, including: S100: Obtain multi-dimensional feature information of the task to be scheduled.

[0018] The multi-dimensional feature information (i.e., task features) includes: the computation type identifier of the task to be scheduled, the estimated computation time, the delay constraint, the cost budget, and the data storage location.

[0019] S200 can obtain the current status information of each node in the heterogeneous computing power resource pool in real time.

[0020] The heterogeneous computing resource pool includes cloud computing centers, edge computing nodes, and terminal devices. Current status information (i.e., node resources) includes node type (cloud computing center node, edge computing node, or terminal device node), CPU / GPU / NPU / FPGA computing unit utilization, peak available computing power, current idle computing power, inter-node network latency, bandwidth utilization, cross-domain packet loss rate, node geographical location, real-time power consumption, unit-time computing cost, and predicted execution time of tasks on each node. Specifically, CPU / GPU / NPU / FPGA computing unit utilization, peak available computing power, and current idle computing power constitute hardware capabilities; inter-node network latency, bandwidth utilization, and cross-domain packet loss rate constitute network metrics; and node geographical location, real-time power consumption, unit-time computing cost, and predicted execution time of tasks on each node constitute dynamic attributes.

[0021] S300: Obtain scheduling environment status information.

[0022] Among them, the scheduling environment status information (i.e., environmental parameters) includes network topology (i.e., routing paths and real-time congestion of links between nodes), energy price of the region where the node is located (i.e., dynamic electricity price or fuel cost of the region where the node is located), and carbon emission intensity (i.e., carbon emission factor) corresponding to the current energy structure of the node.

[0023] S400 obtains historical correlation data based on historical task logs and uses the historical correlation data to train the time series model. It inputs multi-dimensional feature information, the current state information of all nodes and the corresponding scheduling environment state information into the trained time series model to obtain correlation trend information within a future preset time window.

[0024] Among them, the associated trend information includes task load trend information, resource status change trend information (i.e., resource availability), network status change trend information (task arrival rate, node load change trend, network latency fluctuation and predicted available computing power), and the predicted execution time of tasks on each node.

[0025] S500 constructs a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption based on multi-dimensional feature information, current status information of all nodes, scheduling environment status information, and correlation trend information.

[0026] S600 optimizes the multi-objective optimization function based on a preset strategy to obtain the objective decision strategy (i.e., the optimal task-resource allocation decision scheme).

[0027] The preset strategies include applying reinforcement learning, metaheuristic optimization algorithms, or deep learning models.

[0028] S700 analyzes the target decision-making strategy and generates a scheduling instruction set; Obtain task allocation information and corresponding target nodes according to the scheduling instruction set, and then send the tasks to the corresponding target nodes.

[0029] The heterogeneous computing resource dynamic collaborative scheduling method disclosed in this application integrates multi-dimensional real-time perception of task characteristics, resource status, and environmental parameters, and introduces a short-term prediction mechanism based on time series models to effectively improve the foresight and adaptability of scheduling and reduce task latency and resource waste. By constructing a multi-objective comprehensive optimization function that considers latency, default rate, cost, and energy consumption, and supporting flexible weight configuration, the method achieves a dynamic trade-off between service quality, cost control, and energy efficiency in the scheduling scheme. This method supports cross-level collaboration and unified scheduling of heterogeneous resources, enables collaborative management of multiple types of nodes in the cloud, edge, and terminal, supports cross-domain data migration and task distribution, breaks through the limitations of traditional scheduling systems, and significantly improves the overall resource utilization and service response capabilities of the system.

[0030] This method, by acquiring node status in real time, predicting resource trends, and optimizing task allocation, can rationally distribute tasks to nodes with available resources, avoiding resource idleness and waste. It comprehensively considers cost and energy consumption, selecting the optimal task-resource allocation scheme to effectively reduce computing resource usage costs and energy consumption costs. By combining task time requirements with resource usage, scheduling can reduce task execution latency and improve task processing efficiency while meeting time requirements. Simultaneously, it considers scheduling environment information, such as network and energy conditions, enabling scheduling decisions to adapt to different actual situations and improving system stability and reliability. This invention, through a collaborative multi-dimensional perception, prediction module, decision optimization, and coordinated scheduling execution scheme, effectively solves the problems of single-dimensional resource scheduling and lack of predictability and coordination in heterogeneous distributed computing environments. It effectively improves overall resource utilization efficiency, reduces task processing latency, and optimizes system operating costs and energy consumption, making it suitable for high-demand scenarios such as artificial intelligence computing and edge real-time processing.

[0031] Furthermore, the computation type is identified as CPU-intensive, GPU-intensive, I / O-intensive, or AI inference type. The specific methods for obtaining this information include: parsing the dependency library list of the task image, and identifying it as GPU-intensive if it contains CUDA / TensorRT; and detecting the task's input / output data ratio, and identifying it as I / O-intensive if the I / O throughput is greater than a preset threshold.

[0032] The estimated computation time is expressed as floating-point operations (FLOPS) or task execution duration; the latency constraint is the maximum tolerable completion time of the task; the cost budget is the upper limit of the acceptable resource rental cost of the task; and the data storage location is the storage node location of the task input data and the target location of the result output.

[0033] The methods for obtaining the peak available computing power in the current status information of each node include: GPU nodes: obtaining SM utilization and Tensor Core availability through the NVIDIA DCGM tool; NPU nodes: calling the chip manufacturer's SDK to read the depth of the idle queue of computing units.

[0034] Reference Figure 2 The S400 method of "obtaining historical correlation data based on historical task logs and using the historical correlation data to train a time series model," i.e., the method of training a time series model, includes: A100 retrieves historical related data based on historical task logs.

[0035] The historical data includes task computation type identifiers, submission timestamps, estimated computation time, latency constraints, data storage location information, resource usage records, historical node utilization time-series data, historical network latency values, and historical power consumption sequences. Furthermore, data from the past 72 hours is preferred.

[0036] Comprehensive historical correlation data encompasses information on tasks, resources, and the environment, reflecting various characteristics and patterns in past task execution. Acquiring this data provides rich material for training time series models, enabling them to learn the correlations between task execution and various factors, thus allowing for more accurate predictions of future trends.

[0037] A200 divides historical correlation data according to a preset ratio to obtain training and test sets.

[0038] Specifically, 80% of the historical correlation data can be used as the training set, and 20% of the historical correlation data can be used as the test set.

[0039] A300: Determine the time series model and train the time series model based on the training set.

[0040] The time series model can be an LSTM neural network or a Prophet time series model.

[0041] A400 uses a test set to analyze the trained model and dynamically adjusts the model's hyperparameters based on the analysis results until the optimized model meets the preset conditions, thus obtaining a trained time series model.

[0042] Commonly used analytical indicators include mean square error (MSE), root mean square error (RMSE), and mean absolute error (MAE).

[0043] By comprehensively acquiring historical correlation data, rationally dividing the training and test sets, selecting appropriate time series models, and optimizing model hyperparameters, the model can learn the true patterns in the data, thereby improving the accuracy of predicting correlation trends within a preset time window. Dividing the data into training and test sets, and using the test set for model evaluation and hyperparameter tuning, helps avoid overfitting, allowing the model to perform well even with new data and enhancing its generalization ability. Accurate correlation trend information is crucial for the dynamic collaborative scheduling of heterogeneous computing resources. Predictions based on well-trained time series models can help scheduling systems plan ahead, rationally allocate tasks and resources, improve scheduling efficiency and effectiveness, and ultimately enhance the overall system performance.

[0044] Furthermore, regarding the S400's "inputting multi-dimensional feature information, the current state information of all nodes, and the corresponding scheduling environment state information into the trained time series model to obtain the correlation trend information within the future preset time window," specifically, this includes: First, constructing a multi-dimensional time series input feature matrix, including: ① the time series sequence of task arrivals, categorized and statistically analyzed by task type; ② the time series of CPU / GPU / NPU / FPGA utilization rates for each node; ③ the time series of cross-node network latency and bandwidth utilization; ④ the time series of real-time power consumption and computational cost. Subsequently, a multivariate time series prediction method is adopted: for task arrival rate and computational requirements, an attention-enhanced LSTM network is used, with the historical 72-hour task arrival pattern as the training set to capture periodic and bursty features; for resource node state prediction, a GRU network based on a Seq2Seq structure is adopted, with the encoder inputting the node's historical load sequence and the decoder outputting predicted values ​​of load, available computing power, and network latency for multiple future time points. The model performs a prediction every 5 minutes, updating the input data through a sliding window mechanism to ensure prediction timeliness. The final output includes: the arrival distribution of each type of task within the future time window, the estimated interval of computational demand, the load change curve of each node, the probability distribution of available computing power, and the confidence interval of network latency. Specifically, it predicts the arrival rate and computational demand of each type of task, the load change trend of each resource node, the fluctuation of available computing power, and the change of network latency within the next 5-15 minute window; this prediction result serves as correlated trend information, providing forward-looking input for subsequent optimization decisions.

[0045] Reference Figure 3 The method for S500 that "constructs a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption based on multi-dimensional feature information, current state information of all nodes, scheduling environment state information, and correlation trend information," specifically includes the following: S510 obtains the comprehensive temporal and spatial cost index of the task based on the current status information and related trend information of each node.

[0046] S520 obtains the probability of delay constraint violation based on delay constraint and correlation trend information in multi-dimensional feature information.

[0047] The probability of delay constraint violation is : , For the task i The maximum tolerable completion time; the probability of delay constraint violation measures the degree of risk that the scheduling scheme cannot meet the task time limit requirements.

[0048] The probability of a delay constraint violation represents the likelihood that a task cannot be completed within a specified time. In some time-critical applications, such as industrial automation control and real-time monitoring systems, tasks must be completed within a specific timeframe, otherwise serious consequences may result.

[0049] S530 obtains the total task transmission cost based on multi-dimensional feature information, the current status information of each node, and the correlation trend information.

[0050] S540 obtains the carbon emission equivalent of task execution based on the current status information of each node and the scheduling environment status information.

[0051] S550 determines resource fragmentation penalties based on the current status information of each node.

[0052] Resource fragmentation penalty items are : ; 1≤j≤N, Let j be the computing power utilization rate of node j. The minimum computing power utilization rate is calculated; the resource fragmentation penalty term penalizes the uneven distribution of resources by calculating the lowest utilization rate among all nodes, aiming to avoid resource fragmentation and improve overall resource utilization efficiency.

[0053] Resource fragmentation refers to the scattered and discontinuous availability of resources in a system. This leads to low resource utilization and affects the overall performance of the system. For example, memory fragmentation makes it difficult for new tasks to find contiguous memory space, thereby increasing the time and complexity of memory allocation. A resource fragmentation penalty term is used to measure the degree of resource fragmentation and, through a penalty mechanism, encourages the optimization process to minimize the generation of resource fragmentation.

[0054] S560 constructs a multi-objective optimization function based on the task spatiotemporal integrated cost index, the probability of delay constraint violation, the total cost of task transmission, the carbon emission equivalent of task execution, and the resource fragmentation penalty term.

[0055] The multi-objective optimization function is F: F = α· + β· + γ· + δ· + ε· ; α+β+γ+δ+ε=1.

[0056] in, The comprehensive cost index of mission time and space. The total cost of task transmission. The carbon emission equivalent of the task is represented by α, β, γ, δ, and ε, which are the first, second, third, fourth, and fifth coefficients, respectively.

[0057] This multi-objective optimization function comprehensively considers average task latency, constraint default rate, resource leasing and transmission costs, and overall energy consumption. The optimization function supports dynamic weighted configuration, allowing users to flexibly set weights according to business priorities.

[0058] The first coefficient adjusts the importance of the task's combined time and space cost in the overall optimization function. A larger coefficient indicates a greater focus on the task's time and space costs during optimization, prioritizing solutions that reduce latency and geographical distance impacts. The second coefficient adjusts the weight of the probability of delay constraint violation in the optimization. A larger coefficient indicates a greater emphasis on timely task completion, minimizing the probability of delay constraint violations. The third coefficient controls the influence of the total task transmission cost in the optimization function. A larger coefficient suggests a greater preference for solutions with lower resource computation costs and cross-domain data transmission costs.

[0059] The fourth coefficient adjusts the importance of the carbon emission equivalent of task execution in the optimization process. When this coefficient increases, the system will place greater emphasis on selecting task execution schemes with lower carbon emissions to achieve energy conservation and emission reduction goals. The fifth coefficient adjusts the weight of the resource fragmentation penalty term in the optimization function. If the system has high requirements for resource continuity and utilization, this coefficient can be increased to encourage better management and utilization of resources during the optimization process.

[0060] The method disclosed in this embodiment considers multiple factors simultaneously in practical task scheduling and resource allocation problems, including time, cost, environmental impact, and resource utilization efficiency. By incorporating these factors into a multi-objective optimization function, the merits of task execution schemes can be evaluated more comprehensively, thereby finding a comprehensive optimal solution. Different application scenarios may place varying degrees of emphasis on each factor. By adjusting the values ​​of the five coefficients, the weights of each factor can be flexibly adjusted according to specific application scenarios and needs. For example, in scenarios with extremely high real-time requirements, the first and second coefficients can be increased; in scenarios that emphasize energy conservation and emission reduction, the fourth coefficient can be increased. By optimizing task scheduling and resource allocation, task latency, cost, and carbon emissions can be reduced, while resource utilization can be improved, thereby enhancing the overall performance and sustainability of the system. This multi-objective optimization method can achieve a win-win situation for both economic and environmental benefits while meeting business needs.

[0061] The method for S510 to "obtain the comprehensive spatiotemporal cost index of the task based on the current status information and correlation trend information of each node" specifically includes: S511, obtain the network latency between nodes and the predicted execution time of the task on each node from the current status information.

[0062] S512, obtain network state change trend information from the associated trend information.

[0063] S513 obtains the data transmission network latency of the task based on the network latency between nodes and the trend of network status changes.

[0064] S514: Based on the data transmission network latency of the task and the predicted execution time of the task on each node, obtain the predicted average latency of the task.

[0065] The average latency of the predicted task is : , For the task Predicted execution time at the target node For the task Data transmission network latency, The total number of nodes is used to predict the average delay of a task, which reflects the overall time cost from task issuance to completion.

[0066] The predicted average task latency reflects the average time delay from task issuance to completion. The longer the latency, the lower the task execution efficiency, which will affect the response speed and performance of the entire system.

[0067] S515: Obtain the data storage location from the multi-dimensional feature information and the node geographical location from the current state information, and obtain the geographical penalty factor.

[0068] Geographical penalty factor is : , Geographic distance (in kilometers) between the data storage location and the computing node for task i. The geographic penalty factor is used to penalize scheduling decisions where the data is too far from the computing power, in order to reduce data transmission overhead and improve response speed.

[0069] The geographical penalty factor takes into account the impact of geographical distance between the task data storage location and the computing node. The greater the geographical distance, the higher the latency, the greater the bandwidth requirement, and the possible network instability in data transmission, thereby increasing the cost of task execution.

[0070] S516, based on the predicted average mission delay and geographical penalty factor, obtains the mission spatiotemporal comprehensive cost index.

[0071] The mission's spatiotemporal comprehensive cost index is : .

[0072] The method for S530 to "obtain the total task transmission cost based on multi-dimensional feature information, the current status information of each node, and correlation trend information" includes: S531 obtains the node's geographical location, real-time power consumption, and unit time computation cost from the current status information of each node, as well as the predicted execution time of the task on each node from the associated trend information, and obtains the resource computation cost.

[0073] Resource computing cost is ; , Calculate the unit time cost for the node j where the task is located; the resource calculation cost is the cumulative calculation cost of all tasks on the selected resources.

[0074] Resource computing cost refers to the cost of computing resources consumed during task execution, such as the cost of using resources like CPU and memory. Different computing nodes may have different computing capabilities and costs. Rationally allocating tasks to reduce resource computing costs is an important optimization goal.

[0075] S532 determines the cross-domain data transmission cost based on the data storage location in the acquired multi-dimensional feature information, the bandwidth utilization in the current status information of each node, and the network topology in the scheduling environment status information.

[0076] The cost of cross-domain data transfer is : , The amount of data that task i needs to transmit. The cost per GB of data transfer from the source storage location to compute node j; cross-domain data transfer cost quantifies the additional costs incurred due to data migration.

[0077] Because data transmission between different network domains may require additional bandwidth and costs, this cost is known as cross-domain data transmission cost. For example, data transmission between different departments within an enterprise, or data interaction between different data centers, may involve cross-domain data transmission costs.

[0078] S533 determines the total cost of task transmission based on resource computing costs and cross-domain data transmission costs.

[0079] The total cost of task transmission is : .

[0080] The method for S540 to "obtain the carbon emission equivalent of task execution based on the current status information of each node" specifically includes: S541 determines the total energy consumption of task execution based on the real-time power consumption in the acquired current status information and the predicted execution time of the task on each node.

[0081] The total energy consumption for task execution is : , The real-time power consumption of node j; the total energy consumption for task execution is the total power consumed by all tasks.

[0082] S542 determines the carbon emission amplification factor based on the carbon emission intensity corresponding to the current energy structure of the node in the scheduling environment status information.

[0083] In this embodiment, the carbon emission amplification factor is: The carbon emission intensity (i.e., carbon emission factor) corresponding to the current energy structure of the node is directly adopted. This factor reflects the carbon emission intensity corresponding to the energy structure of the node's location and is used to convert energy consumption into carbon emission equivalent.

[0084] S543, determine the carbon emission equivalent of mission execution based on the total energy consumption and carbon emission amplification factor.

[0085] The carbon emission equivalent of the mission is ; , This is the carbon emission amplification factor.

[0086] The carbon emission equivalent of a mission reflects the amount of carbon emissions generated during the execution of a mission. With increasing global emphasis on environmental protection, reducing carbon emissions has become an important goal for many companies and organizations. In calculating the carbon emission equivalent, the energy consumption during mission execution generates carbon emissions, and the impact of a mission on the environment can be measured by calculating the carbon emission equivalent.

[0087] The S600 method of "optimizing a multi-objective optimization function based on a preset strategy to obtain an objective decision strategy" specifically includes: S610 constructs a state space based on multi-dimensional feature information, the current state information of all nodes, the state information of the scheduling environment, and the correlation trend information.

[0088] The S620 is configured with an action space, which is a multi-dimensional allocation decision, including the target node ID, data migration path, and computing power reservation ratio.

[0089] S630 configures the reward function based on the multi-objective optimization function.

[0090] The reward function is R: R = -F + R B ;R B =η·(1-Standard deviation of inter-node utilization),R B The reward is for resource balancing, and η is the weighting coefficient.

[0091] The methods for obtaining the standard deviation of inter-node utilization include: 1) obtaining the resource utilization information of each resource node. The resource utilization can be defined according to the specific resource type. For example, for computing resources, it is the CPU utilization; for storage resources, it is the disk utilization; for network resources, it is the network bandwidth utilization, etc.; 2) calculating the average value of resource utilization; 3) calculating the square of the difference between the resource utilization of each node and the average value, and then calculating the variance. The square root of the variance is the standard deviation.

[0092] In reinforcement learning, the reward function is usually expected to be as large as possible, while the multi-objective optimization function FF is the objective that needs to be minimized (e.g., reducing cost, reducing latency, etc.). Therefore, by taking the negative, the minimization problem is transformed into the problem of maximizing the reward, which prompts the agent to act in the direction of reducing the overall time and space cost of the task, the probability of delay constraint violation, the total cost of task transmission, the carbon emission equivalent of task execution, and resource fragmentation, thereby optimizing the performance of the entire system.

[0093] η is a weighting coefficient used to adjust the importance of resource equalization rewards in the overall reward function. It can be adjusted according to specific application scenarios and needs. A larger η indicates that the system prioritizes balanced resource allocation; a smaller η indicates a greater focus on optimizing the multi-objective optimization function F. By adjusting the value of η, a trade-off can be struck between minimizing the multi-objective optimization function F and improving resource equalization.

[0094] The standard deviation of inter-node utilization is used to measure the dispersion of a set of data. In this context, it reflects the dispersion of the utilization of each resource node. A large standard deviation indicates a large difference in utilization between different nodes, resulting in an unbalanced resource allocation; a small standard deviation indicates that the utilization of each node is relatively close, resulting in a more balanced resource allocation. The value of the standard deviation of inter-node utilization (1 - 1) ranges between [0, 1]. When the standard deviation of inter-node utilization is 0, that is, all nodes have the same utilization, and the resource allocation reaches the most balanced state. At this time, the standard deviation of inter-node utilization (1 - 1) is 1, and the resource balance reward R is 1. B To achieve the maximum value η; the larger the standard deviation of utilization among nodes, the smaller the standard deviation of utilization among nodes, and the greater the resource balance reward R. B The smaller the value, the better. Therefore, this metric incentivizes the agent to strive for a more balanced resource utilization across all nodes. Overall, the reward function in this application is -F + R. B It takes into account both the optimization of the multi-objective optimization function FF and the balanced allocation of resources, guiding the agent to find a suitable balance between the two in order to obtain the optimal task-resource mapping scheme.

[0095] S640 uses a reinforcement learning engine to iteratively learn based on the state space, action space, and reward function to find the task-resource mapping scheme that satisfies the optimal multi-objective optimization function and uses it as the objective decision strategy.

[0096] Furthermore, a prediction-guided genetic algorithm can be used for optimization. Specifically, 1) chromosome encoding is performed first, with the chromosome's gene segments being [task ID, target node ID, data migration flag]; 2) then innovative operators are designed, including: adaptive crossover probability, crossover probability P... cross = 0.7×(1-distance between nodes / maximum distance); Greedy mutation, prioritize replacing high GeoPenalty or high CarbonFactor node tasks; 3) Generate an initial population and inject the Top-K low-load node candidate set from the correlation trend information; Iterative evolution is carried out by a prediction-guided genetic algorithm based on chromosome encoding, innovation operators and the initial population to find the task-resource mapping scheme that makes the multi-objective optimization function optimal as the objective decision strategy.

[0097] Furthermore, the application also includes: continuously optimizing the prediction and decision models based on feedback data from the actual execution results of tasks. Specifically, the execution status of each task is first collected, including key indicators such as actual completion time, energy consumption, cost, and resource utilization. Then, a deviation analysis is performed between this data and the predicted values. If a certain type of model is found to have a prediction error greater than 15% in three consecutive scheduling attempts, a model update process is triggered. By driving the continuous optimization of the prediction model and scheduling strategy through execution result feedback, and combining online updates using reinforcement learning with adaptive evolutionary mechanisms using genetic algorithms, the system achieves rapid response and continuous learning to environmental changes.

[0098] The system supports two paths for policy solving: reinforcement learning and genetic algorithms. The reinforcement learning module uses a PPO policy network to construct task-resource mapping policies and introduces a course learning mechanism during the training phase, gradually expanding from a small-scale environment of 10 nodes to a scheduling scenario of 100 nodes. The genetic algorithm improves the local optimality and overall fitness of the solution through a location-aware crossover mechanism and a greedy mutation strategy.

[0099] Furthermore, the heterogeneous computing power resource dynamic collaborative scheduling method disclosed in this application also includes: when the data storage location of the task does not match the target node, determining whether the migration time and migration cost meet the preset conditions; if so, triggering cross-node data migration.

[0100] The determination of whether the migration time and migration cost meet the preset conditions includes: A100, the first condition is determined based on the delay constraints and estimated calculation time in the multi-dimensional feature information; The first condition includes: ,in, The migration takes time. For task delay constraints, To estimate the calculation time, This is for the safety factor.

[0101] This indicates the migration time, which is the time it takes to migrate data from one region to another. This represents the task delay constraint, which is the maximum allowed time from the start of the task to its mandatory completion. This indicates the estimated computation time, which is a preliminary estimate of the time required for the task to perform computational operations on the computing node.

[0102] A200, determine the second condition; the second condition includes: the migration cost is no more than 10% of the task cost budget.

[0103] A300 determines the matching result of migration time and migration cost based on the first and second conditions.

[0104] The method disclosed in this application consists of five closed-loop stages: perception, prediction, optimization, execution, and feedback. First, the task flow is monitored in real time through the task scheduler interface, and feature data required for scheduling is collected from multiple dimensions. This data includes not only the structural characteristics, computation type, and resource requirements of the task itself, but also external constraints such as node hardware status, network topology, environmental electricity price, and carbon emissions. The collected data can flow into the prediction module in real time through the data pipeline. The prediction module calls historical load and node behavior sequences, and uses time series models such as LSTM and Prophet to predict the trend of task load, resource availability, and network status within the future short window. The prediction results serve as an important reference input for optimized scheduling and are passed to the decision optimization module.

[0105] Next, the decision optimization module solves the scheduling problem. The inputs include perception data, prediction results, and historical execution feedback. This module can select a reinforcement learning path or an improved genetic algorithm path to generate a scheduling strategy based on the system configuration. The generated scheduling instructions will be synchronously transmitted to the collaborative scheduling execution module. This module issues scheduling instructions to cloud, edge, and terminal resource nodes based on the task mapping results, and is responsible for necessary data migration, container scheduling, and task orchestration. Specifically, it achieves cross-domain resource management through cloud management platform API, edge orchestrator, and terminal agent.

[0106] Furthermore, during execution, the running logs, resource usage, and performance metrics of each task are recorded and sent to the feedback learning module. The feedback module is responsible for evaluating the accuracy of the prediction and scheduling models and triggering updates to the relevant models when the error exceeds a threshold. After the model is updated, it is returned to the perception and prediction modules through an interface, thereby achieving closed-loop optimization of the entire scheduling system.

[0107] For multi-dimensional data acquisition, three parallel paths can be employed. The first path involves parsing container dependency libraries to identify whether a task is GPU-intensive, specifically by scanning the image for libraries such as CUDA, cuDNN, or TensorRT. The second path uses eBPF technology to monitor task I / O behavior; if I / O throughput exceeds a threshold, it is identified as an I / O-intensive task. Furthermore, resource status monitoring, as the third path, is deployed across all heterogeneous computing nodes. Its implementation includes NVIDIA DCGM collecting CPU / GPU status data and the heterogeneous chip SDK reading NPU computation queues, enabling precise monitoring of computing power utilization and idle computing units.

[0108] Resource awareness extends beyond hardware load sampling to include perception of network and environmental conditions. By subscribing to terminal devices via the MQTT protocol and actively monitoring network latency, packet loss, and other metrics using the Smokeping tool, a real-time network topology between nodes is constructed. Environmental data is acquired through two layers: firstly, environmental sensing systems collect on-site parameters such as temperature, voltage, and noise interference; secondly, regional electricity price fluctuation data is obtained through API calls to the State Grid Corporation of China, and time-aligned and quantified. This aggregated data is then fed into a data fusion module, where its feature structure is standardized and finally written into the RedisTSDB time-series database for subsequent prediction and optimization.

[0109] Once the scheduling result is generated, the scheduling instructions will be parsed and translated by the collaborative scheduling execution module into a configuration structure recognizable by the infrastructure platform. Finally, they will be distributed to the target node in the form of Terraform scripts or Kubernetes CRDs. Through a unified API interface, the system will maintain synchronous communication with the cloud management platform, edge orchestrator, and terminal agent to ensure that the scheduling instructions can be distributed in milliseconds across different computing power domains. In the event of a mismatch between the task and the computing node, the system will automatically determine whether to trigger a data migration operation and verify whether the migration time and cost meet the scheduling constraints.

[0110] The model updates are divided into two categories: the prediction model uses weighted mean square error (WMSE) based on error weights for backpropagation optimization; the reinforcement learning model updates the policy network parameters through an importance sampling strategy to enhance online learning capabilities. Once the model is updated, it automatically feeds back to the perception and prediction modules, forming a complete closed loop and improving the scheduling accuracy in the next round. Through the system structure, perception process, and feedback mechanism, intelligent, dynamic, and high-precision control of task scheduling can be achieved. It is particularly suitable for cloud-edge-device heterogeneous resource collaboration scenarios, effectively reducing overall system latency, improving resource utilization, and achieving fine-grained control over cost and energy consumption.

[0111] This implementation not only significantly improves scheduling efficiency, but also enables the system to dynamically respond to different scenarios' emphasis on latency, cost, energy consumption, or resource balance through multi-objective function configuration capabilities and strategy algorithm switching mechanisms, achieving true scenario-adaptive computing power orchestration.

[0112] The heterogeneous computing power resource dynamic collaborative scheduling method disclosed in this application comprehensively acquires multi-dimensional characteristics of the task to be scheduled, the current state of the node, and the status information of the scheduling environment. It realizes multi-dimensional perception of tasks, resources, and environment. By comprehensively considering this information, it can more comprehensively understand the task requirements and resource status, avoid the limitations of single-index decision-making, and thus improve the utilization rate of hardware resources. For example, when scheduling tasks, a suitable node can be selected based on the task's GPU requirements and the GPU performance of the node, rather than just considering CPU and memory.

[0113] This application utilizes historical correlation data to train a time series model, enabling it to learn the correlation and changing patterns between task execution and various factors. By inputting current multi-dimensional information, it can predict correlation trends within a preset time window, such as task execution time and resource usage. This allows the scheduling system to plan ahead and cope with sudden traffic surges and unknown tasks, improving the system's adaptability and stability. For example, in the event of a traffic surge, task allocation can be adjusted in advance based on predictive information to avoid scheduling imbalances.

[0114] This application constructs a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption. It can comprehensively consider multiple objectives and optimize the function through preset strategies. It can dynamically weigh different objectives to find the optimal scheduling scheme. For example, under the premise of meeting task delay requirements, it can minimize cost consumption and energy consumption and improve the overall efficiency of the system.

[0115] In this application, the heterogeneous computing resource pool includes cloud computing centers, edge computing nodes, and terminal devices. This application constructs a heterogeneous computing resource pool of a three-level "cloud-edge-device" computing network. By parsing the target decision strategy, a scheduling instruction set is generated, and tasks are distributed to the corresponding target nodes. This achieves cross-level collaborative scheduling of cloud, edge, and device resources. Based on the characteristics of the task and the availability of resources, tasks can be reasonably allocated among nodes at different levels, giving full play to the advantages of resources at each level and improving the utilization rate of computing resources. For example, tasks with low latency requirements can be allocated to edge computing nodes or terminal devices; while computationally intensive tasks can be allocated to cloud computing centers.

[0116] Secondly, this application discloses a heterogeneous computing resource dynamic collaborative scheduling system for executing the heterogeneous computing resource dynamic collaborative scheduling method disclosed in the first aspect of this application. The system includes: The task feature acquisition submodule is used to acquire multi-dimensional feature information of the task to be scheduled; The resource status monitoring submodule is used to obtain the current status information of each node in the heterogeneous computing power resource pool in real time. The heterogeneous computing power resource pool includes cloud computing centers, edge computing nodes and terminal devices. The environmental data acquisition submodule is used to acquire scheduling environment status information; The prediction module is used to obtain historical correlation data based on historical task logs and to train the time series model using the historical correlation data. It inputs multi-dimensional feature information, the current state information of all nodes and the corresponding scheduling environment state information into the trained time series model to obtain correlation trend information within a future preset time window. The decision optimization module is used to construct a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption based on multi-dimensional feature information, current status information of all nodes, scheduling environment status information, and correlation trend information; and to optimize the multi-objective optimization function based on preset strategies to obtain the target decision strategy. The collaborative scheduling and execution module is used to parse the target decision strategy, generate a scheduling instruction set, obtain task allocation information and corresponding target nodes based on the scheduling instruction set, and send the tasks to the corresponding target nodes.

[0117] A computer device according to embodiments of the present disclosure includes a memory and a processor. The memory is used to store non-transitory computer-readable instructions. Specifically, the memory may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may, for example, include random access memory (RAM) and / or cache memory. The non-volatile memory may, for example, include read-only memory (ROM), hard disk, flash memory, etc.

[0118] The processor may be a central processing unit (CPU) or other processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the computer device to perform desired functions. In one embodiment of this disclosure, the processor is used to execute computer-readable instructions stored in the memory, causing the computer device to perform all or part of the steps of the heterogeneous computing resource dynamic collaborative scheduling method described in the foregoing embodiments of this disclosure.

[0119] Those skilled in the art will understand that, in order to solve the technical problem of how to achieve a good user experience, this embodiment may also include well-known structures such as communication buses and interfaces, and these well-known structures should also be included within the protection scope of this disclosure.

[0120] like Figure 4 This is a schematic diagram of a computer device provided for an embodiment of the present disclosure. It illustrates a structural schematic diagram suitable for implementing the computer device in the embodiments of the present disclosure. Figure 4 The computer device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.

[0121] like Figure 4 As shown, a computer device may include a processor (such as a central processing unit, graphics processing unit, etc.), which can perform various appropriate actions and processes based on programs stored in read-only memory (ROM) or programs loaded from storage devices into random access memory (RAM). The RAM also stores various programs and data required for the operation of the computer device. The processor, ROM, and RAM are interconnected via a bus. Input / output (I / O) interfaces are also connected to the bus.

[0122] Typically, the following devices can be connected to the I / O interface: input devices, such as sensors or visual information acquisition devices; output devices, such as displays; storage devices, such as magnetic tapes or hard drives; and communication devices. Communication devices allow the computer device to communicate wirelessly or wiredly with other devices (such as edge computing devices) to exchange data. Although Figure 4 A computer apparatus with various devices is shown, but it should be understood that it is not required to implement or have all of the devices shown. More or fewer devices may be implemented or included alternatively.

[0123] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from a storage device, or installed from a ROM. When the computer program is executed by a processor, all or part of the steps of the heterogeneous computing resource dynamic collaborative scheduling method of embodiments of this disclosure are performed.

[0124] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0125] A computer-readable storage medium according to embodiments of the present disclosure stores non-transitory computer-readable instructions. When the non-transitory computer-readable instructions are executed by a processor, all or part of the steps of the aforementioned heterogeneous computing resource dynamic collaborative scheduling method of the embodiments of the present disclosure are performed.

[0126] The aforementioned computer-readable storage media include, but are not limited to: optical storage media (e.g., CD-ROM and DVD), magneto-optical storage media (e.g., MO), magnetic storage media (e.g., magnetic tape or portable hard drive), media with built-in rewritable non-volatile memory (e.g., memory card), and media with built-in ROM (e.g., ROM cartridge).

[0127] For a detailed description of this embodiment, please refer to the corresponding descriptions in the foregoing embodiments, which will not be repeated here.

[0128] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.

[0129] In this disclosure, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. The block diagrams of devices, apparatuses, devices, and systems involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, devices, and systems can be connected, arranged, and configured in any manner. Words such as "comprising," "including," "having," etc., are open-ended terms meaning "including but not limited to," and are used interchangeably with them. The terms "or" and "and" as used herein refer to the terms "and / or," and are used interchangeably with them unless the context clearly indicates otherwise. The term "such as" as used herein refers to the phrase "such as but not limited to," and is used interchangeably with it.

[0130] Additionally, as used herein, the "or" used in a list of items beginning with "at least one" indicates a separate list, such that a list of, for example, "at least one of A, B, or C" means A or B or C, or AB or AC or BC, or ABC (i.e., A and B and C). Furthermore, the word "exemplary" does not imply that the described example is preferred or better than other examples.

[0131] It should also be noted that in the systems and methods of this disclosure, the components or steps can be decomposed and / or recombined. These decompositions and / or recombinations should be considered as equivalent solutions to this disclosure.

[0132] Various changes, substitutions, and modifications can be made to the technology described herein without departing from the teachings defined by the appended claims. Furthermore, the scope of the claims of this disclosure is not limited to the specific aspects of the processes, machines, manufactures, events, means, methods, and actions described above. Currently existing or later-developed processes, machines, manufactures, events, means, methods, or actions that perform substantially the same function or achieve substantially the same result as the corresponding aspects described herein can be utilized. Therefore, the appended claims include such processes, machines, manufactures, events, means, methods, or actions within their scope.

[0133] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.

[0134] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations therein.

Claims

1. A method for dynamic collaborative scheduling of heterogeneous computing resources, characterized in that, include: Obtain multi-dimensional feature information of the task to be scheduled; The system can acquire the current status information of each node in the heterogeneous computing power resource pool in real time. The heterogeneous computing power resource pool includes cloud computing centers, edge computing nodes, and terminal devices. Obtain scheduling environment status information; Historical correlation data is obtained based on historical task logs, and the time series model is trained using the historical correlation data. The multi-dimensional feature information, the current state information of all nodes, and the corresponding scheduling environment state information are input into the trained time series model to obtain the correlation trend information within the future preset time window. Based on the multi-dimensional feature information, the current state information of all nodes, the scheduling environment state information, and the correlation trend information, a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption is constructed. The multi-objective optimization function is optimized based on a preset strategy to obtain the target decision strategy; The target decision-making strategy is analyzed to generate a scheduling instruction set; The task allocation information and the corresponding target node are obtained according to the scheduling instruction set, and the task is sent to the corresponding target node.

2. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 1, characterized in that, The process of constructing a multi-objective optimization function that integrates task latency, cost consumption, and energy consumption based on the multi-dimensional feature information, the current state information of all nodes, the scheduling environment state information, and the correlation trend information includes: Based on the current status information and the associated trend information of each node, the comprehensive spatiotemporal cost index of the task is obtained; Based on the delay constraint and the associated trend information in the multi-dimensional feature information, the probability of delay constraint violation is obtained; The probability of the delay constraint being violated is : , For the task i Maximum tolerable completion time; Based on the multi-dimensional feature information, the current status information of each node, and the correlation trend information, the total task transmission cost is obtained; The carbon emission equivalent of task execution is obtained based on the current status information of each node and the scheduling environment status information; Based on the current status information of each node, determine the resource fragmentation penalty item; The resource fragmentation penalty item is: : ; 1≤j≤N, Let j be the computing power utilization rate of node j. To achieve the lowest possible utilization of computing power; A multi-objective optimization function is constructed based on the task spatiotemporal integrated cost index, the probability of delay constraint violation, the total cost of task transmission, the carbon emission equivalent of task execution, and the resource fragmentation penalty term; The multi-objective optimization function is F: F = a· + b; + c· + d· + e· ; α+β+γ+δ+ε=1; in, The comprehensive spatiotemporal cost index of the task. The total cost of transmitting the task. The carbon emission equivalent is calculated for the task, where α, β, γ, δ, and ε are the first, second, third, fourth, and fifth coefficients, respectively.

3. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 2, characterized in that, The step of obtaining the task spatiotemporal comprehensive cost index based on the current state information and the correlation trend information of each node includes: Obtain the inter-node network latency and the predicted execution time of the task on each node from the current status information; Obtain the network state change trend information from the associated trend information; The data transmission network latency of the task is obtained based on the network latency between the nodes and the network state change trend information; The average delay of the predicted task is obtained based on the data transmission network latency of the task and the predicted execution time of the task on each node. The average latency of the prediction task is : , For the task Predicted execution time at the target node For the task Data transmission network latency, The total number of nodes; Obtain the data storage location in the multi-dimensional feature information and the node geographical location in the current state information, and obtain the geographical penalty factor; The geographical penalty factor is : , For the task i The geographical distance between the data storage location and the computing node; Based on the predicted average task delay and the geographical penalty factor, the task spatiotemporal comprehensive cost index is obtained; The comprehensive spatiotemporal cost index of the task is: : .

4. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 3, characterized in that, The step of obtaining the total task transmission cost based on the multi-dimensional feature information, the current state information of each node, and the correlation trend information includes: Obtain the node's geographical location, real-time power consumption, and unit time computation cost from the current status information of each node, as well as the predicted execution time of the task on each node from the associated trend information, to obtain the resource computation cost; The resource calculation cost is ; , Calculate the cost per unit time for the node j where the task is located; Based on the data storage location in the acquired multi-dimensional feature information, the bandwidth utilization in the current status information of each node, and the network topology in the scheduling environment status information, the cross-domain data transmission cost is determined. The cost of cross-domain data transmission is : , For the task i The amount of data to be transmitted The cost per GB of data transfer from the source storage location to compute node j; The total task transmission cost is determined based on the resource calculation cost and the cross-domain data transmission cost. The total cost of the task transmission is : .

5. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 4, characterized in that, The step of obtaining the carbon emission equivalent of task execution based on the current state information of each node includes: Based on the real-time power consumption in the obtained current status information and the predicted execution time of the task on each node, the total energy consumption for task execution is determined. The total energy consumption for the task execution is : , The real-time power consumption of node j; The carbon emission amplification factor is determined based on the carbon emission intensity corresponding to the current energy structure of the node in the scheduling environment status information. The carbon emission equivalent of the task is determined based on the total energy consumption of the task execution and the carbon emission amplification factor. The carbon emission equivalent of the task execution is ; , The carbon emission amplification factor is denoted as .

6. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 5, characterized in that, The optimization of the multi-objective optimization function based on a preset strategy to obtain the objective decision strategy includes: A state space is constructed based on the multi-dimensional feature information, the current state information of all nodes, the scheduling environment state information, and the correlation trend information; Configure the action space, which is a multi-dimensional allocation decision, including the target node ID, data migration path, and computing power reservation ratio; Configure the reward function according to the multi-objective optimization function; The reward function is R: R = -F + R B R B For resource-balanced rewards, η is the weighting coefficient; By using a reinforcement learning engine to iteratively learn based on the state space, the action space, and the reward function, a task-resource mapping scheme that satisfies the optimal multi-objective optimization function is found and used as the target decision strategy.

7. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 1, characterized in that, Also includes: When the data storage location of the task does not match the target node, determine whether the migration time and migration cost meet the preset conditions. If so, trigger cross-node data migration.

8. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 7, characterized in that, The determination of whether the migration time and migration cost meet the preset conditions includes: The first condition is determined based on the delay constraints and estimated computation time in the multi-dimensional feature information. The first condition includes: ,in, The migration takes time. For task delay constraints, To estimate the calculation time, For safety factor; The second condition is determined; the second condition includes: the migration cost is no more than 10% of the task cost budget; Based on the first condition and the second condition, determine the matching result of migration time and migration cost.

9. The method for dynamic collaborative scheduling of heterogeneous computing resources according to claim 1, characterized in that, The step of obtaining historical correlation data based on historical task logs and using the historical correlation data to train the time series model includes: Historical associated data is obtained based on historical task logs. The historical associated data includes task calculation type identifier, submission timestamp, estimated calculation time, delay constraints, data storage location information, resource usage records, node historical utilization time series data, network latency historical values, and power consumption historical sequences. The historical correlation data is divided into a training set and a test set according to a preset ratio; Determine the time series model; The time series model is trained based on the training set; The trained model is analyzed using the test set, and the hyperparameters of the model are dynamically adjusted based on the analysis results until the optimized model meets the preset conditions, thus obtaining the trained time series model.

10. A dynamic collaborative scheduling system for heterogeneous computing resources, characterized in that, include: The task feature acquisition submodule is used to acquire multi-dimensional feature information of the task to be scheduled; The resource status monitoring submodule is used to obtain the current status information of each node in the heterogeneous computing power resource pool in real time. The heterogeneous computing power resource pool includes cloud computing centers, edge computing nodes and terminal devices. The environmental data acquisition submodule is used to acquire scheduling environment status information; The prediction module is used to obtain historical correlation data based on historical task logs and to train the time series model using the historical correlation data; the multi-dimensional feature information, the current state information of all nodes and the corresponding scheduling environment state information are input into the trained time series model to obtain the correlation trend information within a future preset time window; The decision optimization module is used to construct a multi-objective optimization function that integrates task delay, cost consumption, and energy consumption based on the multi-dimensional feature information, the current state information of all nodes, the scheduling environment state information, and the correlation trend information. The multi-objective optimization function is optimized based on a preset strategy to obtain the target decision strategy; The collaborative scheduling execution module is used to parse the target decision strategy and generate a scheduling instruction set; The task allocation information and the corresponding target node are obtained according to the scheduling instruction set, and the task is sent to the corresponding target node.

11. A computer device, characterized in that, The computer device includes: At least one processor; and, A memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the heterogeneous computing resource dynamic collaborative scheduling method according to any one of claims 1-9.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing the computer to execute the dynamic collaborative scheduling method for heterogeneous computing resources as described in any one of claims 1-9.

13. A computer program product comprising computer instructions, characterized in that, When executed by a processor, the computer instructions implement the steps of the method according to any one of claims 1-9.

Citation Information

Cited By

  • Green computing power scheduling system and method based on multi-objective optimization

    CN121501463A

  • Cluster fusion feature generation method and system based on feature interaction, and cluster fusion feature application method and system based on feature interaction

    CN121881274A

  • Computing task streaming decomposition method and system for multi-source heterogeneous computing power

    CN122086563A

  • Method and system for multi-source heterogeneous computing task flow decomposition

    CN122086563B

  • Resource scheduling method and device, storage medium and computer program product

    CN122086570A