Heterogeneous computing power unified scheduling method based on electric power global planning
Through real-time monitoring of the multi-tasking parallel processing environment of the power system and intelligent prediction of machine learning models, heterogeneous resource allocation is dynamically adjusted, and the problem of unreasonable resource allocation caused by task load changes is solved, computing performance and power use efficiency are improved, and operational costs and energy consumption are reduced.
Patent Information
- Application Number
- CN202510527796.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
AI Technical Summary
In a multi-task parallel processing environment, the operation mode of the task may change suddenly, especially when some tasks enter the high load stage, their computing needs will increase rapidly. It is difficult for existing scheduling systems to accurately grasp the dynamic load changes in the task execution process, resulting in unreasonable allocation of power resources, insufficient or oversupply of resources, affecting the efficiency of task execution and causing power waste.
By monitoring the multi-task parallel processing environment of the power system in real time, using machine learning models to intelligently predict the task load state, construct task load state feature vectors, generate task stress coefficients, and dynamically adjust heterogeneous resource allocation to ensure that high-load tasks obtain sufficient computing power and avoid excessive resource occupancy of low-load tasks.
It realizes accurate scheduling decisions before sudden load changes, improves computing performance, optimizes power usage efficiency, reduces unnecessary energy consumption in data centers or heterogeneous computing environments, reduces operational costs, and improves computing throughput and task execution stability.
Smart Images

Figure CN120448112A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power dispatching, and in particular to a method for unified dispatching of heterogeneous computing power based on global power planning. Background Art
[0002] Unified scheduling of heterogeneous computing power based on global power planning refers to the coordination and scheduling of different types of computing resources (such as CPUs, GPUs, and FPGAs) during power system operation, taking into account their characteristics, strengths, and weaknesses through global power planning. Its core purpose is to optimize the use of power resources, ensuring that different computing platforms can not only efficiently complete tasks while meeting computing needs, but also reduce energy consumption, thereby achieving the optimal match between power and computing resources. This scheduling method requires real-time monitoring of power load, computing task priorities, and resource availability, enabling efficient resource scheduling on a global scale to ensure stable system operation and sustainable energy utilization.
[0003] The existing technology has the following deficiencies:
[0004] In a multi-tasking parallel processing environment, the operating modes of tasks may change suddenly, especially when certain tasks enter a high-load phase, where their computing requirements increase rapidly. However, these changes are often difficult to predict, especially on heterogeneous computing platforms (such as CPUs, GPUs, and FPGAs), where different computing architectures respond differently to load fluctuations. If the scheduling system cannot accurately grasp the dynamic load changes during task execution, it will be difficult to effectively estimate energy consumption requirements, which may lead to irrational allocation of power resources, either insufficient resources, affecting task execution efficiency, or excessive resources, resulting in power waste.
[0005] The above information disclosed in this Background section is only for enhancement of understanding of the background of the present disclosure and therefore it may contain information that does not form the prior art that is already known to a person of ordinary skill in the art. Summary of the Invention
[0006] The purpose of the present invention is to provide a unified scheduling method for heterogeneous computing power based on global power planning. By real-time monitoring of the operating status of subtasks in a multi-task parallel processing environment and using machine learning models to intelligently predict the task load status, accurate scheduling decisions can be made before sudden load changes occur. This prediction-based dynamic resource adjustment mechanism enables computing resources to be adaptively allocated between different tasks, ensuring that high-load tasks obtain sufficient computing power while avoiding low-load tasks from excessively occupying resources and causing power waste. Compared with traditional static scheduling methods, this solution can not only improve computing performance, but also optimize power utilization efficiency and reduce unnecessary energy consumption in data centers or heterogeneous computing environments, thereby achieving the goals of reducing operating costs, increasing computing throughput, and enhancing task execution stability, in order to solve the problems in the above-mentioned background technology.
[0007] To achieve the above objectives, the present invention provides the following technical solution: a method for unified scheduling of heterogeneous computing power based on global power planning, comprising the following steps:
[0008] Comprehensive and continuous monitoring and data collection of the operating status of all subtasks in the multi-task parallel processing environment of the power system;
[0009] Extract indicators with predictive value for subtask load status from massive data, and construct a feature vector representing the subtask load status based on the analysis of the load indicators to characterize the current load status of the subtask;
[0010] The constructed feature vector is input into the trained machine learning model to complete the intelligent prediction of the load status of each subtask;
[0011] After obtaining the prediction results of the machine learning model for the future load of each subtask, the heterogeneous resource allocation plan for the future period is dynamically adjusted, and the allocation weights of heterogeneous resources are dynamically adjusted according to the load conditions of the subtasks to avoid insufficient or excessive configuration of power resources due to sudden loads.
[0012] Preferably, data collection is performed on the operating conditions of all subtasks in the multi-task parallel processing environment of the power system, and the specific steps are as follows:
[0013] Step 1: Determine the monitoring scope and indicators. Based on the characteristics of the multi-task parallel processing environment, clarify the subtask parameters that need to be monitored and set the corresponding sampling frequency.
[0014] Step 2: Deploy monitoring components to capture the full range of subtask running status;
[0015] Step 3: Transmit the real-time collected monitoring data to a centralized storage system or distributed database via the network or message queue, and perform preliminary format conversion and annotation;
[0016] Step 4: Perform integrity check, denoising, outlier processing, and timestamp alignment on the collected data to provide accurate basic data for subsequent analysis and scheduling.
[0017] Preferably, indicators with predictive value for the subtask load status are extracted from massive data, wherein the extracted indicators include the frequency of subtasks being blocked during execution and the proportion of time that the subtasks are forced to be postponed by the operating system. During the monitoring period, the frequency of subtasks being blocked during execution and the proportion of time that the subtasks are forced to be postponed by the operating system are deeply analyzed, and task blocking quantization values and calculation backoff quantization values are generated respectively. The task blocking quantization values and calculation backoff quantization values are used to construct a characteristic vector representing the subtask load status to characterize the current load status of the subtask.
[0018] Preferably, the constructed feature vector is input into the trained machine learning model, wherein the feature vector is the task blocking quantization value and the calculation backoff quantization value generated after in-depth analysis. The machine learning model generates a task stress coefficient by analyzing the task blocking quantization value and the calculation backoff quantization value, and intelligently predicts the load status of each subtask based on the task stress coefficient.
[0019] Preferably, after obtaining the prediction results of the machine learning model for the future load of each subtask, the heterogeneous resource allocation plan for a period of time in the future is dynamically adjusted. The specific steps are as follows:
[0020] Obtaining the task stress coefficient Task predicted by the machine learning model strain Then, calculate the demand weight of each subtask for different heterogeneous resources. Suppose there are N m subtasks, and the task stress coefficient of each subtask k is The resource demand weight is defined as:
[0021]
[0022] Where: W k is the resource demand weight of subtask k, which determines how much computing resources the subtask should obtain in the next stage, α w is a nonlinear adjustment parameter used to control the nonlinear mapping between task load and resource allocation. j represents the index variable to be summed in all subtasks, that is, the index used to traverse all subtasks in the system. N m Indicates the total number of subtasks in the system, that is, the number of all tasks that are being executed or require resource allocation during the monitoring period.
[0023] Preferably, after calculating the resource requirement weight W of each subtask kFinally, according to the adaptability of different computing architectures, the allocation ratio of heterogeneous resources is dynamically adjusted, and the resource allocation amount of the computing task on the heterogeneous computing resources is calculated using the expression:
[0024]
[0025] Where: R k,p For subtask k in computing resource R p The actual distribution amount on p For computing resources R p The computing performance weight is used to measure the computing power of the resource. M is the number of heterogeneous computing resources in the scheduling system. R total is the total computing resources currently available in the scheduling system, and m represents the index variable used to traverse all heterogeneous computing resources, that is, the variable to be summed in the computing resource set.
[0026] Preferably, within the monitoring period, the specific steps of deeply analyzing the frequency of subtask blocking during execution to generate a task blocking quantification value are as follows:
[0027] During the monitoring cycle, the execution process of each subtask is monitored in real time, and all key events that lead to blocking are collected;
[0028] A cumulative impact factor is introduced to measure the overall impact of all blocking events on task execution. The cumulative impact factor is calculated as follows:
[0029]
[0030] Where: C if is the cumulative impact factor, N is the total number of blocking events detected during the monitoring period, and B i is the duration of the blocking when the i-th blocking occurs, P i is the resource occupancy rate of the task when the i-th blocking occurs, R i is the computational intensity of the task currently being executed when the i-th blocking occurs, D i is the memory bandwidth utilization of the task when the i-th blocking occurs, and α is the adjustment factor used to nonlinearly enhance the impact of extreme cases;
[0031] After obtaining the cumulative impact factor, it is mapped to the task blocking quantification value. The task adaptability factor is introduced and the cumulative impact factor is normalized to ensure the comparability of the blocking index of different tasks. The task blocking quantification value is defined as follows:
[0032]
[0033] Among them: B f is the task blocking quantization value, Taf is the task adaptability factor, which reflects the tolerance of the task to blocking. β is the adaptability adjustment factor, which is used to balance the range of task blocking quantization values of different task categories. γ is the smoothing factor.
[0034] Preferably, within the monitoring period, the specific steps of deeply analyzing the time ratio of the subtask that is forcibly postponed by the operating system to generate and calculate the backoff quantization value are as follows:
[0035] During the monitoring period, the total time of the subtask being forced to postpone execution due to the operating system and the total time of the subtask actually getting the processor and being successfully executed are recorded, respectively as f d and r t At the same time, the number of times the subtask is forcibly preempted during this period and the total number of scheduling fragments are counted, respectively recorded as p c and s c ;
[0036] Introducing two key ratios R x and P x , accurately quantify the degree of subtask constraints and resource competition:
[0037]
[0038] Where: R x Indicates the extent to which the subtask is postponed by the operating system during the monitoring period; P x It reflects the intensity of forced preemption encountered by the subtask;
[0039] Based on the degree to which the subtask is postponed by the operating system during the monitoring period R x and the density of forced preemption encountered by subtasks P x , a nonlinear fusion method is used to construct the calculation backoff quantization value to characterize the load status and degree of limitation of the subtask. The construction expression for the calculation backoff quantization value is:
[0040]
[0041] Among them: back off To calculate the backoff quantization value, To delay the stress-enhancing factor, To seize the interference enhancement factor, R x ·P x is the interaction factor used to capture the synergistic effect of postponement and preemption. λ, μ, and η are the weight coefficients of the postponement pressure enhancement factor, the preemption interference enhancement factor, and the interaction factor, respectively.
[0042] In the above technical solution, the technical effects and advantages provided by the present invention are:
[0043] The present invention monitors the operating status of subtasks in a multi-task parallel processing environment in real time and uses machine learning models to intelligently predict task load status, making accurate scheduling decisions before sudden load changes occur. This prediction-based dynamic resource adjustment mechanism enables computing resources to be adaptively allocated between different tasks, ensuring that high-load tasks have sufficient computing power while preventing low-load tasks from excessively occupying resources and wasting electricity. Compared to traditional static scheduling methods, this solution not only improves computing performance, but also optimizes power usage efficiency and reduces unnecessary energy consumption in data centers or heterogeneous computing environments, thereby achieving the goals of reducing operating costs, increasing computing throughput, and enhancing task execution stability. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction to the drawings required for use in the embodiments will be given below. Obviously, the drawings described below are only some embodiments recorded in the present invention. For ordinary technicians in this field, other drawings can also be obtained based on these drawings.
[0045] Figure 1 This is a flow chart of the method for unified scheduling of heterogeneous computing power based on global power planning of the present invention. DETAILED DESCRIPTION
[0046] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in many forms and should not be construed as limited to the examples set forth herein; rather, these example embodiments are provided so that the description of this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art.
[0047] The present invention provides Figure 1 The unified scheduling method for heterogeneous computing power based on global power planning shown in the figure includes the following steps:
[0048] Comprehensive and continuous monitoring and data collection of the operating status of all subtasks in the multi-task parallel processing environment of the power system;
[0049] Data collection on the operating status of all subtasks in the multi-task parallel processing environment of the power system is carried out. The specific steps are as follows:
[0050] Step 1: Determine the monitoring scope and indicators. Based on the characteristics of the multi-task parallel processing environment, clearly define the subtask parameters that need to be monitored (such as CPU, GPU, and FPGA resource utilization, power consumption, execution time, etc.) and set the corresponding sampling frequency;
[0051] Step 2: Deploy monitoring components and integrate multiple data sources into the system, such as hardware sensors, operating system performance counters, and application-level log modules, to ensure comprehensive capture of subtask running status.
[0052] Step 3: Data transmission and storage management: transmit the real-time collected monitoring data to a centralized storage system or distributed database through the network or message queue, and perform preliminary format conversion and annotation;
[0053] Step 4: Data verification and preprocessing: perform integrity check, denoising, outlier processing, and timestamp alignment on the collected data to provide accurate and reliable basic data for subsequent analysis and scheduling.
[0054] After obtaining the subtask's operating data, we extract indicators with predictive value for the subtask's load status from the massive data. Based on the analysis of the load indicators, we construct a feature vector representing the subtask's load status to characterize the subtask's current load status.
[0055] Indicators with predictive value for subtask load conditions are extracted from massive data. The extracted indicators include the frequency of subtask blocking during execution and the proportion of time that subtask execution is forcibly postponed by the operating system. During the monitoring period, the frequency of subtask blocking during execution and the proportion of time that subtask execution is forcibly postponed by the operating system are deeply analyzed to generate task blocking quantization values and calculation backoff quantization values respectively. The task blocking quantization values and calculation backoff quantization values are used to construct a feature vector representing the subtask load state to characterize the current load condition of the subtask.
[0056] If a subtask is blocked frequently during execution, it usually indicates that its current load is high. This is because task blocking is often related to factors such as resource contention, compute-intensive operations, I / O bottlenecks, or high concurrency competition. For example, when a CPU-intensive task frequently waits for data loading (such as main memory access caused by cache misses) due to excessive computing requirements, the instruction pipeline will be paused, thereby reducing throughput; in multi-threaded tasks, if threads frequently wait due to lock contention or synchronization mechanisms (such as mutexes and semaphores), it indicates that the tasks are highly competitive and under high load; when I / O-intensive tasks read or write large amounts of data, if the storage system cannot respond in time, high-frequency blocking may occur, affecting the efficiency of task execution. Therefore, an increase in blocking frequency is usually a sign that the task load has reached a bottleneck or that resource utilization is too high. If scheduling is not optimized or additional resources are not allocated in a timely manner, the efficiency of task execution may decrease, and even the stability of the entire system may be affected.
[0057] During the monitoring period, the specific steps for deeply analyzing the frequency of subtask blocking during execution and generating task blocking quantification values are as follows:
[0058] During the monitoring cycle, the execution of each subtask must be monitored in real time, and all key events that could cause blocking must be collected. Examples include CPU instruction pipeline pauses (such as data dependencies and cache misses), thread synchronization waits (such as mutexes and conditional variables), and I / O delays (such as storage device response timeouts). To accurately quantify the severity of these blocking events, a cumulative impact factor is introduced to measure the overall impact of all blocking events on task execution. The cumulative impact factor is calculated as follows:
[0059]
[0060] Where: C if is the cumulative impact factor, N is the total number of blocking events detected during the monitoring period, and B i is the duration of the blocking when the i-th blocking occurs, P i is the resource occupancy rate of the task (such as CPU and GPU utilization, ranging from 0 to 1) when the i-th blocking occurs, R i is the computational intensity of the task currently being executed when the i-th blocking occurs, D i is the memory bandwidth utilization of the task when the i-th blocking occurs, α is the adjustment factor used to nonlinearly enhance the impact of extreme cases (typical values are 1.5 to 3);
[0061] Computational intensity refers to the proportion of computational operations (such as arithmetic operations, logical judgments, matrix multiplication, etc.) in the overall instruction stream or operation stream during the execution of a task. It can reflect the level of consumption of processor computing resources by the task. If the computational intensity of a task is high, it means that most of its execution time is spent on heavy arithmetic or logical operations, which puts relatively greater pressure on computing units such as the CPU or GPU; on the contrary, if the computational intensity is low, it means that the task relies more on I / O operations, storage access or network communication, rather than pure computing power input;
[0062] The core goal of this step is to quantify the cumulative blocking impact of tasks during the monitoring period through the cumulative impact factor indicator. The innovation of this formula lies in:
[0063] Considering the interaction between blocking time and resource occupancy (B i ·P i ), if the task is blocked for a long time when the resource usage is high, the performance will be affected more. Introducing the correction factor of computational intensity and memory bandwidth utilization (R i +D iIf the task is computationally intensive or bandwidth-intensive, it may be more susceptible to congestion due to resource contention, so its impact weight can be appropriately reduced. Nonlinear amplification of extreme congestion (α): When severe congestion occurs, the exponential amplification mechanism can more accurately distinguish between minor congestion and severe load issues.
[0064] After obtaining the cumulative impact factor, it is further mapped to a task blocking quantification value to more intuitively represent the load status of the subtask. Since different tasks may have different computational complexity and blocking sensitivity, the task adaptability factor is introduced to normalize the cumulative impact factor to ensure the comparability of the blocking index of different tasks. The task blocking quantification value is defined as follows:
[0065]
[0066] Among them: B f is the task blocking quantification value. The higher the value, the higher the blocking load of the task. af is the task adaptability factor, reflecting the task's tolerance to blocking (value range 0.1 to 10), β is the adaptability adjustment factor (usually 0.5 to 2), used to balance the task blocking quantization value range of different task categories, and γ is the smoothing factor (to prevent the denominator from being too small, which will lead to an explosive growth of the task blocking quantization value, usually 1 to 5);
[0067] The task adaptability factor is a parameter used to measure the tolerance of different types of tasks to blocking situations. It reflects the adaptability of a task to computational, I / O, or synchronization blocking. Obtaining the task adaptability factor requires combining historical task execution data and task type characteristics. It can be calculated using statistical analysis and machine learning methods. For example, the number of blocking events encountered by a task during past execution, their duration, and their contribution to overall execution time can be counted to calculate the task's average blocking tolerance. Furthermore, an initial task adaptability factor value can be set based on the task's category (e.g., compute-intensive, I / O-intensive, or memory-intensive) and dynamically adjusted based on actual execution data. For example, an I / O-intensive task may have a higher tolerance for storage latency and therefore a higher task adaptability factor, while a compute-intensive task may be more sensitive to computational unit blocking and therefore have a lower task adaptability factor. In this way, the task adaptability factor can serve as a dynamically adjusted parameter to ensure that task blocking quantification values remain comparable across tasks, thereby improving the accuracy of the scheduling system.
[0068] The Task Blocking Quantization value shows that, during the monitoring period, the larger the performance value of the Task Blocking Quantization value, generated by in-depth analysis of the frequency of subtask blocking during execution, the higher the current load of the subtask. Specifically, the Task Blocking Quantization value comprehensively considers the frequency of subtask blocking, the duration of blocking, the resource utilization level, and the computational characteristics of the task itself during the monitoring period. When subtask blocking events occur frequently and are accompanied by long blocking durations, high resource utilization, and low adaptability, the Task Blocking Quantization value will increase significantly, reflecting a high load or resource bottleneck on the task. Conversely, a lower Task Blocking Quantization value indicates that the subtask has experienced few blocking events during execution, has low competition for computing resources, or has high adaptability, resulting in a relatively low current load and a relatively stable task operation. Therefore, the Task Blocking Quantization value intuitively and accurately reflects the current load level of the subtask.
[0069] If a subtask is forcibly deferred by the operating system for a high percentage of its execution time, it usually indicates that the subtask is under a high load. The operating system's task scheduling mechanism decides when to execute or defer a task based on the task's priority, resource usage, and overall system load. When the proportion of subtask execution time deferred increases, it may mean that the computing resources in the system (such as the CPU and GPU) are close to saturation, resulting in the scheduler being unable to immediately allocate sufficient processing power to the task. This situation usually occurs when the task requires high computing performance or large amounts of memory access, especially when multiple high-load tasks compete for resources simultaneously. The operating system may be forced to delay the execution of some tasks to ensure the overall stability of the system. In addition, if the subtask involves I / O-intensive operations, access latency to the storage device may also cause the task execution to be delayed. A continuously high proportion of deferred execution time not only affects the response speed of the task, but may also lead to a decline in overall computing performance.
[0070] During the monitoring period, the specific steps for deeply analyzing the proportion of time that the subtask is forced to postpone execution by the operating system to generate the backoff quantization value are as follows:
[0071] During the monitoring period, the total time of the subtask being forced to postpone execution due to the operating system and the total time of the subtask actually getting the processor and being successfully executed are recorded, respectively as f d and r t At the same time, the number of times the subtask is forcibly preempted during this period and the total number of scheduling fragments are counted, respectively recorded as p c and s c On this basis, two key ratios R are introduced x and P x , accurately quantify the degree of subtask constraints and resource competition:
[0072]
[0073] Where: R x Indicates the degree to which the subtask is postponed by the operating system during the monitoring period. The higher the value, the greater the proportion of subtask postponement. x It reflects the intensity of forced preemption encountered by the subtask. The higher the value, the more interference the subtask encounters at the scheduling level.
[0074] The purpose of this step is to accurately quantify the degree of subtask constraints and resource competition, providing key raw metrics for the subsequent calculation of backoff quantization values.
[0075] Based on the degree to which the subtask is postponed by the operating system during the monitoring period R x and the density of forced preemption encountered by subtasks P x , a nonlinear fusion method is used to construct the calculation backoff quantization value to accurately characterize the load status and degree of limitation of the subtask. The construction expression of the calculation backoff quantization value is:
[0076]
[0077] Among them: back off To calculate the backoff quantization value, To delay the pressure enhancement factor, when R x When it approaches 1 (i.e., the task is almost completely postponed), the term increases rapidly, reflecting the extreme constraints imposed on the subtask due to scheduling delays. To seize the interference enhancement factor, when P x When the value is high, the amplification task is frequently preempted, which significantly increases the calculation backoff quantization value of the high preemption task. x ·P x is an interaction factor used to capture the synergy between postponement and preemption. If a subtask is postponed (R x High) and frequently preempted (P x High), this item will greatly improve the calculation backoff quantization value and accurately reflect its degree of limitation. λ, μ, and η are the weight coefficients of the postponement pressure enhancement factor, the preemption interference enhancement factor, and the interaction influence factor, respectively, which can be dynamically adjusted according to different computing environments;
[0078] This step can adaptively amplify the calculation backoff quantization value of high-load tasks while maintaining reasonable suppression of low-load tasks, ensuring that when P x or P x At a moderate level, the backoff quantization value will not be overly inflated, while when the task is severely constrained, the backoff quantization value can quickly reflect abnormal conditions.
[0079] From the calculated backoff quantization value, we can see that during the monitoring period, the calculated backoff quantization value generated by the in-depth analysis of the proportion of time that the subtask is forced to postpone execution by the operating system is larger, indicating that the subtask has experienced more severe system scheduling delays and resource competition during the monitoring period, thus reflecting that its current load condition is higher. Since the calculated backoff quantization value is the degree to which the subtask is postponed by the operating system during the monitoring period, R x and the density of forced preemption encountered by subtasks P x The calculated value comprehensively captures the scheduling obstacles encountered by a task during execution. A large backoff value indicates that the subtask is frequently delayed or preempted due to high competition or limited system resources, hindering its computation and leading to a high-load state. A small backoff value, on the other hand, indicates that the subtask is less likely to be delayed or preempted, running smoothly, with relatively sufficient resource scheduling and a low load. Therefore, the size of the backoff value can serve as an important indicator for determining whether a subtask is experiencing a high-load state.
[0080] The constructed feature vector is input into the trained machine learning model to complete the intelligent prediction of the load status of each subtask;
[0081] The constructed feature vector is input into the trained machine learning model. The feature vector is the task blocking quantization value and the calculation backoff quantization value generated after in-depth analysis. The machine learning model generates the task stress coefficient by analyzing the task blocking quantization value and the calculation backoff quantization value, and intelligently predicts the load status of each subtask based on the task stress coefficient.
[0082] A trained machine learning model is an intelligent model that has completed offline training and reached a usable state through a thorough process of data preparation, algorithm selection, parameter optimization, and verification. It contains all the internal structure and weight information required to learn a specific problem (such as inferring the trend of task load peaks based on task blocking quantization and calculation backoff quantization). In other words, before being officially put into use, the model has gone through several core stages: the first is data collection and feature engineering. Researchers or engineers will iteratively polish key features such as "task blocking quantization value" and "computational backoff quantization value" to ensure that the input vector can fully reflect the main factors affecting the subtask load status; the second is algorithm screening and network structure design. On the basis of considering a variety of machine learning algorithms (such as deep neural networks, random forests or integrated learning methods), combined with actual business needs and data scale, the most suitable model architecture is selected and appropriate hyperparameters are configured; the third is the training process, which often requires iteratively updating model parameters in a large amount of historical task operation data and a simulated test environment to minimize the loss function or improve evaluation indicators such as accuracy and recall rate to prevent overfitting or underfitting; the fourth is verification and testing. The trained model is placed on the verification set and test set for performance evaluation, and cross-validation or independent testing is used to test whether the model can maintain stable and accurate prediction effects on unseen data. When this series of processes is completed, the model will have the feasibility and reliability to be put into actual deployment. That is, in the same or similar distribution data environment, it can give the mapping relationship learned during the training process based on the input feature vector, thereby generating a reasonable prediction output. For this solution, a trained model means that it can effectively infer the relationship between the "task blocking quantization value" and the "computation backoff quantization value", timely capture the potential high load conditions during task execution, and provide valuable reference for the next step of resource scheduling and performance optimization. At the same time, this "completed training" does not mean that the model does not need to be updated at all after deployment. Some systems will also continue to conduct online learning or regular retraining to maintain the model's sensitivity to environmental changes, thereby improving the overall prediction effect and robustness.
[0083] In specific implementations, the emphasis on "trained machine learning models" is mainly to distinguish between the offline training phase and the online inference phase, thereby better explaining the role of the model in actual use. The offline training phase often takes up a lot of computing resources and time, and is usually completed in a test environment or a dedicated data center. The best strategy is selected through repeated trials and model comparisons (for example, comparing different network depths, activation functions, or loss functions). Once the training is stable and verified, the model can be deployed to the runtime environment. At this point, it already has the ability to extract effective feature patterns from the input feature vector (i.e., task blocking quantization values and computational backoff quantization values). When processing new data, the model does not need to go through the tedious and lengthy training steps again. Instead, it directly uses the internally solidified parameters (weight matrix, bias terms, tree structure, or other forms) for inference and outputs a "task stress coefficient" that reflects the task load status. In short, the "trained machine learning model" is not only a distillation of past historical data and patterns, but also a "weapon" for making real-time or quasi-real-time decisions in the current complex system environment. It is "on standby" at any time in the production environment. As long as a new feature vector is input, it can immediately output the prediction results reflecting the subtask load, providing key auxiliary information for system resource allocation, anomaly detection and performance tuning.
[0084] The machine learning model is not specifically limited here, and can achieve the task blocking quantization value B f And calculate the backoff quantization value back off Perform comprehensive analysis to generate task stress coefficient Task strain In order to realize the technical solution of the present invention, the present invention provides a specific implementation method; Task stress coefficient Task strain The generated expression is: Task strain =w p *B f +w q *back off , where w p 、w q They are task blocking quantization value B f And calculate the backoff quantization value back off The preset proportional coefficient, and w p 、w q are all greater than 0. The preset proportional coefficient here refers to the weight coefficient w p and w q , which are used to control the task blocking quantization value B f And calculate the backoff quantization value back off In calculating the task stress coefficient Task strainThe role of these coefficients is to adjust the weight of the two core indicators on the final task stress coefficient according to different computing environments, task types and system requirements. For example, if the high load caused by resource competition (i.e. task blocking) in a computing environment is more serious, then w p It may be set larger so that B f In the calculation task strain If the system is more concerned about the task execution delay caused by the operating system scheduling delay (i.e., backoff), then w q May be set higher, making back off The preset scaling factor is usually set empirically or obtained through data training optimization. Its ultimate goal is to make the calculated task stress coefficient more consistent with the actual system load, allowing the task scheduling system to more accurately allocate computing resources and optimize overall performance.
[0085] It can be seen from the task stress coefficient that during the monitoring period, the greater the performance value of the task blocking quantization value generated by in-depth analysis of the frequency of blocking of subtasks during execution, and the greater the performance value of the calculation backoff quantization value generated by in-depth analysis of the proportion of time when subtasks are forcibly postponed by the operating system, that is, the greater the performance value of the task stress coefficient generated by the intelligent prediction of the load status of each subtask through the trained machine learning model, the higher the current load status of the subtask, and vice versa.
[0086] After obtaining the machine learning model's prediction results for the future load of each subtask, the heterogeneous resource allocation plan for the next period of time is dynamically adjusted. The allocation weights of heterogeneous resources are dynamically adjusted according to the subtask load to avoid insufficient or excessive power resources due to load surges.
[0087] After obtaining the machine learning model's prediction results for the future load of each subtask, dynamically adjust the heterogeneous resource allocation plan for the next period of time. The specific steps are as follows:
[0088] Obtaining the task stress coefficient Task predicted by the machine learning model strain After that, it is necessary to calculate the demand weight of each subtask for different heterogeneous resources (such as CPU, GPU, FPGA) in order to make reasonable allocation. Suppose there are N m subtasks, and the task stress coefficient of each subtask k is (predicted by the machine learning model), the resource demand weight is defined as:
[0089]
[0090] Where: Wk is the resource demand weight of subtask k, which determines how much computing resources the subtask should obtain in the next stage, α w is a nonlinear adjustment parameter used to control the nonlinear mapping between task load and resource allocation. w When α > 1, the resource tilt of high-load tasks is greater, and w <1, the resource allocation is more balanced, j represents the index variable for summing across all subtasks, that is, the index used to traverse all subtasks in the system, N m Indicates the total number of subtasks in the system, that is, the number of all tasks that are being executed or need to allocate resources during the monitoring period;
[0091] The core idea of this step is to dynamically adjust resource allocation weights based on the importance of task load, avoiding resource waste when tasks are underloaded while also preventing high-load tasks from impacting system stability due to insufficient resource allocation. Compared to traditional linear allocation methods, this formula introduces exponential adjustment, significantly increasing the resource share of extremely high-load tasks while reducing the resource share of low-load tasks, thereby more accurately adapting to heterogeneous computing resources.
[0092] After calculating the resource requirement weight W of each subtask k After that, the allocation ratio of heterogeneous resources is dynamically adjusted according to the adaptability of different computing architectures (such as CPU, GPU, FPGA). The resource allocation amount of computing tasks on heterogeneous computing resources (such as CPU, GPU, FPGA) is calculated as follows:
[0093]
[0094] Where: R k,p For subtask k in computing resource R p The actual allocation on (such as CPU, GPU, FPGA), β p For computing resources R p The computing performance weight is used to measure the computing power of the resource (for example, if the GPU computing performance is much higher than that of the CPU, then β GPU >β CPU ), M is the number of heterogeneous computing resources in the scheduling system (if it includes CPU, GPU, and FPGA, then M = 3), R total is the total computing resources currently available to the scheduling system (which can be set based on total computing power consumption, power budget, or total task demand), and m represents the index variable used to traverse all heterogeneous computing resources, that is, the variable to be summed in the set of computing resources (such as CPU, GPU, FPGA, etc.);
[0095] This step implements two layers of dynamic adjustment:
[0096] Task-level dynamic adjustment automatically adjusts the computing resource ratio of tasks according to the stress coefficient of the tasks, ensuring that high-load tasks get more computing resources. Different computing architectures have different computing capabilities, so β is introduced. p As a computing performance weight, tasks can not only obtain resources according to load, but also be reasonably matched to the most suitable computing architecture. For example, highly parallel computing tasks tend to be assigned to the GPU, while logic-intensive tasks are given priority to the CPU.
[0097] Through the above steps, the system can dynamically and adaptively allocate heterogeneous computing resources, while ensuring priority scheduling of high-load tasks, optimizing resource utilization to the greatest extent, and avoiding problems such as over-configuration of computing resources or insufficient power consumption.
[0098] The beneficial effect of this solution is to improve the matching efficiency of power resources and computing resources and achieve efficient energy consumption management. By real-time monitoring of the running status of subtasks in a multi-task parallel processing environment and using machine learning models to intelligently predict the task load status, the system can make accurate scheduling decisions before sudden load changes occur. This prediction-based dynamic resource adjustment mechanism enables computing resources to be adaptively allocated between different tasks, ensuring that high-load tasks have sufficient computing power while avoiding low-load tasks from excessively occupying resources and causing power waste. Compared with traditional static scheduling methods, this solution can not only improve computing performance, but also optimize power usage efficiency and reduce unnecessary energy consumption in data centers or heterogeneous computing environments, thereby achieving the goals of reducing operating costs, increasing computing throughput, and enhancing task execution stability.
[0099] The above formulas are all dimensionless and numerical calculations. The formulas are obtained by collecting a large amount of data and performing software simulation to obtain the most recent real situation. The preset parameters in the formulas are set by technicians in this field according to actual conditions.
[0100] The above description is merely illustrative of certain exemplary embodiments of the present invention. It goes without saying that those skilled in the art will be able to modify the described embodiments in various ways without departing from the spirit and scope of the present invention. Therefore, the above drawings and description are illustrative in nature and should not be construed as limiting the scope of protection of the claims.
[0101] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A unified scheduling method for heterogeneous computing power based on global power planning, characterized by: The following steps are involved: Comprehensive and continuous monitoring and data collection of the operating status of all subtasks in the multi-task parallel processing environment of the power system; Extract indicators with predictive value for subtask load status from massive data, and construct a feature vector representing the subtask load status based on the analysis of the load indicators to characterize the current load status of the subtask; The constructed feature vector is input into the trained machine learning model to complete the intelligent prediction of the load status of each subtask; After obtaining the prediction results of the machine learning model on the future load of each subtask, the heterogeneous resource allocation plan for the future period is dynamically adjusted, and the allocation weights of heterogeneous resources are dynamically adjusted according to the load conditions of the subtasks to avoid insufficient or excessive configuration of power resources due to sudden loads.
2. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 1 is characterized in that: Data collection on the operating status of all subtasks in the multi-task parallel processing environment of the power system is carried out. The specific steps are as follows: Step 1: Determine the monitoring scope and indicators. Based on the characteristics of the multi-task parallel processing environment, clarify the subtask parameters that need to be monitored and set the corresponding sampling frequency. Step 2: Deploy monitoring components to capture the full range of subtask running status; Step 3: Transmit the real-time collected monitoring data to a centralized storage system or distributed database via the network or message queue, and perform preliminary format conversion and annotation; Step 4: Perform integrity check, denoising, outlier processing, and timestamp alignment on the collected data to provide accurate basic data for subsequent analysis and scheduling.
3. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 1 is characterized in that: Indicators with predictive value for subtask load conditions are extracted from massive data. The extracted indicators include the frequency of subtask blocking during execution and the proportion of time that subtask execution is forcibly postponed by the operating system. During the monitoring period, the frequency of subtask blocking during execution and the proportion of time that subtask execution is forcibly postponed by the operating system are deeply analyzed to generate task blocking quantization values and calculation backoff quantization values respectively. The task blocking quantization values and calculation backoff quantization values are used to construct a feature vector representing the subtask load state to characterize the current load condition of the subtask.
4. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 3 is characterized in that: The constructed feature vector is input into the trained machine learning model. The feature vector is the task blocking quantization value and the calculation backoff quantization value generated after in-depth analysis. The machine learning model generates the task stress coefficient by analyzing the task blocking quantization value and the calculation backoff quantization value, and intelligently predicts the load status of each subtask based on the task stress coefficient.
5. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 4 is characterized in that: After obtaining the machine learning model's prediction results for the future load of each subtask, dynamically adjust the heterogeneous resource allocation plan for the next period of time. The specific steps are as follows: Obtaining the task stress coefficient Task predicted by the machine learning model strain Then, calculate the demand weight of each subtask for different heterogeneous resources. Suppose there are N m subtasks, and the task stress coefficient of each subtask k is The resource demand weight is defined as: Where: W k is the resource demand weight of subtask k, which determines how much computing resources the subtask should obtain in the next stage, α w is a nonlinear adjustment parameter used to control the nonlinear mapping between task load and resource allocation. j represents the index variable to be summed in all subtasks, that is, the index used to traverse all subtasks in the system. N m Indicates the total number of subtasks in the system, that is, the number of all tasks that are being executed or require resource allocation during the monitoring period.
6. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 5 is characterized in that: After calculating the resource requirement weight W of each subtask k Finally, according to the adaptability of different computing architectures, the allocation ratio of heterogeneous resources is dynamically adjusted, and the resource allocation amount of the computing task on the heterogeneous computing resources is calculated using the expression: Where: R k,p For subtask k in computing resource R p The actual distribution amount on p For computing resources R p The computing performance weight is used to measure the computing power of the resource. M is the number of heterogeneous computing resources in the scheduling system. R total is the total computing resources currently available in the scheduling system, and m represents the index variable used to traverse all heterogeneous computing resources, that is, the variable to be summed in the computing resource set.
7. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 3 is characterized in that: During the monitoring period, the specific steps for deeply analyzing the frequency of subtask blocking during execution and generating task blocking quantification values are as follows: During the monitoring cycle, the execution process of each subtask is monitored in real time, and all key events that lead to blocking are collected; A cumulative impact factor is introduced to measure the overall impact of all blocking events on task execution. The cumulative impact factor is calculated as follows: Where: C if is the cumulative impact factor, N is the total number of blocking events detected during the monitoring period, and B i is the duration of the blocking when the i-th blocking occurs, P i is the resource occupancy rate of the task when the i-th blocking occurs, R i is the computational intensity of the task currently being executed when the i-th blocking occurs, D i is the memory bandwidth utilization of the task when the i-th blocking occurs, and α is the adjustment factor used to nonlinearly enhance the impact of extreme cases; After obtaining the cumulative impact factor, it is mapped to the task blocking quantification value. The task adaptability factor is introduced and the cumulative impact factor is normalized to ensure the comparability of the blocking index of different tasks. The task blocking quantification value is defined as follows: Among them: B f is the task blocking quantization value, T af is the task adaptability factor, which reflects the tolerance of the task to blocking. β is the adaptability adjustment factor, which is used to balance the task blocking quantization value range of different task categories. γ is the smoothing factor.
8. The method for unified scheduling of heterogeneous computing power based on global power planning according to claim 3 is characterized in that: During the monitoring period, the specific steps for deeply analyzing the proportion of time that the subtask is forced to postpone execution by the operating system to generate the backoff quantization value are as follows: During the monitoring period, the total time of the subtask being forced to postpone execution due to the operating system and the total time of the subtask actually getting the processor and being successfully executed are recorded, respectively as f d and r t At the same time, the number of times the subtask is forcibly preempted during this period and the total number of scheduling fragments are counted, respectively recorded as p c and s c ; Introducing two key ratios R x and P x , accurately quantify the degree of subtask constraints and resource competition: Where: R x Indicates the extent to which the subtask is postponed by the operating system during the monitoring period; P x It reflects the intensity of forced preemption encountered by the subtask; Based on the degree to which the subtask is postponed by the operating system during the monitoring period R x and the density P of subtasks that encounter forced preemption x , a nonlinear fusion method is used to construct the calculation backoff quantization value to characterize the load status and degree of limitation of the subtask. The construction expression for the calculation backoff quantization value is: Among them: back off To calculate the backoff quantization value, To delay the stress-enhancing factor, To seize the interference enhancement factor, R x ·P x is the interaction factor used to capture the synergistic effect of postponement and preemption. λ, μ, and η are the weight coefficients of the postponement pressure enhancement factor, the preemption interference enhancement factor, and the interaction factor, respectively.