Dynamic computing power scheduling method for distributed heterogeneous nodes
By collecting and analyzing data and computing power status in distributed heterogeneous nodes in real time, calculating the matching scores between tasks and nodes, combining data localization and task type adaptation strategies, efficient scheduling and resource utilization of distributed systems are achieved, and the problems of waste of resources and inefficient task execution in the existing technology are solved.
Patent Information
- Application Number
- CN202510562202.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The prior art lacks effective considerations for the dynamic correlation between data and computing power in the computing power scheduling of distributed heterogeneous nodes, resulting in waste of resources and inefficient task execution.
The dynamic scheduling method of distributed heterogeneous nodes is adopted to collect the computing power resource status of nodes, task data characteristics and network transmission status data in real time, calculate the matching scores between tasks and nodes, and generate a scheduling plan based on data localization rules and task type adaptation strategies.
Through deep collaborative data characteristics and computing power status, efficient scheduling and resource utilization of distributed systems are achieved, resource waste is avoided, and task execution efficiency and resource utilization efficiency are significantly improved.
Smart Images

Figure CN120066808A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of computing power scheduling, and particularly to a method for dynamically scheduling the computing power of distributed heterogeneous nodes. Background Art
[0002] In the field of distributed computing, with the wide application of heterogeneous computing resources (such as CPUs, GPUs, FPGAs, etc.), computing power scheduling has become the key to improving system performance and resource utilization. Existing computing power scheduling methods mainly rely on static rules or simple load balancing strategies, and tasks are assigned to different nodes through preset priorities or resource allocation ratios. To a certain extent, these methods can achieve task assignment and execution, but they have obvious limitations when dealing with complex and changing computing environments.
[0003] The main defect of the existing technology is the lack of effective consideration of the dynamic association between data and computing power. Traditional scheduling methods usually regard data storage and computing power allocation as independent processes, and do not fully consider factors such as data access heat, transmission cost, and the heterogeneous performance of nodes for comprehensive scheduling. This leads to situations where hot data may be matched with low-computing-power nodes, and cold data occupies high-computing-power resources in practical applications, resulting in waste of resources and low task execution efficiency. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the defects of waste of resources and low task execution efficiency existing in the prior art. The present invention proposes a method for dynamically scheduling the computing power of distributed heterogeneous nodes.
[0005] To solve the above technical problem, the technical solution adopted by the present invention is a method for dynamically scheduling the computing power of distributed heterogeneous nodes, including the following steps: S1. Data collection and preprocessing: Real-time collect the computing power resource status data, task data characteristics, and network transmission status data of heterogeneous nodes through Prometheus, nvidia-smi, and IP location API, and perform noise filtering, outlier correction, and normalization processing; S2. Feature extraction and model construction: Based on the data collected in step S1, calculate the node load characteristics , heterogeneous performance coefficient , data access heat , and transmission cost , and calculate the matching score of task and node through the following formula, where is the data set associated with task , and is a single data block in set ; is an extremely small constant used to avoid a zero denominator; Indicates the access popularity of a data block, reflecting the frequency and importance of access to the data block; The in is the load characteristic of node reflecting the current idle degree of the node, where a larger value indicates that the node is more idle; is the heterogeneous performance coefficient of node reflecting the computing and processing capabilities of the node; is the transmission cost of data block including costs brought by factors such as data size and transmission distance; is the data localization priority bonus item, used to encourage tasks to be preferentially assigned to nodes storing relevant data; is the task type adaptation weight, which adjusts the matching score according to whether the task is computationally intensive or I / O intensive; S3. Scheduling decision: Based on the matching score calculated by the dynamic affinity model in step S2 and combined with the data localization rule and task type adaptation strategy to generate a scheduling plan; S4. Task execution: Use the Kubernetes container orchestration system to deploy the task to the target node and monitor the task execution status and node resource consumption in real time; S5. Feedback optimization: Dynamically adjust the parameters in the model according to task execution latency, resource utilization, and metrics to optimize subsequent scheduling strategies. Furthermore, the computing power resource status data includes CPU utilization rate, GPU utilization rate, memory occupancy rate, CPU main frequency, number of cores, and network bandwidth; the task data characteristics include the number of accesses, update times, data size, and storage location of the data block; the network transmission status is the geographical distance latency and real-time bandwidth occupancy between data nodes.
[0006] Furthermore, the node load characteristic
[0007] is calculated by the following formula: where is the CPU occupancy rate, is the number of CPU cores, is the memory occupancy rate, is the total memory. is the total memory.
[0008] Furthermore, the heterogeneous performance coefficient is calculated by the following formula: where is the main frequency, is the benchmark CPU performance, is the real-time network bandwidth, is the benchmark network bandwidth.
[0009] Furthermore, the data access popularity is calculated by the following formula: , where is the business type weight, is the number of accesses to the data block in the most recent 1 minute, the number of updates to the data block in the most recent 1 minute.
[0010] Furthermore, the data transmission cost is calculated by the following formula: , where is the weight of data volume and geographical distance, is the data size, the geographical distance delay between the data storage node and the computing power node.
[0011] Furthermore, the has the following value-taking logic: If ≥ 80% of the data in the data set required for the node's local storage task , takes , being the maximum value of the matching scores of all current candidate nodes.
[0012] Furthermore, for compute-intensive tasks, takes ; for I / O-intensive tasks, takes .
[0013] Furthermore, the scheduling decision rule includes: preferentially selecting the node with the highest matching score , and for the first scheduling scenario of cold data, degrading to a computing power priority strategy and selecting the node with the largest ; the feedback optimization step includes: if the task execution delay exceeds 1.5 times the expected value and is higher than the cluster average, increasing the coefficient of this data-node pair by 0.1; if a node is selected 5 times in a row and its load , reducing its load penalty weight to 90% of the original weight.
[0014] Furthermore, the heterogeneous nodes include cloud servers, edge devices, GPU / TPU dedicated computing power nodes, and supercomputing centers; the data access records are updated in real time through a sliding window with a window size of 10 minutes to ensure that only reflects the recent data access trend; real-time monitoring and visualization of resource status are achieved through Prometheus-Grafana, and task containerized deployment and elastic scaling are achieved through Kubernetes.
[0015] Compared with the prior art, the beneficial effects of the present invention include: by calculating the matching score between computing tasks and nodes and combining a real-time feedback optimization mechanism, deep coordination between data features and computing power status is achieved, effectively improving the scheduling efficiency and resource utilization efficiency of the distributed system. This method first collects multi-dimensional information such as data access popularity, data size, storage location, and CPU / GPU utilization rate and network bandwidth of nodes in real time, quantitatively correlates the access characteristics at the data end with the heterogeneous performance and load status at the computing power end to form a unified scheduling matching index, solving the problem of the mutual disconnection between data storage and computing power allocation in traditional solutions. During the scheduling decision-making process, for different task types such as compute-intensive and IO-intensive, the weights of factors such as data volume, geographical distance, and node performance are dynamically adjusted, enabling compute-intensive tasks to preferentially match high-performance computing power nodes and IO-intensive tasks to preferentially select nodes with a high degree of data localization, avoiding the extensive allocation of heterogeneous resources and significantly improving the overall operation efficiency and resource utilization benefit of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them: Figure 1 Schematically shows a method step diagram of a method for dynamically scheduling the computing power of distributed heterogeneous nodes according to an embodiment of the present invention; Figure 2 Schematically shows a flowchart of the operation of a method for dynamically scheduling the computing power of distributed heterogeneous nodes according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0017] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various structural ways and implementation ways that can be mutually replaced. Therefore, the following detailed embodiments and the accompanying drawings are only illustrative descriptions of the technical solution of the present invention and should not be regarded as the whole of the present invention or as a limitation or restriction on the technical solution of the present invention.
[0018] According to the embodiments of the present invention in combination with Figure 1 - Figure 2Shown. The computing power dynamic scheduling method for distributed heterogeneous nodes includes the following steps: S1. Data collection and preprocessing: Real-time collection and processing of the following data through Prometheus, nvidia-smi, and IP location API: Computing power resource status: CPU utilization rate, GPU utilization rate, memory occupancy rate, CPU main frequency, number of cores, network bandwidth, obtained by Prometheus polling nodes every 10 seconds, and nvidia-smi obtains GPU-specific metrics; Task data characteristics: Number of data block accesses in the last 1 minute, number of updates in the last 1 minute, data size, and node ID of the storage location, obtained through task metadata parsing and data access log statistics; Network transmission status: The geographical distance delay between nodes is calculated by the GIS distance through the IP location API, and the real-time bandwidth occupancy is obtained through the Prometheus network monitoring module.
[0019] S2. Feature extraction and model construction: Based on the data collected in step S1, calculate the node load characteristics , heterogeneous performance coefficient , data access heat and transmission cost , reflecting the node resource occupancy, calculation formula: , where is the CPU occupancy rate, is the number of CPU cores, is the memory occupancy rate, is the total memory; Quantify the node computing and transmission capabilities, based on the cluster average performance: , where is the main frequency, is the reference CPU performance, is the real-time network bandwidth, is the reference network bandwidth; Reflects the frequency and importance of data block access: , where is the business type weight, is the number of data block accesses in the last 1 minute, is the number of data block updates in the last 1 minute; The data transmission cost is calculated by the following formula: , where is the data volume and geographical distance weight, is the data size, is the geographical distance delay between the data storage node and the computing power node; Calculate the matching score of task and node : Calculate task and node and node , where is the task associated data set, is a single data block in the set ; is a very small constant used to avoid a zero denominator; reflects the current idle degree of the node, and the larger the value, the more idle the node; is the data locality preference bonus item, used to encourage tasks to be preferentially assigned to nodes storing relevant data; is the task type adaptation weight, which adjusts the matching score according to whether the task is compute-intensive or I / O-intensive; the takes the following value logic: if the node locally stores ≥ 80% of the data in the data set required by the task, take , the maximum value of the matching scores of all current candidate nodes, the for compute-intensive tasks, take ; for I / O-intensive tasks, take .
[0020] S3. Scheduling decision: Based on the matching scores calculated by the dynamic affinity model in step S2 combined with the data locality rule and the task type adaptation strategy to generate a scheduling plan, and preferentially select the node with the highest matching score. For the first scheduling scenario of cold data, it degrades to a computing power priority strategy and selects the node with the largest computing power.
[0021] S4. Task execution: Use the Kubernetes container orchestration system to deploy the task to the target node and monitor the task execution status and node resource consumption in real time.
[0022] S5. Feedback optimization: Dynamically adjust the parameters in the model according to the task execution delay, resource utilization rate, and metrics to optimize the subsequent scheduling strategy. The feedback optimization steps include: if the task execution delay exceeds 1.5 times the expected value and is higher than the cluster average, increase the coefficient of this data-node pair by 0.1; if the node is continuously selected 5 times and the load , reduce its load penalty weight to 90% of the original weight.
[0023] Heterogeneous nodes include cloud servers, edge devices, GPU / TPU dedicated computing power nodes, and supercomputing centers; data access records are updated in real time through a sliding window with a window size of 10 minutes to ensure only reflecting the recent data access trends; real-time monitoring and visualization of resource status are achieved through Prometheus-Grafana, and task containerized deployment and elastic scaling are achieved through Kubernetes.
[0024] The technical scope of the present invention is not limited to the content described above. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.
Claims
1. A method for dynamically scheduling computing power of distributed heterogeneous nodes, characterized in that: The following steps are involved: S1. Data collection and preprocessing: real-time collection of computing resource status data, task data characteristics and network transmission status data of heterogeneous nodes, and noise filtering, outlier correction and normalization processing; S2. Feature extraction and model construction: based on the data collected in step S1, calculate the node load characteristics , Heterogeneous Performance Coefficient , Data access popularity and transmission costs , and calculate the task by the following formula With Node Match score ,in For the task A collection of related data, For collection A single data block in ; is a very small constant; Represents a data block The popularity of visits; In For Node The load characteristics, Reflects the node The current idleness. The larger the value, the more idle the node is. Is a node The heterogeneous performance coefficient reflects the computing and processing capabilities of the node; For data blocks The transmission cost includes the cost caused by factors such as data size and transmission distance; Prioritize data locality to encourage tasks to be assigned to nodes that store relevant data. Adapt the weights for the task types and adjust the matching scores according to whether the tasks are compute-intensive or IO-intensive; S3. Scheduling decision: Based on the matching scores calculated by the dynamic affinity model in step S2 , generate a scheduling plan by combining data localization rules and task type adaptation strategies; S4. Task execution: Use the Kubernetes container orchestration system to deploy tasks to the target node, and monitor the task execution status and node resource consumption in real time; S5. Feedback optimization: Dynamically adjust the parameters in the model according to task execution delay, resource utilization and indicators to optimize subsequent scheduling strategies.
2. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 1, characterized in that: The computing resource status data covers CPU utilization, GPU utilization, memory occupancy, CPU main frequency, number of cores, and network bandwidth; the task data characteristics include the number of data block accesses, the number of updates, the data size, and the storage location; the network transmission status is the geographical distance delay between data nodes and the real-time bandwidth occupancy.
3. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The node load characteristics Calculated by the following formula: ,in is the CPU usage, is the number of CPU cores, is the memory usage, The total amount of memory.
4. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The heterogeneous performance coefficient Calculated by the following formula: ,in is the main frequency, To benchmark CPU performance, is the real-time network bandwidth, The baseline network bandwidth.
5. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The data access popularity Calculated by the following formula: ,in Business type weight, The number of accesses to the data block in the last minute. The number of times the data block has been updated in the last minute.
6. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The data transfer cost Calculated by the following formula: ,in is the weight of data volume and geographical distance, is the data size, It is the geographical distance delay between the data storage node and the computing power node.
7. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: In S2 The value logic is: if the node locally stores the data set required by the task ≥80% of the data, Pick , The maximum matching score of all current candidate nodes.
8. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 1, characterized in that: Said For computationally intensive tasks, Pick ; For IO-intensive tasks, Pick .
9. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 1, characterized in that: The scheduling decision rule includes: giving priority to matching scores The highest node, for the first scheduling scenario of cold data, degenerates to the computing power priority strategy and selects The largest node; the feedback optimization step in S5 includes: if the task execution delay exceeds the expected value by 1.5 times and Higher than the cluster mean, the The coefficient increases by 0.1; if the node is selected 5 times in a row and the load , reducing its load penalty weight to 90% of the original weight.
10. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to any one of claims 1 to 9, characterized in that: The heterogeneous nodes include cloud servers, edge devices, GPU / TPU dedicated computing nodes and supercomputing centers; data access records are updated in real time through sliding windows with a window size of 10 minutes to ensure It only reflects recent data access trends; Real-time monitoring and visualization of resource status is achieved through Prometheus-Grafana, and task container deployment and elastic scaling are achieved through Kubernetes.
Citation Information
Patent Citations
Cloud rendering resource monitoring method and system based on computing power scheduling
CN118093326A
Distributed task scheduling method and system for heterogeneous tasks based on cloud native, and medium
CN119759594A
Dynamic scheduling system and method for container calculation
CN119862037A
KR20220006490A
Cited By
Multi-level forest resource monitoring method and system
CN120258471A
A multi-level forest resource monitoring method and system
CN120258471B
GPU (Graphic Processing Unit) and parallel IO (Input / Output) collaborative optimization method based on mode heterogeneous calculation
CN120295803A
GPU and parallel IO collaborative optimization method based on heterogeneous computing
CN120295803B
Heterogeneous intelligent computing power optimization management scheduling system for accelerating large model reasoning task
CN120743517A