Dynamic computing power scheduling method for distributed heterogeneous nodes

By collecting and analyzing data and computing power status in distributed heterogeneous nodes in real time, calculating the matching scores between tasks and nodes, combining data localization and task type adaptation strategies, efficient scheduling and resource utilization of distributed systems are achieved, and the problems of waste of resources and inefficient task execution in the existing technology are solved.

CN120066808AInactive Publication Date: 2025-05-30ZHONGLIAN YUNGANG DATA TECH CO LTD

Patent Information

Application Number
CN202510562202.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-30
Publication Date
2025-05-30
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art lacks effective considerations for the dynamic correlation between data and computing power in the computing power scheduling of distributed heterogeneous nodes, resulting in waste of resources and inefficient task execution.

Method used

The dynamic scheduling method of distributed heterogeneous nodes is adopted to collect the computing power resource status of nodes, task data characteristics and network transmission status data in real time, calculate the matching scores between tasks and nodes, and generate a scheduling plan based on data localization rules and task type adaptation strategies.

Benefits of technology

Through deep collaborative data characteristics and computing power status, efficient scheduling and resource utilization of distributed systems are achieved, resource waste is avoided, and task execution efficiency and resource utilization efficiency are significantly improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120066808A_ABST
    Figure CN120066808A_ABST
Patent Text Reader

Abstract

The invention provides a computing power dynamic scheduling method for distributed heterogeneous nodes, relates to the technical field of computing power scheduling, and aims to solve the problems of resource waste and low task execution efficiency. According to the method, the matching score between the task and the node is calculated, and a real-time feedback optimization mechanism is combined, so that deep collaboration of the data feature and the computing power state is realized, the scheduling efficiency and the resource utilization efficiency of the distributed system are effectively improved, and firstly, the multi-dimensional information of the data is acquired in real time; according to the method, the access characteristics of the data end are quantitatively associated with the heterogeneous performance and the load state of the computing power end to form a unified scheduling matching index, the problem that data storage and computing power distribution are mutually separated in a traditional scheme is solved, the weight of data volume, geographic distance and node performance factors is dynamically adjusted according to different task types of computation-intensive and IO-intensive tasks, and the scheduling efficiency is improved. Extensive distribution of heterogeneous resources is avoided, and the overall operation efficiency and resource utilization benefits of the system are remarkably improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computing power scheduling, and particularly to a method for dynamically scheduling the computing power of distributed heterogeneous nodes. Background Art

[0002] In the field of distributed computing, with the wide application of heterogeneous computing resources (such as CPUs, GPUs, FPGAs, etc.), computing power scheduling has become the key to improving system performance and resource utilization. Existing computing power scheduling methods mainly rely on static rules or simple load balancing strategies, and tasks are assigned to different nodes through preset priorities or resource allocation ratios. To a certain extent, these methods can achieve task assignment and execution, but they have obvious limitations when dealing with complex and changing computing environments.

[0003] The main defect of the existing technology is the lack of effective consideration of the dynamic association between data and computing power. Traditional scheduling methods usually regard data storage and computing power allocation as independent processes, and do not fully consider factors such as data access heat, transmission cost, and the heterogeneous performance of nodes for comprehensive scheduling. This leads to situations where hot data may be matched with low-computing-power nodes, and cold data occupies high-computing-power resources in practical applications, resulting in waste of resources and low task execution efficiency. Summary of the Invention

[0004] The technical problem to be solved by the present invention is to overcome the defects of waste of resources and low task execution efficiency existing in the prior art. The present invention proposes a method for dynamically scheduling the computing power of distributed heterogeneous nodes.

[0005] To solve the above technical problem, the technical solution adopted by the present invention is a method for dynamically scheduling the computing power of distributed heterogeneous nodes, including the following steps: S1. Data collection and preprocessing: Real-time collect the computing power resource status data, task data characteristics, and network transmission status data of heterogeneous nodes through Prometheus, nvidia-smi, and IP location API, and perform noise filtering, outlier correction, and normalization processing; S2. Feature extraction and model construction: Based on the data collected in step S1, calculate the node load characteristics , heterogeneous performance coefficient , data access heat , and transmission cost , and calculate the matching score of task and node through the following formula, where is the data set associated with task , and is a single data block in set ; is an extremely small constant used to avoid a zero denominator;​ Indicates the access popularity of a data block, reflecting the frequency and importance of access to the data block; The in is the load characteristic of node reflecting the current idle degree of the node, where a larger value indicates that the node is more idle; is the heterogeneous performance coefficient of node reflecting the computing and processing capabilities of the node; is the transmission cost of data block including costs brought by factors such as data size and transmission distance; is the data localization priority bonus item, used to encourage tasks to be preferentially assigned to nodes storing relevant data; is the task type adaptation weight, which adjusts the matching score according to whether the task is computationally intensive or I / O intensive; S3. Scheduling decision: Based on the matching score calculated by the dynamic affinity model in step S2 and combined with the data localization rule and task type adaptation strategy to generate a scheduling plan; S4. Task execution: Use the Kubernetes container orchestration system to deploy the task to the target node and monitor the task execution status and node resource consumption in real time; S5. Feedback optimization: Dynamically adjust the parameters in the model according to task execution latency, resource utilization, and metrics to optimize subsequent scheduling strategies. Furthermore, the computing power resource status data includes CPU utilization rate, GPU utilization rate, memory occupancy rate, CPU main frequency, number of cores, and network bandwidth; the task data characteristics include the number of accesses, update times, data size, and storage location of the data block; the network transmission status is the geographical distance latency and real-time bandwidth occupancy between data nodes.

[0006] Furthermore, the node load characteristic

[0007] is calculated by the following formula: where is the CPU occupancy rate, is the number of CPU cores, is the memory occupancy rate, is the total memory. is the total memory.

[0008] Furthermore, the heterogeneous performance coefficient is calculated by the following formula: where is the main frequency, is the benchmark CPU performance, is the real-time network bandwidth, is the benchmark network bandwidth.

[0009] Furthermore, the data access popularity is calculated by the following formula: , where is the business type weight, is the number of accesses to the data block in the most recent 1 minute, the number of updates to the data block in the most recent 1 minute.

[0010] Furthermore, the data transmission cost is calculated by the following formula: , where is the weight of data volume and geographical distance, is the data size, the geographical distance delay between the data storage node and the computing power node.

[0011] Furthermore, the has the following value-taking logic: If ≥ 80% of the data in the data set required for the node's local storage task , takes , being the maximum value of the matching scores of all current candidate nodes.

[0012] Furthermore, for compute-intensive tasks, takes ; for I / O-intensive tasks, takes .

[0013] Furthermore, the scheduling decision rule includes: preferentially selecting the node with the highest matching score , and for the first scheduling scenario of cold data, degrading to a computing power priority strategy and selecting the node with the largest ; the feedback optimization step includes: if the task execution delay exceeds 1.5 times the expected value and is higher than the cluster average, increasing the coefficient of this data-node pair by 0.1; if a node is selected 5 times in a row and its load , reducing its load penalty weight to 90% of the original weight.

[0014] Furthermore, the heterogeneous nodes include cloud servers, edge devices, GPU / TPU dedicated computing power nodes, and supercomputing centers; the data access records are updated in real time through a sliding window with a window size of 10 minutes to ensure that only reflects the recent data access trend; real-time monitoring and visualization of resource status are achieved through Prometheus-Grafana, and task containerized deployment and elastic scaling are achieved through Kubernetes. ​

[0015] Compared with the prior art, the beneficial effects of the present invention include: by calculating the matching score between computing tasks and nodes and combining a real-time feedback optimization mechanism, deep coordination between data features and computing power status is achieved, effectively improving the scheduling efficiency and resource utilization efficiency of the distributed system. This method first collects multi-dimensional information such as data access popularity, data size, storage location, and CPU / GPU utilization rate and network bandwidth of nodes in real time, quantitatively correlates the access characteristics at the data end with the heterogeneous performance and load status at the computing power end to form a unified scheduling matching index, solving the problem of the mutual disconnection between data storage and computing power allocation in traditional solutions. During the scheduling decision-making process, for different task types such as compute-intensive and IO-intensive, the weights of factors such as data volume, geographical distance, and node performance are dynamically adjusted, enabling compute-intensive tasks to preferentially match high-performance computing power nodes and IO-intensive tasks to preferentially select nodes with a high degree of data localization, avoiding the extensive allocation of heterogeneous resources and significantly improving the overall operation efficiency and resource utilization benefit of the system. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] The disclosure of the present invention will be described with reference to the accompanying drawings. It should be understood that the drawings are only for illustrative purposes and are not intended to limit the scope of protection of the present invention. In the drawings, the same reference numerals are used to refer to the same components. Among them: Figure 1 Schematically shows a method step diagram of a method for dynamically scheduling the computing power of distributed heterogeneous nodes according to an embodiment of the present invention; Figure 2 Schematically shows a flowchart of the operation of a method for dynamically scheduling the computing power of distributed heterogeneous nodes according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0017] It is easy to understand that according to the technical solution of the present invention, without changing the essence of the present invention, those of ordinary skill in the art can propose various structural ways and implementation ways that can be mutually replaced. Therefore, the following detailed embodiments and the accompanying drawings are only illustrative descriptions of the technical solution of the present invention and should not be regarded as the whole of the present invention or as a limitation or restriction on the technical solution of the present invention.

[0018] According to the embodiments of the present invention in combination with Figure 1 - Figure 2Shown. The computing power dynamic scheduling method for distributed heterogeneous nodes includes the following steps: S1. Data collection and preprocessing: Real-time collection and processing of the following data through Prometheus, nvidia-smi, and IP location API: Computing power resource status: CPU utilization rate, GPU utilization rate, memory occupancy rate, CPU main frequency, number of cores, network bandwidth, obtained by Prometheus polling nodes every 10 seconds, and nvidia-smi obtains GPU-specific metrics; Task data characteristics: Number of data block accesses in the last 1 minute, number of updates in the last 1 minute, data size, and node ID of the storage location, obtained through task metadata parsing and data access log statistics; Network transmission status: The geographical distance delay between nodes is calculated by the GIS distance through the IP location API, and the real-time bandwidth occupancy is obtained through the Prometheus network monitoring module.

[0019] S2. Feature extraction and model construction: Based on the data collected in step S1, calculate the node load characteristics , heterogeneous performance coefficient , data access heat and transmission cost , reflecting the node resource occupancy, calculation formula: , where is the CPU occupancy rate, is the number of CPU cores, is the memory occupancy rate, is the total memory; Quantify the node computing and transmission capabilities, based on the cluster average performance: , where is the main frequency, is the reference CPU performance, is the real-time network bandwidth, is the reference network bandwidth; Reflects the frequency and importance of data block access: , where is the business type weight, is the number of data block accesses in the last 1 minute, is the number of data block updates in the last 1 minute; The data transmission cost is calculated by the following formula: , where is the data volume and geographical distance weight, is the data size, is the geographical distance delay between the data storage node and the computing power node; Calculate the matching score of task and node : Calculate task and node and node , where is the task associated data set, is a single data block in the set ; is a very small constant used to avoid a zero denominator; reflects the current idle degree of the node, and the larger the value, the more idle the node; is the data locality preference bonus item, used to encourage tasks to be preferentially assigned to nodes storing relevant data; is the task type adaptation weight, which adjusts the matching score according to whether the task is compute-intensive or I / O-intensive; the takes the following value logic: if the node locally stores ≥ 80% of the data in the data set required by the task, take , the maximum value of the matching scores of all current candidate nodes, the for compute-intensive tasks, take ; for I / O-intensive tasks, take .

[0020] S3. Scheduling decision: Based on the matching scores calculated by the dynamic affinity model in step S2 combined with the data locality rule and the task type adaptation strategy to generate a scheduling plan, and preferentially select the node with the highest matching score. For the first scheduling scenario of cold data, it degrades to a computing power priority strategy and selects the node with the largest computing power.

[0021] S4. Task execution: Use the Kubernetes container orchestration system to deploy the task to the target node and monitor the task execution status and node resource consumption in real time.

[0022] S5. Feedback optimization: Dynamically adjust the parameters in the model according to the task execution delay, resource utilization rate, and metrics to optimize the subsequent scheduling strategy. The feedback optimization steps include: if the task execution delay exceeds 1.5 times the expected value and is higher than the cluster average, increase the coefficient of this data-node pair by 0.1; if the node is continuously selected 5 times and the load , reduce its load penalty weight to 90% of the original weight.

[0023] Heterogeneous nodes include cloud servers, edge devices, GPU / TPU dedicated computing power nodes, and supercomputing centers; data access records are updated in real time through a sliding window with a window size of 10 minutes to ensure only reflecting the recent data access trends; real-time monitoring and visualization of resource status are achieved through Prometheus-Grafana, and task containerized deployment and elastic scaling are achieved through Kubernetes.

[0024] The technical scope of the present invention is not limited to the content described above. Those skilled in the art can make various deformations and modifications to the above embodiments without departing from the technical idea of the present invention, and these deformations and modifications should all fall within the protection scope of the present invention.

Claims

1. A method for dynamically scheduling computing power of distributed heterogeneous nodes, characterized in that: The following steps are involved: S1. Data collection and preprocessing: real-time collection of computing resource status data, task data characteristics and network transmission status data of heterogeneous nodes, and noise filtering, outlier correction and normalization processing; S2. Feature extraction and model construction: based on the data collected in step S1, calculate the node load characteristics , Heterogeneous Performance Coefficient , Data access popularity and transmission costs , and calculate the task by the following formula With Node Match score ,in For the task A collection of related data, For collection A single data block in ; is a very small constant; Represents a data block The popularity of visits; In For Node The load characteristics, Reflects the node The current idleness. The larger the value, the more idle the node is. Is a node The heterogeneous performance coefficient reflects the computing and processing capabilities of the node; For data blocks The transmission cost includes the cost caused by factors such as data size and transmission distance; Prioritize data locality to encourage tasks to be assigned to nodes that store relevant data. Adapt the weights for the task types and adjust the matching scores according to whether the tasks are compute-intensive or IO-intensive; S3. Scheduling decision: Based on the matching scores calculated by the dynamic affinity model in step S2 , generate a scheduling plan by combining data localization rules and task type adaptation strategies; S4. Task execution: Use the Kubernetes container orchestration system to deploy tasks to the target node, and monitor the task execution status and node resource consumption in real time; S5. Feedback optimization: Dynamically adjust the parameters in the model according to task execution delay, resource utilization and indicators to optimize subsequent scheduling strategies.

2. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 1, characterized in that: The computing resource status data covers CPU utilization, GPU utilization, memory occupancy, CPU main frequency, number of cores, and network bandwidth; the task data characteristics include the number of data block accesses, the number of updates, the data size, and the storage location; the network transmission status is the geographical distance delay between data nodes and the real-time bandwidth occupancy.

3. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The node load characteristics Calculated by the following formula: ,in is the CPU usage, is the number of CPU cores, is the memory usage, The total amount of memory.

4. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The heterogeneous performance coefficient Calculated by the following formula: ,in is the main frequency, To benchmark CPU performance, is the real-time network bandwidth, The baseline network bandwidth.

5. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The data access popularity Calculated by the following formula: ,in Business type weight, The number of accesses to the data block in the last minute. The number of times the data block has been updated in the last minute.

6. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: The data transfer cost Calculated by the following formula: ,in is the weight of data volume and geographical distance, is the data size, It is the geographical distance delay between the data storage node and the computing power node.

7. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 2, characterized in that: In S2 The value logic is: if the node locally stores the data set required by the task ≥80% of the data, Pick , The maximum matching score of all current candidate nodes.

8. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 1, characterized in that: Said For computationally intensive tasks, Pick ; For IO-intensive tasks, Pick .

9. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to claim 1, characterized in that: The scheduling decision rule includes: giving priority to matching scores The highest node, for the first scheduling scenario of cold data, degenerates to the computing power priority strategy and selects The largest node; the feedback optimization step in S5 includes: if the task execution delay exceeds the expected value by 1.5 times and Higher than the cluster mean, the The coefficient increases by 0.1; if the node is selected 5 times in a row and the load , reducing its load penalty weight to 90% of the original weight.

10. The method for dynamically scheduling computing power of distributed heterogeneous nodes according to any one of claims 1 to 9, characterized in that: The heterogeneous nodes include cloud servers, edge devices, GPU / TPU dedicated computing nodes and supercomputing centers; data access records are updated in real time through sliding windows with a window size of 10 minutes to ensure It only reflects recent data access trends; Real-time monitoring and visualization of resource status is achieved through Prometheus-Grafana, and task container deployment and elastic scaling are achieved through Kubernetes.

Citation Information

Patent Citations

  • Cloud rendering resource monitoring method and system based on computing power scheduling

    CN118093326A

  • Distributed task scheduling method and system for heterogeneous tasks based on cloud native, and medium

    CN119759594A

  • Dynamic scheduling system and method for container calculation

    CN119862037A

  • KR20220006490A

Cited By

  • Multi-level forest resource monitoring method and system

    CN120258471A

  • A multi-level forest resource monitoring method and system

    CN120258471B

  • GPU (Graphic Processing Unit) and parallel IO (Input / Output) collaborative optimization method based on mode heterogeneous calculation

    CN120295803A

  • GPU and parallel IO collaborative optimization method based on heterogeneous computing

    CN120295803B

  • Heterogeneous intelligent computing power optimization management scheduling system for accelerating large model reasoning task

    CN120743517A