Distributed resource dynamic scheduling system based on adaptive load balancing

CN122802505APending Publication Date: 2026-09-22CHINA SOUTHERN POWER GRID COMPREHENSIVE ENERGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610801854.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-06-04
Publication Date
2026-09-22

AI Technical Summary

Technical Problem

然而,这些方法存在明显的局限性:一方面,多数调度系统未充分考虑边缘计算的地理分布特性,任务与节点间的地理位置关系往往被忽略,导致网络传输延迟较大;另一方面,计算节点的资源状态是动态变化的,静态调度策略无法实时适应这种变化,容易导致某些节点过载,而其他节点闲置

Benefits of technology

通过节点资源状态标记模块实时采集并评估CPU、内存、磁盘I/O、网络带宽等多维度资源数据,动态地将节点标记为空闲、正常、繁忙、过载等状态,并结合任务数据获取模块赋予任务的权重值,使得任务分配模块能够进行精细化、智能化的调度决策。这避免部分节点过载而其他节点闲置的资源利用不均问题,实现负载在各边缘网关节点间的自适应均衡,从而显著提高整个分布式系统的资源利用率和整体处理能力;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122802505A_ABST
    Figure CN122802505A_ABST
Patent Text Reader

Abstract

The application discloses a distributed resource dynamic scheduling system based on adaptive load balancing and relates to the technical field of computer networks, and comprises the following: a task data acquisition module is used for acquiring task data including physical geographic position information of tasks and assigning a weight value to each task according to the task data; a node resource state marking module is used for collecting physical geographic position information and resource state data of each edge gateway node in a distributed system and marking the resource state of each edge gateway node according to the resource state data; and a task allocation module generates a task allocation strategy according to the physical geographic position information of tasks, the physical geographic position information of edge gateway nodes, the resource state data and the weight value and allocates tasks to edge gateway nodes according to the task allocation strategy. The application can effectively avoid node overload or idling through a load balancing strategy, thereby improving the stability and response speed of the system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer network technology, and in particular to a distributed resource dynamic scheduling system based on adaptive load balancing. Background Technology

[0002] In modern distributed computing environments, with the continuous increase in the number and complexity of tasks, how to efficiently schedule computing tasks and rationally allocate resources has become a critical issue. Traditional resource scheduling methods mainly rely on static rules, such as allocation based on task priority or node load rate. However, these methods have obvious limitations: on the one hand, most scheduling systems do not fully consider the geographical distribution characteristics of edge computing, and the geographical relationship between tasks and nodes is often ignored, resulting in large network transmission latency; on the other hand, the resource status of computing nodes changes dynamically, and static scheduling strategies cannot adapt to this change in real time, which can easily lead to some nodes being overloaded while others are idle. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a distributed resource dynamic scheduling system based on adaptive load balancing.

[0004] This invention relates to a distributed resource dynamic scheduling system based on adaptive load balancing, comprising the following: Task data acquisition module (11): used to acquire task data including the physical geographical location information of the task, and assign weight values ​​to each task according to the task data; Node resource status marking module (12): used to collect the physical geographical location information and resource status data of each edge gateway node in the distributed system, and mark the resource status of each edge gateway node according to the resource status data; Task allocation module (13): Generates a task allocation strategy based on the physical geographic location information of the task, the physical geographic location information of the edge gateway node, the resource status data and the weight value, and allocates the task to the edge gateway node according to the task allocation strategy. Task progress monitoring module (14): Used to monitor the execution status data of each task in real time, including task start time, execution progress, task end time and resource usage; Fault handling module (15): Used to handle faults when a task execution fault or an edge gateway node fault is detected.

[0005] Preferably, the task data acquisition module (11) includes, Task feature extraction unit (111): used to extract task data from task requests, including task data volume, memory requirements, task execution time and task priority, as well as the physical geographical location information of the task; Weight allocation unit (112): Assigns a weight value to each task based on the task data.

[0006] Preferably, the node resource status marking module (12) includes, Resource data acquisition unit (121): used to collect physical location information and resource status data of each edge gateway node in the distributed system in real time. The resource status data includes CPU utilization, disk I / O, memory usage, network bandwidth and storage utilization. Resource status assessment unit (122): It is used to set corresponding resource status thresholds according to the hardware configuration parameters of each edge gateway node, including the number of CPU cores, memory capacity, disk capacity and network bandwidth limit, and mark the resource status of each node according to the assessment results. The resource status includes idle status, normal status, busy status, overload status and fault status. Resource status update unit (123): It is used to dynamically update the resource status of each edge gateway node according to the real-time data of the resource data acquisition unit (121), and pass the updated resource status data to the task allocation module (13) to support task allocation decision. Node anomaly handling unit (124): If one or more edge gateway nodes are marked as faulty or overloaded, the fault handling module (15) will be triggered to handle the fault.

[0007] Preferably, the task allocation module (13) includes, Candidate node filtering unit (131): used to filter out candidate nodes suitable for executing tasks based on the physical geographical location information of the task, the physical geographical location information of the node, the task weight value and resource status data; Task allocation unit (132): used to allocate tasks to candidate nodes in order of their weight values; Update unit (133): Used to update the resource status data of the edge gateway node and transmit the updated data to the node resource status marking module.

[0008] Preferably, the task progress monitoring module (14) includes, Task status acquisition unit (141): used to collect the execution status data of each task in real time, including task start time, execution progress, task end time and resource usage, including CPU usage, memory usage, disk space usage and network bandwidth. Task progress evaluation unit (142): Based on the execution status data, evaluate the execution progress of the task and calculate the remaining execution time of the task; Resource occupancy status transmission unit (143): used to transmit the resource occupancy status to the node resource status marking module to support dynamic updates of resource status; Task execution anomaly detection unit (144): used to detect abnormal situations during task execution, including task execution timeout, abnormal resource usage or task execution failure, and trigger the fault handling module (15) to handle the situation when an anomaly is detected; Task status update unit (145): Used to dynamically update the execution status of the task based on the real-time data of the task status acquisition unit (141), and pass the updated status data to the task allocation module (13) to support dynamic task scheduling. Task completion notification unit (146): When a task is completed, it generates a task completion notification and passes the task execution result and resource release status to the node resource status marking module (12) and the task allocation module (13).

[0009] Preferably, the fault handling module (15) includes, Abnormal data detection unit (151): used to detect abnormal data of task abnormal detection unit (144) and node abnormal processing unit (124); Task retry unit (152): used to reassign the task and retry execution when a task execution failure is detected. The retry execution includes reassigning the task to other available edge gateway nodes, adjusting the task resource requirements and then re-executing. Node recovery unit (153): used to restore the normal operation of the edge gateway node when a failure of the edge gateway node is detected. The recovery operation includes restarting the failed edge gateway node; migrating the tasks on the failed edge gateway node to other available edge gateway nodes; and releasing the resources occupied by the failed edge gateway node. Fault notification unit (154): Used to notify relevant personnel when a fault is detected and cannot be handled.

[0010] This invention discloses a distributed resource dynamic scheduling system based on adaptive load balancing, which has the following beneficial effects: The node resource status marking module collects and evaluates multi-dimensional resource data such as CPU, memory, disk I / O, and network bandwidth in real time, dynamically marking nodes as idle, normal, busy, or overloaded. Combined with the weight values ​​assigned to tasks by the task data acquisition module, this enables the task allocation module to make refined and intelligent scheduling decisions. This avoids the problem of uneven resource utilization where some nodes are overloaded while others are idle, achieving adaptive load balancing among edge gateway nodes, thereby significantly improving the resource utilization and overall processing capacity of the entire distributed system. The scheduling decision comprehensively considers the physical geographic location information of the task and the edge gateway node, tending to assign tasks to geographically proximate nodes for execution, effectively reducing the data transmission distance and time in the network. Combined with consideration of the real-time network bandwidth of the nodes, the overall latency of task execution is further reduced, improving the user experience, and making it particularly suitable for latency-sensitive edge computing applications. Attached Figure Description

[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0012] Figure 1 The system architecture diagram of the distributed resource dynamic scheduling system based on adaptive load balancing provided by the present invention; Figure 2 A unit structure diagram of the task data acquisition module provided by the present invention; Figure 3 A unit structure diagram of the node resource status marking module provided by the present invention; Figure 4 A unit structure diagram of the task allocation module provided by the present invention; Figure 5 A unit structure diagram of the task progress monitoring module provided by the present invention; Figure 6 This is a unit structure diagram of the fault handling module provided by the present invention. Detailed Implementation

[0013] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention are described clearly and completely. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0014] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.

[0015] refer to Figure 1 The distributed resource dynamic scheduling system based on adaptive load balancing provided by this invention includes the following: Task data acquisition module (11): used to acquire task data including task data volume, memory requirements, task priority and task execution time, as well as the physical geographical location information of the task, and assign weight values ​​to each task according to the task data; Node resource status marking module (12): Used to collect the physical geographical location information and resource status data of each edge gateway node in the distributed system. The resource status data includes CPU utilization, disk I / O, memory utilization, network bandwidth and storage utilization. Based on the resource status data, the resource status of each edge gateway node is marked, including idle status, normal status, busy status, overload status and fault status. Task allocation module (13): Based on the physical geographic location information of the task, the physical geographic location information of the edge gateway node, resource status data and weight value, generate a task allocation strategy and allocate the task to the edge gateway node according to the task allocation strategy. Task progress monitoring module (14): Used to monitor the execution status data of each task in real time, including task start time, execution progress, task end time and resource usage; Fault handling module (15): Used to handle faults when a task execution fault or an edge gateway node fault is detected.

[0016] The distributed resource dynamic scheduling system based on adaptive load balancing provided by this invention can collect and evaluate multi-dimensional resource data such as CPU, memory, disk I / O, and network bandwidth in real time through a node resource status marking module. It dynamically marks nodes as idle, normal, busy, or overloaded, and combines this with the weight values ​​assigned to tasks by a task data acquisition module, enabling the task allocation module to make refined and intelligent scheduling decisions. This avoids the problem of uneven resource utilization where some nodes are overloaded while others are idle, achieving adaptive load balancing among edge gateway nodes, thereby significantly improving the resource utilization and overall processing capacity of the entire distributed system. The scheduling decision comprehensively considers the physical geographic location information of the task and the edge gateway node, tending to assign tasks to geographically proximate nodes for execution, effectively reducing the data transmission distance and time in the network. Combined with consideration of the real-time network bandwidth of the nodes, the overall latency of task execution is further reduced, improving the user experience, and making it particularly suitable for latency-sensitive edge computing applications.

[0017] refer to Figure 2 In a preferred embodiment, the task data acquisition module 11 includes, Task feature extraction unit 111: used to extract task data from task requests, including task data volume, memory requirements, task execution time and task priority, and geographical information of the task; Weight allocation unit 112: Assigns weight values ​​to each task based on the task data.

[0018] Specifically, the weight value of each task is obtained by weighted summation of the task's data volume, memory requirements, priority, and execution time; the formula for calculating the weight value X is as follows: , Where D, M, T, and P represent the data volume, memory requirements, estimated execution time, and priority of the current task, respectively.

[0019] D avg M avg ,T avg These are the average values ​​of the characteristics corresponding to all tasks to be scheduled in the system, used for normalization.

[0020] L represents the geographical proximity between the task and the candidate edge gateway node. Its value can be calculated based on the geographical location attributes of the task and the node. In edge computing scenarios, proximity L is one of the key factors in optimizing task allocation.

[0021] The formula for calculating geographical proximity is as follows: , Where d represents the Euclidean geographical distance between the task and the edge gateway node. This formula satisfies the following condition: the closer the distance, the greater the proximity (the maximum value approaches 1), thus prioritizing the selection of closer nodes during scheduling. In practical applications, d can be calculated based on latitude and longitude coordinates.

[0022] α, β, γ, δ, ε are adjustable weight coefficients, where ε is the geographical proximity weight coefficient, which is usually assigned a higher value in edge gateway scheduling scenarios to prioritize task processing in the nearest location and reduce network latency.

[0023] refer to Figure 3 In a preferred embodiment, the node resource status marking module 12 includes, Resource data acquisition unit 121: used to collect physical location information and resource status data of each edge gateway node in the distributed system in real time. The resource status data includes CPU utilization, disk I / O, memory usage, network bandwidth and storage utilization. The physical location information of each edge gateway node is determined when the node is registered and remains basically unchanged during the operation cycle.

[0024] Resource status assessment unit (122): It is used to set corresponding resource status thresholds according to the hardware configuration parameters of each edge gateway node, including the number of CPU cores, memory capacity, disk capacity and network bandwidth limit, and mark the resource status of each node according to the assessment results. The resource status includes idle status, normal status, busy status, overload status and fault status. Specifically, the hardware configurations of different gateway nodes may vary, such as the number of CPU cores, memory size, disk capacity, and network bandwidth limit. This system sets differentiated thresholds for idle, normal, busy, and overloaded states for each node based on its hardware capabilities: for high-configuration nodes, a higher CPU utilization threshold, such as 90%, can be set, while for low-configuration nodes, a relatively lower threshold, such as 70%, is set, thus more accurately reflecting the actual carrying capacity of different nodes. Resource states include idle, normal, busy, overloaded, and fault states, and the specific determination methods are as follows: A node is marked as idle when its CPU utilization, memory usage, and network bandwidth utilization are all below their respective low thresholds. A node is marked as normal when any one or more of its CPU utilization, memory usage, or network bandwidth utilization are between their respective low and medium thresholds. A node is marked as busy when any one or more of its CPU utilization, memory usage, or network bandwidth utilization are between their respective medium and high thresholds. A node is marked as overloaded when any one or more of its CPU utilization, memory usage, or network bandwidth utilization exceeds its corresponding high threshold. When a node fails to respond to resource monitoring requests or experiences abnormal fluctuations in resource utilization, such as CPU utilization approaching 100% and continuing to rise, it is marked as a fault state.

[0025] Resource status update unit 123: It is used to dynamically update the resource status of each node based on the real-time data of resource data acquisition unit 121, and transmit the updated resource status data to task allocation module 13 to support task allocation decision. Node anomaly handling unit 124: If one or more nodes are marked as faulty or overloaded, the fault handling module 15 will be triggered to handle the fault.

[0026] refer to Figure 4 In a preferred embodiment, the task allocation module 13 includes, Candidate node filtering unit 131: used to filter out candidate nodes suitable for executing tasks based on task weight values, resource status data, and the geographical location attributes of tasks and nodes; Specifically, the screening process adopts a hierarchical screening strategy: First, the geographical proximity between the task and all nodes is calculated, and nodes with proximity values ​​higher than a first preset threshold are selected to form a set of neighboring nodes; then, the nodes in this set are sorted a second time according to their resource status. Based on this sorting result, the task allocation unit (132) prioritizes assigning high-weight tasks to the edge gateway node with the best geographical location and the best resource status.

[0027] Task allocation unit 132: used to allocate tasks to candidate nodes in order of their weight values; The specific allocation strategy is as follows: For tasks with weight values ​​higher than the preset high weight threshold, priority is given to assigning them to idle nodes in the "neighboring node set" to ensure low-latency execution; if there are no idle nodes in the neighboring node set, they are assigned to the node in the set that is in a normal state and has the highest proximity ranking.

[0028] For tasks whose weight values ​​are within the preset medium weight range, they are assigned to nodes in the neighboring node set that are in a normal state; if this condition is not met, they are assigned to nodes in an idle state.

[0029] Tasks with weight values ​​lower than the preset low weight threshold can be flexibly assigned to nodes in idle or normal states; if none of the aforementioned resource state nodes are available, they can be assigned to nodes in busy states.

[0030] Update unit 133: Used to update the resource status data of the node and transmit the updated data to the node resource status marking module.

[0031] refer to Figure 5 In a preferred embodiment, the task progress monitoring module 14 includes, Task status acquisition unit 141: used to collect the execution status data of each task in real time, including task start time, execution progress, task end time and resource usage, including CPU usage, memory usage, disk space usage and network bandwidth. Task progress evaluation unit 142: Based on the execution status data, evaluate the execution progress of the task and calculate the remaining execution time of the task; Resource occupancy transmission unit 143: used to transmit resource occupancy information to the node resource status marking module to support dynamic updates of resource status; Task execution anomaly detection unit 144: used to detect abnormal situations during task execution, including task execution timeout, abnormal resource usage, or task execution failure, and to trigger the fault handling module 15 to handle the situation when an anomaly is detected. Task status update unit 145: Used to dynamically update the execution status of the task based on the real-time data of the task status acquisition unit 141, and transmit the updated status data to the task allocation module 13 to support dynamic task scheduling. Task completion notification unit 146: When a task is completed, it generates a task completion notification and transmits the task execution result and resource release status to the node resource status marking module 12 and the task allocation module 13 so that the system can reallocate resources and schedule new tasks.

[0032] refer to Figure 6 In a preferred embodiment, the fault handling module 15 includes, Abnormal data detection unit 151: used to detect abnormal data of task abnormality detection unit 144 and node abnormality handling unit 124; Task retry unit 152: Used to reassign and retry tasks when a task execution failure is detected. Retry execution includes: reassigning the task to other available nodes, adjusting task resource requirements, and then re-executing. For example, if the task retry unit detects that task A failed to execute on node 1, since task A is a high-priority task, the task retry unit decides to reassign and retry the task. Based on the resource status data provided by the node resource status marking module, node 2 (idle state) is selected as the retry target. Task A is reassigned to node 2 and execution begins. The task progress monitoring module monitors the execution status of task A in real time to ensure successful task completion.

[0033] Node recovery unit 153: Used to restore the normal operating state of a node when a node failure is detected. The recovery operation includes: restarting the failed node; migrating tasks on the failed node to other available nodes; and releasing the resources occupied by the failed node. Fault notification unit 154: When a fault is detected and cannot be handled, it notifies the relevant personnel so that further measures can be taken in a timely manner.

Claims

1. A distributed resource dynamic scheduling system based on adaptive load balancing, characterized in that, Including the following: Task data acquisition module (11): used to acquire task data including the physical geographical location information of the task, and assign weight values ​​to each task according to the task data; Node resource status marking module (12): used to collect the physical geographical location information and resource status data of each edge gateway node in the distributed system, and mark the resource status of each edge gateway node according to the resource status data; Task allocation module (13): Generates a task allocation strategy based on the physical geographic location information of the task, the physical geographic location information of the edge gateway node, the resource status data and the weight value, and allocates the task to the edge gateway node according to the task allocation strategy. Task progress monitoring module (14): Used to monitor the execution status data of each task in real time, including task start time, execution progress, task end time and resource usage; Fault handling module (15): Used to handle faults when a task execution fault or an edge gateway node fault is detected.

2. The distributed resource dynamic scheduling system based on adaptive load balancing according to claim 1, characterized in that, The task data acquisition module (11) includes, Task feature extraction unit (111): used to extract task data from task requests, including task data volume, memory requirements, task execution time and task priority, as well as the physical geographical location information of the task; Weight allocation unit (112): Assigns a weight value to each task based on the task data.

3. The distributed resource dynamic scheduling system based on adaptive load balancing according to claim 1, characterized in that, The node resource status marking module (12) includes, Resource data acquisition unit (121): used to collect physical location information and resource status data of each edge gateway node in the distributed system in real time. The resource status data includes CPU utilization, disk I / O, memory usage, network bandwidth and storage utilization. Resource status assessment unit (122): It is used to set corresponding resource status thresholds according to the hardware configuration parameters of each edge gateway node, including the number of CPU cores, memory capacity, disk capacity and network bandwidth limit, and mark the resource status of each node according to the assessment results. The resource status includes idle status, normal status, busy status, overload status and fault status. Resource status update unit (123): It is used to dynamically update the resource status of each edge gateway node according to the real-time data of the resource data acquisition unit (121), and pass the updated resource status data to the task allocation module (13) to support task allocation decision. Node anomaly handling unit (124): If one or more edge gateway nodes are marked as faulty or overloaded, the fault handling module (15) will be triggered to handle the fault.

4. The distributed resource dynamic scheduling system based on adaptive load balancing according to claim 1, characterized in that, The task allocation module (13) includes, Candidate node filtering unit (131): used to filter out candidate nodes suitable for executing tasks based on the physical geographical location information of the task, the physical geographical location information of the node, the task weight value and resource status data; Task allocation unit (132): used to allocate tasks to candidate nodes in order of their weight values; Update unit (133): Used to update the resource status data of the edge gateway node and transmit the updated data to the node resource status marking module.

5. The distributed resource dynamic scheduling system based on adaptive load balancing according to claim 1, characterized in that, The task progress monitoring module (14) includes, Task status acquisition unit (141): used to collect the execution status data of each task in real time, including task start time, execution progress, task end time and resource usage, including CPU usage, memory usage, disk space usage and network bandwidth. Task progress evaluation unit (142): Based on the execution status data, evaluate the execution progress of the task and calculate the remaining execution time of the task; Resource occupancy status transmission unit (143): used to transmit the resource occupancy status to the node resource status marking module to support dynamic updates of resource status; Task execution anomaly detection unit (144): used to detect abnormal situations during task execution, including task execution timeout, abnormal resource usage or task execution failure, and trigger the fault handling module (15) to handle the situation when an anomaly is detected; Task status update unit (145): Used to dynamically update the execution status of the task based on the real-time data of the task status acquisition unit (141), and pass the updated status data to the task allocation module (13) to support dynamic task scheduling. Task completion notification unit (146): When a task is completed, it generates a task completion notification and passes the task execution result and resource release status to the node resource status marking module (12) and the task allocation module (13).

6. The distributed resource dynamic scheduling system based on adaptive load balancing according to claims 3 and 5, characterized in that, The fault handling module (15) includes, Abnormal data detection unit (151): used to detect abnormal data of task abnormal detection unit (144) and node abnormal processing unit (124); Task retry unit (152): used to reassign the task and retry execution when a task execution failure is detected. The retry execution includes reassigning the task to other available edge gateway nodes, adjusting the task resource requirements and then re-executing. Node recovery unit (153): used to restore the normal operation of the edge gateway node when a fault is detected, the recovery operation includes restarting the faulty edge gateway node; Migrate tasks on the failed edge gateway node to other available edge gateway nodes; Release the resources occupied by the faulty edge gateway node; Fault notification unit (154): Used to notify relevant personnel when a fault is detected and cannot be handled.