Resource collaborative awareness-based Storm flow calculation dynamic scheduling method

Through real-time monitoring and historical data-driven resource collaborative optimization model, the scheduling performance problems of the Storm flow computing system under dynamic resource fluctuations and node abnormalities are solved, efficient task instance allocation and topology division are achieved, and the system's throughput and latency performance is improved.

CN120390015APending Publication Date: 2025-07-29CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 4 Cited by

Patent Information

Application Number
CN202510541490.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

The existing Storm flow computing system has problems of degradation in scheduling performance under dynamic resource fluctuations, node exceptions and multi-dimensional resource constraints, especially in heterogeneous clusters, which fails to effectively utilize node characteristics and optimize communication overhead between tasks, resulting in insufficient system delay and throughput.

Method used

By monitoring the node status in real time, a resource collaborative optimization model is built, a greedy algorithm is used to prioritize the allocation of highly correlated task instances to nodes with excellent historical performance, and dynamically adjust load balancing parameters through smooth processing through time windows, optimize task instance allocation and topological division, and reduce communication costs and delays.

Benefits of technology

It significantly reduces system latency by 37.5%, improves throughput to more than 95%, ensures stable CPU and memory utilization, and realizes rapid recovery and stable operation of heterogeneous clusters in high concurrency scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120390015A_ABST
    Figure CN120390015A_ABST
Patent Text Reader

Abstract

The invention discloses a resource collaborative awareness-based Storm flow calculation dynamic scheduling method, and provides the following technical scheme aiming at the problem that the performance of a Storm default scheduling algorithm is reduced in node abnormity, resource fluctuation and rescheduling scenes: firstly, dynamically sensing node busy, downtime and data flow fluctuation events through a real-time monitoring module; triggering a rescheduling process; secondly, constructing a resource collaborative optimization model based on historical task instance resource requirements, node performance indexes and communication overhead, and generating a task instance allocation scheme by taking minimization of inter-node communication cost as a target and combining CPU / memory dynamic threshold constraints; further, a greedy algorithm is adopted to sort high-relevance task instances, the high-relevance task instances are preferentially distributed to nodes with the optimal historical performance, and it is ensured that the node resource utilization rate does not exceed a dynamic threshold value; meanwhile, node load balancing parameters are corrected in real time through time window smoothing processing, and the remaining resource state is updated; and finally, outputting an optimized topology division result, so that the system delay after rescheduling is remarkably reduced, and the throughput is improved. According to the method, rapid recovery and stable operation of the heterogeneous cluster are realized through resource collaborative modeling, historical data driven dynamic scheduling and load balancing optimization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of distributed stream computing, and specifically relates to a dynamic scheduling method for Storm stream computing based on resource collaborative perception, which is used to solve the problems of node resource fluctuations, degraded rescheduling performance of task instances, and collaborative optimization under multi-dimensional resource constraints in heterogeneous clusters. Background Art

[0002] With the rapid development of scenarios such as the Internet of Things, financial transactions, and real-time monitoring, distributed stream computing systems have become the core infrastructure for processing high-speed and continuous data streams. As a representative of open-source stream processing frameworks, Apache Storm is widely used in scenarios such as real-time data analysis and complex event processing due to its low latency and high reliability. Storm abstracts computing task instances into topologies, and splits data streams into multiple task instances for parallel processing. However, its default scheduling strategy is based on static resource allocation logic, and only distributes task instances according to the preset number of task instance slots, lacking the ability to perceive dynamic resource requirements in real time. In actual production environments, the throughput of data streams often fluctuates significantly. For example, there may be a surge in traffic during e-commerce promotions or sudden data reports from Internet of Things devices, and the node resource status (such as CPU overload, memory exhaustion, network congestion) may also change suddenly due to hardware failures or uneven loads. Such dynamic changes make it difficult for static scheduling strategies to maintain system stability. When a node crashes due to resource overload, the system needs to re-partition the task instance topology. However, because the scheduler does not combine key information such as historical resource utilization and task instance communication patterns, the rescheduling process takes a long time, and the newly allocated task instances may trigger resource bottlenecks again due to unreasonable node selection, forming a vicious cycle.

[0003] Existing scheduling strategies have significant deficiencies in dynamic resource perception and multi-dimensional constraint coordination. Usually, they use a single resource metric (such as CPU utilization) as the decision basis, resulting in unbalanced resource allocation. For example, although a certain node has idle CPU but exhausted memory, the scheduler may still allocate task instances with high memory requirements to this node, exacerbating memory competition and even triggering frequent container restarts or task instance migrations, further reducing system throughput. Taking a certain securities trading system as an example, the Storm cluster needs to process 100,000 order data per second within 1 millisecond. However, because the default scheduler does not distinguish the characteristics of GPU nodes, the risk control model inference task is allocated to a node without GPU resources, resulting in a delay exceeding the limit. In addition, existing methods (such as Kubernetes VPA) rely on historical data to predict resource requirements, have a lag in response in scenarios of sudden traffic, and do not optimize the communication overhead between tasks, resulting in an increase in cross-rack data transmission delay of more than 30%.

[0004] The complexity of heterogeneous clusters further amplifies the scheduling challenges. Modern data centers often adopt a hybrid hardware architecture. For example, some nodes are equipped with high-performance GPUs for machine learning inference, while ordinary nodes only provide basic computing capabilities. The default scheduler of Storm does not distinguish node characteristics and may misallocate GPU-dependent task instances to nodes without GPU resources, resulting in task instance execution failures or significant performance degradation. At the same time, the data dependencies between task instances are not effectively utilized. For example, if upstream and downstream task instances are scattered across nodes in different racks, multiple layers of network switching are required, introducing additional communication delays. Existing methods do not optimize such topologies, making it difficult for the overall system latency to meet real-time requirements. These problems are particularly prominent in scenarios such as financial high-frequency trading and industrial real-time control, where millisecond-level latency fluctuations can trigger major business risks.

[0005] The current scheduling methods of stream computing systems have significant deficiencies in aspects such as dynamic resource awareness, multi-dimensional constraint coordination, and heterogeneous environment adaptation. There is an urgent need for a dynamic scheduling scheme that can respond to resource changes in real time, make intelligent decisions by combining historical data, and optimize the utilization rate of multi-dimensional resources to support high-concurrency and strong-real-time business requirements. Summary of the Invention

[0006] Aiming at the problem of degraded scheduling performance in the existing Storm stream computing system under dynamic resource fluctuations, node anomalies, and multi-dimensional resource constraints, the present invention proposes a dynamic scheduling method for Storm stream computing based on resource collaborative awareness. Through real-time resource monitoring, history data-driven collaborative modeling, and dynamic threshold adjustment, efficient task instance allocation and rapid recovery in heterogeneous clusters are achieved. The specific steps are as follows:

[0007] S1. Monitor the resource status of cluster nodes in real time. When it is detected that node busyness, downtime, or data flow fluctuations trigger rescheduling, start the dynamic scheduling process;

[0008] S2. Based on the resource requirements of task instances, node performance metrics, and communication overhead between task instances in historical data, construct a resource collaborative optimization model and generate an allocation scheme for task instances;

[0009] S3. Use the greedy algorithm to sort the task instances to be allocated, and preferentially allocate highly correlated task instances to the nodes with the best historical performance, ensuring that the CPU and memory utilization rates of the nodes do not exceed the dynamic threshold;

[0010] S4. Dynamically adjust the load balancing parameters of the nodes according to the monitoring data smoothed by the time window, and update the remaining resource status of the nodes;

[0011] S5. Output the optimized topology partitioning result, enabling the system to quickly recover to a stable state after rescheduling, reducing system latency and improving throughput.

[0012] Furthermore, the objective function of the resource collaborative optimization model is to minimize the communication cost between nodes, expressed as:

[0013]

[0014] where E ij represents the communication overhead between task instances task i and task j , m represents a set of m machines with CPU capacity of and memory capacity of , n represents a list of n task instances T = {task1, task2,..., task n}, and γ i,k represents whether task instance task i is allocated to machine v k .

[0015] The constraint conditions include:

[0016] (1) Each task instance is allocated to only a single node;

[0017] (2) The resource requirements need to be satisfied:

[0018]

[0019] where is the CPU resource requirement of task instance task i , and is the memory resource requirement of task instance task i .

[0020] Furthermore, the calculation formula for the node performance metric described in step S2 is:

[0021]

[0022] where is the CPU resource used by machine v k , is the memory resource used by machine v k , u is the number of CPU cores of the node, is the load usage of machine v k , and its calculation formula is:

[0023]

[0024] where μ1 and μ2 respectively represent the weight coefficients of the CPU and memory utilization rates in measuring the node load.

[0025] Further, in step S3, the relevance of task instances is evaluated through the communication overhead E in historical data. Tasks with high E are preferentially assigned to the same node to reduce the communication latency between nodes. ij evaluated, and tasks with high E ij are preferentially assigned to the same node to reduce the communication latency between nodes.

[0026] The specific dynamic threshold adjustment in step S4 includes:

[0027] Based on a preset time window W, the sliding window mean of the CPU utilization rate and memory utilization rate of the node is calculated to filter out instantaneous fluctuation noise. When the average CPU utilization rate of the node exceeds or the average memory utilization rate exceeds in continuous W time periods, the values of α and β are reduced to limit resource overload.

[0028] If the fluctuation range of the node resource utilization rate within the W window is less than the preset threshold, α and β are increased to enhance the resource utilization rate.

[0029] Further, the specific metrics for performance recovery in step S5 are: the system latency is reduced by more than 37.5% (compared with the default scheduling algorithm); the fluctuation ranges of CPU and memory utilization rates are reduced to ±5% of the preset threshold; the throughput is increased to more than 95% of that before the data stream fluctuation.

[0030] The beneficial effects of the present invention are as follows: By introducing a dynamic scheduling mechanism based on resource collaborative perception, the real-time monitoring and historical data analysis are organically combined to achieve a comprehensive perception and response to the node resource status, communication overhead between task instances, and data stream fluctuation in the Storm stream computing environment. This method can not only quickly trigger the rescheduling process when detecting node busyness, downtime, or data stream anomalies, but also achieve an efficient matching of task instances and node resources by constructing a resource collaborative optimization model, fundamentally reducing the communication cost and latency between nodes. Using the greedy algorithm, tasks with high relevance are preferentially assigned to nodes with excellent historical performance, thereby significantly optimizing the task instance scheduling efficiency and topology division results on the premise that the CPU and memory utilization rates of each node do not exceed the dynamic threshold. The sliding window average method is used to smooth the monitoring data, which not only effectively filters out instantaneous fluctuation noise but also can adjust the node load balancing parameters in real time, enabling the system to quickly schedule resources to meet sudden demands when the load surges and release excess resources in a timely manner when the load decreases, avoiding resource waste and over-allocation. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 is a flowchart of the Storm scheduling algorithm based on resource collaborative perception in the present invention

[0032] Figure 2Structural diagram of the improved D-Storm system in the present invention Detailed implementation manners

[0033] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0034] The present invention provides a dynamic scheduling method for Storm stream computing based on resource collaborative perception, as Figure 1 shown, including the following steps:

[0035] S1. Monitor the resource status of cluster nodes in real time. When detecting that the nodes are busy, down, or the data stream fluctuates and triggers rescheduling, start the dynamic scheduling process;

[0036] Specifically, in the real-time monitoring and dynamic scheduling mechanism, the system collects key indicators such as CPU utilization rate, memory occupancy rate, and network bandwidth of each node in real time through a lightweight monitoring daemon process, and uses a time window (such as a 60-second sliding window) to perform mean smoothing processing on instantaneous fluctuation data to filter out noise to reflect the stable load trend. When detecting that the CPU or memory utilization rate of a node continuously exceeds the preset threshold, the node heartbeat is lost and it is determined to be down, or the data stream rate suddenly changes (such as exceeding 3 times the standard deviation of the historical mean), a rescheduling event is triggered. The objective function of the resource collaborative optimization model is to minimize the communication cost between nodes, expressed as:

[0037]

[0038] where E ij represents the communication overhead between task instances task i and task j , m represents a set of m machines with a given CPU capacity of and a memory capacity of , n represents a list of n task instances T = {task1, task2,..., task n}, γ i,k represents whether the task instance task i is assigned to the machine v k .

[0039] Among them, the constraint conditions include:

[0040] (1) Each task instance is only assigned to a single node;

[0041] (2) The resource requirements need to be met:

[0042]

[0043] Among them, is the CPU resource requirement of task instance task i of. is the memory resource requirement of task instance task i of.

[0044] S2. Based on the task instance resource requirements, node performance metrics, and communication overhead between task instances in historical data, construct a resource collaborative optimization model and generate an allocation plan for task instances;

[0045] Specifically, the node performance metrics described in step S2

[0046]

[0047] Among them, is the CPU resource used by machine v k of. is the memory resource used by machine v k of. u is the number of CPU cores of the node, is the load usage of machine v k of, and the calculation formula is:

[0048]

[0049] Among them, μ1 and μ2 respectively represent the weight coefficients of the CPU and memory utilization rates when measuring the node load.

[0050] S3. Use the greedy algorithm to sort the tasks to be assigned, and preferentially assign high-correlation task instances to the node with the best historical performance, ensuring that the CPU and memory utilization rates of the node do not exceed the dynamic threshold;

[0051] Specifically, in the task instance allocation stage, the greedy algorithm first sorts the task instance groups in descending order according to the communication overhead between task instances and the historical collaboration frequency, and preferentially selects high-correlation task instances. Subsequently, traverse the sorted queue and assign each task instance group to the node with the best historical performance score, which comprehensively considers historical data such as node CPU stability, memory utilization, and communication latency. During the assignment, the remaining resources of the node are checked in real time to ensure that the CPU and memory occupancy rates after the task instances are superimposed are respectively lower than the dynamic thresholds α and β. If the limit is exceeded, skip to the second-best node until the task instance is successfully deployed or a resource exhaustion warning is issued.

[0052] S4. Dynamically adjust the load balancing parameters of the node according to the monitoring data smoothed by the time window, and update the remaining resource status of the node;

[0053] Specifically, the monitoring module performs mean smoothing on the CPU / memory utilization rate of the node based on a sliding time window (such as 60 seconds). After eliminating the instantaneous jitter interference, it dynamically calculates the load balancing weights μ1 and μ2. If the CPU usage variance within the window is significantly higher than that of the memory (such as in a compute-intensive task instance), the proportion of μ1 is automatically increased; otherwise, the weight of μ2 is increased to adapt to the memory-intensive scenario. The remaining resource status of the node is refreshed in real time in combination with the smoothed data. If the CPU utilization rate of a certain node exceeds α (such as 0.8) or the memory exceeds β for three consecutive windows, the allocation priority of its subsequent task instances is reduced until the resource pressure drops back within the threshold to ensure dynamic load balancing.

[0054] S5. Output the optimized topology partitioning result, enabling the system to quickly return to a stable state after rescheduling, reducing system latency and improving throughput.

[0055] Specifically, the optimized topology partitioning aggregates highly correlated task instances into communication-intensive subgraphs, and preferentially deploys them to the same node or cooperative nodes connected by low-latency links, significantly reducing the communication overhead of cross-node data transmission. A dynamic resource buffer space is reserved during partitioning to avoid the risk of secondary overload caused by sudden local load increases.

[0056] After rescheduling, the system quickly binds task instances to node resources, significantly reducing the end-to-end latency of the processing pipeline by minimizing serialization and network interaction losses. The efficient cooperation and load balancing mechanism of task instances within the subgraph effectively improve the overall throughput performance of the data stream, enabling the cluster to quickly converge to a stable state after topology reconstruction, ensuring the high responsiveness and continuity of real-time computing.

[0057] Figure 2 Shows the improved D-Storm architecture design, whose core function is to optimize cluster resource allocation through dynamic topology partitioning. D-Storm integrates a resource analysis engine, uses ThreadMXBean to track the CPU occupancy rate of task instance execution threads in real time, and jointly models with historical task instance execution logs (such as resource consumption patterns, communication overheads, etc.) to form a multi-dimensional task instance demand profile.

[0058] When a node resource bottleneck or a decline in the communication efficiency among task instances is detected, the system automatically triggers a rescheduling mechanism to achieve resource rebalancing through the intelligent merging and migration of task instances. At the implementation level, D-Storm is embedded in the Storm framework in the form of a pluggable component. As an extended module of the Nimbus node, it works in coordination with the Storm native scheduler by implementing the IScheduler interface. After each scheduling decision is generated, the system calls the RESTful API to obtain actual running metrics (such as throughput and end-to-end latency) for effect verification, forming a closed-loop optimization of the scheduling strategy.

[0059] The supporting distributed monitoring system adopts a two-level architecture: lightweight monitoring agents are deployed at the worker node layer to continuously collect basic metrics such as CPU / memory; the central analysis layer introduces a sliding time window technology to perform mean filtering and variance analysis on the original monitoring data, eliminating the interference of instantaneous peaks on load determination, thereby accurately capturing the long-term trend of resource usage and providing stable data support for dynamic scheduling.

[0060] The present invention aims at the scheduling efficiency problem of the Storm cluster in a dynamic load scenario and proposes a dynamic scheduling mechanism with collaborative resource awareness. This mechanism uses a multi-dimensional resource monitoring module to continuously collect the utilization rates of node CPU, memory, and network bandwidth, combines the sliding time window smoothing technology to eliminate instantaneous noise interference, and accurately identifies node overload, downtime, or data flow mutation events. At the task instance scheduling layer, based on the greedy algorithm, priority sorting is performed on task instance groups with high communication relevance, and the optimal deployment node is dynamically selected by combining historical performance scores and real-time resource margins. A load balancing coefficient constraint is introduced to ensure that the CPU / memory utilization rate is maintained within a safe threshold. Through topology partition optimization, strongly associated task instances are aggregated and deployed to minimize the cross-node communication overhead, and at the same time, an elastic resource buffer space is reserved to cope with sudden load fluctuations. Compared with traditional static scheduling strategies, this solution achieves low-latency reconfiguration during task instance migration, significantly improves the throughput performance and stability of the cluster during the data flow peak period, reduces the risk of resource fragmentation while ensuring service quality, and provides efficient dynamic resource adaptation capabilities for the streaming computing scenario.

[0061] The above-mentioned embodiments further elaborate on the purpose, technical solutions, and advantages of the present invention. It should be understood that the above-mentioned embodiments are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made to the present invention within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A dynamic scheduling method for Storm stream computing based on resource collaborative perception, characterized in that, It includes the following steps: S1. Monitor the resource status of cluster nodes in real time. When detecting that a node is busy, down, or data flow fluctuations trigger rescheduling, start the dynamic scheduling process; S2. Based on the resource requirements of task instances, node performance metrics, and communication overhead between task instances in historical data, construct a resource collaborative optimization model and generate an allocation plan for task instances; S3. Use the greedy algorithm to sort the task instances to be allocated, and preferentially allocate highly correlated task instances to the node with the best historical performance, ensuring that the CPU and memory utilization rates of the node do not exceed the dynamic threshold; S4. Dynamically adjust the load balancing parameters of the node according to the monitored data smoothed by the time window, and update the remaining resource status of the node; S5. Output the optimized topology partition result, enabling the system to quickly recover to a stable state after rescheduling, reducing system latency and increasing throughput.

2. According to the method described in claim 1, the objective function of the resource collaborative optimization model is to minimize the communication cost between nodes, expressed as: Among them, E ij represents the communication overhead between the task instance task i and task j . m represents a set of m machines with a given CPU capacity of and a memory capacity of . n represents a list of n task instances T = {task1, task2,..., task n}. γ i,k indicates whether the task instance task i is allocated to the machine v k . where the constraint conditions include: (1) Each task instance is allocated to only one node; (2) The resource requirements need to satisfy: Among them, is the CPU resource requirement of task instance task i , and is the memory resource requirement of task instance task i .

3. The node performance metrics described in step S2 of the method according to claim 1 have the following calculation formula: Among them, is the CPU resources used by machine v k The is the memory resources used by machine v k where u is the number of CPU cores of the node is the load usage of machine v k The calculation formula is: where μ1 and μ2 respectively represent the weight coefficients of the CPU and memory utilization rates when measuring node load.

4. According to the method described in claim 1, in step S3, the relevance of task instances is evaluated by the communication overhead E in historical data, and task instances with high E are preferentially assigned to the same node to reduce the communication latency between nodes. ij is evaluated, and task instances with high E ij are preferentially assigned to the same node to reduce the communication latency between nodes.

5. According to the method described in claim 1, the specific dynamic threshold adjustment in step S4 includes: Based on the preset time window W, calculate the moving window average of the CPU utilization rate and memory utilization rate of the node to filter out instantaneous fluctuation noise; When the average CPU utilization rate of a node exceeds within W consecutive time periods or the average memory utilization rate exceeds reduce the values of α and β to limit resource overload; If the fluctuation range of the node resource utilization rate within the W window is less than the preset threshold, increase α and β to enhance the resource utilization rate.

6. According to the method described in claim 1, the specific performance recovery indicators in step S5 are: the system latency is reduced by more than 37.5% (compared with the default scheduling algorithm); the fluctuation range of the CPU and memory utilization rates is reduced to ±5% of the preset threshold; the throughput is increased to more than 95% before the data flow fluctuation.

Citation Information

Cited By

  • Hardware resource allocation method and device applied to edge server

    CN120973543A

  • Load balancing optimization method and system for cloud computing resources

    CN121750654A

  • Task scheduling system of data stream processor

    CN121807572A

  • A data flow processor task scheduling system

    CN121807572B