Dynamic task scheduling resource optimization method and device for distributed system
By monitoring task dependencies and hardware resource status, dynamically quantifying the task dependency release amount R and resource efficiency index A, and adopting a three-stage scheduling strategy, the problem of resource allocation imbalance in distributed systems is solved, CPU parallel utilization and memory resource utilization are improved, and task processing efficiency is optimized.
Patent Information
- Application Number
- CN202511104355.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-07
- Publication Date
- 2025-12-05
AI Technical Summary
Traditional static scheduling strategies cannot dynamically perceive the topological relationships between tasks and the real-time resource status, leading to resource imbalances in distributed systems when the load fluctuates, low utilization of CPU parallel computing resources, or memory shortages, which affects system performance.
By monitoring task dependencies and hardware resource status, the dependency release amount R and resource efficiency index A of tasks are dynamically quantified. A three-stage scheduling strategy is adopted: under low load, tasks with the largest R are allocated first; under stable load, tasks are allocated by weighted average of R and A; and under high load, tasks with the largest A are allocated first. Combined with real-time load status and resource status updates, efficient and coordinated utilization of resources is achieved.
By improving CPU parallel utilization during low-load phases and alleviating memory pressure during high-load phases, long-task latency is prevented, thus achieving overall optimization of system resource utilization and task processing efficiency and significantly improving the overall performance of the distributed system.
Smart Images

Figure CN121070540A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of task scheduling of distributed systems, and particularly relates to a dynamic task scheduling resource optimization method and device for distributed systems. BACKGROUND
[0002] With the wide application of distributed computing systems in enterprise service bus (ESB) and microservice architecture, task schedulers need to efficiently coordinate a large number of heterogeneous tasks with complex dependency relationships. The traditional static scheduling strategy faces significant technical challenges in actual operation: when the system load fluctuates, the fixed priority scheduling method cannot dynamically perceive the topological relationship between tasks and the real-time resource state, resulting in serious imbalance in the allocation of internal resources of the computer system. Specifically, at the low load stage, the scheduler fails to prioritize the execution of key nodes that can activate multiple downstream tasks, resulting in low utilization of CPU parallel computing resources; while at the high load stage, short tasks excessively preempt resources, causing long tasks to be continuously delayed, triggering memory resource shortage and even overflow risk. Although the existing improvement scheme such as dynamic weight adjustment can partially alleviate the resource contention problem, due to the lack of coordinated judgment on the task structure value and system load state, it still cannot fundamentally solve the resource optimization problem in the high concurrency environment. This technical limitation makes the hardware resource utilization rate of the existing system decrease sharply when the task dependency depth exceeds 5 layers or the concurrency increases by 200%, which seriously restricts the overall performance of the distributed system. SUMMARY
[0003] The present application aims to solve the imbalance problem of CPU and memory resource allocation in distributed systems caused by static scheduling strategies, and to realize the optimal allocation of resources under system load fluctuations by dynamically quantifying the task structure value and execution efficiency.
[0004] To achieve the above-mentioned purpose, the first aspect of the present application provides a dynamic task scheduling resource optimization method for distributed systems, comprising the following steps:
[0005] Step one: monitoring the task dependency relationship data and hardware resource state data of the distributed system;
[0006] Step two: generating the dependency release amount R of each task based on the task dependency relationship data, wherein R represents the number of downstream tasks that can be activated after the task is completed;
[0007] Step three: generating the resource efficiency indicator A of each task based on the hardware resource state data, wherein A represents the efficiency of improving the utilization of hardware resources by the task;
[0008] Step four: according to the real-time load state, selecting one of the following scheduling strategies to allocate tasks to computing nodes:
[0009] When the load state is low load, the task with the largest R is preferentially assigned;
[0010] When the load state is stable load, the task is assigned according to the weighted value of R and A;
[0011] When the load state is high load, the task with the largest A is preferentially assigned;
[0012] Step five: driving the computing node to execute the assigned task and updating the hardware resource state.
[0013] The second aspect of the application provides a dynamic task scheduling resource optimization device of a distributed system, comprising:
[0014] A data monitoring module is configured to monitor task dependency relationship data and hardware resource state data of the distributed system;
[0015] A task dependency analysis module is connected to the data monitoring module and is configured to generate a dependency release amount R of each task based on the task dependency relationship data, wherein R represents the number of downstream tasks that can be activated after the task is completed;
[0016] A resource efficiency evaluation module is connected to the data monitoring module and is configured to generate a resource efficiency indicator A of each task based on the hardware resource state data, wherein A represents the promotion efficiency of the task on the utilization rate of the hardware resource;
[0017] A scheduling decision module is connected to the task dependency analysis module and the resource efficiency evaluation module and is configured to select a scheduling strategy according to a real-time load state, comprising:
[0018] A first strategy unit is configured to select the task with the largest dependency release amount R for assignment when the load state is low load;
[0019] A second strategy unit is configured to select the task for assignment according to the weighted value of R and A when the load state is stable load;
[0020] A third strategy unit is configured to select the task with the largest resource efficiency indicator A for assignment when the load state is high load;
[0021] A task execution driving module is connected to the scheduling decision module and the data monitoring module and is configured to drive the computing node to execute the assigned task and feed back the execution result to the data monitoring module to update the hardware resource state.
[0022] The dynamic task scheduling resource optimization method and device of the distributed system provided by the application have at least the following beneficial effects:
[0023] The application realizes efficient collaborative utilization of hardware resources in a computer system by monitoring task dependency and system resource state in real time, combining dynamic priority calculation and three-stage scheduling strategy.
[0024] Specifically, high-dependence release-amount tasks are preferentially scheduled in a low-load stage to improve CPU parallel utilization, and the system automatically switches to a high-activation task priority mode in a high-load stage to reduce memory pressure, while a minimum-efficiency guarantee mechanism prevents long tasks from being starved, so that the overall optimization of system resource utilization and task processing efficiency is achieved. BRIEF DESCRIPTION OF DRAWINGS
[0025] Figure 1 A distributed system dynamic task scheduling resource optimization method flowchart is provided for the embodiments of the application.
[0026] Figure 2 A load state determination flowchart is provided for the embodiments of the application.
[0027] Figure 3 A dynamic task scheduling device structure diagram based on multi-module collaboration is provided for the embodiments of the application.
[0028] Figure 4 An ESB-Adaptive scheduling system architecture diagram is provided for the embodiments of the application.
[0029] Figure 5 An ESB-Adaptive scheduling core pseudocode diagram is provided for the embodiments of the application.
[0030] REFERENCE NUMERALS
[0031] Data monitoring module-100, task dependency analysis module-200, resource efficiency evaluation module-300, scheduling decision module-400, first strategy unit-401, second strategy unit-402, third strategy unit-403, task execution driving module-500, runtime adapter-510, service dependency analysis module-520, task execution monitoring module-530, priority evaluator-540, and scheduling controller-550. DETAILED DESCRIPTION
[0032] The technical solutions in the embodiments of the application will be described clearly and completely below with reference to the drawings in the embodiments of the application. Obviously, the described embodiments are only part of the embodiments of the application, rather than all the embodiments of the application. Based on the embodiments in the application, all other embodiments obtained by those skilled in the art without creative work fall within the scope of protection of the application.
[0033] Reference to an "embodiment" in this document means that a particular feature, structure, or characteristic described in connection with the embodiment can be included in at least one embodiment of the application. The appearances of the phrase in various places in the specification are not necessarily all referring to the same embodiment, nor are they necessarily mutually exclusive of other embodiments. It is expressly understood that the embodiments described herein can be combined with other embodiments in any way deemed useful.
[0034] Embodiment one
[0035] The embodiment of the application provides a dynamic task scheduling resource optimization method of a distributed system. The method realizes real-time sensing of task topological relation and hardware resource state, dynamically quantifies task structure value and execution efficiency, and adaptively selects a scheduling strategy according to system load fluctuation. As shown in the figure, the method comprises the following sequentially executed steps: task dependency relation and hardware resource state monitoring, dependency release quantity calculation, resource efficiency index calculation, dynamic strategy scheduling, task execution and state updating. The steps are cooperatively used to realize resource optimization, and specifically comprise the following steps. Figure 1
[0036] S100: monitoring task dependency relation data and hardware resource state data of the distributed system
[0037] The task dependency relation data refers to directed graph structure information describing the topological connection relation between tasks, and specifically includes task node identification, parent-child dependency relation between nodes and dependency strength weight. The hardware resource state data represents real-time resource occupation of a computing node, for example, can include central processing unit core utilization rate, memory occupation rate, network bandwidth usage and other physical resource indexes. In this step, the monitoring agent deployed on the computing node periodically collects the above data, and the data is summarized to the scheduling controller to form a global view.
[0038] S200: generating dependency release quantity R of the task
[0039] The dependency release quantity R is defined as the number of directly downstream tasks that can be activated immediately after the task is completed, and is used to quantify the pivotal value of the task in the topological structure. In a possible implementation manner, the calculation process comprises the following steps.
[0040] (1) analyzing the task dependency relation data, and constructing a task dependency graph;
[0041] (2) traversing each task node, and counting the number of directly connected child tasks;
[0042] (3) taking the number of child tasks as the R value of the task and outputting the R value.
[0043] For example, if the completion of task X can trigger the simultaneous start of tasks A, B and C, the R value of task X is 3. This index preferentially schedules the key node that can release more parallel tasks.
[0044] S300: Generating resource efficiency indicator A of the task
[0045] The resource efficiency indicator A represents the efficiency of the task in improving the utilization of hardware resources per unit time. Preferably, the calculation of A takes into account the execution time and resource consumption characteristics of the task, specifically:
[0046] (1) Predict the task duration based on historical execution records, for example, using exponential smoothing method;
[0047] (2) Analyze the task resource demand pattern, for example, calculate the average processor instruction throughput and memory access efficiency per unit time;
[0048] (3) Output A value through normalization and weighting formula, which may include the combination of processor utilization improvement coefficient and memory occupation optimization coefficient.
[0049] For example, a task can complete 1000-1500 instructions per millisecond and the memory page hit rate is as high as 85%-95% in typical execution, so its A value is significantly better than that of a long-time inefficient task. Those skilled in the art can understand that the above numerical range is only exemplary, and the actual execution efficiency may be adjusted according to the hardware architecture or task type.
[0050] S400: Dynamic policy scheduling
[0051] Select the scheduling policy according to the real-time load state, and determine the load state by calculating the total resource utilization of the node cluster:
[0052] Low load state determination criteria: the average resource utilization of the cluster is lower than the preset threshold (e.g. lower than 40%). At this time, tasks with the largest R value are preferentially allocated to improve CPU parallel computing resource utilization by activating more downstream tasks.
[0053] Stable load state determination criteria: resource utilization is in the middle interval (e.g. 40%-80%). Calculate the task priority according to the formula P = α·R + β·A, where α and β are configurable weight coefficients (e.g. α = 0.6, β = 0.4), and allocate tasks with the highest P value to balance resource efficiency and topology value.
[0054] High load state determination criteria: resource utilization exceeds the critical threshold (e.g. >80%). Preferentially allocate tasks with the largest A value to reduce memory pressure by quickly releasing resources and avoid overflow risk caused by resource contention.
[0055] Special scenario adaptation: When there are tasks to be allocated with the same weight, the default is to use the first-ready-first-served strategy; if a computing node fails, the scheduler automatically skips the node until its state is restored, ensuring the robustness of the method.
[0056] S500: Task execution and state update
[0057] The scheduler drives the target computing node to execute the assigned task through a remote procedure call. After the task is started, the monitoring agent continuously collects the hardware resource state data of the node (such as the newly added CPU computing load and the memory allocation amount) and updates the global resource pool in real time. When the task is completed or abnormally terminated, the monitoring data refresh in step one is triggered, forming a closed loop control.
[0058] In the embodiment of the application, each key mechanism cooperates to realize dynamic scheduling.
[0059] The load state sensing mechanism divides the load stage through a dynamic resource utilization rate threshold, for example, a plurality of configurable threshold values are adapted to different scale clusters. When the system suddenly changes from low load to high load, the scheduling strategy is automatically switched within a single monitoring period, ensuring real-time response.
[0060] The minimum efficiency guarantee mechanism prevents long tasks from being continuously delayed in a high load state. When the task waiting time exceeds the preset tolerance value, the system temporarily increases the weight of the A value, avoiding the task starvation phenomenon. This mechanism extends the exception handling capability of the method without modifying the core scheduling logic.
[0061] The dynamic task scheduling resource optimization method of a distributed system in the exemplary embodiment, by continuously monitoring the task dependency relationship and the hardware resource state, combining the dynamically calculated dependency release amount R and the resource efficiency indicator A, prioritizes high R value tasks in the low load stage to improve CPU parallel resource utilization, automatically switches to high A value task priority mode in the high load stage to relieve memory pressure, and maintains the execution rights and interests of long tasks through the minimum efficiency guarantee mechanism. The three strategies are dynamically switched based on the load state, and the real-time resource state is updated, realizing the cooperative optimization configuration of CPU and memory resources, significantly improving the overall resource utilization and task processing efficiency of the distributed system.
[0062] Embodiment two
[0063] On the basis of the above-mentioned embodiment one, the application embodiment further provides another dynamic task scheduling resource optimization method of a distributed system, comprising the following steps:
[0064] Step one: monitoring the task dependency relationship data and the hardware resource state data of the distributed system;
[0065] Step two: generating the dependency release amount R of each task based on the task dependency relationship data, wherein R represents the number of downstream tasks that can be activated after the task is completed;
[0066] Step three: generating resource efficiency index A of each task based on the hardware resource state data, wherein A represents the promotion efficiency of the task on the hardware resource utilization;
[0067] Step four: according to the real-time load state, selecting one of the following scheduling strategies to distribute the tasks to the computing nodes:
[0068] When the load state is low load, the task with the largest R is preferentially distributed;
[0069] When the load state is stable load, the task is distributed according to the weighted value of R and A;
[0070] When the load state is high load, the task with the largest A is preferentially distributed;
[0071] Step five: driving the computing nodes to execute the distributed tasks, and updating the hardware resource state.
[0072] In the embodiment, the specific implementation is please refer to the above embodiment one, the following preferred other exemplary embodiments are described:
[0073] In a preferred embodiment, as Figure 2 as shown in the flowchart of the load state determination provided in this embodiment, the load state is determined by the following steps:
[0074] S410, real-time statistics of the number of ready tasks N in the task pool;
[0075] S420, comparing N with the preset threshold value:
[0076] S421, when N < the first threshold value T, it is determined as low load state, which is used to trigger the high parallelism optimization mode;
[0077] S422, when T ≤ N ≤ the second threshold value T, it is determined as stable load state, which is used to balance the resource utilization;
[0078] S423, when N > T, it is determined as high load state, which is used for emergency memory release.
[0079] The above-mentioned load state determination method of step four is described in detail, and the determination process is used as the basis for load state determination. In a possible implementation manner, the load state is determined by the following sequential operations:
[0080] (1) Real-time statistics of the number of ready tasks N in the task pool:
[0081] The task monitoring module of the scheduler periodically scans the task pool, and counts the total number of tasks in the ready state (i.e., the dependent conditions have been met, and only the computing nodes are to be allocated), denoted as N. For example, the ready tasks include task instances that have completed the pre-dependence check but have not been allocated by the scheduler.
[0082] (2) Determine the load state based on the comparison of N and the preset threshold value:
[0083] Compare the real-time statistical N value with the preset threshold value, and divide the load stage according to the comparison result:
[0084] When N < the first threshold value T1, it is determined to be a low load state. This state triggers a high parallelism optimization mode, that is, a task with the largest dependence release amount R is preferentially allocated to improve CPU resource utilization.
[0085] When T1≤N≤the second threshold value T2, it is determined to be a stable load state. This state enables a resource utilization balancing mode, and executes a strategy of allocating tasks according to a weighted value (P=α·R+β·A).
[0086] When N>T2, it is determined to be a high load state. This state activates an emergency memory release mode, and executes a strategy of preferentially allocating a task with the largest resource efficiency indicator A.
[0087] It should be noted that T1 and T2 are integer values dynamically configured according to the size of the distributed system, for example, in a typical scenario, T1=50 and T2=150 can be set, but the actual value can be adjusted according to the number of cluster computing nodes;
[0088] Among them, the threshold value updating mechanism is: when the cluster node is scaled up or down, the scheduler automatically resets T1 and T2 in proportion (for example, if the number of nodes increases by 50%, the threshold value is also increased by 50%), to ensure that the load determination matches the system size.
[0089] The output of the determination step directly drives the strategy selection in step four:
[0090] Low load determination result→trigger the "preferentially allocate R largest task" strategy of step four;
[0091] Stable load determination result→trigger the "allocate according to R and A weighted value" strategy of step four;
[0092] High load determination result→trigger the "preferentially allocate A largest task" strategy of step four.
[0093] In a preferred embodiment, the dependence release amount R is generated by the following steps:
[0094] 1. Analyze the service call chain between tasks and extract explicit dependency relationships;
[0095] 2. Detecting shared resource contention, identifying implicit dependency relationship;
[0096] 3. Constructing a directed acyclic graph (DAG), traversing to count the number of direct downstream tasks of each task node as the R value, which is used to quantify the contribution of the task to the parallelism of the system.
[0097] The present embodiment details the dependency release amount R generation method, which is the specific implementation of the above step two. Specifically, the R value is generated through the following sequential operations:
[0098] (1) Analyze the service call chain to extract explicit dependency relationship:
[0099] Collect the call logs between tasks through a distributed tracing system (such as an Open Telemetry-based call chain monitoring tool), and analyze the parent-child task relationship in the service call path. For example, when task A calls the service interface of task B, record the explicit dependency link A→B, which directly determines the topological order of task execution.
[0100] (2) Detecting shared resource contention to identify implicit dependency relationship:
[0101] Monitor the access conflict of tasks to shared resources, specifically including:
[0102] Database lock contention: detect other tasks blocked when a task holds a table-level lock or a row-level lock;
[0103] Memory resource contention: identify multiple tasks simultaneously applying for memory blocks exceeding the preset limit;
[0104] Hardware device preemption: record the task queuing state of special devices such as GPU / NPU.
[0105] For example, when tasks C and D concurrently access the same database table, even if there is no direct call relationship between them, an implicit dependency is established.
[0106] (3) Constructing DAG and counting the number of direct downstream tasks:
[0107] Integrate explicit and implicit dependency relationships to construct a task directed acyclic graph (DAG), where nodes represent tasks and directed edges represent dependency directions;
[0108] Traverse each task node and count the number of its directly connected child nodes, which is the R value. For example, task E directly points to tasks F, G, and H in the DAG, so R = 3.
[0109] Technical feature association explanation
[0110] Explicit / implicit dependency synergy: explicit dependency guarantees the correctness of the underlying topology, while implicit dependency supplements the indirect constraints generated by resource competition. The combination of the two makes the DAG more realistic and reflects the system parallelism bottleneck.
[0111] R-value quantification logic: the number of direct downstream tasks directly determines the number of parallel tasks that can be activated after the task is completed. For example, the release of a high R-value task (e.g., R≥5) can increase CPU parallel utilization by 30%-50% (example data), achieving precise quantification of the "task contribution to system parallelism".
[0112] Integration with core method: the R-value generated in this step will be directly input into Step Four above to participate in scheduling decisions:
[0113] Low-load stage: high R-value tasks are preferentially allocated to maximize parallel resource utilization;
[0114] Stable load stage: R-value participates in weighted calculation to balance topology value and resource efficiency.
[0115] It should be noted that when implicit dependency causes a loop in the DAG, the lowest priority task dependency edge in the loop is automatically interrupted to ensure that the graph structure always meets the acyclic constraint; R-value calculation ignores downstream tasks in a failed state to avoid invalid scheduling. This embodiment realizes precise quantification and evaluation of task parallel value by constructing an explicit and implicit dual dependency detection mechanism combined with a multi-dimensional topology feature extraction algorithm for directed acyclic graphs (DAG).
[0116] In a preferred embodiment, the resource efficiency indicator A is calculated as follows:
[0117] 1. Set the standard execution time window T according to the task type s :
[0118] For compute-intensive tasks, where k is the computational complexity coefficient;
[0119] For I / O-intensive tasks, where k is the data volume coefficient;
[0120] 2. Obtain the actual execution time Ta of the task;
[0121] 3. Calculate A value in segments:
[0122] When Ta≤Ts, representing the activation capacity per unit time under ideal conditions, where R is the dependency release amount;
[0123] When Ta>Ts, where λ is the decay coefficient, used to punish tasks that exceed the timeout.
[0124] The embodiment details the calculation method of resource efficiency indicator A, which is the specific implementation of the above step three. Specifically, the A value is generated by the following sequential operations:
[0125] (1) Set the standard execution time window Ts
[0126] According to the task type, dynamically configure the reference time window:
[0127] Computing-intensive tasks: CPU frequency, where k is a configurable coefficient reflecting the computational complexity of the task (for example, a matrix operation task can set k = 1.2 x 10 instruction number), and the CPU frequency is obtained from the real-time computing node hardware information;
[0128] I / O-intensive tasks: Where k is a configurable coefficient based on the amount of task data (for example, a log processing task can set k = 500 MB), and the disk throughput is collected in real time through the storage system performance monitoring interface.
[0129] (2) Obtain the actual execution time Ta of the task
[0130] Through the recording module of the task execution engine, the complete time consumption of the task from the start of allocation to the computing node to the end of execution is counted, with a precision of milliseconds (for example, based on the taskstats mechanism of the Linux kernel).
[0131] (3) Calculate the A value in segments
[0132] When Ta≤Ts (ideal execution efficiency): This formula represents the ability of the task to activate downstream tasks in unit time, where R is the dependency release amount. For example, a task R = 4 and Ts = 200 ms, then A = 20 tasks / s;
[0133] When Ta>Ts (timeout penalty scenario): Where λ is a configurable decay coefficient (for example, 0.05), which punishes the execution delay through an exponential decay mechanism.
[0134] Technical meaning of the formula: for every 1 unit of time increase in the timeout part, the A value is attenuated by e -λ Proportional decay (for example, when λ = 0.05, a delay of 100 ms causes the A value to decay to about 0.6 times the original value).
[0135] The present embodiment constructs a complete technical system for task resource efficiency evaluation through the synergistic design of multi-dimensional technical features. At the task type classification level, the system clearly distinguishes between compute-intensive tasks (with CPU main frequency as the driving parameter) and I / O-intensive tasks (with disk throughput as the benchmark index), laying the foundation for differentiated calculation of the resource efficiency index A. The configurable coefficients k1 and k2 are introduced to construct an adaptive evaluation model, and through typical example values (such as k1 = 1.2 x 10 9 ) to realize decoupling design of hardware characteristics and algorithm parameters, ensuring the portability of index calculation.
[0136] For the timeliness of task execution, an exponential decay function is introduced to quantify the efficiency decay of timeout tasks with λ = 0.05. This mechanism forms a technical linkage with the R value of claim 3, which increases the weight of A value to accelerate resource release under high load, and the priority of timeout tasks decreases exponentially with ΔT, avoiding long-tail task blocking at the algorithm level.
[0137] To enhance environmental adaptability, the present embodiment sets a dynamic standard time window Ts, which is adjusted in real time with CPU main frequency / disk throughput to ensure the consistency of heterogeneous environment indexes; new tasks (such as network communication type) use a hybrid calculation mode Ts = (k1 / CPU main frequency + k2 / disk throughput) / 2 to expand applicability; and the λ value supports adaptive optimization based on historical timeout delay rate (such as the average of the last 10 times) to enhance environmental matching degree.
[0138] At the mathematical expression level, the present embodiment uses an exponential decay function as the preferred solution, while compatible with linear decay and other alternative forms to ensure that the core inventive concept of the punishment mechanism is not limited by specific functions.
[0139] In a preferred embodiment, the decay coefficient λ is dynamically adjusted according to the following rule: λ n = λ0·[1+log(n+1)], where λ is the initial decay coefficient, n is the number of consecutive timeouts, and log(n+1) is the punishment gain term, used to gradually reduce the priority of repeated timeout tasks.
[0140] The present embodiment proposes a dynamic decay coefficient adjustment mechanism, which adaptively adjusts the resource efficiency evaluation parameter λ by monitoring task timeout behavior to optimize distributed task scheduling. The core process includes:
[0141] (1) Initially set the reference coefficient λ (typical value 0.03);
[0142] (2) Real-time statistics of consecutive timeout times n;
[0143] (3) Adjust λ n= λ0·[1+ln(n+1)] dynamically calculates the punishment strength.
[0144] The mechanism uses a natural logarithm function to achieve gradual gain, so that the first timeout λ value is increased by 57% (λ≈0.051), and the increase reaches 140% (λ≈0.072) after three timeouts, but the increase slows down as n increases, avoiding excessive punishment.
[0145] In terms of technical coordination, the dynamic λ value and the exponential decay formula linkage, forming a step-by-step priority adjustment: short-term timeout tasks retain scheduling opportunities, and long-term timeout task A values decay exponentially (for example, when λ=0.07, a 100ms delay causes the A value to drop to 30%). The system provides double protection: the task is completed on time, and the n value is reset, and the maximum value of λ (such as 0.15) is limited to prevent task starvation.
[0146] The mechanism precisely suppresses inefficient tasks through sub-linear punishment gain, and significantly improves the adaptability of heterogeneous environments by combining cross-node timeout count synchronization and dynamic calibration of initial coefficients (optimizing λ based on historical data). Experiments show that it can make the priority of repeated timeout tasks decrease step by step, and the system resource idle rate is reduced by 27%, forming a closed-loop control with the technical solution of claim 4.
[0147] In a preferred embodiment, it also includes protecting critical tasks by the following steps:
[0148] 1. Set the lower limit value of the activation rate: where ∈ is the lower limit coefficient, and ∈ is in the range of 0<∈≤0.3, R is the dependent release amount, T s is the standard execution time window;
[0149] 2. When a task is marked as a critical path node, the ∈ value is increased by 50%-100% to ensure the execution right of the core task of the business process.
[0150] The present embodiment proposes a critical task protection mechanism, which builds a business continuity protection system for distributed systems through dynamic lower limit control and priority enhancement technology, and the core innovations are as follows:
[0151] Double-stage activation rate protection strategy
[0152] After calculating the resource efficiency indicator A, a forced lower limit calibration is added: where ∈ is a configurable coefficient (default 0.15, range 0<∈≤0.3). The mechanism uses a theoretical benchmark value to establish a priority bottom line, for example, when a long task decays to 0.05 due to timeout A value, it can be forced to increase to 3.0 through ∈=0.2 calibration, avoiding scheduling starvation.
[0153] Key path privilege enhancement
[0154] Dynamic protection for DAG critical path nodes:
[0155] ∈ value promotion 50%-100%(e.g., from 0.15 to 0.3), making the lower limit of critical task A value expand 1 times;
[0156] ∈ value synchronization propagation across system-dependent nodes through distributed transaction coordinator, ensuring core business chain integrity.
[0157] Conflict resolution and elasticity design
[0158] Dynamic competition isolation: when multiple critical tasks compete, resources are allocated according to secondary sorting based on business priority;
[0159] Intelligent backoff mechanism: after task completion, automatically restore the original ∈ value to prevent long-term occupation of protection allowance;
[0160] False node filtering: "critical task" automatically unmarked if not invoked for 3 consecutive periods, avoiding resource abuse.
[0161] System-level synergistic effect
[0162] Form a closed-loop control with the basic scheduling mechanism:
[0163] Basic protection layer: Formula prevents ordinary task priority from collapsing;
[0164] Strengthen the protection layer: critical task A lower limit is increased by 100%, ensuring that resources are executed preferentially during contention period;
[0165] High-load adaptation: non-critical task ∈ value can be dynamically adjusted down(e.g., from 0.3 to 0.1), concentrating resources to ensure core link.
[0166] This mechanism realizes task classification protection without changing the core scheduling logic through configurable parameters and business priority linkage. Experimental data shows that it can reduce the critical business interruption rate by 63%, and improve the overall resource utilization of the system by 19%, making it particularly suitable for distributed scenarios with high reliability requirements such as financial transactions and real-time control.
[0167] In a preferred embodiment, the task allocation priority Score in a stable load state is calculated as follows:
[0168] 1. Real-time monitoring of CPU idle rate U idle ;
[0169] 2. Dynamically adjust the weight coefficient: Where: η is the adjustment factor; U threshold is the CPU idle rate threshold;
[0170] 3. Calculate Score = a * R + (1-a) * A, to balance CPU and memory resource occupation.
[0171] This embodiment proposes a task scheduling method based on dynamic weight optimization under stable load, which realizes the collaborative optimization of computing and memory resources by real-time sensing of CPU idle rate and adaptive adjustment of weight ratio of dependent release amount R and resource efficiency A. The core mechanism is as follows:
[0172] Dynamic weight calculation model
[0173] Real-time collection of CPU idle rate U idle (e.g. through Linux / proc / stat interface, updated every 200ms), combined with preset threshold U threshold (default 15%-25%) and adjustment factor η (0.8-1.2), according to Generate dynamic weight. For example, when U idle From 10% to 30%, α is linearly adjusted between 0.5 and 1.0.
[0174] Priority decision engine
[0175] Build composite index Score = a * R + (1-a) * A:
[0176] High α scenario (α≥0.8): focus on R value, prefer to activate high parallel tasks to consume idle CPU resources;
[0177] Low α scenario (α≤0.3): focus on A value, select short tasks with fast memory release to prevent resource competition.
[0178] Adaptive collaborative mechanism
[0179] Linkage with load determination module: only activate this strategy in stable load interval (T1≤N≤T2), avoid strategy conflict;
[0180] Threshold dynamic calibration: automatically adjust U threshold (e.g. 24-hour mean ± 5% fluctuation);
[0181] Boundary protection: the lower limit of α value is forced to be 0.1, to ensure that A value always participates in decision-making, to prevent memory optimization failure.
[0182] Resource optimization effect
[0183] Experimental data shows that CPU utilization is improved by 10%-30% (when U idle >25%), memory peak occupation is reduced by 12%-18% (when U idle <15%);
[0184] Different business characteristics are adapted by η value differentiation configuration (such as scientific computing cluster η = 1.2, transaction system η = 0.9).
[0185] Robustness enhanced design
[0186] Heterogeneous cluster adaptation: partition statistics U according to CPU architecture idle And weighted calculation is performed.
[0187] Anti-shake processing: to U idle Mutation implementation 3 period smoothing filter;
[0188] Cold start protection: the default alpha = 0.5 in the initialization stage, and the dynamic mode is switched after the data is stable.
[0189] The mechanism quantifies the hardware state into dynamic weight parameters, constructs a resource-aware scheduling decision model, effectively solves the collaborative optimization problem of CPU and memory resource allocation under stable load, and experimentally verifies that the comprehensive resource utilization of the system can be improved by 18%-25%.
[0190] In a preferred embodiment, the first threshold T and the second threshold T are dynamically set as follows:
[0191] 1. Obtain the maximum capacity C of the task pool;
[0192] 2, calculate T = δ × C, T = γ × C, wherein: δ, γ are proportional coefficients, satisfy 0 < δ < γ < 1; for making the scheduling strategy automatically adapt to different scale distributed system.
[0193] The embodiment proposes a task pool capacity adaptive based load threshold configuration method, which realizes intelligent adaptation of scheduling strategy by dynamically associating system scale. The method first calculates the maximum capacity C of the task pool according to the ratio of the total available memory Mtotal of the cluster to the average memory occupancy M task of a single task. When the cluster is scaled up or down or the task memory characteristics change, the value of C is automatically updated. Based on the value of C, the dynamic load threshold T1 = δ × C and T2 = γ × C are generated by presetting the proportional coefficients δ and γ, wherein 0 < δ < γ < 1 form the stable load interval. Under typical configuration, δ takes 0.15-0.25, and γ takes 0.45-0.65, for example, when C = 500, T1 = 100 triggers the low load mode, and T2 = 300 activates the high load protection.
[0194] The mechanism builds a scale-adaptive closed loop: the task pool capacity C scales linearly with the cluster resources, driving the thresholds T1 / T2 to scale synchronously, ensuring that the load determination criteria strictly match the system scale. By decoupling the delta / gamma configuration from the physical resources, small-scale test clusters and super-large-scale production clusters share the same set of determination logic. To enhance environmental adaptability, the system supports time-periodized dynamic optimization of delta / gamma, such as using different coefficient combinations for daytime trading periods and nighttime batch processing periods, while setting a strong constraint of gamma-delta >= 0.3 to prevent strategy oscillation.
[0195] In terms of technical effects, the method enables different-scale clusters to automatically obtain load thresholds that adapt to their capacities, such as T1 = 40 for a 10-node cluster and T1 = 400 for a 100-node cluster, completely eliminating the risk of misjudgment with fixed thresholds. When the cluster is expanded, the T2 value automatically floats to extend the stable load interval length by about 35%, ensuring that the weighted scheduling strategy continues to take effect. At the operation and maintenance level, once the proportion coefficient is set, it can adapt to any scale changes without the need for manual threshold adjustment. For special scenarios, the system implements a memory fluctuation compensation mechanism that delays C value updates when a single task memory suddenly increases by 20%, and forces small clusters (C < 50) to maintain a minimum stable interval width, ensuring the robustness of load determination under various environments.
[0196] In a preferred embodiment, step five implements task allocation by the following steps:
[0197] 1. Configure a plug-in runtime adapter for compatibility with heterogeneous enterprise service bus protocols;
[0198] 2. Select an interface driver that matches the target bus:
[0199] For WSO2 ESB, call the SOAP / REST service interface;
[0200] For Mule ESB, access the task event bus message queue.
[0201] The present embodiment proposes a plug-in protocol adaptation scheme that dynamically matches enterprise service bus interfaces at runtime to achieve efficient driving of tasks to computing nodes in a heterogeneous distributed environment. The scheme builds an extensible adapter framework, including two core components: a protocol abstraction layer and a plug-in loader. The protocol abstraction layer encapsulates a unified ESB communication interface specification, covering connection management, message serialization, and exception retrying, among other basic functions. The plug-in loader supports hot deployment of bus protocol driver plug-ins, dynamically registering protocol implementations in JAR package form through configuration files. The scheduler automatically scans the specified directory when starting, such as detecting WSO2 and Mule ESB plug-ins, and then loading the SOAP / REST service interface driver and AMQP message queue adapter, respectively.
[0202] For different ESB types, the adapter automatically selects the driving mode: for WSO2 ESB nodes, task instructions are delivered through SOAP 1.2 protocol or RESTful API, an XML message body containing task ID and resource constraint parameters is constructed, and the WS-Addressing standard interface is called to receive the execution status; for Mule ESB nodes, the AMQP 0-9-1 protocol is used to publish the JSON format task object to the specified exchange, and the result feedback is obtained by listening to the exclusive queue. This design realizes protocol conversion without feeling, automatically adapts the unified task instruction to the format required by the target ESB, and maintains independent connection pools for various protocols to avoid resource competition.
[0203] In terms of technical cooperation, the responses of different ESBs are converted into standard structures through a unified callback interface. When a 5-second timeout is detected without reply, task redistribution to the standby node is automatically triggered. Tests show that this scheme supports six types of mainstream enterprise bus protocols such as SOAP / REST / AMQP, covering 92% of heterogeneous environment needs, and reduces task allocation delay from 120ms to 35ms. New protocol support only requires the development of 200 lines of plug-in code without modifying the core scheduling logic. Special scenario processing includes automatically switching to the REST interface when detecting WSO2 ESB version upgrade, implementing primary and backup dual bus driving for high-availability tasks, and ensuring plug-in security through digital signature verification to comprehensively guarantee reliable execution of tasks in a mixed bus environment.
[0204] In a preferred embodiment, a dynamic task scheduling method for enterprise service bus is disclosed, which is characterized by constructing a three-stage cooperative scheduling mechanism to achieve load adaptive optimization through dynamic balance of dependence release (R) and unit activation rate (A). The method includes three technical modules: activation rate calculation model, phased scheduling strategy and closed-loop execution process. In particular, the dimensional imbalance problem of multi-index cooperative scheduling is solved through the normalized weighting mechanism in the load stabilization period.
[0205] At the activation rate quantization level, this embodiment uses a segmented calculation model to realize timeliness compensation. The basic calculation formula is where t s According to the dynamic setting of the task type, t a is the real-time execution time. To prevent excessive priority decay of timeout tasks, a lower limit protection mechanism is specially set The configurable parameter ∈ ∈ (0, 0.3] ensures the basic scheduling opportunity. Experimental verification shows that this model can retain at least 10% of the original priority for long-tail tasks.
[0206] As a core innovation, the three-stage scheduling mechanism divides the system running period into traffic rise period, load stabilization period and traffic decline period, and adopts differentiated priority strategies in each stage: when the number of ready tasks Nr When entering the rising period of traffic at T1, the P = R strategy is adopted to preferentially activate high-parallel tasks, and it is measured that the critical path duration can be shortened by 28%; when T1 ≤ N r ≤ T2, it enters the load-stable period, and an innovative formula is proposed. Through the dynamic normalization operator to eliminate the difference in task scale, and cooperate with the weight coefficient to achieve resource-aware scheduling; when N r > T2, the traffic decline period is started, and it is switched to the P = A strategy to focus on resource release. It is measured that the task backlog can be reduced by 35% - 52%.
[0207] This design shows a significant technical synergy effect: the normalization processing mechanism enables tasks of different magnitudes to obtain a fair competition environment, effectively avoiding the phenomenon of "big R tasks monopolizing resources"; the dynamic weight adjustment constructs a direct mapping from the resource state to the scheduling decision through the ratio relationship between the CPU idle rate U idle and the threshold U threshold In the Alibaba Cloud ESB test environment, this method increases the task throughput by 22% - 37%, reduces the risk of memory overflow by 40%, and fully supports runtime self-adjustment of all parameters. It has been verified to be applicable to heterogeneous scenarios such as financial transactions and e-commerce promotions.
[0208] The scheduling execution process forms a complete closed loop: the system continuously monitors N r to determine the stage it is in. After calculating the priority P according to the corresponding formula, the task is distributed to the ESB node through the plug-in protocol adapter. After the task is completed, the execution duration t a is updated and the ready queue is refreshed, thereby constructing a complete control loop of resource state awareness, scheduling strategy adaptation, and execution result feedback.
[0209] Example Three
[0210] This embodiment discloses a dynamic task scheduling device based on multi-module collaboration, which realizes real-time matching of task allocation and resource state through a closed-loop control system. As Figure 3 shown is a schematic structural diagram of a dynamic task scheduling device based on multi-module collaboration provided by way of example in this embodiment. The device consists of a data monitoring module 100, a task dependency analysis module 200, a resource efficiency evaluation module 300, a scheduling decision module 400, and a task execution drive module 500. Each module is interconnected via a system bus to form a complete control loop, specifically including:
[0211] The data monitoring module 100 is used to monitor the task dependency relationship data and hardware resource state data of the distributed system;
[0212] The task dependency analysis module 200 is connected to the data monitoring module 100, and is configured to generate a dependency release amount R of each task based on the task dependency relationship data, where R represents a number of downstream tasks that can be activated after the task is completed.
[0213] The resource efficiency evaluation module 300 is connected to the data monitoring module 100, and is configured to generate a resource efficiency indicator A of each task based on the hardware resource state data, where A represents an efficiency of improving utilization of hardware resources by the task.
[0214] The scheduling decision module 400 is connected to the task dependency analysis module 200 and the resource efficiency evaluation module 300, and is configured to select a scheduling strategy according to a real-time load state, including:
[0215] The first strategy unit 401 is configured to select a task with a maximum dependency release amount R for distribution when the load state is a low load.
[0216] The second strategy unit 402 is configured to select a task for distribution according to a weighted value of R and A when the load state is a stable load.
[0217] The third strategy unit 403 is configured to select a task with a maximum resource efficiency indicator A for distribution when the load state is a high load.
[0218] The task execution driving module 500 is connected to the scheduling decision module 400 and the data monitoring module 100, and is configured to drive a computing node to execute a distributed task, and feed back an execution result to the data monitoring module 100 to update a hardware resource state.
[0219] Optionally, the data monitoring module 100 is deployed on each computing node, and periodically collects task dependency relationship data and hardware resource state data through a monitoring agent. Specifically, the data monitoring module 100 is configured to perform the following operations: first, collect task node identifiers, parent-child dependency relationships, and dependency strength weights, and synchronously acquire resource indicators such as CPU utilization and memory occupancy; then, aggregate the foregoing data to a central processor to generate a resource state global view; finally, receive task execution result feedback and update a hardware resource state in real time to provide data support for scheduling decisions.
[0220] Optionally, the task dependency analysis module 200 is connected to the data monitoring module 100, and its core functions include dependency graph construction and R value calculation. The task dependency analysis module 200 parses task dependency relationship data and generates a directed acyclic graph, and counts a number of direct downstream tasks by traversing DAG nodes to output a dependency release amount R representing parallel activation capability. For example, when a task activates three downstream tasks, the R value of the task is marked as 3, and the parameter is pushed to the scheduling decision module 400 in real time to participate in priority calculation.
[0221] Optionally, the resource efficiency evaluation module 300 performs multi-dimensional efficiency analysis based on the resource state data provided by the data monitoring module 100. It first uses the exponential smoothing method to predict the task execution time, and then identifies the task resource demand characteristics, such as CPU instruction throughput and memory access efficiency. The resource efficiency index A is generated through a normalized weighting formula, and the minimum efficiency guarantee mechanism is integrated to temporarily increase the A value weight when the task is waiting for timeout to prevent long tail task scheduling from being starved.
[0222] Optionally, the scheduling decision module 400 includes three strategy units and a load sensor to realize dynamic strategy switching. When the load sensor determines that the cluster resource utilization is less than 40%, the first strategy unit 401 is activated to preferentially select the task with the maximum R value to improve CPU parallel utilization; when the utilization is in the interval of 40%-80%, the second strategy unit 402 calculates the priority according to the formula P = a R + b A, and balances the dependence activation and resource efficiency through configurable weight coefficients; when the utilization exceeds 80%, the third strategy unit 403 switches to the A value priority mode to significantly reduce the memory peak pressure. The aforementioned strategy switching is completed within a single monitoring period to ensure that the scheduling strategy strictly matches the real-time load state.
[0223] Optionally, the task execution driving module 500 realizes task instruction distribution through an RPC interface and configures a fault tolerance mechanism. When a node fault is detected, it automatically redirects the task to a backup node to ensure system availability. After execution is completed, the module returns the result data such as CPU load and memory allocation to the data monitoring module 100, forming a complete closed-loop control link.
[0224] The device works through five modules. In the initialization stage, it completes dependency relationship analysis and efficiency index generation. In the scheduling stage, it dynamically selects the optimal strategy according to the load state. Finally, it realizes task landing through the execution driving module. Its technical advantages are: through the multi-strategy switching mechanism driven by load sensing, it simultaneously considers parallel optimization in low load scenarios and resource pressure relief in high load scenarios; the lower limit protection design of the resource efficiency evaluation module 300 effectively avoids scheduling starvation caused by long task index decay; the closed-loop feedback architecture ensures real-time updating of system state, providing continuous data support for dynamic scheduling decisions.
[0225] In a preferred embodiment, the embodiment discloses a dynamic task scheduling system for enterprise service bus, whose modular architecture and core scheduling logic are shown in Figure 4 and Figure 5 The system operation mechanism is described in detail below in conjunction with the drawings:
[0226] As shown in Figure 4An ESB-Adaptive scheduling system architecture schematic diagram provided by an embodiment example of the application is shown, and the system constructs an end-to-end closed loop scheduling architecture through five core modules. The runtime adapter 510 acts as the hub for the system to interact with the external ESB environment, and adopts a bidirectional arrow design to realize protocol decoupling: on the one hand, it receives the API scheduling instructions transmitted by external buses such as WSO2 / MuleESB, and on the other hand, it returns the instructions generated by the scheduling controller 550 to the target node. This module supports plug-in protocol drivers, and can dynamically load JAR package plug-ins conforming to the unified interface specification, to ensure seamless access to heterogeneous bus environments.
[0227] The service dependency analysis module 520 undertakes the task static feature analysis responsibility, and extracts the dependency release amount R and the standard execution time t s by constructing a task directed acyclic graph (DAG). Specifically, this module traverses the task nodes to count the number of direct downstream tasks, for example, when a task activates three parallel sub-tasks, its R value is marked as 3, and this parameter, together with the preset t s value, constitutes the basic data for priority evaluation.
[0228] The task execution monitoring module 530 collects dynamic running parameters in real time, including the actual execution time t a of the task and the ready queue length N r . This module deploys a monitoring agent on the computing node, synchronizes indicators such as CPU utilization and memory occupancy at a period of 200 ms, and pushes the dynamic feature data to the priority evaluator 540.
[0229] As shown in Figure 5 , an ESB-Adaptive scheduling core pseudocode diagram provided by an embodiment example of the application is shown, and the priority evaluator 540 acts as the scheduling decision core, and its processing logic contains two major innovative algorithms. First, the activation rate segmented calculation algorithm (see Figure 5 ) implements differentiated processing according to the relationship between t a and t s : when t a ≤ t s , the basic formula A=R / t s is used to quantify the task execution efficiency; when t a >t s , an exponential decay term exp(-k*(t a -t s )) is introduced to implement timeout punishment, and a lower limit protection mechanism is used to prevent long-tail task scheduling starvation. Second, the three-stage priority decision algorithm dynamically switches the scheduling strategy according to the N r value: in the traffic rising period (N r<T1) prefer large R value task to accelerate parallel link construction; load stable period (T1≤N r ≤T2) adopt normalized weighted formula Eliminate large R value task monopoly problem by dynamic normalization; flow decline period (N r >T2) switch to A value priority mode to quickly release resources.
[0230] The scheduling controller 550 generates scheduling instructions based on the priority evaluation results, which contains two major innovative features: on the one hand, it establishes a low-latency link (decision latency <10ms) with the priority evaluator 540 through a direct connection channel, and on the other hand, it dynamically adjusts the threshold T1 / T2 according to the weight 8 embodiment, ensuring that the scheduling strategy strictly matches the real-time load state. The final generated scheduling instructions are issued to the ESB nodes for execution through the runtime adapter 510, forming a complete control loop.
[0231] The system verifies significant technical effects in the Mule ESB test environment through modular division and algorithm coordination: the normalization mechanism improves small-scale task scheduling opportunities by 40%, reduces memory peak pressure by 18%, and controls the whole process latency from instruction access to scheduling decision within 50ms. The bidirectional adapter architecture and dynamic priority decision model designed by the system provide an efficient solution for task scheduling in a heterogeneous distributed environment.
[0232] The above only describes the preferred embodiments of the present application, and does not limit the patent scope of the present application, and any equivalent structure or equivalent process transformation using the content of the present application specification and drawings, or direct or indirect application in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A dynamic task scheduling resource optimization method of a distributed system, characterized in that, Comprising: Step one: monitor the task dependency data and hardware resource state data of the distributed system; Step two: generate the dependency release amount R of each task based on the task dependency data, wherein R represents the number of downstream tasks that can be activated after the task is completed; Step three: generate the resource efficiency indicator A of each task based on the hardware resource state data, wherein A represents the efficiency of task in improving hardware resource utilization; Step four: according to the real-time load state, select one of the following scheduling strategies to allocate tasks to computing nodes: When the load state is low load, preferentially allocate the task with the largest R; When the load state is stable load, allocate tasks according to the weighted value of R and A; When the load state is high load, preferentially allocate the task with the largest A; Step five: drive the computing node to execute the allocated task, and update the hardware resource state.
2. The method of claim 1, wherein, The load state is determined by the following steps: (1) Real-time statistics of the number of ready tasks N in the task pool; (2) Compare N with the preset threshold value: When N < the first threshold value T, it is determined as low load state, which is used to trigger high parallel optimization mode; When T ≤ N ≤ the second threshold value T, it is determined as stable load state, which is used to balance resource utilization; When N > T, it is determined as high load state, which is used for emergency memory release.
3. The method of claim 2, wherein, The dependency release amount R is generated by the following steps: (1) Analyze the service call chain between tasks to extract explicit dependency relationships; (2) Detect shared resource competition to identify implicit dependency relationships; (3) Build a directed acyclic graph DAG, traverse and count the number of direct downstream tasks of each task node as R value, which is used to quantify the contribution of the task to the parallelism of the system.
4. The method of claim 3, wherein, The resource efficiency indicator A is calculated according to the following steps: (1) Set the standard execution time window Ts according to the task type: for computationally intensive tasks, where k is a computational complexity coefficient; for I / O intensive tasks, where k is a data volume coefficient; (2) Get the actual execution time Ta of the task; (3) Calculate A value in segments: When Ta≤ Ts, Characterizes the activation capacity per time under ideal conditions, where R is the dependency on the release amount. When Ta > Ts, where λ is a decay coefficient used to penalize overdue tasks.
5. The method of claim 4, wherein, The attenuation coefficient λ is dynamically adjusted according to the following rule: λ n = λ0·[1+log(n+1)], where: λ is the initial attenuation coefficient; n is the number of consecutive timeouts of the task; log(n+1) is a penalty gain term for gradually reducing the priority of the repeated timeout task.
6. The method of claim 4, wherein, Further comprising, protecting critical tasks by the following steps: (1) Set the lower limit of the activation rate: where ∈ is the lower limit coefficient, 0 < ∈ ≤ 0.3, R is the dependent release amount, T s is the standard execution time window; (2) When the task is marked as a critical path node, the value of ∈ is increased by 50%-100%, which is used to protect the execution right of the core task of the business process.
7. The method of claim 2, wherein, The task allocation priority Score of the stable load state is calculated according to the following steps: (1) Real-time monitoring of CPU idle rate U idle ; (2) dynamically adjusting the weight coefficient: wherein: η is an adjustment factor; U threshold is a CPU idle rate threshold value; (3) Calculate Score = α·R + (1-α)·A, which is used to balance CPU and memory resource occupation.
8. The method of claim 2, wherein, The first threshold value T and the second threshold value T are dynamically set according to the following steps: (1) Get the maximum capacity C of the task pool; (2) Calculate T = δ × C, T = γ × C, where: δ, γ are proportional coefficients, satisfying 0 < δ < γ < 1; for making the scheduling strategy automatically adapt to different sizes of distributed system.
9. The method of claim 1, wherein, Step five realizes task allocation by the following steps: (1) Configure plug-in runtime adapter to be compatible with heterogeneous enterprise service bus protocol; (2) Select the interface driver matching the target bus: For WSO2 ESB, call SOAP / REST service interface; For Mule ESB, access task event bus message queue.
10. A dynamic task scheduling resource optimization apparatus of a distributed system, characterized in that, Comprising: A data monitoring module for monitoring the task dependency data and hardware resource state data of the distributed system; a task dependency analysis module connected to the data monitoring module, configured to generate a dependency release amount R of each task based on the task dependency relationship data, wherein R represents the number of downstream tasks that can be activated after the task is completed; a resource efficiency evaluation module connected to the data monitoring module, configured to generate a resource efficiency indicator A of each task based on the hardware resource state data, wherein A represents the promotion efficiency of the task on the utilization of hardware resources; a scheduling decision module connected to the task dependency analysis module and the resource efficiency evaluation module, configured to select a scheduling strategy according to the real-time load state, including: a first strategy unit configured to select a task with the largest dependency release amount R for distribution when the load state is low load; a second strategy unit configured to select a task according to the weighted value of R and A for distribution when the load state is stable load; a third strategy unit configured to select a task with the largest resource efficiency indicator A for distribution when the load state is high load; a task execution driving module connected to the scheduling decision module and the data monitoring module, configured to drive the computing node to execute the distributed task, and feed back the execution result to the data monitoring module to update the hardware resource state.
Citation Information
Cited By
Communication link intelligent sharing method of edge communication gateway
CN121509354A
Multi-core NPU task scheduling method, system and equipment based on bandwidth awareness
CN122086576A
Dynamic scheduling and resource allocation optimization method for multi-base collaborative observation platform
CN122088955A