Intelligent dynamic big data platform resource scheduling method and system

By constructing DAG and using graph convolutional networks to optimize task priorities, the problem of insufficient modeling of complex task dependencies and resource heterogeneity in resource scheduling on big data platforms is solved, real-time adaptive scheduling of dynamic task flows and resource pools is achieved, and the system's scheduling efficiency and resource utilization are improved.

CN120780429AActive Publication Date: 2025-10-14BEIJING ZHONGYU TAINUO DATA TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510889008.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-30
Publication Date
2025-10-14
Estimated Expiration
2045-06-30

AI Technical Summary

Technical Problem

Existing big data platform resource scheduling technologies are unable to adapt to the complex dependencies between tasks and the changing resource supply when faced with large-scale dynamic and heterogeneous computing environments. They are unable to identify critical paths and bottleneck nodes in a timely manner, resulting in limited global execution efficiency. In addition, they lack the ability to comprehensively model data communication delays, resource heterogeneity, and task execution dynamics.

Method used

By constructing a directed acyclic graph (DAG), task splitting and topological sorting are performed, and graph convolutional networks are used to optimize task priorities. Incremental rescheduling is triggered when tasks are added or fail to execute. Real-time optimization is combined with graph neural networks to achieve efficient scheduling of complex task dependencies and resource environments.

Benefits of technology

It significantly improves the ability to perceive and optimize global bottlenecks of tasks, shortens the total task completion time, improves resource utilization, and supports dynamic addition and deletion of tasks and resources and incremental scheduling under dependency adjustment, with good scalability and real-time response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120780429A_ABST
    Figure CN120780429A_ABST
Patent Text Reader

Abstract

The invention discloses an intelligent dynamic big data platform resource scheduling method and system, and relates to the technical field of intelligent resource management, and the method comprises the steps: collecting the information of a to-be-scheduled task, constructing a directed acyclic graph (DAG), and carrying out the task splitting according to a load balancing principle; performing topological sorting on the DAG, determining a task priority sequence, and allocating available resources to the tasks according to the priority sequence; converting the DAG into a graph neural network input, outputting a priority enhancement value by using a graph convolutional network, and fusing the priority enhancement value with the sequence according to a preset weight to readjust the task priority; when a new task is added and task execution fails, incremental rescheduling is triggered, the graph convolutional network calculates task priorities of affected sub-graph nodes, and local redistribution is carried out in an original dispatch table. According to the method, the perception and optimization capability of the global bottleneck of the task is remarkably improved, intelligent decision and adaptive adjustment of resource allocation are realized, the total completion time of the task is effectively shortened, and the resource utilization rate is improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of intelligent resource management, in particular to an intelligent dynamic big data platform resource scheduling method and system. BACKGROUND

[0002] With the rapid development of big data and cloud computing technology, more and more industries begin to rely on large-scale distributed computing platforms to efficiently process and analyze multi-source heterogeneous data. Especially in the financial, telecommunications, transportation, manufacturing and other scenarios with high requirements for data timeliness and computing resource utilization, task scheduling and resource management have become the core problem in the design of platform architecture. In recent years, the emergence of YARN, Kubernetes, ApacheMesos and other resource scheduling frameworks has made cluster management and task scheduling automated and flexible, supporting the efficient operation of multi-tenant and multi-type tasks. At the same time, with the increasing application of artificial intelligence and machine learning in big data platform scheduling decisions, scheduling strategies are gradually evolving from traditional static allocation to intelligent, adaptive and predictive direction. For example, some studies propose to combine historical load characteristics for predictive scheduling, or use deep learning to realize dynamic task grading and resource optimal allocation, further improving the overall throughput and resource utilization efficiency of the system.

[0003] However, the existing big data platform scheduling technology still has many deficiencies when faced with large-scale dynamic and heterogeneous computing environments. First, traditional scheduling methods based on rules or static priorities are difficult to adapt to complex dependencies between tasks and changing resource supplies, often failing to identify and optimize key paths and bottleneck nodes in DAG (Directed Acyclic Graph) task flows in a timely manner, resulting in limited global execution efficiency. Second, current resource scheduling frameworks generally lack comprehensive modeling capabilities for data communication delays, resource heterogeneity, and task execution dynamics. In the face of runtime task additions, deletions, and dependency relationship adjustments, the scheduling strategy adjustment response is slow and lacks precision. In addition, although some intelligent scheduling methods have introduced neural networks or heuristic algorithms, they are mostly offline trained or single-point optimized, unable to achieve real-time adaptation to dynamic task flows and resource pools, and the scheduling results lack explainability, making it difficult to make fine-grained adjustments to actual bottleneck problems. SUMMARY

[0004] In view of the above problems, the present application is proposed.

[0005] Therefore, the technical problem solved by the present application is that the existing big data platform resource scheduling technology generally lacks modeling of complex task dependencies and resource heterogeneity, has limited adaptability to runtime dynamic changes, and is difficult to meet the efficient scheduling needs in large-scale dynamic scenarios.

[0006] To solve the above technical problems, the present application provides the following technical solutions: a kind of intelligent dynamic big data platform resource scheduling method, comprising:

[0007] The information of the task to be scheduled is collected, a directed acyclic graph (DAG) is constructed, and the task is split according to the load balancing principle;

[0008] The DAG is topologically sorted to determine the priority order of the tasks, and the available resources are allocated to the tasks according to the priority order;

[0009] The DAG is converted into a graph neural network input, and a priority enhancement value is output using a graph convolution network; the priority enhancement value and the ranking are fused according to a preset weight to readjust the priority of the tasks;

[0010] When a new task is added or a task fails, incremental rescheduling is triggered, the graph convolution network locates and calculates the priority of the task of the affected subgraph node, and local reallocation is performed in the original scheduling table.

[0011] As a preferred scheme of the intelligent dynamic big data platform resource scheduling method of the present application, wherein: the task splitting according to the load balancing principle includes calculating the average execution level of the task to be scheduled, setting a splitting threshold based on the average execution level; traverse the tasks in the directed acyclic graph (DAG), and determine whether the average execution time of the task on all resources exceeds the threshold; if not, keep the original task; if yes, split the task into multiple balanced sub-tasks;

[0012] For each split task, remove the original node in the graph, and create multiple new nodes according to the number of balanced sub-tasks; each new node inherits all the predecessor dependencies of the original node; if the sub-tasks must be executed in series, add a sequential dependency relationship between the sub-tasks; otherwise, keep parallel; set the corresponding execution time information for each sub-task;

[0013] After splitting, the DAG is detected in a loop to remove the dependency loop and update the predecessor and successor list of the newly added node;

[0014] The information of the task to be scheduled includes the execution time, dependency relationship and data transmission cost of each task.

[0015] As a preferred scheme of the intelligent dynamic big data platform resource scheduling method of the present application, wherein: the topological sorting of the DAG includes performing a traversal search on the DAG to generate a topological sequence that satisfies the dependency constraint;

[0016] According to the topological sequence, the scheduler calculates the upward rank of each node from bottom to top starting from the head of the sequence; for the exit node without successor, the upward rank is equal to the average execution time of the task itself; for other nodes, the upward rank is equal to the average execution time of the task plus the data transmission time between the successor nodes, plus the upward rank of the successor nodes, and the maximum value of all successor node path cases is taken;

[0017] The key task affecting the global completion time is identified, that is, the task with the highest cumulative upward rank from the entrance to the exit; the scheduler arranges all tasks in descending order of upward rank to obtain the task priority order.

[0018] As a preferred scheme of the intelligent dynamic big data platform resource scheduling method, the method comprises the following steps: increment scheduling decision is made, the earliest start time of the current task is determined, the running time of the current task on the to-be-allocated resource is added, and the completion time is evaluated; the current task is compared with the allocatable resource, and the current task is allocated to the resource that can be completed earliest;

[0019] The idle time of the selected resource is updated to the actual completion time of the task, and the start and end times of the current task on the resource are recorded; the scheduling table is obtained by distributing all tasks according to the priority order.

[0020] As a preferred scheme of the intelligent dynamic big data platform resource scheduling method, the method comprises the following steps: increment scheduling decision is made, the earliest start time of the current task is determined, the running time of the current task on the to-be-allocated resource is added, and the completion time is evaluated; the current task is compared with the allocatable resource, and the current task is allocated to the resource that can be completed earliest;

[0021] A multi-layer graph neural network is built, each layer including three kinds of mappings: mapping one acting on the node itself feature, mapping two acting on the feature from the predecessor node, and mapping three acting on the edge feature; all mappings are linear transformations;

[0022] In each layer, the node first performs a linear transformation on the upper layer representation of itself;

[0023] All predecessor nodes are traversed, the upper layer representation of each predecessor node is linearly transformed, and the linear transformation of the corresponding edge feature is added; the transformation result of the node itself is added to the transformation results of all predecessors, and the new feature representation of the node in the current layer is obtained through ReLU; the iteration is repeated to the maximum iteration number, and each node obtains a high-dimensional representation containing the calculation characteristics and topological information of itself and neighbors;

[0024] After the last layer of the GNN, a set of readout parameters is added to linearly map the high-dimensional representation of each node to a priority offset value;

[0025] A fusion weight is set to balance the upward ranking and the priority offset value output by the GNN; for each node, the upward ranking of the node is proportionally weighted and summed with the priority offset value to obtain a final scheduling priority score; all tasks are reordered using the new score to obtain a scheduling sequence that can dynamically respond to bottleneck and key nodes in the graph.

[0026] As a preferred scheme of the intelligent dynamic big data platform resource scheduling method, the trigger incremental rescheduling includes deploying the trained GNN model for scheduling service, and real-time monitoring of new tasks, canceled tasks, and modification events of task dependencies; starting from the task nodes and dependency relationships designed in the events, bounded traversal is performed in the predecessor and successor directions to find tasks that have not started execution to obtain an affected task set;

[0027] For each affected task, execution stability, task connectivity, remaining workload, and edge feature information are collected; the subgraph corresponding to the affected task is input into the GNN as a single graph sample; the priority offset value and resource affinity score of the affected task node are obtained by inference output;

[0028] The resource affinity and the priority offset value are fused to obtain a comprehensive priority; the affected tasks are incrementally scheduled according to the comprehensive priority.

[0029] As a preferred scheme of the intelligent dynamic big data platform resource scheduling method, the local reallocation includes using the difference between the predicted completion time and the actual completion time and the incremental subgraph as a new sample to perform a small amount of gradient update on the GNN; after each new DAG is formed and updated, the GNN is forward run to update the comprehensive priority of all nodes, and the task start time and resource allocation obtained by incremental scheduling are input into the execution layer.

[0030] As a preferred scheme of the intelligent dynamic big data platform resource scheduling system, it includes a task management module, a priority calculation module, a HEFT scheduling module, and a dynamic optimization module.

[0031] The task management module is used to collect task information and construct and maintain the DAG structure.

[0032] The priority calculation module is used for topological analysis, upward ranking, and priority sorting of the DAG.

[0033] The HEFT scheduling module is used for mapping tasks to specific resources according to priorities and processing communication overheads.

[0034] The dynamic optimization module is used for runtime monitoring, GNN feature extraction and online model fine-tuning.

[0035] A computer device comprises a memory and a processor, the memory stores a computer program, and the processor executes the computer program to implement the steps of the intelligent dynamic big data platform resource scheduling method.

[0036] A computer readable storage medium stores a computer program, and the computer program is executed by a processor to implement the steps of the intelligent dynamic big data platform resource scheduling method.

[0037] The intelligent dynamic big data platform resource scheduling method provided by the present application can realize efficient and intelligent scheduling for complex task dependencies and variable resource environments. Through DAG task dependency modeling and critical path identification, the perception and optimization ability of the global bottleneck of the task are significantly improved. Combined with the dynamic priority sorting enhanced by HEFT and GNN, intelligent decision-making and adaptive adjustment of resource allocation are realized, which effectively shortens the total task completion time and improves the resource utilization rate. In addition, the system supports dynamic addition and deletion of tasks and resources, incremental scheduling under dependency relationship adjustment, and has good scalability and real-time response capability. BRIEF DESCRIPTION OF DRAWINGS

[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings needed in the embodiment description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and for those skilled in the art, other drawings can be obtained without creative labor.

[0039] Figure 1 The overall flowchart of the intelligent dynamic big data platform resource scheduling method provided by the first embodiment of the present application. DETAILED DESCRIPTION

[0040] In order to make the above-mentioned purposes, features and advantages of the present application more apparent and easy to understand, the specific embodiments of the present application will be described in detail below with reference to the drawings of the specification. Obviously, the described embodiments are part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor should be within the scope of protection of the present application.

[0041] Embodiment 1, refer to Figure 1, as one embodiment of the present invention, provides an intelligent dynamic big data platform resource scheduling method, comprising:

[0042] S1: Collect information about tasks to be scheduled, construct a directed acyclic graph (DAG), and split tasks according to the load balancing principle.

[0043] Furthermore, the initialization input is to read all task lists {T1, T2, ..., T n} and resource list {R1, R2, …, R K}.

[0044] For each task T i Read: In each resource R k Execution time estimation on Dependency relationships with other tasks {(T p ,T i )}, data transmission time between tasks Calculate the average execution time with reference Makespan.

[0045] For each task T i , calculate its average execution time:

[0046]

[0047] Set Global Reference Used to determine the "large task" threshold. The benchmark can be the sum of the average execution time of all tasks or the longest completion time in historical scheduling.

[0048] Initially construct the DAG and use the task set directly as the node set of the graph:

[0049] V←{T1,T2,…,T n}

[0050] For each dependency (T p ,T i ), add the edge set E by T p Pointing to T i The directed edge of .

[0051] Task splitting according to the load balancing principle includes calculating the average execution level of the tasks to be scheduled and setting the splitting threshold based on the average execution level; traversing the tasks in the directed acyclic graph (DAG) to determine whether the average execution time of the task on all resources exceeds the threshold; if not, retaining the task as it is; if exceeded, splitting the task into multiple balanced subtasks.

[0052] Large task detection and splitting, traversing each node T i , if its average execution time satisfies:

[0053]

[0054] it is considered as a "big task" and needs to be split.

[0055] For each split task, remove the original node in the graph, create multiple new nodes according to the number of subtasks after balanced division; each new node inherits all the predecessor dependencies of the original node, if the subtasks must be executed in series, then add the precedence relationship between the subtasks in turn, otherwise keep parallel; set the corresponding execution time information for each subtask.

[0056] The information of the task to be scheduled includes the execution time, dependency relationship and data transmission cost of each task.

[0057] Calculate the number of small tasks after splitting:

[0058]

[0059] Let the execution time of each subtask be about

[0060] Remove the original task T from the node set i , add m subtask nodes {T i,1 ,…,T i,m}. Point all predecessors of the original T i to each T i,j ; point each T i,j to all successors to inherit data dependencies. If it is necessary to ensure that the subtasks are executed in order, add serial dependencies between T i,1 →T i,2 →…→T i,m ; otherwise, all subtasks remain parallel.

[0061] After splitting, perform cyclic detection on the DAG, remove the dependency loop, and update the predecessor and successor lists of the newly added nodes.

[0062] Update the transmission time and the graph structure for the new nodes introduced by splitting. If the subsequent scheduling will deploy them on different resources, still use the between the original tasks to estimate the transmission between subtasks (which can be allocated or conservatively take the original value). Finally, output the new node set V' and edge set E' after splitting and dependency inheritance, that is, get the preprocessed DAG.

[0063] It should be noted that by modeling the dependency relationship between tasks as a directed acyclic graph (DAG), the execution order and data flow of complex business processes can be accurately described, effectively avoiding deadlocks and conflicts during scheduling. At the same time, the task automatic splitting mechanism splits large-scale computing units into fine-grained sub-tasks, making the utilization of computing resources more sufficient and improving the parallelism and overall throughput of the system. The intelligent splitting criterion of "execution time occupancy Makespan ratio" is introduced to realize dynamic granularity adjustment. Task splitting not only optimizes load balancing, but also lays the foundation for subsequent adaptive scheduling strategies (such as GNN-based learning enhancement). Unlike traditional static task configuration methods, the system flexibility and scheduling accuracy are greatly improved.

[0064] S2: Topologically sorting the DAG to determine the priority order of tasks, and assigning available resources to the tasks according to the priority order.

[0065] Further, the topological sorting of the DAG includes performing a traversal search on the DAG to generate a topological sequence that satisfies the dependency constraints.

[0066] According to the topological sequence, the scheduler starts from the head of the sequence and calculates the upward ranking of each node from bottom to top; for the exit node without successors, the upward ranking is equal to the average execution time of the task itself; for other nodes, the upward ranking is equal to the average execution time of the task plus the data transmission time between the task and its successors, plus the upward ranking of the successors, and the maximum value of all successor paths is taken.

[0067] Identify the key tasks that affect the global completion time, i.e., the task path with the highest cumulative upward ranking from the entrance to the exit; the scheduler arranges all tasks in descending order of upward ranking to obtain the priority order of tasks.

[0068] Further, the assignment of available resources to tasks includes making incremental scheduling decisions, determining the earliest start time of the current task, adding the expected runtime of the current task on the available resources, and evaluating the completion time; comparing the current task with the available resources, and assigning the current task to the earliest available resource.

[0069] Update the idle time of the selected resource to the actual completion time of the task, and record the start and end times of the current task on the resource; distribute all tasks according to the priority order to obtain a scheduling table.

[0070] Further, initialize the data structure for each task node T i Preparation: pred(T i ): contains all tasks directly pointing to T i ; succ(T i) : contains all tasks starting from T i i state(T

[0071] Depth-first search (DFS) generates a topological sequence by traversing all task nodes in turn. If state(T i ) is "unvisited", the recursive visit is started from this node.

[0072] Recursive visit procedure (for the current node T i ) : state(T i ) is marked as "in visit"; for each successor T j ∈ succ(T i ) : if state(T j ) is still "unvisited", recursively visit T j ; after all successors are visited, T i is added to the "end of topological sequence"; state(T i ) is marked as "finished".

[0073] The final "topological sequence" is the legal scheduling precedence order in reverse order of tasks being added.

[0074] Upward ranking calculation. Based on the generated topological sequence, the sequence head (exit node) is processed in turn: if task T i has no successor in the computation graph (i.e. succ(T i ) is empty), its upward ranking is directly set as the average execution time of the task:

[0075]

[0076] Otherwise, the maximum value in the following formula is taken for all successors T j ∈ succ(T i ) :

[0077]

[0078] where, is the average execution time of task T i on all resources, is the data transmission time of T i → T j .

[0079] Determination of critical path and task priority. The critical path is identified. The critical path of the entire DAG is the path from a certain entry task to a certain exit task that can accumulate the maximum ​The value of the path from the source node to the sink node is the global minimum possible makespan lower bound.

[0080] Task ordering, the value of each task is arranged in descending order, get the priority list of tasks. In the subsequent resource allocation phase, the scheduler will first schedule higher ranking (greater impact on makespan) task.

[0081] Enumerate the resource set to list all the computing resources available in the system as a set:

[0082] R = {R1, R2, …, R K}

[0083] Set the initial available time to each resource R k , record its current "available time":

[0084] avail(R k ) = 0

[0085] The value represents the earliest idle time when no task is allocated on the resource.

[0086] Get the task scheduling list, extract the ordered list from the task priority order obtained:

[0087]

[0088] Among them has the highest priority (the largest ), in turn, decrease.

[0089] Assign tasks according to priority, for each task T i in the list L, perform the following operations:

[0090] Prepare the predecessor completion time, if T i has a predecessor set pred(T i ), then for each predecessor T p , record its completion time finish(T p ).

[0091] Calculate the EST and EFT on each resource, for each resource R k ∈R, calculate. The earliest start time formula is expressed as:

[0092]

[0093] Among them avail(R k ) is the idle time of the resource itself; If the predecessor T p is allocated on the same resource as T i , the corresponding communication time ​

[0094] The earliest completion time formula is:

[0095]

[0096] in It's T i In R k Execution time estimation on .

[0097] Select the optimal resource to find the resource index that minimizes the completion time. The formula is:

[0098]

[0099] T i Assigned to Record the scheduling results and update the resource status; record the start time of the task. The formula is:

[0100]

[0101] And the completion time formula is expressed as:

[0102]

[0103] Update resource availability time:

[0104]

[0105] Output schedule table Repeat the above iteration until all T i ∈L are all allocated.

[0106] Finally, output each task T i Allocation of resources Start time start(T i ) and completion time finish(T i ).

[0107] It should be noted that the combination of the critical path method and topological sorting can accurately determine system bottlenecks and scheduling priorities from a global perspective, ensure that critical tasks are executed first, effectively shorten the overall completion time (Makespan), and improve overall scheduling efficiency. Topological sorting also ensures that dependency constraints are strictly met, avoiding unnecessary waiting and conflicts. On the basis of traditional topological sorting, the "upward ranking" algorithm of the critical path method is integrated to achieve the quantification of the global contribution of tasks to the final completion time. This hierarchical recursive ranking mechanism provides a theoretically rigorous dynamic basis for priority allocation in DAG scheduling scenarios for the first time, and provides a structured input feature foundation for the introduction of machine learning enhancements in subsequent algorithms.

[0108] S3: converting the DAG into a graph neural network input, outputting a priority enhancement value using a graph convolution network, and fusing the priority enhancement value with the ranking according to a preset weight to readjust the task priority.

[0109] Further, outputting the priority enhancement value using the graph convolution network includes extracting, for each node, a single average execution time on a computing resource, a number of predecessor tasks, a number of successor tasks, and a remaining workload ranking, and combining the normalized values as a node feature.

[0110] A multi-layer graph neural network is built, each layer including three mappings: mapping one acting on the node feature, mapping two acting on the feature from the predecessor node, and mapping three acting on the edge feature; all mappings are linear transformations.

[0111] In each layer, the node first performs a linear transformation on the upper layer representation. All predecessor nodes are traversed, and the upper layer representation of each predecessor node is linearly transformed, and then the linear transformation of the corresponding edge feature is added; the node transformation result and all predecessor transformation results are added, and the new feature representation of the node in the current layer is obtained through ReLU; the iteration is repeated to the maximum iteration number, and each node obtains a high-dimensional representation containing the calculation characteristics and topological information of itself and neighbors.

[0112] After the last layer of the graph neural network, a set of readout parameters is added to linearly map the high-dimensional representation of each node to obtain a priority offset value.

[0113] A fusion weight is set to balance the upward ranking and the priority offset value output by the GNN; for each node, the upward ranking of the node is proportionally weighted and summed with the priority offset value to obtain a final scheduling priority score; the new score is used to reorder all tasks to obtain a scheduling sequence that can dynamically respond to bottlenecks and key nodes in the graph.

[0114] Further, the node feature vector for each task node T i , the initial feature vector is constructed as:

[0115]

[0116] wherein, represents the average execution time of the task on all resources. |pred(T i )|, |succ(T i )| represents the number of predecessor / successor tasks (discrete value, which can be normalized), represents the upward ranking.

[0117] The edge feature vector for each directed dependency edge (T p →Ti ), build:

[0118]

[0119] denotes data transmission time (standardized), "1" indicates the weight of a single dependent place.

[0120] GNN encoder designs network topology. Wherein the number of layers L, the hidden dimension d of each layer. Use graph convolution or graph attention (GAT) layer with edge features.

[0121] Message passing rules, let the k-th layer node be represented as The 0th layer is For each layer k = 1…L, execute:

[0122]

[0123] wherein, is the weight matrix to be learned. σ(·) is the ReLU activation. The output readout layer is linearly mapped after the Lth layer, and the "priority score offset" is obtained:

[0124]

[0125] wherein, Node priority score prediction. Offset meaning δ i denotes the GNN's task T i "Priority increment relative to traditional ranking".

[0126] δ i > 0 indicates that the GNN believes that the priority of this task should be raised δ i <0 indicates that its priority can be moderately reduced. Normalization processing maps all {δ i} to the [-1, 1] interval, ensuring numerical stability.

[0127] Priority fusion and scheduling reordering, the fusion formula introduces a fusion coefficient α ∈ [0, 1], which combines the classic upward ranking with the GNN output:

[0128]

[0129] wherein, denotes the standardized δ i 2 . Rearrange the schedule according to the priority (T i ) descending to generate a new task queue, which is passed into the HEFT allocation process to dynamically respond to bottlenecks and topological criticality.

[0130] Model training and online inference, training set construction samples multiple labeled DAG instances from historical DAG scheduling logs.

[0131] Each node gives the real scheduling "precedence order" or the contribution to the global Makespan (as a regression label y i ).

[0132] The loss function is expressed as:

[0133]

[0134] The first term is the mean square error, and the second term is the L2 regularization (weight decay coefficient λ).

[0135] When the validation set error no longer decreases, the trained GNN encoder is loaded into the scheduling service. Each time a DAG is constructed or updated, a forward propagation is performed to calculate all δ i , and a new scheduling priority queue is generated by fusion

[0136] It should be noted that the earliest start and earliest completion time of the task on each resource is evaluated in parallel, and the dynamic available resource pool information is combined in real time to realize the adaptive optimal matching of tasks and resources. Unlike pure greedy or static allocation strategies, this method effectively compatible with heterogeneous environment and task dependency complexity, improves the intelligence and adaptability of scheduling.

[0137] S4: Trigger incremental rescheduling when new tasks join and task execution fails, and the graph convolution network locates and calculates the task priority of the affected subgraph nodes, and performs local reallocation in the original scheduling table.

[0138] Further, the trigger incremental rescheduling includes deploying the trained GNN model for scheduling service, and real-time monitoring of new tasks, canceled tasks, and modification events of task dependencies; starting from the task nodes and dependency relationships designed in the event, performing bounded traversal in the predecessor and successor directions, finding tasks that have not started execution, and obtaining an affected task set.

[0139] For each affected task, collect execution stability, task connectivity, remaining workload, and edge feature information; input the subgraph corresponding to the affected task into the GNN; output the priority offset value and resource affinity score of the affected task node through inference; the resource affinity score represents the preference score of the task resource combination under simulation execution.

[0140] According to the resource affinity and priority offset value, a comprehensive priority is obtained by fusion; and according to the comprehensive priority, an incremental scheduling decision is made for the affected tasks.

[0141] The local re-allocation includes a difference between a predicted completion time and an actual completion time, and an increment subgraph as a new sample, and a small gradient update of the GNN; after each new DAG formation and update, the GNN is forwardly run, the comprehensive priority of all nodes is updated, and the task start time and resource allocation input execution layer derived from the increment scheduling are executed according to the task metadata {ET, pred, succ}.

[0142] Further, the preliminary subgraph positioning processes three cases, a new task is added, the task metadata {ET, pred, succ} is captured, a task execution fails, and a task is canceled, and the failed task T f and the allocated resource R fail .

[0143] The relevant node set AV = {T new} or {T f} is changed as a seed, a breadth-first search in a successor direction of the DAG is performed once with a limited depth d, all affected tasks that have not started or are queuing are collected to form a subgraph node set A.

[0144] Meanwhile, the edge set E A of these nodes in the original DAG is retained.

[0145] The node feature vector x i includes an original feature, an average execution time , a predecessor / successor degree |pred(T i )|, |succ(T i )|, and an original directed ranking

[0146] An event identifier, a newly added task is marked with a vector “+1”, a failed task is marked with “1”, and the rest is 0.

[0147] An edge feature, a data transmission time or a simple assignment 1 indicates that there is a dependency.

[0148] GCN forward inference and priority reevaluation. An adjacency matrix A A is constructed, for the subgraph G[A], a normalized adjacency matrix with a self-loop is constructed:

[0149]

[0150] Multi-layer graph convolution is performed, and an initial hidden representation The formula of the lth layer is represented as:

[0151]

[0152] Wherein, σ is ReLU, W (l) is a trainable weight. The priority score output takes the last layer representation H(L) The new priority enhancement value of each node is calculated by the fully connected readout layer:

[0153]

[0154] The delta i is normalized to [0, 1] as the GCN score.

[0155] The traditional ranking is fused to calculate the fusion priority of each affected node T i ∈ A, and the formula is:

[0156]

[0157] Wherein, α ∈ [0, 1] controls the proportion of CPM heuristic and GCN score.

[0158] Local sorting, according to priority ′ Reorder the tasks in A from large to small, generate a new local scheduling sequence S A Local resource reallocation and schedule table update, lock area reservation keeps the resource allocation and start / complete time of the tasks that have been started or completed in the original schedule table unchanged.

[0159] Perform HEFT local scheduling, according to the order in S A Enumerate the resources R i for each task T k , calculate the incremental earliest start / finish time, and the formula is:

[0160]

[0161] Select R k that makes EFT minimum and assign, update avail(R k ). Schedule table local merging, write the new assignment result of S A back to the global schedule table, overwrite the old A area entries. The unaffected task entries remain unchanged

[0162] The scheduling system supports real-time changes of tasks and dependencies at runtime, greatly enhancing the flexibility and robustness of the platform. It can adaptively handle task addition, cancellation or dependency adjustment, maintain the consistency and optimality of resource allocation and task execution, and is suitable for large-scale distributed and dynamic business scenarios.

[0163] The DAG dynamic scheduling scheme with "incremental adjustment" and "online GNN fine-tuning" as the core. Unlike global recalculation, the system only rearranges the affected subgraph locally, significantly reducing the scheduling overhead; the GNN model supports small-step online learning, quickly adapting to topology changes. This combination not only guarantees the real-time and optimality of scheduling, but also promotes the development of intelligent scheduling systems towards self-adaptation and self-evolution.

[0164] It should be noted that by automatically learning the potential scheduling rules in the DAG structure through the graph neural network (GNN), the bottleneck nodes and critical path tasks can be dynamically identified, and the intelligent enhancement of task priority can be achieved. With the accumulation of historical data, the model continuously evolves to adapt to different types of task structures and resource states, making the scheduling strategy more refined and personalized. By modeling the DAG scheduling as the input of the graph neural network, combining the multi-dimensional features of tasks and dependencies, and outputting the enhanced priority through end-to-end training, this hybrid scheduling scheme that combines deep learning and classical heuristics breaks through the limitations of previous scheduling methods that rely solely on static rules or a single indicator, achieving a deep integration of structure recognition and decision optimization.

[0165] Embodiment 2, as an embodiment of the present application, provides an intelligent dynamic big data platform resource scheduling method. In order to verify the beneficial effects of the present application, economic benefit calculation and simulation experiments are carried out for scientific demonstration.

[0166] Firstly, this embodiment takes the scientific computing task scheduling scene of a big data computing center as the background, and the system contains 20 heterogeneous computing nodes, which need to complete a batch of scientific tasks containing various dependency relationships. In actual operation of the platform, high-throughput and high-concurrency multi-task automatic scheduling needs to be realized to adapt to the dynamic changes of task distribution, load and resource state. In order to objectively compare the advantages and disadvantages of the intelligent dynamic scheduling method and the traditional heuristic scheduling, the existing earliest finish time priority (EFT) method and the "DAG-load balancing-topology GNN enhancement-incremental scheduling" combined method of the present application are respectively used to complete the scheduling and execution of the same batch of tasks.

[0167] The detailed process of the experiment is as follows:

[0168] The platform receives 40 scientific computing tasks in batches, respectively gives the task execution time interval ([3, 45] minutes), and collects the dependency relationship and data transmission overhead (0-5 minutes) between each task. The task information is used to automatically build a DAG, forming a task network.

[0169] The DAG after splitting is topologically sorted, and the upward ranking (average execution time plus maximum data transmission delay accumulation) of each task node is calculated according to the bottom-up principle, forming a global priority sequence, and determining the bottleneck nodes and critical paths.

[0170] The platform encodes the global DAG by using a three-layer graph convolution network (GCN) based on node features (average duration, dependency, remaining workload, etc.) and edge features (data transmission delay), and outputs a node priority enhancement value. The GNN output and upward ranking are fused with a weight of 0.6:0.4, the task priority is reordered, and the bottleneck task scheduling is dynamically strengthened.

[0171] During execution, the system randomly injects disturbances such as task addition and task failure, triggering the incremental scheduling process. The GNN quickly locates the affected subgraph and calculates its priority and resource affinity, and the platform only reallocates resources locally to related tasks, while unaffected tasks maintain their existing allocation.

[0172] The overall completion time, CPU / GPU utilization rate, task waiting time, and other key indicators caused by sudden changes are recorded and compared with the traditional EFT method.

[0173] In terms of total completion time, the traditional EFT algorithm is difficult to cope with the load concentration of bottleneck nodes or critical path tasks, and is prone to cause resource idleness and critical task blocking. After using the DAG-GNN method of the present application, the tasks are reasonably split, the critical path identification is more accurate, the GNN further enhances the perception and optimization of dependent and bottleneck nodes, and the global task completion time is shortened from 138 minutes to 122 minutes (improved by 17.2%). When new tasks or execution failures occur, the method of the present application maintains the global Makespan improvement through subgraph positioning and local rearrangement, while the traditional scheme is affected by global rescheduling, and the completion time increases significantly.

[0174] In terms of resource utilization, the traditional EFT allocation strategy fails to fully consider task splitting and dynamic resource affinity. The present application actively splits large tasks based on the load balancing principle and relies on the feature learning of GNN to effectively improve the global and local parallelism, and the resource utilization rate is stably improved to more than 78%.

[0175] The average task waiting time is significantly reduced, especially under dynamic disturbance events. The intelligent scheduling method shortens the waiting period of affected tasks from 28 minutes to about 14 minutes through priority rearrangement and local resource reallocation, improving efficiency by nearly 50%.

[0176] The increase in the number of critical path tasks reflects that after GNN optimization, the scheduling system can actively identify and strengthen the parallelism of the bottleneck path. The GNN output priority adjustment number reaches 13-14 times, indicating that the system can dynamically identify bottleneck and acceleration tasks based on graph topology and node characteristics, effectively making up for the shortcoming of the classic algorithm that "has limited global dependency and parallelism perception ability".

[0177] Example 3 is an embodiment of the present invention, which provides an intelligent dynamic big data platform resource scheduling system, including a task management module, a priority calculation module, a HEFT scheduling module, and a dynamic optimization module.

[0178] The task management module is used to collect task information, build and maintain the DAG structure.

[0179] The priority calculation module is used to perform topological analysis, upward ranking and priority sorting on the DAG.

[0180] The HEFT scheduling module is used to map tasks to specific resources according to priority and handle communication overhead.

[0181] The dynamic optimization module is used for runtime monitoring, GNN feature extraction and model online fine-tuning.

[0182] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of each embodiment of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.

[0183] The logic and / or steps represented in the flowcharts or otherwise described herein, for example, can be considered as an ordered list of executable instructions for implementing the logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (e.g., a computer-based system, a system including a processor, or other system that can fetch and execute instructions from an instruction execution system, apparatus, or device). For purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by, or in conjunction with, an instruction execution system, apparatus, or device.

[0184] More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection (electronic) having one or more wires, a portable computer diskette (magnetic), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can also be paper or another suitable medium upon which the program is printed, as the program can be electronically captured, for example, via an optical scanner, then compiled, interpreted, or otherwise processed, and stored in a computer memory in a form that is then reproducible into a computer readable medium.

[0185] It should be understood that aspects of the application can be implemented in hardware, software, firmware or combinations thereof. In the embodiments described above, various steps or methods can be implemented, for example in software or firmware, stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, and in another embodiment, any of the following technologies, known in the art, or combinations thereof, can be used: discrete logic circuitry having logic gates for implementing logic functions upon an application of data signals, application specific integrated circuits having appropriate combinational logic gates, programmable gate arrays (PGA), field programmable gate arrays (FPGA), and so forth. It should be appreciated that the foregoing examples have been given for illustrative purposes only and are not intended to limit the techniques of the present application, as described herein, in that as persons skilled in the art will recognize from this disclosure that other configurations comprising substitutions, combinations and / or equivalents of those described herein can be utilized, and that the scope of the present application is not limited to the specific embodiments disclosed.

[0186] It should be understood that the foregoing examples have been given for illustrative purposes only and are not intended to limit the techniques of the present application, as described herein, in that as persons skilled in the art will recognize from this disclosure that other configurations comprising substitutions, combinations and / or equivalents of those described herein can be utilized, and that the scope of the present application is not limited to the specific embodiments disclosed.

Claims

1. Intelligent dynamic big data platform resource scheduling method, characterized in that: include: Collect information about tasks to be scheduled, construct a directed acyclic graph (DAG), and split tasks according to the load balancing principle; Topologically sort the DAG, determine the priority order of tasks, and allocate available resources to tasks according to the priority order; Convert the DAG into a graph neural network input, use the graph convolutional network to output a priority enhancement value, and fuse the priority enhancement value with the ranking according to the preset weight to readjust the task priority; When new tasks are added or task execution fails, incremental rescheduling is triggered. The graph convolutional network locates and calculates the task priorities of the affected subgraph nodes and performs local redistribution in the original scheduling table.

2. The intelligent dynamic big data platform resource scheduling method according to claim 1, characterized in that: The task splitting according to the load balancing principle includes calculating the average execution level of the tasks to be scheduled and setting a splitting threshold according to the average execution level; Traverse the tasks in the directed acyclic graph (DAG) and determine whether the average execution time of the tasks on all resources exceeds the threshold; If it does not exceed the limit, the task will be kept as it is; If it exceeds, the task will be divided into multiple balanced subtasks; For each split task, remove the original node from the graph and create multiple new nodes based on the number of evenly divided subtasks. Each new node inherits all predecessor dependencies of the original node. If subtasks must be executed serially, add dependencies between them in sequence; otherwise, keep them in parallel. Set the corresponding execution duration information for each subtask. After the split is completed, the DAG is checked for cycles, dependency cycles are removed, and the predecessor and successor lists of the newly added nodes are updated; The information of the tasks to be scheduled includes the execution time, dependency and data transmission overhead of each task.

3. The intelligent dynamic big data platform resource scheduling method according to claim 2, characterized in that: The topological sorting of the DAG includes performing a traversal search on the DAG to generate a topological sequence that satisfies dependency constraints; Based on the topological sequence, the scheduler starts from the sequence head and calculates the upward ranking of each node from the bottom up; for an exit node without successors, the upward ranking is equal to the average execution time of the task itself; For other nodes, the upward ranking is equal to the average execution time of the task plus the data transmission time between the successor node and the upward ranking of the successor node, and the maximum value among all successor node paths is taken; Identify the critical tasks that affect the global completion time, that is, the task path with the highest cumulative upward ranking from the entry to the exit; the scheduler arranges all tasks from large to small according to the upward ranking to obtain the task priority order.

4. The intelligent dynamic big data platform resource scheduling method according to claim 3, characterized in that: Allocating available resources to tasks includes making incremental scheduling decisions to determine the earliest possible start time for the current task, adding the estimated runtime of the current task on the resources to be allocated, and evaluating the completion time; Compare the current task with the available resources and assign the current task to the resource that can complete it the earliest; The idle time of the selected resource is updated to the actual completion time of the task, and the start and end time of the current task on the resource are recorded; all tasks are assigned in order of priority to obtain a schedule.

5. The intelligent dynamic big data platform resource scheduling method according to claim 4, characterized in that: The use of the graph convolutional network to output the priority enhancement value includes extracting the average execution time of each node on the computing resources, the number of predecessor tasks, the number of successor tasks, and the ranking of the remaining workload, and combining them into node features after normalization; For each dependency edge, record the data transmission time and constant mark between the corresponding tasks, and use it as the edge feature after normalization; Build a multi-layer graph neural network, with each layer including three mappings: Mapping 1 acts on the node's own features, Mapping 2 acts on the features from the predecessor nodes, and Mapping 3 acts on the edge features; All mappings are linear transformations; In each layer, the node first performs a linear transformation on its upper layer representation; Traverse all predecessor nodes, perform a linear transformation on the upper layer representation of each predecessor node, and then add the linear transformation of the corresponding edge feature; add the node's own transformation result to all predecessor transformation results, and use ReLU to obtain the new feature representation of the node layer; repeat the iteration to the maximum number of times, and each node obtains a high-dimensional representation that contains the computational characteristics and topological information of itself and its neighbors; After the last layer of the graph neural network, a set of readout parameters is added to perform linear mapping on the high-dimensional representation of each node to obtain the priority offset value; Set the fusion weight to balance the upward ranking and the priority offset value of the GNN output; For each node, the node's upward ranking is proportionally weighted and summed with the priority offset value to obtain the final scheduling priority score; all tasks are reordered using the new score to obtain a scheduling sequence that can dynamically respond to bottlenecks and key nodes in the graph.

6. The intelligent dynamic big data platform resource scheduling method according to claim 5, characterized in that: Triggering incremental rescheduling includes deploying the trained GNN model for scheduling services and monitoring new tasks, task cancellations, and task dependency modification events in real time; Starting from the task nodes and dependencies designed in the event, perform bounded traversal along the predecessor and successor directions to find tasks that have not started execution and obtain the set of affected tasks. Task execution failure includes task cancellation and task dependency modification. For each affected task, execution stability, task connectivity, remaining workload, and edge feature information are collected. The subgraph corresponding to the affected task is used as a single graph sample and input into the GNN. The priority offset value and resource affinity score of the affected task node are obtained through inference output. The resource affinity score represents the preference score of the task resource combination under simulated execution. The resource affinity and priority offset values ​​are integrated to obtain a comprehensive priority; incremental scheduling decisions are made for the affected tasks based on the comprehensive priority.

7. The intelligent dynamic big data platform resource scheduling method according to claim 6, characterized in that: The local reallocation includes taking the difference between the predicted completion time and the actual completion time and the incremental subgraph as new samples and performing a small amount of gradient update on the GNN; After each new DAG is formed and updated, the GNN is run forward to update the comprehensive priority of all nodes, and the task start time and resource allocation obtained by incremental scheduling are input into the execution layer.

8. A system using the intelligent dynamic big data platform resource scheduling method according to any one of claims 1 to 7, characterized in that: Including task management module, priority calculation module, HEFT scheduling module, dynamic optimization module; The task management module is used to collect task information, build and maintain the DAG structure; The priority calculation module is used to perform topological analysis, upward ranking and priority sorting on the DAG; The HEFT scheduling module is used to map tasks to specific resources according to priority and handle communication overhead; The dynamic optimization module is used for runtime monitoring, GNN feature extraction and model online fine-tuning.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the intelligent dynamic big data platform resource scheduling method described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the intelligent dynamic big data platform resource scheduling method described in any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • Cloud edge-end resource scheduling optimization method based on double-layer graph neural network

    CN118134029A

  • Cloud computing parallel task optimization scheduling method based on priority dependency graph

    CN119806776A

  • DAG task scheduling method and apparatus, device, and storage medium

    WO2023241000A1

Cited By

  • Robot production scheduling method and system based on task priority

    CN121543999A