Complex equipment sensing data calculation task scheduling method
By adopting a directed acyclic graph task scheduling method based on Top-p sampling strategy on the cloud computing platform, combined with graph neural network and multi-head attention mechanism, the problem of low scheduling of complex equipment perceived data calculation tasks in the existing technology is solved, and more efficient task scheduling and resource utilization are achieved.
Patent Information
- Application Number
- CN202510115766.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-24
- Publication Date
- 2025-05-06
AI Technical Summary
When handling complex equipment-aware data computing tasks, existing cloud computing platforms are difficult to make full use of the structured characteristics and dependencies of the tasks, resulting in low online task execution efficiency and waste of computing resources for offline tasks.
Directed acyclic graph task scheduling method based on Top-p sampling strategy is adopted, and the p-value is dynamically adjusted to adapt to different computing task types, and DAG is encoded in combination with graph neural network, and the multi-head attention mechanism is used to explore potential relationships and optimize the task scheduling order.
It improves the scheduling efficiency of complex equipment-aware data computing tasks, can handle online and offline tasks more efficiently, reduce waste of computing resources, and improve the overall operating efficiency of the system.
Smart Images

Figure CN119938279A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of complex equipment perception data, and more specifically to a method for scheduling complex equipment perception data computing tasks. Background Art
[0002] As a typical skeleton-type assembly product, complex equipment has complex structural composition and constraints, and is assembled from tens of thousands of detachable parts. High-end energy equipment is a typical product of this type of equipment. Taking its bogie subsystem as an example, it needs to be closely monitored by installing various sensors during the operation and maintenance stage. These sensors and other data acquisition devices will continue to collect complex sensory data such as temperature, pressure, and vibration throughout the life cycle of the bogie. With the in-depth operation and maintenance of energy equipment such as wind power, these sensory data are growing explosively, and the processing of massive data has become a major challenge that the industry needs to solve urgently.
[0003] Faced with such a large and complex sensory data resource, cloud computing has become the preferred solution due to its powerful data processing capabilities. However, with the rapid growth of data scale, the computing tasks undertaken by cloud computing platforms are becoming increasingly heavy. At present, cloud computing platforms present significant characteristics such as highly heterogeneous data center clusters, high volatility of cluster operating environments and task submission numbers, and diversity of optimization targets, which puts unprecedented high demands on the task scheduling capabilities of cloud computing platforms. When solving directed acyclic graph (DAG) scheduling tasks in sensory data computing tasks, current task scheduling algorithms are difficult to fully utilize their structured characteristics and dependencies, and fail to effectively combine the characteristics of online and offline tasks, resulting in low efficiency when executing online tasks and waste of computing resources for offline tasks. Summary of the invention
[0004] In order to overcome the defects existing in the above-mentioned prior art, the present invention discloses a method for scheduling complex equipment perception data computing tasks. The present invention establishes a directed acyclic graph task scheduling method based on the Top-p sampling strategy, uses a strategy of dynamically adjusting the p value to adapt to different computing task types, and effectively improves the scheduling efficiency of complex equipment perception data computing tasks.
[0005] In order to achieve the above objectives, the technical solution adopted by the present invention is:
[0006] A method for scheduling complex equipment perception data computing tasks, comprising the following steps:
[0007] 1. Directed Acyclic Graph Coding
[0008] Step S1, encoding the directed acyclic graph to obtain the embedding of each node;
[0009] Preferably, step S1 comprises the following steps:
[0010] Step S11, using Topoformer to encode the directed acyclic graph DAG to obtain multiple subgraphs;
[0011] In the present invention, step S11 encodes the DAG using Topoformer, which is similar to the traditional Transformer. The propagation rule of Topoformer is the same as that of Transformer encoder, but a mask is added to ensure that each attention head only focuses on its assigned graph. Step S11 specifically includes steps S111-S115.
[0012] Preferably, step S11 includes the following steps:
[0013] Step S111, generating a directed acyclic graph DAG(G) of the tasks formed by the dependency relationship between intelligent computing tasks, and a transfer specification subgraph TR(G) of DAG(G);
[0014] In step S111, a transitive reduction subgraph TR(G) of the original graph DAG(G) is generated. The transitive reduction is a subgraph with the least edges that can lead to the same partial order relationship. In the partial order relationship, if a and b have a partial order relationship, and b and c have a partial order relationship, then a and c also have a partial order relationship.
[0015] Step S112, after removing TR(G) from DAG(G), a subgraph G\ETR(G) is generated;
[0016] Step S113, generating a subgraph TC(G)\E by removing the edges of the directed acyclic graph from the transitive closure TC of DAG(G), wherein the transitive closure is a graph containing all edges of potential partial orders;
[0017] Step S114, reverse the directions of the subgraph edges obtained from steps S111 to S113 to generate three subgraphs;
[0018] Step S115: bidirectionally connect all nodes without partial order relations in DAG (G) to generate a subgraph.
[0019] In the present invention, steps S111 to S113 obtain three sub-graphs, step S114 obtains three sub-graphs, and step S115 obtains one sub-graph, so seven sub-graphs are obtained in total.
[0020] Step S12: construct a multi-head attention mechanism layer based on the directed acyclic graph DAG and multiple subgraphs, and perform scaled dot product attention on each subspace in the multi-head attention mechanism layer;
[0021] Preferably, step S12 includes:
[0022] In step S11, 7 subgraphs are obtained by encoding the directed acyclic graph DAG. Figure 1 It is mapped to 8 subspaces, which is the number of subspaces in the multi-head attention mechanism;
[0023] In the multi-head attention layer, we create an input sequence X = [x1, x2, ..., x n ], where n is the length of the sequence, each x i is a vector that maps the input sequence to h = 8 subspaces, each of which has a different weight matrix W. Q ,W K ,W V , where Q represents query, which represents the vector that generates attention weights, K represents key, which represents the vector that calculates the degree of association between query and input; V represents value, which represents the input vector of weighted summation;
[0024] For each subspace i, calculate the query vector Q i =XW i Q , key vector K i =XW i K Sum value vector V i =XW i V ; Then, by performing a dot product between the query vector and the key vector, we get the attention score of the i subspace And scale the attention score; where S i is the attention score of subspace i, T is the set of all scheduling orders that satisfy the given DAG (G) priority and resource constraints, d k is the dimension of the key vector K.
[0025] Step S13: feature connect the scaled dot product attention results and perform linear transformation to obtain the embedding of each node.
[0026] 2. Node decoding and sequence generation distribution
[0027] Step S2: decode each node and generate sequence distribution based on the Top-p sampling strategy;
[0028] Preferably, step S2 comprises the following steps:
[0029] Step S21, decode each node, and the decoder uses the chain rule of conditional probability to gradually generate the distribution of the target sequence according to the generated sequence and node embedding information;
[0030] In the present invention, S21 is a random strategy derived from the effective topological order of the graph, and each factor selects the probability of the next node according to the previous node and the embedding of the graph, which will generate multiple optional paths. S22 selects the next node based on the optional nodes generated by S21 using the Top-p sampling strategy.
[0031] Preferably, the distribution of the target sequence in step S21 includes:
[0032]
[0033] Among them, p(σ|G) is the distribution of the sequence obtained by using the sampling strategy according to the effective topological order of the graph, V represents all nodes, G represents the directed acyclic graph, σ represents the sequence element, | represents the probability of the previous variable appearing under the latter condition, t represents the tth node, and p θ represents the node probability considering the network weight, σ t represents the tth node element, σ 1:t-1 represents the preceding node element of t, from 1 to t-1 elements, h represents the output embedding dimension of each node, and σ1 represents the first node element.
[0034] In step S21, each node is decoded, and the decoder uses the chain rule of conditional probability to gradually generate the distribution of the target sequence based on the generated sequence and node embedding information, that is, it calculates the conditional probability of each possible sequence element in the current state in an iterative manner, and constructs the distribution of the entire sequence based on this, as shown in the above formula. The above formula decomposes the strategy into a product form, and each factor is the probability of selecting the next node based on the previous node and graph embedding.
[0035] Step S22, input the h-dimensional embedding of each node into a separate multilayer perceptron, output the priority logits and network weight θ of each node, use the Top-p sampling strategy to select important nodes from the output priority logits and generate a sequence distribution of the nodes;
[0036] Preferably, step S22 outputs the priority logits and network weight θ of each node, including:
[0037]
[0038] Among them, logits is the priority, θ is the network weight, v is any node in DAG, G is a directed acyclic graph, := means that the previous variable is defined by the following structure, MLP θ Represents GNN based on multi-layer perceptron θ represents the random strategy set of the output of step S21, represents the reward, which is used to indicate the action taken by the algorithm, and V represents the set of DAG nodes.
[0039] Preferably, the Top-p sampling strategy of step S2 comprises the following steps:
[0040] Step S221, calculating the probability distribution of the node through the softmax function;
[0041] Step S222, sort the probability distribution from high to low to obtain a sorted probability distribution;
[0042] Step S223: Select a minimum subset whose cumulative probability does not exceed a given threshold p from the sorted probability distribution as a sampling candidate set, where the subset includes the sequence when the probability accumulates to p;
[0043] Step S224, re-sampling from the candidate set according to probability;
[0044] Step S225: Control the diversity of the generated sequence by adjusting the threshold p.
[0045] Preferably, in the Top-p sampling strategy of step S2, when scheduling online computing tasks of complex equipment perception data, a threshold p close to 0 is selected, and when scheduling offline computing tasks of complex equipment perception data, a threshold p close to 1 is selected; after selecting the threshold p according to different scenarios, a sequence distribution of nodes is generated.
[0046] Step S23: Perform logits norm regularization on a single directed acyclic graph DAG and optimize the training process.
[0047] In step S23, logits norm regularization is used for a single DAG to improve the generalization ability of the model and optimize the training process.
[0048] 3. Sorting
[0049] Step S3, using a sorting algorithm on the sequence distribution to obtain the priority of each node;
[0050] In step S3, the priority of each node is obtained by using a sorting algorithm based on the node weight for the sequence distribution of the nodes.
[0051] 4. List Scheduling
[0052] Step S4, using the list scheduling algorithm for the priority of each node to obtain the scheduling order of the tasks;
[0053] Preferably, step S4 comprises the following steps:
[0054] Step S41, input the node priority list, clear the decision time, the decision time is the execution time of the task;
[0055] Step S42: Find nodes that are currently in a schedulable state. The schedulable nodes are nodes whose previous tasks have been completed. Set these nodes as ready nodes.
[0056] Step S43: Schedule the ready nodes according to the node priority order until all the ready nodes are scheduled or there are no available resources;
[0057] Step S44: For the currently unfinished nodes, change their decision time to the earliest completion time, and then repeat steps S42 to S44 until all nodes are scheduled and the scheduling order of all tasks is obtained.
[0058] 5. Improvement of model loss function
[0059] Step S5: Calculate the total scheduling execution time and improve the loss function of the model.
[0060] Preferably, in step S5, the loss function of the model is:
[0061]
[0062] Where LS represents the list scheduling algorithm of step S4, C LS (t, G) represents the completion time of the directed acyclic graph G in order t, t~π(·G) represents the probability that the directed acyclic graph G is scheduled in order t, θ is the neural network weight, is the expected value, G~GS is the probability of G in the DAG set GS.
[0063] In the present invention, the above-mentioned model refers to the implementation structure of the DAG task scheduling method proposed in the present invention using a Topoformer encoder and introducing a Top-p strategy in the probability sampling process, including steps S1 to S5, constituting a graph neural network model structure.
[0064] Beneficial effects of the present invention:
[0065] 1. The fault diagnosis task in the complex equipment perception computing task is a typical directed acyclic graph computing task. The fault diagnosis algorithm relies on the input data of multiple sensors, and these input data are based on the warning signals of the fault warning algorithm, that is, a multi-level dependency relationship is formed, and the execution conditions have a sequence. For different tasks in this type of DAG structure, the current scheduling algorithm is often unable to handle the complexity and dependency of the task, resulting in poor scheduling effect of the algorithm in structured tasks. The present invention uses a graph neural network to encode the DAG, and uses a multi-head attention mechanism through transitive reduction, transitive closure and analysis of partial order relations to fully explore the potential relationship between DAGs.
[0066] 2. Perception data computing tasks contain both online tasks and offline tasks, and their different characteristics can be utilized to improve the efficiency of task scheduling. Based on the Top-p sampling strategy, the present invention adjusts the p value according to different task types, controls the diversity of generated sequences to obtain more scheduling sequences, uses a smaller p value for online tasks to obtain more stable output results, and uses a larger p value for offline tasks to find a better task scheduling strategy as much as possible at the cost of a higher algorithm execution time. BRIEF DESCRIPTION OF THE DRAWINGS
[0067] Figure 1 The overall structure of the network model of the present invention is a flow chart;
[0068] Figure 2 It is a multi-head attention mechanism diagram of the directed acyclic graph of the present invention;
[0069] Figure 3 Generating a weighted sum graph for the multi-head attention of the directed acyclic graph of the present invention;
[0070] Figure 4 Schematic diagram of logits and network weights of nodes of the present invention;
[0071] Figure 5 Schematic diagram of the list scheduling algorithm of the present invention. DETAILED DESCRIPTION
[0072] The concept, specific structure and technical effects of the present invention will be clearly and completely described below in conjunction with the embodiments and drawings to fully understand the purpose, characteristics and effects of the present invention.
[0073] The present invention discloses a scheduling method for complex equipment perception data computing tasks. Specifically, the method is based on the existing Topoformer model and proposes a DAG task scheduling algorithm based on the Top-p sampling strategy. The algorithm combines the Top-p sampling strategy with the Topoformer model, fully considers the different characteristics of online and offline tasks in complex equipment perception data computing, and performs targeted optimization. The algorithm first uses the Topoformer model to encode the directed acyclic graph, and generates an embedded representation for each node through a multi-head attention mechanism. Subsequently, the embedded representation of each node is input into a separate multi-layer perceptron to generate a sequence distribution. According to different task types, the algorithm selects an appropriate p value and determines the priority of each node through the Top-p sampling strategy. Finally, the order of task scheduling is obtained by combining the sorting algorithm and the list scheduling algorithm. Since a larger p value can generate a richer sequence distribution, the DAG task scheduling algorithm based on the Top-p sampling strategy can more efficiently find the optimization strategy for task scheduling. In addition, the algorithm is also optimized for a single DAG, and the total completion time of the task is further reduced by improving the model loss function.
[0074] like Figure 1 As shown, the present invention comprises the following steps:
[0075] Step S1: Encode the directed acyclic graph to obtain the embedding of each node.
[0076] Step S2: Decode each node and generate sequence distribution.
[0077] Step S3: Use a sorting algorithm on the sequence distribution to obtain the priority of each node.
[0078] Step S4: Use the list scheduling algorithm for the priority of each node to obtain the scheduling order of the tasks.
[0079] Step S5: Calculate the total scheduling execution time and improve the loss function of the model.
[0080] In this embodiment, step S1 is to encode the directed acyclic graph to obtain the embedding of each node, which specifically includes steps S11-S13.
[0081] In this embodiment, step S11 is to encode the DAG using Topoformer, which specifically includes steps S111-S115. Figure 2 shown.
[0082] Step S111: Generate a transfer reduction subgraph TR(G) of the original graph DAG(G), wherein the original graph DAG(G) refers to a task DAG graph formed by the dependency relationship between intelligent computing tasks, which is represented by the variable G in subsequent formulas;
[0083] Step S112: After removing TR(G) from G, a subgraph G\ETR(G) is generated;
[0084] Step S113: Remove the edges of G from the transitive closure TC of G and generate a subgraph TC(G)\E;
[0085] Step S114: Reverse the direction of the graph edges obtained from steps S111 to S113 to generate three subgraphs;
[0086] Step S115: bidirectionally connect all nodes without partial order relationships to generate a subgraph.
[0087] In this embodiment, step S12 is to construct a multi-head attention mechanism layer, and establish an input sequence X=[x1,x2,...,x n ], where n is the length of the sequence, each x i is a vector that maps the input sequence into h = 8 subspaces, each with a different weight matrix W Q ,W K ,W V , where Q stands for Query, which represents the vector that generates attention weights, K stands for Key, which represents the vector that calculates the degree of association between Query and input; V stands for Value, which represents the weighted sum of the input vector. For each subspace i, the query vector Q is calculated. i =XW i Q , the key vector K i =XW i K Sum value vector V i =XW i V ; Then, by taking the dot product of the query vector and the key vector, we get the attention score of the i subspace And scale the attention score, such as Figure 3 shown.
[0088] Step S13 is: feature-connecting the scaled dot product attention result of step S12 and performing a linear transformation to obtain the embedding of each node.
[0089] In this embodiment, step S2 includes steps S21-S23;
[0090] In this embodiment, step S21 is to decode each node. The decoder uses the chain rule of conditional probability to gradually generate the distribution of the target sequence according to the generated sequence and node embedding information. That is, the conditional probability of each possible sequence element in the current state is calculated in an iterative manner, and the distribution of the entire sequence is constructed accordingly, as shown in formula (1);
[0091]
[0092] Where p(σ|G) is the distribution of the sequence obtained by using the sampling strategy according to the effective topological order of the graph, V represents all nodes, G represents the directed acyclic graph, and formula (1) decomposes the strategy into a product form, where each factor is the probability of selecting the next node based on the previous node and the embedding of the graph.
[0093] In this embodiment, step S22 inputs the h-dimensional embedding of each node into a separate multi-layer perceptron (MLP), and outputs the priority logits and network weight θ of each node through formula (2), such as Figure 4 As shown, the Top-p sampling strategy is used to select important nodes from the output priority logits;
[0094]
[0095] In this embodiment, the Top-p sampling strategy of step S22 specifically includes steps S221-S225;
[0096] Step S221: Calculate the probability distribution of the node through the softmax function;
[0097] Step S222: sort the probability distribution from high to low to obtain a sorted probability distribution;
[0098] Step S223: Select the smallest subset whose cumulative probability does not exceed a given threshold value p as the sampling candidate set. This subset includes the sequence when the probability accumulates to p.
[0099] Step S224: re-sampling from the candidate set according to their probabilities;
[0100] Step S225: Control the diversity of the generated sequences by adjusting the threshold p. When scheduling online computing tasks of complex equipment perception data, select a smaller p value, and the candidate set will be limited to nodes with higher probabilities. When performing offline computing tasks, select a p value close to 1, and the candidate set includes almost all possible sequences, and the generated sequences will be more diverse. After selecting the p value according to different scenarios, the sequence distribution of the generated nodes is generated;
[0101] In this embodiment, step S23 uses logits norm regularization for a single DAG to improve the generalization ability of the model and optimize the training process.
[0102] Step S3 is: for the sequence distribution of the nodes, a sorting algorithm is used according to the weight of the nodes to obtain the priority of each node;
[0103] Step S4 is to use the list scheduling algorithm to obtain the scheduling order of tasks for each node priority, and S4 specifically includes steps S41-S44, such as Figure 5 As shown;
[0104] Step S41 is: input the node priority list and clear the decision time, which is the execution time of the task;
[0105] Step S42 is: finding nodes whose current status is schedulable, that is, nodes whose previous tasks have been completed, and setting these nodes as ready nodes;
[0106] Step S43 is: scheduling the ready nodes according to the node priority order until all the ready nodes are scheduled or there are no available resources;
[0107] Step S44 is: for the currently unfinished nodes, change their decision time to the earliest completion time, and then repeat steps S42 to S44 until all nodes are scheduled, that is, the scheduling order of all tasks is obtained;
[0108] In this embodiment, step S5 uses formula (3) as the loss function of the model, where LS represents the list scheduling algorithm of step S4, C LS (t, G) represents the completion time of the directed acyclic graph G under order t, t~π(·G) represents the probability that the directed acyclic graph G is scheduled in order t;
[0109]
[0110] The optimization goal of formula (3) emphasizes the consideration of time factors. Its core goal is to minimize the average task turnaround time to improve the overall operating efficiency and response speed of the system.
[0111] In order to test the method proposed in the present invention, the present invention selects some competitive sampling strategies for comparison, mainly including the following strategies:
[0112] Greedy strategy: Based on the value evaluation of the current state, select the action with the highest value for sampling to maximize the immediate reward.
[0113] Beam search strategy: It efficiently searches for the optimal solution by maintaining a fixed-size candidate list and selecting the sequence with the highest score to expand at each step.
[0114] In this experiment, the total completion time of the scheduling task and the execution time of the algorithm are used as two evaluation indicators to comprehensively evaluate the performance of the proposed algorithm. The total completion time of the scheduling task corresponds to the optimization objective of the network (3). The comparison of the online task scheduling results of different sampling strategies is shown in the following table, where J represents the number of tasks in JSSP, the number of computing resource virtual machines in the experiment is 20, and p represents the threshold in Top-p.
[0115]
[0116] After analyzing the table, when the p value is small and the number of tasks (J) is 500, the Top-p strategy shows certain advantages over the greedy and beam search strategies, and the greedy strategy has a greater advantage when J is 100. It is verified that the Top-p sampling strategy can generate scheduling sequences more stably and efficiently when the p value is small, taking advantage of the small number of single-batch computing tasks and high execution time requirements in online computing tasks.
[0117] The comparison of offline task scheduling results of different sampling strategies is shown in the following table.
[0118]
[0119] Through the analysis of the table, when the p value takes a larger value and after multiple inferences, in the face of complex scenarios with 1000 tasks, the Top-p strategy can still discover better solutions that have not been discovered before. Although the execution time of the algorithm is relatively long in this case, it fits the computing scenario of offline computing tasks of complex equipment perception data. In such scenarios, computing time is not the primary concern, but there are extremely high requirements for the quality of the solution. Therefore, while pursuing high-quality solutions, the Top-p strategy sacrifices some computing efficiency, but its excellent performance in complex task scheduling still has extremely high practical application value.
[0120] The above is a specific description of the implementation mode of the present invention, but the present invention is not limited to the described embodiments. Those skilled in the art may make various equivalent modifications or substitutions without violating the spirit of the present invention, and these equivalents or substitutions are all included in the scope defined by the claims of the present invention.
Claims
1. A method for scheduling complex equipment perception data computing tasks, characterized in that: The following steps are involved: Step S1, encoding the directed acyclic graph to obtain the embedding of each node; Step S2: decode each node and generate sequence distribution based on the Top-p sampling strategy; Step S3, using a sorting algorithm on the sequence distribution to obtain the priority of each node; Step S4, using the list scheduling algorithm for the priority of each node to obtain the scheduling order of the tasks; Step S5: Calculate the total scheduling execution time and improve the loss function of the model.
2. A method for scheduling complex equipment perception data computing tasks as claimed in claim 1, characterized in that: Step S1 includes the following steps: Step S11, using Topoformer to encode the directed acyclic graph DAG to obtain multiple subgraphs; Step S12: construct a multi-head attention mechanism layer based on the directed acyclic graph DAG and multiple subgraphs, and perform scaled dot product attention on each subspace in the multi-head attention mechanism layer; Step S13: feature connect the scaled dot product attention results and perform linear transformation to obtain the embedding of each node.
3. A method for scheduling complex equipment perception data computing tasks as claimed in claim 2, characterized in that: Step S11 includes the following steps: Step S111, generating a directed acyclic graph DAG(G) of the tasks formed by the dependency relationship between intelligent computing tasks, and a transfer specification subgraph TR(G) of DAG(G); Step S112, after removing TR(G) from DAG(G), a subgraph G\ETR(G) is generated; Step S113, generating a subgraph TC(G)\E by removing the edges of the directed acyclic graph from the transitive closure TC of DAG(G), wherein the transitive closure is a graph containing all the edges of the potential partial order; Step S114, reverse the directions of the subgraph edges obtained from Step S111 to Step S113 to generate three subgraphs; Step S115: bidirectionally connect all nodes without partial order relations in DAG (G) to generate a subgraph.
4. A method for scheduling complex equipment perception data computing tasks as claimed in claim 2, characterized in that: Step S12 includes: In step S11, 7 subgraphs are obtained by encoding the directed acyclic graph DAG, and mapped to 8 subspaces together with the original graph, which is the number of subspaces in the multi-head attention mechanism; In the multi-head attention layer, we create an input sequence X = [x1, x2, ..., x n ], where n is the length of the sequence, each x i is a vector that maps the input sequence to h = 8 subspaces, each of which has a different weight matrix W. Q ,W K ,W V , where Q represents query, which represents the vector that generates attention weights, K represents key, which represents the vector that calculates the degree of association between query and input; V represents value, which represents the input vector of weighted summation; For each subspace i, calculate the query vector Q i =XW i Q , key vector K i =XW i K Sum value vector V i =XW i V ; Then, by performing a dot product between the query vector and the key vector, we get the attention score of the i subspace And scale the attention score; where S i is the attention score of subspace i, T is the set of all scheduling orders that satisfy the given DAG (G) priority and resource constraints, d k is the dimension of the key vector K.
5. The method for scheduling complex equipment perception data computing tasks according to claim 1, characterized in that: Step S2 includes the following steps: Step S21, decode each node, and the decoder uses the chain rule of conditional probability to gradually generate the distribution of the target sequence according to the generated sequence and node embedding information; Step S22, input the h-dimensional embedding of each node into a separate multilayer perceptron, output the priority logits and network weight θ of each node, use the Top-p sampling strategy to select important nodes from the output priority logits and generate a sequence distribution of the nodes; Step S23: Perform logits norm regularization on a single directed acyclic graph DAG and optimize the training process.
6. A method for scheduling complex equipment perception data computing tasks as claimed in claim 5, characterized in that: The distribution of the target sequence in step S21 includes: Among them, p(σ|G) is the distribution of the sequence obtained by using the sampling strategy according to the effective topological order of the graph, V represents all nodes, G represents the directed acyclic graph, σ represents the sequence element, | represents the probability of the previous variable appearing under the latter condition, t represents the tth node, and p θ represents the node probability considering the network weight, σ t represents the tth node element, σ 1:t-1 represents the preceding node element of t, from 1 to t-1 elements, h represents the output embedding dimension of each node, and σ1 represents the first node element; Step S22 outputs the priority logits and network weight θ of each node including: Among them, logits is the priority, θ is the network weight, v is any node in DAG, G is a directed acyclic graph, := means that the previous variable is defined by the following structure, MLP θ Represents GNN based on multi-layer perceptron θ represents a random strategy set, represents the reward, which is used to indicate the action taken by the algorithm, and V represents the set of DAG nodes.
7. A method for scheduling complex equipment perception data computing tasks as claimed in claim 1, characterized in that: The Top-p sampling strategy of step S2 includes the following steps: Step S221, calculating the probability distribution of the node through the softmax function; Step S222, sort the probability distribution from high to low to obtain a sorted probability distribution; Step S223: Select a minimum subset whose cumulative probability does not exceed a given threshold p from the sorted probability distribution as a sampling candidate set, where the subset includes the sequence when the probability accumulates to p; Step S224, re-sampling from the candidate set according to probability; Step S225: Control the diversity of the generated sequence by adjusting the threshold p.
8. The method for scheduling complex equipment perception data computing tasks according to claim 1, characterized in that: In the Top-p sampling strategy of step S2, when scheduling online computing tasks of complex equipment perception data, a threshold p close to 0 is selected, and when scheduling offline computing tasks of complex equipment perception data, a threshold p close to 1 is selected; After selecting the threshold p according to different scenarios, the sequence distribution of nodes is generated.
9. A method for scheduling complex equipment perception data computing tasks as claimed in claim 1, characterized in that: Step S4 includes the following steps: Step S41, input the node priority list, clear the decision time, the decision time is the execution time of the task; Step S42: Find nodes that are currently in a schedulable state. The schedulable nodes are nodes whose previous tasks have been completed. Set these nodes as ready nodes. Step S43: Schedule the ready nodes according to the node priority order until all the ready nodes are scheduled or there are no available resources; Step S44: For the currently unfinished nodes, change their decision time to the earliest completion time, and then repeat steps S42 to S44 until all nodes are scheduled and the scheduling order of all tasks is obtained.
10. A method for scheduling complex equipment perception data computing tasks as claimed in claim 1, characterized in that: In step S5, the loss function of the model is: Where LS represents the list scheduling algorithm of step S4, C LS (t, G) represents the completion time of the directed acyclic graph G in order t, t~π(·G) represents the probability that the directed acyclic graph G is scheduled in order t, θ is the neural network weight, is the expected value, G~GS is the probability of G in the DAG set GS.