Automatic scheduling method for experimental tasks

By combining deep neural networks and Monte Carlo tree search methods to construct graph neural networks and policy value networks, we solved the experimental task scheduling problem in highly dynamic and deadlock-prone scenarios in life science laboratories, achieved efficient task scheduling decision sequence generation, and improved scheduling performance.

CN120671722APending Publication Date: 2025-09-19CHENGDU QINGSOFT QINGZHI SOFTWARE CO LTD

Patent Information

Application Number
CN202510684670.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-26
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve high-quality experimental task scheduling in the highly dynamic, flexible environment and deadlock-prone complex task scenarios of life science laboratories. Traditional methods such as mathematically precise solvers, rule-based heuristic algorithms and reinforcement learning methods are difficult to meet the requirements in terms of time and effect.

Method used

A method combining deep neural networks and Monte Carlo tree search is used to construct graph neural networks and policy value networks. Experimental task scheduling is performed through the Monte Carlo tree search module. The network parameters are optimized by combining the initialization of the graph neural network and the gradient descent method to generate an efficient task scheduling decision sequence.

Benefits of technology

In highly dynamic and deadlock-prone laboratory scenarios, high-quality experimental task scheduling is achieved, scheduling performance is improved, no expert experience is required, and scheduling efficiency and accuracy are improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120671722A_ABST
    Figure CN120671722A_ABST
Patent Text Reader

Abstract

The invention relates to an automatic scheduling method for experimental tasks, which comprises the following steps of: firstly, constructing a deep neural network comprising a graph neural network and a strategy value network, and randomly initializing parameters; a deep neural network is adopted to replace selection, expansion and simulation processes in a Monte Carlo tree search (MCTS) algorithm; constructing a to-be-scheduled task into a directed acyclic graph, inputting the directed acyclic graph into a scheduling model, and executing an MCTS fused with the deep neural network to generate a decision tag; the deep neural network is trained by adopting labels, and the optimization target is that the similarity between the action selection probability distribution of the strategy value network and the action selection probability distribution of the MCTS is maximum in the same task state, and the difference between the value estimation and the complete task scheduling time of the MCTS is minimum; and alternately executing MCTS tag generation and neural network training optimization, and combining the MCTS with the trained neural network to obtain a final scheduling model. According to the method, high-quality scheduling can be completed in a laboratory complex task scene which is high in dynamic performance, flexible in environment and easy to deadlock.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of artificial intelligence and automated scheduling optimization, and in particular to an automated scheduling method for experimental tasks. Background Art

[0002] Life science automated laboratories utilize modern technologies and equipment to automate laboratory operations and processes. Through the collaboration of laboratory information systems and laboratory automation systems, they can collaboratively perceive experimental tasks and equipment, controlling experimental equipment to efficiently coordinate and complete complex experimental tasks.

[0003] Due to the complexity of tasks in life science laboratories, automation relies heavily on efficient experimental scheduling algorithms. Scheduling in automated laboratories is an NP-hard combinatorial optimization problem. As the problem size increases, the search space grows exponentially, making it impossible to find solutions to large-scale optimization scheduling problems in reasonable polynomial time.

[0004] Current approaches to solving these problems can be categorized into mathematically precise solvers, rule-based heuristic algorithms, meta-heuristic algorithms based on biological evolution, and reinforcement learning. Due to the complexity of the problem and the dynamic scheduling requirements common in this field, precise solvers and meta-heuristic algorithms struggle to achieve feasible solutions within an acceptable timeframe. Heuristic algorithms often require domain expertise and struggle to guarantee effective scheduling. Many advances in artificial intelligence utilize supervised learning systems, which are trained to replicate the decisions of human experts. However, scheduling optimization algorithms are a type of decision-making optimization problem, for which reliable supervised data is not readily available.

[0005] Reinforcement Learning (RL) systems train from their own experience, in principle allowing them to surpass human capabilities and operate in domains where human expertise is lacking.

[0006] However, the extreme complexity of scheduling problems makes deep reinforcement learning training difficult to converge, and generalization has long been a difficult problem. Furthermore, the stochastic nature of pure learning algorithms often makes it difficult to guarantee effective solutions in real-world production. In the field of life science laboratory automation, current scheduling algorithms are particularly unable to achieve high-quality scheduling in complex laboratory scenarios characterized by high dynamics, flexible environments, and the risk of deadlock. Summary of the Invention

[0007] In order to solve the technical problems existing in the above-mentioned prior art, the purpose of the present invention is to provide an automated scheduling method for experimental tasks, which can complete high-quality scheduling in complex laboratory task scenarios with high dynamics, flexible environments, and prone to deadlock.

[0008] To achieve the above-mentioned object of the invention, the present invention provides a method for automated scheduling of experimental tasks, comprising the following steps:

[0009] The experimental task including multiple subtasks is constructed as a directed acyclic graph to obtain input data samples;

[0010] We construct a deep neural network consisting of a graph neural network and a policy value network, and randomly initialize the network parameters. We use a deep neural network to replace the selection, expansion, and simulation operations of Monte Carlo search, and build a Monte Carlo Tree Search (MCTS) module that integrates deep neural networks.

[0011] Input the experimental task into MCTS to train the deep neural network;

[0012] The scheduling neural network is configured to maximize the similarity between the action selection probability distribution generated by the neural network and the action selection distribution of the Monte Carlo search tree under the same task state, and minimize the difference between the value estimate generated by the neural network and the complete task scheduling time of the Monte Carlo search tree;

[0013] The Monte Carlo search tree is iterated based on the action selection probability distribution generated by the neural network and the value estimate generated by the neural network until all subtasks are covered;

[0014] Starting from the root node of the Monte Carlo search tree, a child node in the next layer is selected based on the number of node visits until the child node of the last layer is selected; then the corresponding subtasks are arranged in the order in which the child nodes are selected to obtain the task scheduling decision sequence.

[0015] According to a technical solution of the present invention, the experimental task is represented by a directed acyclic graph G(V,E);

[0016] The node V of the directed acyclic graph G(V,E) represents a subtask, and the information of the node V includes the device type, the number of devices, and the execution time required for the corresponding subtask;

[0017] The directed edge E of the directed acyclic graph G(V, E) represents the dependency relationship between the two terminal subtasks.

[0018] According to a technical solution of the present invention, the scheduling neural network is a pre-trained graph neural network, and the training process of the scheduling neural network is as follows:

[0019] S01. Initialize the graph neural network f θ (s);

[0020] Among them, θ is the neural network parameter, s is the task state;

[0021] S02, combined with graph neural network f θ (s) and Monte Carlo tree search algorithm to construct a Monte Carlo search tree;

[0022] S03, calculate the current task state s according to the number of edge visits to the current root node recorded in the Monte Carlo search tree i The probability of selecting the next scheduling action a N is the number of subtasks in the experimental task;

[0023] S04: Execute a scheduling action, set the child node corresponding to the scheduling action as a new root node, and delete the nodes at the same level and the child nodes at the same level of the new root node;

[0024] S05. Repeat S03 to S04 N times to obtain a subtask sequence of length N; and record all task status s i , and the task status s i The maximum task completion time z under i ;

[0025] S06, according to the task status s i , maximum task completion time z i and the scheduling action selection distribution function The gradient descent method is used to train the graph neural network f θ (s) is updated by updating the neural network parameters θ to minimize the loss function L and obtain the updated graph neural network f θ (s′);

[0026] The minimization loss function L is:

[0027]

[0028] Among them, v i is the graph neural network f θ (s) in task state s i The estimated value of is the graph neural network f θ (s) in task state s i The action selection probability distribution under , c is the L2 regularization loss term parameter;

[0029] S07. Use the updated graph neural network f θ (s′) replaces the original neural network and repeats S2 to S6 until the preset number of neural network training rounds is reached to obtain a pre-trained graph neural network.

[0030] According to a technical solution of the present invention, the process of constructing a Monte Carlo search tree is as follows:

[0031] S021. Starting from the initial root node of the Monte Carlo tree, search downwards for the child node with the largest UCB value until a leaf node is reached;

[0032] The USB value is obtained by the following formula:

[0033]

[0034] Among them, N(s,a) represents the number of visits, Q(s,a) represents the average value, and p σ (a|s) represents the selection probability of scheduling action a under task state s, that is, the prior probability of the selected edge;

[0035] S022. When the leaf node is a non-terminal node, expand all child nodes of the leaf node;

[0036] And initialize the edge (s,a) of any leaf node to {N(s,a)=0,W(s,a)=0,Q(s,a)=0,P(s,a)=p a};

[0037] Where W(s,a) is the sum of the values ​​in all scheduling actions;

[0038] S023,

[0039] In the verification step, a deep neural network f θ (s) Output of the node’s value estimation parameter v;

[0040] S024, back propagation, update the current node and all its parent nodes along the way in the search tree using the following formula:

[0041] N(s,a)=N(s,a)+1

[0042] W(s,a)=W(s,a)+V(v)

[0043]

[0044] Where V(v) = e*(1 / v); e is a hyperparameter used to control the return reward coefficient;

[0045] S025. Repeat S021 to S024 until a preset number of search rounds is reached to obtain a Monte Carlo search tree.

[0046] According to a technical solution of the present invention, the method for automated scheduling of experimental tasks further includes: after obtaining the task scheduling decision sequence, scheduling the experimental tasks according to the task scheduling decision sequence. The specific process is as follows:

[0047] Get the earliest available time for all devices;

[0048] Traverse the subtasks in the task scheduling decision sequence in order;

[0049] Query all devices required for the subtask, assign the subtask to the device with the earliest available time sorted first, and update the earliest available time of the device;

[0050] Traverse the task scheduling decision sequence until all subtasks are assigned to the corresponding devices;

[0051] Calculate the start time EST(n) of any subtask on the corresponding device i ,d j ) and end time EFT(n i ,d j );

[0052]

[0053] EFT(n i ,d j )=EST(n i ,d j )+w i,j

[0054] Among them, pred(n i ) represents subtask n m All predecessor subtasks of AFT(n m ) represents subtask n i The latest completed subtask n among all subtasks m The final completion time; avail[j] indicates the time when the device d is traversed to the current subtask j The earliest available time; w i,j Represents subtask n i On device d j Execution time on ;

[0055] According to the start time EST(n i ,d j ) and end time EFT(n i ,d j ), calculate the maximum task completion time makespan of the scheduling scheme.

[0056] According to a technical solution of the present invention, the maximum task completion time, makespan, is obtained by the following formula:

[0057] z i =makespan=max{AFT(n i )}.

[0058] According to a technical solution of the present invention, starting from the root node of the Monte Carlo search tree, a child node in the next layer is selected based on the number of node visits. The specific process is as follows:

[0059] Calculate the action selection distribution function p for scheduling action a under task state s through the number of node visits N(s,a) σ (a|s):

[0060]

[0061] Among them, τ is the temperature parameter;

[0062] Action selection distribution function p σ (a|s) is sampled, a scheduling action is selected for execution, and then a child node in the next layer is selected.

[0063] According to a technical solution of the present invention, the method for automated scheduling of experimental tasks further includes:

[0064] According to the scheduling order and start time EST(n i ,d j ) and end time EFT(n i ,d j ) and generate a task scheduling Gantt chart.

[0065] According to one aspect of the present invention, an electronic device is provided, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the above-mentioned one or more computer programs are stored in the memory, and when the electronic device is running, the processor executes the one or more computer programs stored in the memory, so that the electronic device performs an automated scheduling method for experimental tasks as described in any one of the above-mentioned technical solutions.

[0066] According to one aspect of the present invention, a computer-readable storage medium is provided for storing computer instructions. When the computer instructions are executed by a processor, an automated scheduling method for experimental tasks as described in any one of the above technical solutions is implemented.

[0067] Compared with the prior art, the present invention has the following beneficial effects:

[0068] The present invention proposes an automated scheduling method for experimental tasks, which is used to generate an automated execution scheduling scheme for tasks based on experimental task processes and equipment.

[0069] This method combines the learning capabilities of deep neural networks with the search capabilities of Monte Carlo tree search to design a universal laboratory scheduling model. This approach, which requires no expert knowledge, uses supervised data generated by a neural network-based Monte Carlo tree search to guide deep neural network optimization, further improving the Monte Carlo tree search's performance. By iteratively executing these two approaches, the resulting model achieves efficient scheduling performance, enabling high-quality scheduling in highly dynamic, flexible, and deadlock-prone laboratory scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] To more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the following briefly introduces the drawings required for the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be derived from these drawings without inventive effort.

[0071] Figure 1 A system framework diagram schematically showing a method for automated scheduling of experimental tasks according to an embodiment of the present invention;

[0072] Figure 2 A flowchart schematically showing a method for automated scheduling of experimental tasks according to an embodiment of the present invention;

[0073] Figure 3 Schematically showing a graph neural network framework diagram in a method for automated scheduling of experimental tasks according to one embodiment of the present invention;

[0074] Figure 4 The figure schematically shows a schematic diagram of the root node conversion in the Monte Carlo search tree in the method for automatic scheduling of experimental tasks according to one embodiment of the present invention. DETAILED DESCRIPTION

[0075] The description of the embodiments in this specification should be combined with the corresponding drawings, which should be considered a complete part of this specification. In the drawings, the shapes and thicknesses of the embodiments may be exaggerated and indicated for simplicity or convenience. Furthermore, the various structural components in the drawings will be described separately. It is worth noting that components not shown in the drawings or not described in words are known to those of ordinary skill in the art.

[0076] The description of the embodiments herein and any references to directions and orientations are for ease of description only and are not to be construed as limiting the scope of the present invention. The following description of the preferred embodiments may involve combinations of features, which may exist independently or in combination. The present invention is not specifically limited to the preferred embodiments. The scope of the present invention is defined by the claims.

[0077] like Figures 1 to 4 As shown, a method for automated scheduling of experimental tasks of the present invention includes the following steps:

[0078] Step 1: Input the experimental task into the Monte Carlo tree search;

[0079] The experimental task includes multiple subtasks, and the initial Monte Carlo search tree is a Monte Carlo tree with only one root node;

[0080] Step 2: Input the experimental task into the scheduling neural network to obtain the action selection probability distribution and value estimation generated by the scheduling neural network;

[0081] The scheduling neural network is configured to maximize the similarity between the action selection probability distribution generated by the neural network and the action selection distribution of the Monte Carlo search tree under the same task state, and to minimize the difference between the value estimate generated by the neural network and the complete task scheduling time of the Monte Carlo search tree; the task state includes subtasks, that is, the same task state at least satisfies the same subtasks;

[0082] Step 3: The Monte Carlo search tree is iterated based on the action selection probability distribution generated by the neural network and the value estimate generated by the neural network until all subtasks are covered;

[0083] An iteration is usually considered complete when the number of layers is equal to the number of subtasks in the experimental task;

[0084] Step 4: Starting from the root node of the Monte Carlo search tree, select a child node in the next layer based on the number of node visits, until the child node of the last layer is selected; then arrange the corresponding subtasks in the order in which the child nodes are selected to obtain the task scheduling decision sequence.

[0085] In this embodiment, the specific working process is as follows:

[0086] Before scheduling the experimental task, obtain the pre-trained deep neural network (p, v) = f θ (s).

[0087] For the scheduled instance S, the optimized Monte Carlo tree search algorithm is executed, and the pre-trained neural network is used to guide the Monte Carlo tree search process. In the expansion and verification process, (p, v) = f θ The estimated value v obtained in (s) replaces the simulation process of the traditional MCTS algorithm, and (p,v)=f is used in the selection process. θ (s) The obtained prior action selection distribution function p optimizes the scheduling action selection.

[0088] Furthermore, the actual choice of p to be executed during inference σ(a|s) no longer uses the sampling strategy, but selects the action with the maximum value for scheduling. The final task scheduling decision sequence is obtained by step-by-step iterative scheduling.

[0089] In the experimental task automation scheduling method, the experimental task is represented by a directed acyclic graph G(V,E);

[0090] The node V of the directed acyclic graph G(V,E) represents a subtask. The information of the node V includes the device type, number of devices, and execution time required for the corresponding subtask.

[0091] The directed edge E of the directed acyclic graph G(V,E) represents the dependency relationship between the two terminal tasks.

[0092] In this embodiment, the experimental tasks and equipment resources of the automated laboratory are specifically modeled.

[0093] Step 1: This embodiment first models the experimental tasks and equipment resources of the automated laboratory through a modeling calculation module.

[0094] An experimental process in an automated laboratory can be represented as a directed acyclic graph, G = (V, E), V = {n1, n2, ..., n m} represents the nodes of a directed acyclic graph, and V={n1,n2,…,n m} represents the operation subtask, which represents the set of m task nodes in the experimental task, E={e1,e2,…,e s} represents the directed edge set of the task node. The edge e(i,j)∈E represents the dependency constraint relationship of the task, that is, task n i Must be in task n j Completed before.

[0095] The experimental tasks include the scheduled experimental tasks and subtask relationship dependencies. The equipment resources include processing machines, processing time, equipment constraint information, machine equipment and attribute information.

[0096] The automated scheduling method for experimental tasks uses a pre-trained graph neural network as the scheduling neural network. The training process of the scheduling neural network is as follows:

[0097] S01. Initialize the graph neural network f θ (s);

[0098] Among them, θ is the neural network parameter, s is the task state;

[0099] S02, combined with graph neural network f θ (s) and Monte Carlo tree search algorithm to construct a Monte Carlo search tree;

[0100] S03, calculate the current task state s according to the number of edge visits to the current root node recorded in the Monte Carlo search tree i The probability of selecting the next scheduling action a N is the number of subtasks in the experimental task;

[0101] S04: Execute a scheduling action, set the child node corresponding to the scheduling action as the new root node, and delete the nodes on the same layer and the child nodes on the same layer of the new root node;

[0102] S05. Repeat S03 to S04 N times to obtain a subtask sequence of length N; and record all task status s i , and the task status s i The maximum task completion time z under i ;

[0103] S06, according to the task status s i , maximum task completion time z i and the scheduling action selection distribution function Using gradient descent method and graph neural network f θ The neural network parameters θ of (s) are updated to minimize the loss function L, and the updated graph neural network f is obtained. θ (s′);

[0104] The minimized loss function L is:

[0105]

[0106] Among them, v i is the graph neural network f θ (s) in task state s i The estimated value of is the graph neural network f θ (s) in task state s i The action selection probability distribution under , c is the L2 regularization loss term parameter;

[0107] S07. Use the updated graph neural network f θ (s ′ ) replaces the original neural network and repeats S2 to S6 until the preset number of neural network training rounds is reached to obtain a pre-trained graph neural network.

[0108] In this embodiment, the training process of the graph neural network is described in detail.

[0109] Step 2: Build and initialize the deep neural network (p, v) = f θ(s), which is used to learn the probability distribution and state value of the scheduling action under the current scheduling state composed of tasks and resource conditions. The graph attention network is used in the deep neural network framework to learn the embedded representation of the node information of the directed acyclic graph. The output layer includes a strategy selection component and a value evaluation component. The strategy selection function p π (a|s) means outputting the probability p of all feasible scheduling actions a under the scheduling state s, and the value head v(s) is used to output the evaluation value of the scheduling state s. The number of initial training scheduling instances is And the deep neural network f θ The parameters are randomly initialized.

[0110] Step 3: Combine deep neural network f θ The Monte Carlo tree search algorithm of (s) constructs a search tree. The process of constructing a Monte Carlo search tree can be divided into the process of selection, expansion and verification, and backpropagation. Repeat k iterations to construct a complete Monte Carlo search tree.

[0111] Step 4: Calculate the computational action selection distribution function p of the current node based on the number of visits recorded in the multiple edges of the current root node in the constructed Monte Carlo search tree (MCTS). σ (a|s). Then the distribution function p is selected for the action σ (a|s) is sampled and an actual scheduling action is selected to be executed by the environment. σ (a|s) and the neural network policy network p π (a|s) constitutes a data pair (p σ (a|s),p π (a|s)) is stored and used for training the policy network in the subsequent deep neural network. Its purpose is to provide guidance on the exploration direction for the initial neural network through the experience value of Monte Carlo tree search.

[0112] Step 5: Set the node reached after the actual selection of the scheduling action as the root node, retain all statistical data of the subtree under the node to be reused in subsequent time steps, and at the same time delete the remaining nodes and branch data.

[0113] Step 6: Repeat steps 3 to 5 for N times, where N represents the total number of currently scheduled subtasks, until all experimental tasks are scheduled and a scheduling decision sequence of length N is obtained. The actual scheduling is performed using the list scheduling method. Record the maximum completion time makespan value z after the execution is completed, and record the supervision data pairs (s t ,z t ), where s t Indicates the current scheduling status, z tIndicates the maximum completion time makespan of this scheduling instance solution.

[0114] Step 7: Generate the search tree according to MCTS and perform the supervised training data collected during the complete scheduling process (s,p σ ,z), train the neural network, and update the initial strategy network (p,v) = f θ (s) parameter θ. Where s represents the current scheduling state, p σ represents the probability distribution of all edges under the root node in each actual scheduling in the Monte Carlo tree search, and z represents the maximum completion time makespan obtained under state s. The goal of training the neural network is to maximize the neural network f θ (s) The probability distribution p of selecting all edges in state s π (a|s) and the probability p of Monte Carlo tree search σ The similarity between (a|s) minimizes the difference between the predicted value v and the makespan value z in the scheduling result of this instance. Specifically, the parameter θ is used to minimize the loss function L through gradient descent as follows:

[0115]

[0116] Among them, L includes the sum of the mean square error and the cross entropy loss, and c is used as a parameter to control the L2 regularization loss term to avoid overfitting.

[0117] In step 7, the supervised data generated by MCTS is used to train the deep neural network. The Monte Carlo tree search can balance exploration and utilization, making the choice p σ (a|s) Compared to the action p made by the randomly initialized policy network π (a|s) is better. After the supervision signal training optimization generated by Monte Carlo tree search, p π (a|s) will be further enhanced. During the iteration into the next round of search execution, since the MCTS selection function depends on the neural network p π (a|s), thus better selection results can be obtained, which can further generate better supervision signals to guide the training of deep neural networks.

[0118] Step 8: Repeat steps 3 through 7, alternating between Monte Carlo tree search and deep neural network training optimization, iteratively training the current scheduled task instance or different task instances. The supervised labels generated by the Monte Carlo tree search algorithm guide the deep neural network optimization, further improving the MCTS search performance in this process, until the pre-set number of training rounds is reached.

[0119] Step 9: Perform inference testing on the initial task or unknown task instance. A pre-trained neural network is used to assist in Monte Carlo tree search for scheduling decisions, obtaining a task scheduling decision sequence. A list scheduling method is used to generate a Gantt chart for the task scheduling decision sequence. Based on the Gantt chart, the subtasks to be assigned at the current decision time are actually scheduled to the device for execution according to the earliest available principle.

[0120] In the experimental task automation scheduling method, the process of constructing the Monte Carlo search tree is as follows:

[0121] S021. Starting from the initial root node of the Monte Carlo tree, search downwards for the child node with the largest UCB value until a leaf node is reached;

[0122] The USB value is obtained by the following formula:

[0123]

[0124] Among them, N(s,a) represents the number of visits, Q(s,a) represents the average value, and p σ (a|s) represents the selection probability of scheduling action a under task state s, that is, the prior probability of the selected edge;

[0125] S022. When the leaf node is a non-terminal node, expand all child nodes of the leaf node;

[0126] And initialize the edge (s,a) of any leaf node to {N(s,a)=0,W(s,a)=0,Q(s,a)=0,P(s,a)=p a};

[0127] Where W(s,a) is the sum of the values ​​in all scheduling actions;

[0128] S023. In the verification step, the deep neural network f θ The output of (s) obtains the makespan estimate v of the node;

[0129] S024, back propagation, update the parameter values ​​of the node and all its parent nodes using the following formula:

[0130] N(s,a)=N(s,a)+1

[0131] W(s,a)=W(s,a)+V(v)

[0132]

[0133] Where V(makespan) = e*(1 / v); e is a hyperparameter used to control the return reward coefficient;

[0134] S025. Repeat S021 to S024 until a preset number of search rounds is reached to obtain a Monte Carlo search tree.

[0135] In this embodiment, the construction process of the Monte Carlo search tree is described in detail.

[0136] The Monte Carlo tree search method combined with the policy network in step 3 above includes the following steps:

[0137] Step 31: Initialize the Monte Carlo search tree. Construct the initial state node as the root node, initialize the variable set [Q, N, D] for the root node, and set the number of search rounds k. The value of k depends on the size of the task. If the search time is sufficient, k should be as large as possible.

[0138] A Monte Carlo search tree consists of nodes and edges. Nodes represent scheduling states (including the subtask itself, completed tasks preceding it, the equipment required for the subtask and its execution time, and equipment constraints). Edges between nodes represent scheduling actions, which are sets of executable tasks. Each node s in the search tree contains edges (s, a) for all legal actions a∈A(s). Each edge stores the following statistics: {N(s, a), W(s, a), Q(s, a), P(s, a)}. N(s, a) represents the number of visits, W(s, a) is the sum of the values ​​of all visited actions, Q(s, a) represents the average value, and P(s, a) represents the prior probability of selecting an edge.

[0139] Step 32: Select. Starting from the current root node, search downwards for the child node with the largest upper confidence bound value UCB until reaching the leaf node. If the leaf node is a non-terminal node, perform the expansion operation. The present invention improves the selection function in the original Monte Carlo tree search algorithm and adopts a deep neural network (p, v) = f θ (s) The obtained policy network p π (a|s) is used to guide the selection operation of Monte Carlo tree search, as shown in the following formula.

[0140]

[0141] where p π (a|s) represents the probability of selecting action a for a scheduled task under the scheduling state s of the task and machine, and C is the hyperparameter that determines the exploration registration. The selection process of the Monte Carlo tree search in this embodiment relies on the prior probability of the policy network in the deep neural network.

[0142] Step 33: Expansion and verification. When the selected node is a leaf node and the current MCTS layer number is less than the total number of subtasks, the expansion operation is performed. All legal leaf nodes are generated on this node. For any leaf node sL , the edge (s L ,a) initialized to {N(s,a)=0,W(s,a)=0,Q(s,a)=0,P(s,a)=p a}, and use deep neural network (p,v) = f θ The evaluation result v of (s) obtains the evaluation value of each expandable node.

[0143] Step 34: Backpropagation. After the extended node is evaluated using the deep neural network, the evaluation value v generated by the neural network is backpropagated.

[0144] Specifically, this implementation propagates the value evaluation value from the terminal leaf node back to the root node based on the makespan value obtained in the verification phase, and updates the parameter values ​​of the nodes along the path, including the number of visits N(s,a), the sum of the visit action values ​​W(s,a), the average value Q(s,a), and the prior probability P(s,a). The update formulas for each item are as follows:

[0145] N(s,a)=N(s,a)+1

[0146] W(s,a)=W(s,a)+V(v)

[0147]

[0148] The optimization goal of the scheduling method of the present invention is to minimize the maximum completion time. Therefore, the score evaluation function after the Monte Carlo tree search completes a verification is defined by the following formula:

[0149] V(v)=e*(1 / v)

[0150] Where e is a hyperparameter used to control the coefficient of the return reward.

[0151] Step 35: Iterate steps 32 to 34 until a preset number of search rounds is reached.

[0152] The method for automated scheduling of experimental tasks further includes: after obtaining the task scheduling decision sequence, scheduling the experimental tasks according to the task scheduling decision sequence. The specific process is as follows:

[0153] Get the earliest available time for all devices;

[0154] Traverse the subtasks in the task scheduling decision sequence in order;

[0155] Query all devices required for the subtask, assign the subtask to the device with the highest earliest available time, and update the device's earliest available time;

[0156] Traverse the task scheduling decision sequence until all subtasks are assigned to the corresponding devices;

[0157] Calculate the start time EST(n) of any subtask on the corresponding device i ,d j ) and end time EFT(n i ,d j );

[0158]

[0159] EFT(n i ,d j )=EST(n i ,d j )+w i,j

[0160] Among them, pred(n i ) represents subtask n m All predecessor subtasks of AFT(n m ) represents subtask n i The latest completed subtask n among all subtasks m Final completion time; c m,i Represents subtask n i and subtask n m The time gap between them; avail[j] indicates that when traversing to the current subtask, the device d j The earliest available time; w i,j Represents subtask n i On device d j Execution time on ;

[0161] Sequentially according to subtask n i The start time EST(n i ,d j ) and end time EFT(n i ,d j ) to schedule the experimental tasks and calculate the maximum task completion time makespan of the experimental task scheduling.

[0162] In this embodiment, the process of obtaining the task scheduling decision sequence and the maximum task completion time makespan is obtained.

[0163] The steps for generating the scheduling scheme S from the decision sequence D are as follows:

[0164] Step a) Input the scheduling decision sequence result obtained by Monte Carlo tree search method, and set the current decision time (traversal to the current task n i time) is set to 0;

[0165] Step b) construct and maintain an array of the earliest available time of all devices in the device pool, avail[j];

[0166] Step c) looping through the operation subtasks to be assigned at the current decision moment according to the decision sequence;

[0167] Step d) queries all devices required for the task, assigns the task to the device with the earliest available start time according to the earliest available time principle, and updates the earliest available start time of the device.

[0168] Repeat steps c) to d) until all nodes have completed the scheduling assignment. Obtain the processing start time EST (n) of all task flows on their assigned equipment. i ,d j ) and end time EFT(n i ,d j ), and generate the Gantt chart G of the experimental process plan and the maximum completion time makespan of the scheduling plan.

[0169] The maximum completion time makespan (in step 5) is calculated as follows: by calculating the earliest start time EST and the earliest end time EFT of the task and using the list scheduling method, the makespan value of the execution scheduling decision sequence can be obtained.

[0170] This embodiment calculates task n in the following way i On the device p i The earliest start time on the i ,p j )express:

[0171]

[0172] Among them, pred(n i ) represents subtask n m All predecessor node tasks, AFT(n m ) represents task n i The latest completed task n among all predecessor nodes m The final completion time of avail[j] is the time when the device p j The earliest available time of task n i On the device p j The earliest start time of task n is determined by the maximum value of the set of the latest completion time of all its predecessor nodes and the earliest available time of the device. i On the device p j The earliest completion time on the i ,pj ):

[0173] EFT(n i ,p j )=EST(n i ,p j )+w i,j

[0174] in, Represents task n i In p j The execution time on .

[0175] In the experimental task automation scheduling method, the maximum task completion time makespan is obtained by the following formula:

[0176] z i =makespan=max{AFT(n i )}.

[0177] In the automated scheduling method for experimental tasks, starting from the root node of the Monte Carlo search tree, a child node in the next layer is selected based on the number of node visits. The specific process is as follows:

[0178] Calculate the action selection distribution function p for scheduling action a under task state s through the number of node visits N(s,a) σ (a|s):

[0179]

[0180] Among them, τ is the temperature parameter;

[0181] Action selection distribution function p σ (a|s) is sampled, a scheduling action is selected for execution, and then a child node in the next layer is selected.

[0182] In this embodiment, the action selection distribution function p is described in detail. σ The calculation process of (a|s).

[0183] The action selection distribution function p used in step 4 above is σ The calculation formula for (a|s) is as follows:

[0184]

[0185] Here, τ is a temperature parameter used to control the degree of exploration. The larger τ is, the greater the exploration ratio is. Conversely, it is more inclined to the option with the best current value.

[0186] In step 4, the action selection distribution function p σ (a|s) and the neural network policy network pπ (a|s) constitutes a data pair (p σ (a|s),p π (a|s)) is stored and used for subsequent training of the policy network in the deep neural network. Its purpose is to provide guidance for the initial neural network to explore the direction through the experience of Monte Carlo tree search.

[0187] The automated scheduling method for experimental tasks also includes:

[0188] According to the scheduling order and start time EST(n i ,d j ) and end time EFT(n i ,d j ) and generate a task scheduling Gantt chart.

[0189] In summary, the experimental task automation scheduling method based on deep neural network and Monte Carlo tree search self-reinforcement learning of the present invention includes the following steps. The overall algorithm flow chart is as follows: Figure 2 As shown:

[0190] 1. Task modeling and deep neural network initialization:

[0191] 1. Model the experimental tasks and equipment resources of the automated laboratory through the modeling and calculation module.

[0192] 2. Neural network construction and initialization

[0193] Specifically, given a directed acyclic graph (DAG) G = (V, E), it is fed into a Graph Attention Network (GAT) for K iterations to compute the p-dimensional embedding of each node v∈V. The node features of the first layer are updated to obtain the node features of the bottom l+1 layers according to the following steps:

[0194] (1) Perform linear transformation on the original eigenvector:

[0195]

[0196] in Represents the i-th node information of the l-th layer input, After linear table transformation, the node feature embedding

[0197] (2) Calculate the additive attention score between neighbor nodes:

[0198]

[0199] in, Represents node embedding and The original attention score is calculated, || represents the concatenation operation, and the LeakReLU function is the activation function.

[0200] (3) Softmax the attention scores of all adjacent edges of a single node embedding vector to calculate the attention weight

[0201]

[0202] (4) Finally, the node feature update value of GCN is obtained, and the weighted sum of all adjacent nodes is performed according to the attention score to obtain the embedding representation of the original node feature after passing through the graph attention network at the first layer of the graph:

[0203]

[0204] (5) Finally, the network with a combined architecture of a policy network and a value network is used as the final output layer. The output layer is composed of a vector and a scalar, which represent the evaluation of the current scheduling state and the probability of the next scheduling task, respectively. The following formula is shown:

[0205]

[0206] Where W p , W v , b p , b v is the neural network parameter matrix and bias vector.

[0207] 2. Alternately perform Monte Carlo tree search and deep neural network updates.

[0208] 1. Monte Carlo tree search technology integrated with the policy network to construct training samples:

[0209] (1) A Monte Carlo tree search algorithm combined with a policy network is used to gradually construct a search tree. The process of constructing a search tree can be divided into the process of selection, expansion and verification, and backpropagation. The complete search tree is constructed by repeating k iterations. The setting of the k value depends on the size of the task. If the search time is sufficient, k is taken as large as possible. Specifically, the Monte Carlo tree search process includes the following steps:

[0210] a) Initialize the Monte Carlo search tree:

[0211] b) Selection: Starting from the current root node, search downwards for the child node with the largest UCB value until a leaf node is reached. If the leaf node is a non-terminal node, an expansion operation is performed. In the original Monte Carlo tree search algorithm, the selection function is as follows:

[0212]

[0213] Compared with the UCB method used in the original MCTS, the selection function of the present invention converts the policy network p of the deep neural network into σ (a|s) is fused and the probability distribution generated by MCTS is used to select the appropriate task to allocate the machine. The item representing the number of times the current node is selected in the original MCTS is modified and replaced with the initially defined policy network p π (a|s). As shown in the following formula.

[0214]

[0215] where p σ (a|s) represents the probability distribution of selecting each executable scheduling action a under the scheduling state s, and C is the hyperparameter value that determines the exploration registration. In this process, the choice of MCTS depends on the policy network p in the deep neural network. π The prior probability distribution of (a|s).

[0216] c) Expansion and Verification: Compared with the expansion algorithm in the existing Monte Carlo tree search, the present invention does not use dynamic expansion when executing the expansion strategy, but expands all child nodes. The expansion operation will generate all legal leaf nodes on the node. For any leaf node s i , the edge (s i ,a) initialized to {N(s,a)=0,W(s,a)=0,Q(s,a)=0,P(s,a)=p a}.

[0217] When performing the verification operation, compared with the existing Monte Carlo tree search that adopts the Rollout strategy, that is, randomly executing the scheduling steps downward from the expanded node to the terminal node to obtain the maximum completion time Makespan, the present invention uses the value network (p,v)=f in the deep neural network θ The evaluation result v of (s) is replaced by the estimated value, and the estimated value is used to obtain and feedback rewards, which effectively avoids the longest simulation process in the original MCTS and can significantly improve the search efficiency and accuracy.

[0218] d) Backpropagation: Backward propagation, updating the parameters of the node and all parent nodes along the way in the search tree.

[0219] Specifically, the present invention backpropagates from the terminal leaf node to the root node based on the node value evaluation value v obtained in the verification phase, and updates the parameter values ​​of the nodes along the path, including the number of visits N(s,a), the sum of the visit action values ​​W(s,a), the average value Q(s,a), and the prior probability P(s,a). The update formulas for each item are as follows:

[0220] N(s,a)=N(s,a)+1

[0221] W(s,a)=W(s,a)+V(v)

[0222]

[0223] The optimization goal of the scheduling method of the present invention is to minimize the maximum completion time. Therefore, the score evaluation function returned after the Monte Carlo tree search completes a verification can be defined by the following formula:

[0224] V(v)=e*(1 / v)

[0225] Where e is a hyperparameter used to control the coefficient of the return reward.

[0226] e) Repeat steps b)->c) until the preset number of search rounds T is reached, at which point a biased search tree is constructed.

[0227] (2) According to the number of selections recorded in the multiple edges of the current root node in the constructed Monte Carlo search tree, the action selection distribution function p is calculated using the following formula: σ (a|s):

[0228]

[0229] Where τ is a temperature parameter that controls the degree of exploration. The larger the τ, the greater the exploration ratio. Conversely, the more biased towards the option with the best current value. This search strategy will initially tend to prefer actions with higher prior probabilities and fewer visits, and then gradually tend to choose high-value actions. Using the action selection distribution function p σ (a|s) is sampled and an actual scheduling action is selected for execution by the environment. After executing the scheduling step, the actually selected node is set as the root node, and all statistics of the subtree under the node are retained for reuse in subsequent time steps. At the same time, the remaining nodes and branches are deleted.

[0230] The present invention uses the distribution function p σ (a|s) and the neural network policy network p π (a|s) constitutes a data pair (p σ (a|s),p π The (a|s) is stored and subsequently used for training the policy network in the deep neural network. Its purpose is to provide guidance for the initial neural network's exploration direction through the empirical value of the Monte Carlo tree search. Compared to existing Monte Carlo tree search algorithms that select the action with the largest UCB value, and general reinforcement learning methods that use a policy network to generate action selections, the method used in this paper uses the probability distribution generated by MCTS to select appropriate tasks for machine allocation.

[0231] (3) Repeat (1) and (2) for N times, where N represents the total number of currently scheduled subtasks, until all experimental processes of the experimental task are scheduled. A scheduling decision sequence of length N can be obtained, where N is equal to the number of task nodes in the task graph and also equal to the depth of the constructed search tree. The task scheduling execution sequence obtained by the above algorithm is actually scheduled using the list scheduling method. For task n i , obtain all available devices for the current task and assign the task to the available devices according to the earliest available principle. Specifically, the steps for generating the scheduling solution S from the decision sequence D are as follows:

[0232] Step a) inputting the scheduling decision sequence result obtained by the Monte Carlo tree search method and setting the current decision time to 0;

[0233] Step b) construct and maintain an array of the earliest available time of all devices in the device pool, avail[j];

[0234] Step c) looping through the operation subtasks to be assigned at the current decision moment according to the decision sequence;

[0235] Step d) queries all devices required for the task, assigns the task to the device with the earliest available start time according to the earliest available time principle, and updates the earliest available start time of the device.

[0236] Repeat steps c) to d) until all nodes have completed the scheduling assignment. Obtain the processing start time EST (n) of all task flows on their assigned equipment. i ,p j ) and end time EFT(n i ,p j ), and generate the Gantt chart G of the experimental process plan and the maximum completion time makespan of the scheduling plan.

[0237] The present invention calculates task n in the following way: i On the device p i The earliest start time on the i ,p j )express:

[0238]

[0239] Among them, pred(n i ) represents task n m All predecessor node tasks, AFT(n m ) represents task n i The latest completed task n among all predecessor nodes mThe final completion time of avail[j] is the time when the device p j The earliest available time of task n i On the device p j The earliest start time of task n is determined by the maximum value of the set of the latest completion time of all its predecessor nodes and the earliest available time of the device. i On the device p j The earliest completion time on the i ,p j ):

[0240] EFT(n i ,p j )=EST(n i ,p j )+w i,j

[0241] in, Represents task n i In p j The execution time on .

[0242] (4) After the execution is completed, the maximum completion time makespan value obtained is recorded as z, where the calculation formula of the maximum completion time makespan is: makespan=max{AFT(n m )}. The data of each scheduling execution step t of this scheduling is stored as a data pair (s t ,π t ,z t ), where s t Represents the current state of the scheduling system, π t represents the current policy network distribution, z t The data pairs obtained in this process will be used to train and optimize the deep neural network.

[0243] 2. Train the deep neural network, initialize the deep neural network (p, v) = f θ (s) parameter θ is updated:

[0244] The present invention generates a search tree based on MCTS and collects supervised training data (s, p σ ,z) train the neural network, where s represents the current scheduling state, p σ represents the probability distribution of all edges under the root node selected by MCTS each time it is actually scheduled, and z represents the maximum completion time makespan obtained under state s. The goal of training the neural network is to maximize the neural network f θ(s) The probability distribution p of selecting all edges in state s π (a|s) and the probability p of Monte Carlo tree search σ The similarity of (a|s) minimizes the difference between the predicted value v and the makespan value z in the scheduling result of this instance.

[0245] Specifically, the present invention optimizes the parameter θ by minimizing the loss function L through the gradient descent method:

[0246]

[0247] Among them, L includes the sum of the mean square error and the cross entropy loss, and c is used as a parameter to control the L2 regularization loss term to avoid overfitting.

[0248] Since the selection function of MCTS depends on the neural network p π (a|s), in the initial stage, MCTS makes the choice p due to its own ability to balance exploration and utilization. σ (a|s) Compared to the action p made by the randomly initialized policy network π (a|s) is better. After the supervision signal training and optimization generated by MCTS, p π (a|s) will be further enhanced. In the process of iterating into the next round of MCTS execution, due to the policy network p that MCTS relies on π (a|s) is optimized, resulting in better selection results, which in turn generates a better supervisory signal to guide the training of deep neural networks. By alternating between Monte Carlo tree search and neural network training optimization, the network gradually converges and performance improves.

[0249] 3. Repeat 1 and 2, alternating between Monte Carlo tree search and deep neural network training (p, v) = f θ (s) process until the preset number of training rounds is reached. The main idea of ​​the present invention is to repeatedly use these supervisory signals generated by Monte Carlo tree search to update the neural network f during the policy iteration process. θ The parameter θ of (s) makes the neural network action selection distribution function p π (a|s) is closer to the action selection distribution function p generated by Monte Carlo tree search σ (a|s), while making the neural network f θ The output v of (s) is closer to the maximum completion time evaluation z under the current state s. The neural network process composed of these new parameters is used in the next iteration process, making the MCTS search more enhanced.

[0250] 3. Inference Scheduling Using Trained Neural Networks and Monte Carlo Tree Search

[0251] After step 1 and step 2 are performed alternately K times, the pre-trained deep neural network (p, v) = f is obtained. θ (s), the K value depends on the scale of the problem to be solved, the hardware performance of CPU and GPU, the training time limit and other conditions. The larger the K value, the better the scheduling model effect. For the instance to be scheduled S, the above optimized Monte Carlo tree search algorithm is executed, and the pre-trained neural network is used to guide the Monte Carlo tree search process. In the expansion and verification process, (p, v) = f θ The estimated value v obtained in (s) replaces the simulation process of the traditional MCTS algorithm, and (p,v)=f is used in the selection process. θ (s) The obtained prior action selection distribution function p optimizes the scheduling action selection. In addition, the p that is actually selected to be executed during the inference process σ (a|s) no longer uses a sampling strategy, but instead selects the action with the maximum value for scheduling. The final scheduling sequence is obtained through step-by-step iterative scheduling, and machines are assigned using a list scheduling method, generating the final scheduling Gantt chart G.

[0252] The present invention has the following advantages:

[0253] (1) This paper combines the search capability of the Monte Carlo tree search algorithm with the learning capability of deep neural networks to propose a powerful laboratory automation scheduling model, which can obtain strong performance in the reasoning stage of actual scenarios through pre-training and is suitable for large-scale and complex automated laboratory scheduling scenarios.

[0254] (2) The method of the present invention replaces the simulation process and selection process in the Monte Carlo tree search algorithm by adopting a pre-trained deep neural network, thereby realizing width pruning and depth pruning of MCTS respectively, further improving the search efficiency of MCTS, and being able to effectively meet the frequent dynamic rescheduling requirements that occur in actual production scenarios.

[0255] The present invention is an intelligent scheduling method for life science laboratory automation based on deep neural network and Monte Carlo tree search self-reinforcement learning, which is mainly used in large-scale and complex automated laboratory scheduling scenarios. The overall framework of the system is as follows Figure 1 shown.

[0256] Based on Monte Carlo tree search and deep neural networks, this paper proposes a scheduling algorithm with self-reinforcement training to achieve automated scheduling of large-scale tasks in life science laboratories.

[0257] The present invention combines the powerful feature learning ability of deep learning with the powerful search ability of Monte Carlo tree search. It uses the supervisory signal generated by MCTS to train the deep neural network, and then updates the neural network parameters by alternating between Monte Carlo tree search and deep neural network. The deep neural network is used to replace the Rollout and selection process of MCTS, which implements another kind of pruning for the Monte Carlo tree search algorithm. The alternating execution of generating data through MCTS and using data to enhance the deep learning network makes the model gradually improve as the training progresses, solving the problem of automated scheduling in life science laboratories. At the same time, in the present invention, the supervised training data used to improve the performance of the policy network is collected autonomously during the operation process, solving the problem of the difficulty of obtaining supervised data in combinatorial optimization problems.

[0258] According to one aspect of the present invention, an electronic device is provided, comprising: one or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the above-mentioned one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory so that the electronic device performs the automated scheduling method for experimental tasks as in the above-mentioned technical solution.

[0259] According to one aspect of the present invention, a computer-readable storage medium is provided for storing computer instructions. When the computer instructions are executed by a processor, the method for automated scheduling of experimental tasks as in the above technical solution is implemented.

[0260] Computer-readable storage media may include any medium capable of storing or transmitting information. Examples of computer-readable storage media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, and the like. The code segments may be downloaded via a computer network such as the Internet, an intranet, and the like.

[0261] The present invention provides an automated scheduling method and system for experimental tasks, and the method includes: the present invention provides an automated scheduling method and system for experimental tasks, which relates to the field of information retrieval, and the method is as follows: based on multiple telemetry parameters and operation symbols, a satellite telemetry parameter logical operation expression is constructed; according to the screening conditions of each telemetry parameter, a single parameter valid time period set corresponding to each telemetry parameter is obtained; the single parameter valid time period set is: a set of time periods when the telemetry parameter meets the corresponding screening conditions; the single parameter valid time period set corresponding to the telemetry parameter is substituted into the satellite telemetry parameter logical operation expression and solved to obtain the effective time of multiple telemetry parameters.

[0262] Furthermore, it should be noted that the present invention may be provided as a method, apparatus, or computer program product. Thus, embodiments of the present invention may take the form of entirely hardware embodiments, entirely software embodiments, or embodiments combining software and hardware aspects. Furthermore, embodiments of the present invention may take the form of a computer program product embodied on one or more computer-usable storage media containing computer-usable program code.

[0263] The embodiments of the present invention are described with reference to the flowcharts and / or block diagrams of the methods, terminal devices (systems), and computer program products according to the embodiments of the present invention. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, an embedded processor, or other programmable data processing terminal device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing terminal device generate instructions for implementing the process in the flowchart and / or block diagram. Figure 1 a process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0264] These computer program instructions may also be stored in a computer readable memory that can direct a computer or other programmable data processing terminal device to operate in a specific manner, so that the instructions stored in the computer readable memory produce a manufactured product including an instruction device, which implements the process Figure 1 a process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing terminal device, so that a series of operation steps are executed on the computer or other programmable terminal device to produce a computer-implemented process, thereby providing instructions for implementing the process in the process. Figure 1 a process or multiple processes and / or boxes Figure 1 The steps for the function specified in one or more boxes.

[0265] It should also be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or terminal device comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or terminal device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or terminal device comprising the element.

[0266] Finally, it should be noted that the above is a preferred embodiment of the present invention. It should be noted that although the preferred embodiment of the present invention has been described, it is clear that those skilled in the art, once they understand the basic inventive concept of the present invention, can make various improvements and modifications without departing from the principles of the present invention. Such improvements and modifications should also be considered as within the scope of protection of the present invention. Therefore, the appended claims are intended to be interpreted as including the preferred embodiment and all changes and modifications that fall within the scope of the embodiments of the present invention.

Claims

1. A method for automated scheduling of experimental tasks, characterized in that: The following steps are involved: The experimental task including multiple subtasks is constructed as a directed acyclic graph to obtain input data samples; We construct a deep neural network consisting of a graph neural network and a policy value network, and randomly initialize the network parameters. We use a deep neural network to replace the selection, expansion, and simulation operations of Monte Carlo search, and build a Monte Carlo Tree Search (MCTS) module that integrates deep neural networks. Input the experimental task into MCTS to train the deep neural network; The scheduling neural network is configured through iterative training to maximize the similarity between the action selection probability distribution generated by the neural network and the action selection distribution of the Monte Carlo search tree under the same task state, and to minimize the difference between the value estimate generated by the neural network and the complete task scheduling time of the Monte Carlo search tree; The Monte Carlo search tree is iterated based on the action selection probability distribution generated by the neural network and the value estimate generated by the neural network until all subtasks are covered; Starting from the root node of the Monte Carlo search tree, a child node in the next layer is selected based on the number of node visits until the child node of the last layer is selected; then the corresponding subtasks are arranged in the order in which the child nodes are selected to obtain the task scheduling decision sequence.

2. The method for automated scheduling of experimental tasks according to claim 1, characterized in that: The experimental task is represented by a directed acyclic graph G(V,E); The node V of the directed acyclic graph G(V,E) represents a subtask, and the information of the node V includes the device type, the number of devices, and the execution time required for the corresponding subtask; The directed edge E of the directed acyclic graph G(V, E) represents the dependency relationship between the two terminal subtasks.

3. The method for automated scheduling of experimental tasks according to claim 2, characterized in that: The scheduling neural network is a pre-trained graph neural network. The training process of the scheduling neural network is as follows: S01. Initialize the graph neural network f θ (s); Among them, θ is the neural network parameter, s is the task state; S02, combined with graph neural network f θ (s) and Monte Carlo tree search algorithm to construct a Monte Carlo search tree; S03, calculate the current task state s according to the number of edge visits to the current root node recorded in the Monte Carlo search tree i The probability of selecting the next scheduling action a s i ∈s,i=1,2,3……N; N is the number of subtasks in the experimental task; S04: Execute a scheduling action, set the child node corresponding to the scheduling action as a new root node, and delete the nodes at the same level and the child nodes at the same level of the new root node; S05. Repeat S03 to S04 N times to obtain a subtask sequence of length N; and record all task status s i , and the task status s i The maximum task completion time z under i ; S06, according to the task status s i , maximum task completion time z i and the scheduling action selection distribution function Using gradient descent method and graph neural network f θ The neural network parameters θ of (s) are updated to minimize the loss function L, and the updated graph neural network f is obtained. θ (s′); The minimization loss function L is: Among them, v i is the graph neural network f θ (s) in task state s i The estimated value of is the graph neural network f θ (s) in task state s i The action selection probability distribution under , c is the L2 regularization loss term parameter; S07. Use the updated graph neural network f θ (s′) replaces the original neural network and repeats S2 to S6 until the preset number of neural network training rounds is reached to obtain a pre-trained graph neural network.

4. The method for automated scheduling of experimental tasks according to any one of claims 1 to 3, characterized in that: The process of constructing a Monte Carlo search tree is as follows: S021. Starting from the initial root node of the Monte Carlo tree, search downwards for the child node with the largest UCB value until a leaf node is reached; The UCB value is obtained by the following formula: Among them, N(s,a) represents the number of visits, Q(s,a) represents the average value, and p σ (a|s) represents the selection probability of scheduling action a under task state s, that is, the prior probability of the selected edge; S022. When the leaf node is a non-terminal node, expand all child nodes of the leaf node; And initialize the edge (s,a) of any leaf node to {N(s,a)=0,W(s,a)=0,Q(s,a)=0,P(s,a)=p a }; Where W(s,a) is the sum of the values ​​in all scheduling actions; S023. In the verification step, a deep neural network f θ (s) outputs the value estimation parameter v of the node and assigns a value to each node; S024, back propagation, update the current node and all its parent nodes along the way in the search tree using the following formula: N(s,a)=N(s,a)+1 W(s,a)=W(s,a)+V(v) Where V(v) = e*(1 / v); e is a hyperparameter used to control the return reward coefficient; S025. Repeat S021 to S024 until a preset number of search rounds is reached to obtain a Monte Carlo search tree.

5. The method for automated scheduling of experimental tasks according to claim 4, characterized in that: Also includes: After obtaining the task scheduling decision sequence, the experimental task scheduling is performed using the task scheduling decision sequence. The specific process is as follows: Get the earliest available time for all devices; Traverse the subtasks in the task scheduling decision sequence in order; Query all devices required for the subtask, assign the subtask to the device with the earliest available time sorted first, and update the earliest available time of the device; Traverse the task scheduling decision sequence until all subtasks are assigned to the corresponding devices; Calculate the start time EST(n) of any subtask on the corresponding device i ,d j ) and end time EFT(n i ,d j ); EFT(n i ,d j )=IS(n i ,d j )+w i,j Among them, pred(n i ) represents subtask n m All predecessor subtasks of AFT(n m ) represents subtask n i The latest completed subtask n among all subtasks m The final completion time; avail[j] indicates the time when the device d is traversed to the current subtask j Earliest available time; w i,j Represents subtask n i On device d j Execution time on ; According to the start time EST(n i ,d j ) and end time EFT(n i ,d j ), calculate the maximum task completion time makespan of the scheduling scheme.

6. The method for automated scheduling of experimental tasks according to claim 4, characterized in that: The maximum task completion time makespan is obtained by the following formula: z i =makespan=max{AFT(n i )}。 7. The method for automated scheduling of experimental tasks according to claim 1, 2, 3 or 5, characterized in that: Starting from the root node of the Monte Carlo search tree, a child node in the next layer is selected based on the number of node visits. The specific process is as follows: Calculate the action selection distribution function p for scheduling action a under task state s through the number of node visits N(s,a) σ (a|s): Among them, τ is the temperature parameter; Action selection distribution function p σ (a|s) is sampled, a scheduling action is selected for execution, and then a child node in the next layer is selected.

8. The method for automated scheduling of experimental tasks according to claim 7, characterized in that: Also includes: According to the scheduling order and start time EST(n i ,d j ) and end time EFT(n i ,d j ) and generate a task scheduling Gantt chart.

9. An electronic device, characterized in that: include: One or more processors, one or more memories, and one or more computer programs; wherein the processor is connected to the memory, and the above-mentioned one or more computer programs are stored in the memory. When the electronic device is running, the processor executes the one or more computer programs stored in the memory to enable the electronic device to perform the automated scheduling method for experimental tasks as described in any one of claims 1 to 8.

10. A computer-readable storage medium, characterized in that Used to store computer instructions, which, when executed by a processor, implement the method for automated scheduling of experimental tasks as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Automatic driving speed planning method, electronic equipment, storage medium and computer program product

    CN118439053A

  • Automatic laboratory deadlock-free dynamic scheduling method based on Monte Carlo tree search

    CN118860610A

  • Unmanned aerial vehicle cluster task chain scheduling method based on deep reinforcement learning and product

    CN118863366A

  • MCTS and self-game-based confrontation effectiveness evaluation method, system and equipment and medium

    CN119849310A

Cited By

  • Task execution method and device with semantic comprehension and intelligent optimization capability

    CN121599426A

  • A method and system for implementing multi-task scheduling of chemical experiments, and a storage medium

    CN122347319A