High-bay warehouse entry and exit joint scheduling method based on heterogeneous graph neural network
Patent Information
- Application Number
- CN202611064877.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-17
- Publication Date
- 2026-09-18
- Estimated Expiration
- 2046-07-17
AI Technical Summary
然而,若直接采用普通向量或矩阵形式对高架库系统状态进行建模,往往难以准确刻画出入库任务、堆垛机与货位之间复杂的异构关联关系,也不利于策略网络充分提取反映局部交互与全局状态的高层次结构特征
[0015] Compared with existing technologies, the advantages of this invention are that it models the scheduling of inbound and outbound tasks and the allocation of storage locations in a high-bay warehouse as a multi-stage, multi-action sequential decision-making problem. Addressing the complex relationships between tasks, stacker cranes, and storage locations, a heterogeneous graph composed of three types of nodes is constructed. Node attributes and relational edges jointly characterize task priority, equipment status, and spatial location information. Based on this, a three-stage embedded heterogeneous graph neural network (HGNN) is designed to gradually aggregate the features of tasks, stacker cranes, and storage locations, obtaining a graph representation that can characterize the local relationships and global state of the system. Furthermore, a multi-action Actor-Critic agent based on the Proximal Policy Optimization (PPO) algorithm is combined to sequentially complete task-stacker allocation and stacker crane-storage location allocation under feasible action constraints, achieving efficient joint scheduling in complex dynamic environments.
Smart Images

Figure CN122573366B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent warehouse optimization scheduling technology, and in particular to a method for joint scheduling of inbound and outbound operations in elevated warehouses based on heterogeneous graph neural networks. Background Technology
[0002] High-bay warehouses are automated storage and retrieval systems widely used in manufacturing, logistics, and storage. These systems achieve high-density storage and automated retrieval of goods through the coordinated operation of high-rise racks, stacker cranes, and conveyor systems. Under conditions of mixed product storage and order-driven demand fluctuations, high-bay warehouse inbound and outbound operations exhibit significant dynamism and complexity. In actual operation, the sequencing of inbound and outbound tasks affects the stacker crane's operation and waiting time, while the stacker crane's status further impacts task response efficiency. Furthermore, the selection of storage locations not only determines the current operational path but also affects the connection of subsequent inbound and outbound tasks and the spatial distribution within the warehouse. Therefore, the high-bay warehouse inbound and outbound problem is essentially a joint decision-making problem involving the coupling of "inbound / outbound tasks, stacker cranes, and storage locations." While traditional static optimization methods can obtain feasible solutions in specific scenarios, they often struggle to balance solution efficiency and decision quality in complex environments with continuously changing system states and multiple constraints.
[0003] In recent years, deep reinforcement learning methods have demonstrated good adaptive decision-making capabilities in dynamic scheduling problems, providing a new solution for the joint scheduling of inbound and outbound operations in high-bay warehouses. However, directly modeling the state of a high-bay warehouse system using ordinary vectors or matrices often fails to accurately depict the complex heterogeneous relationships between inbound and outbound tasks, stacker cranes, and storage locations. It also hinders the policy network from fully extracting high-level structural features reflecting local interactions and global states. Furthermore, high-bay warehouse scheduling typically needs to simultaneously satisfy business rules such as partition matching, equipment service range limitations, storage location occupancy constraints, and priority exit for long-term storage. This poses significant challenges to traditional reinforcement learning methods in terms of action space construction, state representation, and constraint handling, thus affecting the learning effectiveness and practical application performance of the scheduling strategy. Therefore, there is currently a lack of a unified method for joint scheduling of inbound and outbound operations in high-bay warehouses that can represent the complex relationships between inbound and outbound tasks, stacker cranes, and storage locations, and achieve efficient joint decision-making in dynamic environments. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to provide a joint scheduling method for inbound and outbound operations of elevated warehouses based on heterogeneous graph neural networks, which can improve the scheduling efficiency and adaptive decision-making ability of elevated warehouses in complex operation scenarios.
[0005] The technical solution adopted by this invention to solve the above-mentioned technical problems is: a joint scheduling method for inbound and outbound operations of elevated warehouses based on heterogeneous graph neural networks, comprising the following specific steps: S1. Establish a mathematical model for joint scheduling of inbound and outbound operations in high-bay warehouses, considering inbound and outbound tasks, stacker crane resource allocation, storage location occupancy status, storage location zoning matching, and complex operation conditions. S2. Map the inbound / outbound tasks, stacker cranes, and storage locations in the mathematical model to task nodes, stacker crane nodes, and storage location nodes in the heterogeneous graph, respectively, and establish the heterogeneous graph; S3. Establish a heterogeneous graph neural network, extract features from the heterogeneous graph through the heterogeneous graph neural network, obtain node-level embedding representations of task nodes, stacker crane nodes and storage location nodes, and then aggregate the node-level embedding representations to obtain graph-level state embedding vectors and state value functions used to characterize the overall operation status of the high-bay warehouse. S4. Input the graph-level state embedding vector and state value function into the Actor-Critic network trained in the near-end policy optimization algorithm. The Actor network outputs two-stage scheduling actions: task-stall crane and stacker crane-cargo location, until a complete high-bay warehouse inbound and outbound joint scheduling scheme is generated.
[0006] Furthermore, in step S1, the storage locations are divided into zones based on the turnover rate of the goods; The composite operation conditions are: when the stacker crane performs an inbound task, after the inbound task is completed, there is an executable outbound task that has arrived but not yet been completed in the outbound scheduling system, and there is an inventory location within the current service range of the stacker crane that matches the category of goods for the outbound task.
[0007] Furthermore, in step S1, the mathematical model includes an objective function and constraints. The objective function aims to minimize the sum of the total waiting time and the total operation time of the inbound / outbound scheduling system, and its relationship is as follows: , , , , Where Z represents the objective function. This represents the total waiting time for all tasks. This represents the total job time when all tasks are executed in single-job mode. Let p represent the time saved when the i-th inbound task is paired with the j-th outbound task to form a composite operation. p is the index of the task. Tasks are divided into inbound and outbound tasks. The number of tasks is the sum of the number of inbound and outbound tasks. O is the task set, O = {O1, O2, ..., O...} p}, where i is the index of the inbound task, O in For the collection of inbound tasks, j is the index of the outbound task, O out For the set of outbound tasks, u ij This variable represents whether the i-th inbound task and the j-th outbound task form a compound operation. If they are arranged as a compound operation to be executed consecutively by the same stacker crane, then u ij =1, otherwise 0; S represents a single task, C represents a compound task. The time for executing a single job for the i-th inbound task. The time for executing the order for the j-th outbound task. Let r represent the waiting time for the p-th task. k Let a be the earliest available time for the k-th stacker crane, where k is the stacker crane index. p Let $\mathbf{p}$ be the arrival time of the $p$-th task. When the $p$-th task arrives, the $k$-th stacker crane assigned to it is already in an idle state. At this time, the p-th task can start execution immediately. When the p-th task arrives, the k-th stacker crane is still performing the preceding operation, i.e. When the p-th task needs to wait in the queue, its waiting time is the difference between the available time of the k-th stacker crane and the arrival time of the task; Let x be the time for executing a single job for the p-th task. p y p Let V be the coordinates of the storage location corresponding to the p-th task. x and V y These refer to the stacker crane's travel speeds in the lateral and longitudinal directions, respectively. The time required to execute the composite operation after pairing the i-th inbound task with the j-th outbound task to form a composite operation, (x i y i Let (x) be the coordinates of the storage location corresponding to the i-th inbound task. j y j Let be the coordinates of the storage location corresponding to the j-th outbound task; The constraints are: , , , , , , Among them: B p is the actual start time of the p-th task; M is a positive number used to implement linearization of logical constraints, and its value is not less than the maximum time difference occurring within the scheduling period; x pkLet x be a binary decision variable, representing whether the p-th task is executed by the k-th stacker crane. If the p-th task is assigned to the k-th stacker crane, then x... pk =1, otherwise 0; W is the set of stacker cranes, W={W1, W2, ..., W...} k}, where l is the location index, L is the set of locations, L={L1, L2, ..., L l}, y pl Let y be a binary decision variable, representing whether the p-th task matches the l-th storage location. If the p-th task matches the l-th storage location, then y... pl =1, otherwise 0; Let z be the classification attribute of the goods corresponding to the p-th task. l This refers to the category area to which the l-th storage location belongs. τ represents the storage time of goods in the l-th storage location; j Let be the in-stock time requirement for the j-th outbound task.
[0008] Furthermore, in step S2, the heterogeneous graph is defined as: G = (O', W', L', E') OO E OW E WL (Start, End), where: the task node set O' consists of all task nodes The stacker crane node set W' consists of all stacker crane nodes. The set of storage location nodes L' consists of all storage location nodes. Composition, E OO Let E be a set of task-task edges. Edges between inbound tasks are directed conjunctive arcs, representing the first-come, first-served time sequence constraint between them. Edges between outbound tasks are undirected disjunctive arcs, and edges between inbound and outbound tasks are undirected disjunctive arcs, representing potential pairing relationships in composite operations. OW This is a set of task-stacking machine edges, where each edge is an undirected disjunctive arc, used to connect task nodes and their executable stacker machine nodes. E WL This is a set of stacker crane-cargo location edges, where each edge is an undirected disjunctive arc. These edges represent candidate cargo locations that the stacker crane can reach when performing a transportation task, and reflect the travel distance relationship between the stacker crane and the corresponding cargo location. Start and End are two virtual nodes, representing the start and end of the scheduling process, respectively. The original characteristics of task nodes, stacker crane nodes, storage location nodes, and various edges are defined as follows: The original characteristics of a task node include task type, task assignment status, task arrival time, and cargo category information; the original characteristics of a stacker crane node include its aisle, current location, and remaining busy time; the original characteristics of a storage location node include its aisle, coordinate location, occupancy status, inventory category, cargo storage duration, and static travel time; the original characteristics of the edge connecting the stacker crane and the storage location are defined as the travel time from the current location of the stacker crane to the coordinates of the target storage location.
[0009] Furthermore, in step S3, the embedded representations of the storage location node, the stacker crane node, and the task node are obtained sequentially through a three-stage embedding process, specifically as follows: Phase 1: Propagate the feature vectors of the task node and stacker crane node to the connected storage location node to obtain the updated storage location node embedding representation; The second stage involves propagating the feature vectors of the task node and the storage location node to the first-order adjacent stacker crane nodes to obtain the updated stacker crane node embedding representation. The third stage involves propagating the feature vectors of the location node, stacker crane node, and task node to the task node to obtain an updated task node embedding representation.
[0010] Furthermore, in step S3, the updated location node embedding representation The relation is: , , in: For the activation function, m pk For task nodes via stacker crane node The aggregated intermediate embedding representation, u kl This refers to the original features of the connection edge between the stacker crane and the storage location. Let represent the set of stacker crane nodes connected to the storage location nodes, and α be the attention weight coefficient during the storage location node embedding phase. h represents the set of task nodes connected to the stacker crane node. p For the original characteristics of the task node, m k These are the original features of the stacker crane node; For the storage location node, the stacker crane node is embedded in the stage normalized. For cargo location nodes Attention coefficient Embed the phase-normalized task node for the storage location node. For stacker crane nodes Attention coefficient For the storage location node, the stacker crane node is embedded in the stage normalized. For stacker crane nodes Attention coefficient; Updated stacker node embedding representation The relation is: , , in: For stacker crane nodes Through task nodes The aggregated intermediate embedding representation, where β is the attention weight coefficient during the stacker crane node embedding stage. This represents the set of storage location nodes connected to the stacker crane node. For the stacker crane node embedding stage, the normalized storage location node For stacker crane nodes Attention coefficient For the stacker crane node embedding stage, the normalized task node For stacker crane nodes Attention coefficient; Updated task node embedding representation The relation is: , Among them, MLP0, MLP1, MLP2, and MLP3 are all multilayer perceptrons. Each MLP contains two 128-dimensional hidden layers, and ELU is an exponential linear unit used as the activation function. This represents the set of stacker crane nodes connected to the task node.
[0011] Furthermore, in step S3, the joint inbound and outbound scheduling is divided into two stages: the task-staller allocation stage and the stacker-location allocation stage. The task node embedding representation, the stacker node embedding representation, and the location node embedding representation are aggregated to obtain the graph-level state embedding vector for the task-staller allocation stage. Graph-level state embedding vectors for stacker crane-location phase The relationship is as follows: , , Embed the graph-level state of the task-stacker assignment phase into a vector. Graph-level state embedding vectors for stacker crane-location phase Further concatenation yields the state value function, which serves as the input to the subsequent Critic network. The state value function's equation is: , in: Represents the Critic network, V(s)t ) represents the state s at time step t. t Value estimate.
[0012] Furthermore, in step S4, an environment model based on a Markov decision process is first constructed in the proximal policy optimization algorithm. The proximal policy optimization algorithm trains an Actor-Critic network based on the environment model. The environment model includes: State space: The scheduling state of the high-bay warehouse inbound and outbound operations is represented by a heterogeneous graph consisting of task nodes, stacker crane nodes, and storage location nodes. The state is denoted as s, which is a set of the original features of task nodes, stacker crane nodes, and storage location nodes. Action Space: Let A be the set of actions for scheduling inbound and outbound operations in the high-bay warehouse, and let a be the action taken by the agent at time step t. t ,and The action is broken down into: task - stacker machine action. With stacker crane - storage position operation ,Right now: , in: This represents a task-stacking machine sub-action, specifically: assigning an executable task to the stacker crane at time step t. O-W This represents the set of possible actions during the task-stacker allocation phase. This indicates a stacker crane-position sub-action, meaning: selecting a target storage location for the task assuming the stacker crane has already been selected at time step t. (A) W-L This represents the set of possible actions during the stacker crane-location phase. State transition function: based on the state s at time step t. t and the action chosen by the agent at time step t. After the inbound / outbound scheduling system executes this action, it updates the stacker crane position, remaining busy time, storage location occupancy status, task completion status, and task queue information. Based on the updated status, it regenerates the connection edges between candidate tasks and stacker cranes, as well as the connection edges between stacker cranes and storage locations. The heterogeneous graph is then completed from state s. t to state s t+1 The transfer; Reward function: Let A be the set of actions corresponding to time step t. t When performing a single operation, A t It contains only one task; when performing compound operations, A t If a task contains one inbound task and one outbound task, then the single-step scheduling cost function C is... t Represented as: , in, This represents the waiting time incurred by the task at time step t before it begins execution. This represents the stacker crane's running time at time step t. When a single operation is executed at time step t... Single job time When a composite operation actually occurs at time step t, For composite operation time ; This indicates the duration of time the goods have been in storage for the corresponding outbound task. This represents the average duration of inventory in storage, and η is a penalty coefficient, set to 0.002, used to adjust the degree of influence of inventory storage time on rewards. This represents the set of actions for outbound tasks at time step t. When time step t does not contain any outbound tasks, then... Therefore, the reward function is defined as R. t = -C t .
[0013] Furthermore, in step S4, the Actor-Critic network includes two independent Actor networks and one Critic network. The two graph-level state embedding vectors are used as inputs to the two Actor networks, and the state value function is used as input to the Critic network.
[0014] Furthermore, in step S4, during the training of the Actor-Critic network, the agent generates actions using a two-stage sequential allocation method. The first stage is based on the local state of the task-stacker allocation. After completing the task and stacker crane matching, the graph-level state corresponding to this stage is embedded into a vector and input into the first Actor network. The Actor network contains a single-layer LH hidden layer and is equipped with a tanh activation function. It outputs the original probability distribution of task-stacker sub-actions in the current state, which is then normalized to obtain the first-stage allocation strategy. Based on the determined task-stacking machine matching results, construct the local state of the stacker crane-warehouse allocation stage. After completing the matching of the stacker crane and the storage location, the graph-level state corresponding to this stage is embedded into a vector and input into a second Actor network with the same structure. The original probability distribution of the stacker crane-storage location sub-actions is output and normalized to obtain the second-stage allocation strategy. π represents the policy function; Simultaneously, the graph-level state embedding vectors of the two stages are concatenated, and the resulting state value function is input into the Critic network, which outputs the value estimate V(s) of the current state. tDuring training, the reward function is used to obtain the reward corresponding to the current scheduled action. This reward is then used in conjunction with the value estimate V(s). t The difference between the two actors is used to evaluate the merits of the current two-stage scheduling actions, and the parameters of the two actor networks are updated accordingly to adjust the selection probabilities of the sub-actions in the two stages; simultaneously, based on the reward and value estimate V(s)... t The error between the two is used to update the parameters of the Critic network, so that the value estimate output by the Critic network gradually approaches the reward corresponding to the current scheduling action; through multiple rounds of iterative training, a well-trained Actor-Critic network is obtained.
[0015] Compared with existing technologies, the advantages of this invention are that it models the scheduling of inbound and outbound tasks and the allocation of storage locations in a high-bay warehouse as a multi-stage, multi-action sequential decision-making problem. Addressing the complex relationships between tasks, stacker cranes, and storage locations, a heterogeneous graph composed of three types of nodes is constructed. Node attributes and relational edges jointly characterize task priority, equipment status, and spatial location information. Based on this, a three-stage embedded heterogeneous graph neural network (HGNN) is designed to gradually aggregate the features of tasks, stacker cranes, and storage locations, obtaining a graph representation that can characterize the local relationships and global state of the system. Furthermore, a multi-action Actor-Critic agent based on the Proximal Policy Optimization (PPO) algorithm is combined to sequentially complete task-stacker allocation and stacker crane-storage location allocation under feasible action constraints, achieving efficient joint scheduling in complex dynamic environments. Attached Figure Description
[0016] Figure 1 This is a flowchart of the present invention; Figure 2 This is a high-bay warehouse logistics layout diagram of the present invention; Figure 3 This is a heterogeneous diagram of the present invention; Figure 4 This is a diagram of the heterogeneous graph neural network architecture of the present invention; Figure 5 This is a training curve graph of the present invention with a scale of 40×2×100; Figure 6 This is a diagram of the initial state of the shelving system of the present invention, with a scale of 40×2×100. Figure 7 This is a diagram of the final state of the shelving system of the present invention, with a scale of 40×2×100. Figure 8 This is a diagram of the initial state of the shelving unit of the present invention, with a scale of 60×3×100. Figure 9 This is a diagram of the final state of the 60×3×100 scale shelving of the present invention. Detailed Implementation
[0017] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments.
[0018] like Figure 1-4 As shown, the high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks includes the following specific steps: S1. A mathematical model for joint scheduling of inbound and outbound operations in high-bay warehouses is established, considering inbound and outbound tasks, stacker crane resource allocation, storage location occupancy status, storage location zoning matching, and complex operation conditions. Storage locations are divided according to the turnover rate of goods. The complex operation conditions are: when a stacker crane performs an inbound task, after the inbound task is completed, there is an executable outbound task that has arrived but not yet been completed in the inbound and outbound scheduling system, and there are inventory storage locations within the service range of the current stacker crane that match the goods category of the outbound task. The mathematical model includes an objective function and constraints. The objective function aims to minimize the sum of the total waiting time and the total operation time of the inbound / outbound scheduling system. The relationship is as follows: , , , , Where Z represents the objective function. This represents the total waiting time for all tasks. This represents the total job time when all tasks are executed in single-job mode. Let p represent the time saved when the i-th inbound task is paired with the j-th outbound task to form a composite operation. p is the index of the task. Tasks are divided into inbound and outbound tasks. The number of tasks is the sum of the number of inbound and outbound tasks. O is the task set, O = {O1, O2, ..., O...} p}, where i is the index of the inbound task, O in For the collection of inbound tasks, j is the index of the outbound task, O out For the set of outbound tasks, u ij This variable represents whether the i-th inbound task and the j-th outbound task form a compound operation. If they are arranged as a compound operation to be executed consecutively by the same stacker crane, then u ij =1, otherwise 0; S represents a single task, C represents a compound task. The time for executing a single job for the i-th inbound task. The time for executing the order for the j-th outbound task. Let r represent the waiting time for the p-th task. k Let a be the earliest available time for the k-th stacker crane, where k is the stacker crane index.p Let $\mathbf{p}$ be the arrival time of the $p$-th task. When the $p$-th task arrives, the $k$-th stacker crane assigned to it is already in an idle state. At this time, the p-th task can start execution immediately. When the p-th task arrives, the k-th stacker crane is still performing the preceding operation, i.e. When the p-th task needs to wait in the queue, its waiting time is the difference between the available time of the k-th stacker crane and the arrival time of the task; Let x be the time for executing a single job for the p-th task. p y p Let V be the coordinates of the storage location corresponding to the p-th task. x and V y These refer to the stacker crane's travel speeds in the lateral and longitudinal directions, respectively. The time required to execute the composite operation after pairing the i-th inbound task with the j-th outbound task to form a composite operation, (x i y i Let (x) be the coordinates of the storage location corresponding to the i-th inbound task. j y j Let be the coordinates of the storage location corresponding to the j-th outbound task; The constraints are: , , , , , , Among them: B p is the actual start time of the p-th task; M is a positive number used to implement linearization of logical constraints, and its value is not less than the maximum time difference occurring within the scheduling period; x pk Let x be a binary decision variable, representing whether the p-th task is executed by the k-th stacker crane. If the p-th task is assigned to the k-th stacker crane, then x... pk =1, otherwise 0; W is the set of stacker cranes, W={W1, W2, ..., W...} k}, where l is the location index, L is the set of locations, L={L1, L2, ..., L l}, y pl Let y be a binary decision variable, representing whether the p-th task matches the l-th storage location. If the p-th task matches the l-th storage location, then y... pl =1, otherwise 0; Let z be the classification attribute of the goods corresponding to the p-th task. lThis refers to the category area to which the l-th storage location belongs. τ represents the storage time of goods in the l-th storage location; j The in-stock time requirement for the j-th outbound task is defined by this constraint, which states that outbound tasks can only be allocated from locations that have already met the in-stock time requirement. S2. Map the inbound / outbound tasks, stacker cranes, and storage locations in the mathematical model to task nodes, stacker crane nodes, and storage location nodes in the heterogeneous graph, respectively, and establish the heterogeneous graph, defined as: G = (O', W', L', E'). OO E OW E WL (Start, End), where: The task node set O' consists of all task nodes The stacker crane node set W' consists of all stacker crane nodes. The set of storage location nodes L' consists of all storage location nodes. Composition, E OO Let E be a set of task-task edges. Edges between inbound tasks are directed conjunctive arcs, representing the first-come, first-served time sequence constraint between them. Edges between outbound tasks are undirected disjunctive arcs, and edges between inbound and outbound tasks are undirected disjunctive arcs, representing potential pairing relationships in composite operations. OW Let E be a set of task-stacking machine edges, where each edge is an undirected disjunctive arc, used to connect task nodes and their executable stacker machine nodes, indicating which stacker machines can complete a given task. WL This is a set of stacker crane-cargo location edges, where each edge is an undirected disjunctive arc. These edges represent candidate cargo locations that the stacker crane can reach when performing a transportation task, and reflect the travel distance relationship between the stacker crane and the corresponding cargo location. Start and End are two virtual nodes, representing the start and end of the scheduling process, respectively. The original characteristics of task nodes, stacker crane nodes, storage location nodes, and various edges are defined as follows: The original characteristics of a task node include task type, task allocation status, task arrival time, and goods category information. Among them, task type is used to distinguish between inbound and outbound tasks, task allocation status is used to identify whether the current task has been assigned, and task arrival time is used to describe the timing information of the task entering the scheduling system. The original characteristics of a stacker crane node include its associated aisle, current location, and remaining busy time. The associated aisle defines the stacker crane's service area, the current location reflects its current position, and the remaining busy time can be estimated based on the remaining path corresponding to its current operation type: when performing an inbound task, it is calculated based on the remaining travel time from the current location to the target storage location; when performing an outbound task, if a combined operation is triggered, it is calculated based on the total travel time of the uncompleted paths from the current location to the inbound storage location, from the inbound storage location to the outbound storage location, and from the outbound storage location to the entrance / exit; otherwise, it is calculated based on the remaining travel time from the current location to the entrance / exit. The original characteristics of a storage location node include its associated aisle, coordinates, occupancy status, inventory category, storage duration, and static walking time. Among these, the inventory category and storage duration are used to characterize the attribute information of the goods stored in the current storage location, providing a basis for storage location selection and long-term storage priority strategy learning during the outbound stage. The static walking time is used to characterize the standard running time from the entrance / exit to the storage location. The original feature of the connection edge between the stacker crane and the storage location is defined as the travel time from the current position of the stacker crane to the coordinates of the target storage location; S3. Establish a heterogeneous graph neural network. The heterogeneous graph neural network extracts features from the heterogeneous graph through a three-stage embedding process, successively obtaining the embedding representations of the storage location nodes, the stacker crane nodes, and the task nodes, specifically: Phase 1: Propagate the feature vectors of the task node and stacker crane node to the connected storage location nodes to obtain updated storage location node embeddings. The relation is: , , in: For the activation function, m pk For task nodes via stacker crane node The aggregated intermediate embedding representation, u kl This refers to the original features of the connection edge between the stacker crane and the storage location. Let represent the set of stacker crane nodes connected to the storage location nodes, and α be the attention weight coefficient during the storage location node embedding phase. h represents the set of task nodes connected to the stacker crane node. p For the original characteristics of the task node, m k These are the original features of the stacker crane node; For the storage location node, the stacker crane node is embedded in the stage normalized. For cargo location nodes Attention coefficient Embed the phase-normalized task node for the storage location node. For stacker crane nodes Attention coefficient For the storage location node, the stacker crane node is embedded in the stage normalized. For stacker crane nodes Attention coefficient; The second stage involves propagating the feature vectors of the task nodes and storage location nodes to the first-order adjacent stacker crane nodes to obtain updated stacker crane node embeddings. The relation is: , , in: For stacker crane nodes Through task nodes The aggregated intermediate embedding representation, where β is the attention weight coefficient during the stacker crane node embedding stage. This represents the set of storage location nodes connected to the stacker crane node. For the stacker crane node embedding stage, the normalized storage location node For stacker crane nodes Attention coefficient For the stacker crane node embedding stage, the normalized task node For stacker crane nodes Attention coefficient; The third stage involves propagating the feature vectors of the storage location node, stacker crane node, and task node to the task node to obtain an updated task node embedding representation. The relation is: , Among them, MLP0, MLP1, MLP2, and MLP3 are all multilayer perceptrons. Each MLP contains two 128-dimensional hidden layers, and ELU is an exponential linear unit used as the activation function. This represents the set of stacker crane nodes connected to the task node; The joint inbound and outbound scheduling is divided into two stages: the task-staller allocation stage and the stacker crane-location stage. The embedded representations of task nodes, stacker crane nodes, and location nodes are aggregated to obtain the graph-level state embedding vector for the task-staller allocation stage. Graph-level state embedding vectors for stacker crane-location phase The relationship is as follows: , , Embed the graph-level state of the task-stacker assignment phase into a vector. Graph-level state embedding vectors for stacker crane-location phase Further concatenation yields the state-value function, which serves as the input to the subsequent Critic (value) network. The state-value function's equation is: , in: Represents the Critic network, V(s) t ) represents the state s at time step t. t Value estimation; S4. In the near-end policy optimization algorithm, first construct an environment model based on Markov decision process. The environment model includes: State space: The scheduling state of the high-bay warehouse inbound and outbound operations is represented by a heterogeneous graph consisting of task nodes, stacker crane nodes, and storage location nodes. The state is denoted as s, which is a set of the original features of task nodes, stacker crane nodes, and storage location nodes. Action Space: Let A be the set of actions for scheduling inbound and outbound operations in the high-bay warehouse, and let a be the action taken by the agent at time step t. t ,and The action is broken down into: task - stacker machine action. With stacker crane - storage position operation ,Right now: , in: This represents a task-stacking machine sub-action, specifically: assigning an executable task to the stacker crane at time step t. O-W This represents the set of possible actions during the task-stacker allocation phase. This indicates a stacker crane-position sub-action, meaning: selecting a target storage location for the task assuming the stacker crane has already been selected at time step t. (A) W-L This represents the set of possible actions during the stacker crane-location phase. State transition function: based on the state s at time step t. t and the action chosen by the agent at time step t. After the inbound / outbound scheduling system executes this action, it updates the stacker crane position, remaining busy time, storage location occupancy status, task completion status, and task queue information. Based on the updated status, it regenerates the connection edges between candidate tasks and stacker cranes, as well as the connection edges between stacker cranes and storage locations. The heterogeneous graph is then completed from state s. t to state s t+1 The transfer; Reward function: Let A be the set of actions corresponding to time step t. t When performing a single operation, A t It contains only one task; when performing compound operations, A tIf a task contains one inbound task and one outbound task, then the single-step scheduling cost function C is... t Represented as: , in, This represents the waiting time incurred by the task at time step t before it begins execution. This represents the stacker crane's running time at time step t. When a single operation is executed at time step t... Single job time When a composite operation actually occurs at time step t, For composite operation time ; This indicates the duration of time the goods have been in storage for the corresponding outbound task. This represents the average duration of inventory in storage, and η is a penalty coefficient, set to 0.002, used to adjust the degree of influence of inventory storage time on rewards. This represents the set of actions for outbound tasks at time step t. When time step t does not contain any outbound tasks, then... Therefore, the reward function is defined as R. t = -C t ; The proximal policy optimization algorithm trains an Actor-Critic network based on an environment model. The Actor-Critic network consists of two independent Actor (policy) networks and one Critic (value) network. During the training of the Actor-Critic network, the agent generates actions using a two-stage sequential allocation method. The first stage is based on the local state of the task-stacker allocation. Complete the matching of tasks with stacker cranes, and embed the graph-level state corresponding to that stage into a vector. Input the first Actor network, which contains a single LH hidden layer and a tanh activation function. Output the original probability distribution of the task-stacking machine sub-actions in the current state. After normalization, the first-stage allocation strategy is obtained. Based on the determined task-stacking machine matching results, construct the local state of the stacker crane-warehouse allocation stage. Complete the matching of the stacker crane and the storage location, and embed the corresponding graph-level state into a vector. Input a second Actor network with the same structure, and output the original probability distribution of stacker crane-cargo location sub-actions. After normalization, the second-stage allocation strategy is obtained. π represents the policy function; simultaneously, the graph-level states of the two stages are embedded into a vector. and The concatenation is performed, and the resulting state-value function is input into the Critic network. The Critic network then outputs the value estimate V(s) of the current state. t During training, the reward function is used to obtain the reward corresponding to the current scheduled action. This reward is then used in conjunction with the value estimate V(s). t The difference between the two actors is used to evaluate the merits of the current two-stage scheduling actions, and the parameters of the two actor networks are updated accordingly to adjust the selection probabilities of the sub-actions in the two stages; simultaneously, based on the reward and value estimate V(s)... t The error between the two is used to update the parameters of the Critic network, so that the value estimate output by the Critic network gradually approaches the reward corresponding to the current scheduling action; through multiple rounds of iterative training, a well-trained Actor-Critic network is obtained. After training, the graph-level state embedding vector and state value function of the instance are input into the trained Actor-Critic network in the near-end policy optimization algorithm. The Actor network outputs two-stage scheduling actions: task-stacker and stacker-location, until a complete high-bay warehouse inbound and outbound joint scheduling scheme is generated.
[0019] In the above embodiments, the location node embedding representation in step S3 is... The attention coefficient in the equation is obtained through the following process: Since task nodes, stacker crane nodes, and storage location nodes differ in attribute space and semantic information in heterogeneous graphs, the linear transformation concept of Graph Attention Network (GAT) in isomorphic graphs is adopted, defining independent linear transformation matrices for task nodes, stacker crane nodes, and storage location nodes respectively: M O M W and M L The original features of the three types of nodes are first mapped to a unified feature space through corresponding linear transformation matrices. Then, the transformed features are concatenated and input into a single-layer feedforward neural network. The unnormalized attention coefficients between nodes are calculated using the trained parameter vector 'a' and the LeakyReLU activation function. Since the node neighborhood in GAT typically includes the node itself, the attention coefficients of the nodes themselves also need to be calculated simultaneously. Finally, the task nodes can be obtained separately. For stacker crane nodes Unnormalized attention coefficient Stacker node For stacker crane nodes Unnormalized attention coefficient and stacker crane nodes For cargo location nodes Unnormalized attention coefficient The calculation formula is as follows: , , , Where: M O M is the linear transformation matrix of the task node. W M is the linear transformation matrix of the stacker crane node. L This is the linear transformation matrix for the storage location nodes; Then, the Softmax function is used to normalize the unnormalized attention coefficients, resulting in the following: , and .
[0020] Stacker crane node embedded representation The attention coefficients in the representation are obtained by embedding them with the storage location nodes. The attention coefficients are obtained using the same method.
[0021] The optimization effect of the high-bay warehouse inbound and outbound joint scheduling method of the present invention was experimentally verified.
[0022] This invention constructs a warehouse scheduling experimental environment using a random instance generation method. Each row of shelves is divided into four areas: A, B, C, and D, with a unified area-goods category mapping relationship. The proportion of storage locations in areas A, B, C, and D are approximately 16%, 33%, 32%, and 19%, respectively. Initially, each area independently generates inventory with a 70% fill rate, meaning approximately 70% of storage locations within the same area are pre-occupied. In a single instance, the ratio of inbound to outbound tasks is set to 1:1. The goods categories for inbound tasks are randomly generated evenly from the four categories, resulting in each category accounting for approximately 25% of inbound tasks. The category distribution of outbound tasks is determined by the initial number of storage locations occupied in each area. Both inbound and outbound tasks are generated using the same random dynamic arrival mechanism: first, task intervals are generated based on an exponential distribution, and an original arrival sequence is constructed using consistent baseline parameters. Then, local fluctuations are introduced by switching between busy and idle periods, causing the task flow to exhibit dynamic characteristics of density or sparseness at different time periods. Finally, the arrival sequence is mapped to a given time window to obtain the final arrival time of each task.
[0023] In step S4, during training, instances of size 40×2×100 are selected for training, i.e., the number of tasks is 40, the number of stacker cranes is 2, and the number of storage locations is 100. The training converges as follows: Figure 5As shown, the trained model is directly applied to solve examples of 50×2×100, 60×2×100, 50×3×100, 60×3×120, 70×3×120 and 80×3×120 to verify the generalization ability of the training strategy. The instances used for training, validation and testing are all constructed according to the same random generation mechanism. During the training phase, instances are generated online. At the beginning of training, 100 additional instances of the same size as the training set are generated as a validation set. The test set consists of 100 instances generated independently and randomly according to the same generation mechanism at each size.
[0024] In comparison, the present invention, the HGNN-PPO-G deep reinforcement learning algorithm based on a greedy strategy, the RC-RS random stacker crane and random storage location allocation strategy, the RC-NS random stacker crane and nearest storage location allocation strategy, the EAC-NS earliest idle stacker crane and nearest storage location allocation strategy, and the EAC-RS earliest idle stacker crane and random storage location allocation strategy were used to test 100 instances of different sizes. The index C, which is the sum of the total system waiting time and the total running time, was used. total An evaluation was conducted, where Gap is the ratio of the difference between each method and the present invention to the present invention. The experimental results are shown in the table below:
[0025] The results in the table show that the present invention and the HGNN-PPO-G method achieve superior results in various warehouse scheduling instances, and are generally better than the other four methods. Traditional methods can still achieve certain results in small-scale instances, but their solutions gradually deteriorate as the scale increases. For example, the gap in RC-RS increases from 19.36% to 21.34%, and the gap in RC-NS increases from 16.42% to 18.28%, indicating that scheduling methods based on fixed rules are difficult to effectively cope with collaborative decision-making problems in complex dynamic warehouse systems. In contrast, the method proposed in this invention can better learn the joint optimization strategy of task scheduling and location allocation.
[0026] To further verify the optimization effect of the method of the present invention in dynamic inbound and outbound operations, the present invention selected 40×2×100 and 60×3×100 scale instances to compare and analyze the shelf status before and after task execution, such as... Figure 6-9As shown, different colors represent occupied storage locations in zones I, II, III, and IV, respectively. White areas represent empty storage locations, and darker areas represent storage locations reused during task execution. The numbers in the cells indicate the time the corresponding goods have been in the warehouse. By comparing the initial and final states of the shelving, it can be seen that after completing a batch of inbound and outbound tasks, the storage location occupancy structure of the shelving on both sides of each aisle has been significantly adjusted, the overall distribution of goods is more orderly, and the zoning storage characteristics are well maintained. Further analysis of the final state of the shelving reveals that the method of this invention can select better target storage locations for inbound tasks. At the same time, some empty storage locations released by outbound operations are reused in subsequent task execution. In example 40×2×100, the empty storage locations released by outbound operations are reused in subsequent task execution. Figure 6 , 7 As shown, the number of reusable storage locations corresponding to the darker color areas is 4; in example 60×3×100, by Figure 8 , 9 As shown, the number of reused storage locations is 9. Furthermore, outbound operations prioritize the removal of goods with longer storage times, indicating that this method, while meeting constraints, comprehensively considers subsequent operational connections and overall space utilization efficiency, prioritizing the accessibility of reused storage locations, thereby reducing unnecessary round-trip travel.
[0027] The scope of protection of this invention includes, but is not limited to, the above embodiments. The scope of protection is defined by the claims. Any substitutions, modifications, or improvements to this technology that are easily conceived by those skilled in the art fall within the scope of protection of this invention.
Claims
1. A method for joint scheduling of inbound and outbound operations in elevated warehouses based on heterogeneous graph neural networks, characterized in that... The specific steps include the following: S1. A mathematical model for the joint scheduling of inbound and outbound operations in a high-bay warehouse is established, considering inbound and outbound tasks, stacker crane resource allocation, storage location occupancy status, storage location zoning matching, and complex operational conditions. The mathematical model includes an objective function and constraints. The objective function aims to minimize the sum of the total waiting time and total operation time of the inbound and outbound scheduling system. The relationship is as follows: , , , , Where Z represents the objective function. This represents the total waiting time for all tasks. This represents the total job time when all tasks are executed in single-job mode. Let p represent the time saved when the i-th inbound task is paired with the j-th outbound task to form a composite operation. p is the index of the task. Tasks are divided into inbound and outbound tasks. The number of tasks is the sum of the number of inbound and outbound tasks. O is the task set, O = {O1, O2, ..., O...} p }, where i is the index of the inbound task, O in For the collection of inbound tasks, j is the index of the outbound task, O out For the set of outbound tasks, u ij This variable represents whether the i-th inbound task and the j-th outbound task form a compound operation. If they are arranged as a compound operation to be executed consecutively by the same stacker crane, then u ij =1, otherwise 0; S represents a single task, C represents a compound task. The time for executing a single job for the i-th inbound task. The time for executing the order for the j-th outbound task. Let r represent the waiting time for the p-th task. k Let a be the earliest available time for the k-th stacker crane, where k is the stacker crane index. p Let $\mathbf{p}$ be the arrival time of the $p$-th task. When the $p$-th task arrives, the $k$-th stacker crane assigned to it is already in an idle state. At this time, the p-th task can start execution immediately. When the p-th task arrives, the k-th stacker crane is still performing the preceding operation, i.e. When the p-th task needs to wait in the queue, its waiting time is the difference between the available time of the k-th stacker crane and the arrival time of the task; Let x be the time for executing a single job for the p-th task. p y p Let V be the coordinates of the storage location corresponding to the p-th task. x and V y These refer to the stacker crane's travel speeds in the lateral and longitudinal directions, respectively. The time required to execute the composite operation after pairing the i-th inbound task with the j-th outbound task to form a composite operation, (x i y i Let (x) be the coordinates of the storage location corresponding to the i-th inbound task. j y j Let be the coordinates of the storage location corresponding to the j-th outbound task; S2. Map the inbound / outbound tasks, stacker cranes, and storage locations in the mathematical model to task nodes, stacker crane nodes, and storage location nodes in the heterogeneous graph, respectively, and establish the heterogeneous graph; S3. Establish a heterogeneous graph neural network, extract features from the heterogeneous graph through the heterogeneous graph neural network, obtain node-level embedding representations of task nodes, stacker crane nodes and storage location nodes, and then aggregate the node-level embedding representations to obtain graph-level state embedding vectors and state value functions used to characterize the overall operation status of the high-bay warehouse. S4. Input the graph-level state embedding vector and state value function into the Actor-Critic network trained in the near-end policy optimization algorithm. The Actor network outputs two-stage scheduling actions: task-stall crane and stacker crane-cargo location, until a complete high-bay warehouse inbound and outbound joint scheduling scheme is generated.
2. The high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S1, the storage locations are divided into zones based on the turnover rate of the goods. The composite operation conditions are: when the stacker crane performs an inbound task, after the inbound task is completed, there is an executable outbound task that has arrived but not yet been completed in the outbound scheduling system, and there is an inventory location within the current service range of the stacker crane that matches the category of goods for the outbound task.
3. The high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S1, the constraint condition is: , , , , , , Among them: B p is the actual start time of the p-th task; M is a positive number used to implement linearization of logical constraints, and its value is not less than the maximum time difference occurring within the scheduling period; x pk Let x be a binary decision variable, representing whether the p-th task is executed by the k-th stacker crane. If the p-th task is assigned to the k-th stacker crane, then x... pk =1, otherwise 0; W is the set of stacker cranes, W={W1, W2, ..., W...} k }, where l is the location index, L is the set of locations, L={L1, L2, ..., L l }, y pl Let y be a binary decision variable, representing whether the p-th task matches the l-th storage location. If the p-th task matches the l-th storage location, then y... pl =1, otherwise 0; Let z be the classification attribute of the goods corresponding to the p-th task. l This refers to the category area to which the l-th storage location belongs. τ represents the storage time of goods in the l-th storage location; j Let be the in-stock time requirement for the j-th outbound task.
4. The high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S2, the heterogeneous graph is defined as: G = (O', W', L', E) OO E OW E WL (Start, End), where: the task node set O' consists of all task nodes The stacker crane node set W' consists of all stacker crane nodes. The set of storage location nodes L' consists of all storage location nodes. Composition, E OO Let E be a set of task-task edges. Edges between inbound tasks are directed conjunctive arcs, representing the first-come, first-served time sequence constraint between them. Edges between outbound tasks are undirected disjunctive arcs, and edges between inbound and outbound tasks are undirected disjunctive arcs, representing potential pairing relationships in composite operations. OW This is a set of task-stacking machine edges, where each edge is an undirected disjunctive arc, used to connect task nodes and their executable stacker machine nodes. E WL This is a set of stacker crane-cargo location edges, where each edge is an undirected disjunctive arc. These edges represent candidate cargo locations that the stacker crane can reach when performing a transportation task, and reflect the travel distance relationship between the stacker crane and the corresponding cargo location. Start and End are two virtual nodes, representing the start and end of the scheduling process, respectively. The original characteristics of task nodes, stacker crane nodes, storage location nodes, and various edges are defined as follows: The original characteristics of a task node include task type, task assignment status, task arrival time, and cargo category information; the original characteristics of a stacker crane node include its aisle, current location, and remaining busy time; the original characteristics of a storage location node include its aisle, coordinate location, occupancy status, inventory category, cargo storage duration, and static travel time; the original characteristics of the edge connecting the stacker crane and the storage location are defined as the travel time from the current location of the stacker crane to the coordinates of the target storage location.
5. The high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S3, the embedded representations of the storage location node, the stacker crane node, and the task node are obtained sequentially through a three-stage embedding process, specifically as follows: Phase 1: Propagate the feature vectors of the task node and stacker crane node to the connected storage location node to obtain the updated storage location node embedding representation; The second stage involves propagating the feature vectors of the task node and the storage location node to the first-order adjacent stacker crane nodes to obtain the updated stacker crane node embedding representation. The third stage involves propagating the feature vectors of the location node, stacker crane node, and task node to the task node to obtain an updated task node embedding representation.
6. The high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks as described in claim 5, characterized in that: In step S3, the updated location node embedding representation The relation is: , , in: For the activation function, m pk For task nodes via stacker crane node The aggregated intermediate embedding representation, u kl This refers to the original features of the connection edge between the stacker crane and the storage location. Let represent the set of stacker crane nodes connected to the storage location nodes, and α be the attention weight coefficient during the storage location node embedding phase. h represents the set of task nodes connected to the stacker crane node. p For the original characteristics of the task node, m k These are the original features of the stacker crane node; For the storage location node, the stacker crane node is embedded in the stage normalized. For cargo location nodes Attention coefficient Embed the phase-normalized task node for the storage location node. For stacker crane nodes Attention coefficient For the storage location node, the stacker crane node is embedded in the stage normalized. For stacker crane nodes Attention coefficient; Updated stacker node embedding representation The relation is: , , in: For stacker crane nodes Through task nodes The aggregated intermediate embedding representation, where β is the attention weight coefficient during the stacker crane node embedding stage. This represents the set of storage location nodes connected to the stacker crane node. For the stacker crane node embedding stage, the normalized storage location node For stacker crane nodes Attention coefficient For the stacker crane node embedding stage, the normalized task node For stacker crane nodes Attention coefficient; Updated task node embedding representation The relation is: , Among them, MLP0, MLP1, MLP2, and MLP3 are all multilayer perceptrons. Each MLP contains two 128-dimensional hidden layers, and ELU is an exponential linear unit used as the activation function. This represents the set of stacker crane nodes connected to the task node.
7. The high-bay warehouse inbound / outbound joint scheduling method based on heterogeneous graph neural networks as described in claim 5, characterized in that: In step S3, the joint inbound and outbound scheduling is divided into two stages: the task-staller allocation stage and the stacker crane-location stage. The task node embedding representation, the stacker crane node embedding representation, and the location node embedding representation are aggregated to obtain the graph-level state embedding vector for the task-staller allocation stage. Graph-level state embedding vectors for stacker crane-location phase The relationship is as follows: , , Embed the graph-level state of the task-stacker assignment phase into a vector. Graph-level state embedding vectors for stacker crane-location phase Further concatenation yields the state value function, which serves as the input to the subsequent Critic network. The state value function's equation is: , in: Represents the Critic network, V(s) t ) represents the state s at time step t. t Value estimate.
8. The method for joint scheduling of inbound and outbound operations in a high-bay warehouse based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S4, an environment model based on a Markov decision process is first constructed in the proximal policy optimization algorithm. The proximal policy optimization algorithm trains an Actor-Critic network based on the environment model. The environment model includes: State space: The scheduling state of the high-bay warehouse inbound and outbound operations is represented by a heterogeneous graph consisting of task nodes, stacker crane nodes, and storage location nodes. The state is denoted as s, which is a set of the original features of task nodes, stacker crane nodes, and storage location nodes. Action Space: Let A be the set of actions for scheduling inbound and outbound operations in the high-bay warehouse, and let a be the action taken by the agent at time step t. t ,and The action is broken down into: task - stacker machine action. With stacker crane - storage position operation ,Right now: , in: This represents a task-stacking machine sub-action, specifically: assigning an executable task to the stacker crane at time step t. O-W This represents the set of possible actions during the task-stacker allocation phase. This indicates a stacker crane-position sub-action, meaning: selecting a target storage location for the task assuming the stacker crane has already been selected at time step t. (A) W-L This represents the set of possible actions during the stacker crane-location phase. State transition function: based on the state s at time step t. t and the action chosen by the agent at time step t. After the inbound / outbound scheduling system executes this action, it updates the stacker crane position, remaining busy time, storage location occupancy status, task completion status, and task queue information. Based on the updated status, it regenerates the connection edges between candidate tasks and stacker cranes, as well as the connection edges between stacker cranes and storage locations. The heterogeneous graph is then completed from state s. t to state s t+1 The transfer; Reward function: Let A be the set of actions corresponding to time step t. t When performing a single operation, A t It contains only one task; when performing compound operations, A t If a task contains one inbound task and one outbound task, then the single-step scheduling cost function C is... t Represented as: , in, This represents the waiting time incurred by the task at time step t before it begins execution. This represents the stacker crane's running time at time step t. When a single operation is executed at time step t... Single job time When a composite operation actually occurs at time step t, For composite operation time ; This indicates the duration of time the goods have been in storage for the corresponding outbound task. This represents the average duration of inventory in storage, and η is a penalty coefficient, set to 0.002, used to adjust the degree of influence of inventory storage time on rewards. This represents the set of actions for outbound tasks at time step t. When time step t does not contain any outbound tasks, then... Therefore, the reward function is defined as R. t = -C t .
9. The method for joint scheduling of inbound and outbound operations in a high-bay warehouse based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S4, the Actor-Critic network includes two independent Actor networks and one Critic network. The two graph-level state embedding vectors are used as inputs to the two Actor networks, and the state value function is used as input to the Critic network.
10. The method for joint scheduling of inbound and outbound operations in a high-bay warehouse based on heterogeneous graph neural networks as described in claim 1, characterized in that: In step S4, during the training of the Actor-Critic network, the agent generates actions using a two-stage sequential allocation method. The first stage is based on the local state of the task-stacker allocation. After completing the task and stacker crane matching, the graph-level state corresponding to this stage is embedded into a vector and input into the first Actor network. The Actor network contains a single-layer LH hidden layer and is equipped with a tanh activation function. It outputs the original probability distribution of task-stacker sub-actions in the current state, which is then normalized to obtain the first-stage allocation strategy. Based on the determined task-stacking machine matching results, construct the local state of the stacker crane-warehouse allocation stage. After completing the matching of the stacker crane and the storage location, the graph-level state corresponding to this stage is embedded into a vector and input into a second Actor network with the same structure. The original probability distribution of the stacker crane-storage location sub-actions is output and normalized to obtain the second-stage allocation strategy. π represents the policy function; Simultaneously, the graph-level state embedding vectors of the two stages are concatenated, and the resulting state value function is input into the Critic network, which outputs the value estimate V(s) of the current state. t During training, the reward function is used to obtain the reward corresponding to the current scheduled action. This reward is then used in conjunction with the value estimate V(s). t The difference between the two actors is used to evaluate the merits of the current two-stage scheduling actions, and the parameters of the two actor networks are updated accordingly to adjust the selection probabilities of the sub-actions in the two stages; simultaneously, based on the reward and value estimate V(s)... t The error between the two is used to update the parameters of the Critic network, so that the value estimate output by the Critic network gradually approaches the reward corresponding to the current scheduling action; through multiple rounds of iterative training, a well-trained Actor-Critic network is obtained.
Citation Information
Patent Citations
Warehousing dynamic scheduling method and system based on digital twinning and deep learning, and storage medium
CN121212744A
System and method for ride order dispatching and vehicle repositioning
US20200074353A1