Stereo warehouse goods location allocation and picking scheduling method and system based on digital twinning

By optimizing the allocation of storage locations in automated warehouses using digital twin technology and deep reinforcement learning algorithms, and combining virtual potential field theory to guide the collaborative scheduling of stacker cranes, the dynamic optimization and path conflict problems of storage location allocation and picking scheduling in automated warehouses have been solved, thereby improving warehousing efficiency and emergency order response capabilities.

CN120851774BActive Publication Date: 2025-11-21XIAMEN SINOSERVICES INFORMATION TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511339817.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-19
Publication Date
2025-11-21
Estimated Expiration
2045-09-19

AI Technical Summary

Technical Problem

Automated warehouses suffer from problems in location allocation and picking scheduling, such as static zoning strategies failing to adapt to market changes, path conflicts during multi-machine collaborative operations, and insufficient response to urgent orders.

Method used

A three-dimensional digital model of an automated warehouse is constructed using digital twin technology. Deep reinforcement learning algorithms are used for warehouse location allocation, and virtual potential field theory is used to guide the collaborative scheduling of stacker cranes, achieving dynamic optimization and intelligent obstacle avoidance.

Benefits of technology

It improved warehouse space utilization and order picking efficiency, resolved path conflict issues, and enhanced the ability to respond to urgent orders.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120851774B_ABST
    Figure CN120851774B_ABST
Patent Text Reader

Abstract

The application provides a stereoscopic warehouse goods location allocation and picking scheduling method and system based on digital twinning, relates to the technical field of digital twinning, and comprises the following steps: constructing a three-dimensional digital model of a stereoscopic warehouse; allocating goods locations based on a deep reinforcement learning algorithm, taking commodity storage frequency, associated purchase probability and seasonal demand prediction as input parameters, taking the shortest picking path and associated commodity centralized storage as optimization objectives to generate goods location allocation instructions; performing warehouse entry operations and receiving picking orders; constructing a virtual potential field in the model, determining task priorities based on order urgency and triggering obstacle avoidance when repulsive force between stackers exceeds a threshold; calculating an obstacle avoidance trajectory and generating a collaborative scheduling instruction to perform picking operations. The application realizes efficient goods location allocation and intelligent multi-stacker collaborative scheduling.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to digital twin technology, and more particularly to a method and system for the allocation and picking scheduling of storage locations in an automated warehouse based on digital twins. Background Technology

[0002] With the rapid development of e-commerce and the increasing demands of consumers for fast logistics response, traditional flat warehouses can no longer meet the needs of modern logistics systems for efficient storage and rapid picking. Automated storage and retrieval systems (AS / RS), as a modern form of warehousing that efficiently utilizes three-dimensional space, are gradually becoming core facilities in large logistics centers due to their high-density storage capacity and automated operation advantages.

[0003] However, the operation and management of automated warehouses face numerous technical challenges and shortcomings: Location allocation often employs a static zoning strategy, relying on simple fixed rules or rules of thumb, which cannot adapt to seasonal changes in product sales characteristics and related purchasing patterns. It is difficult to dynamically optimize based on product characteristics and market demand, resulting in low storage and retrieval efficiency. When multiple stacker cranes operate collaboratively in a high-density track space, the scheduling algorithm is mainly based on the shortest path principle, lacking an effective mechanism for conflict avoidance during multi-machine collaborative operations. This frequently leads to path conflicts and scheduling congestion, making it impossible to respond to urgent order demands in real time, thus affecting overall picking efficiency. Furthermore, the assessment of order urgency is overly simplistic, and there is a lack of effective collaborative decision-making mechanisms between equipment status and order priority, making it difficult to make optimal task priority decisions in complex environments and easily creating bottlenecks under high load conditions. Summary of the Invention

[0004] To address the shortcomings of existing technologies, this invention provides a method and system for the allocation and picking scheduling of storage locations in an automated warehouse based on digital twins, which can solve the problems in existing technologies.

[0005] A first aspect of this invention provides a method for location allocation and picking scheduling in an automated warehouse based on digital twins, comprising:

[0006] A three-dimensional digital model of the automated warehouse was constructed using digital twin technology;

[0007] In the three-dimensional digital model, storage location allocation is performed based on a deep reinforcement learning algorithm. The frequency of goods access, probability of related purchases, and seasonal demand forecast results are used as input parameters. The optimization objectives are to minimize the picking path and maximize the centralized storage of related goods. Based on the iterative calculation results, storage location allocation instructions are generated.

[0008] Execute the goods receiving operation according to the warehouse location allocation instruction, and receive picking task orders;

[0009] A virtual potential field is constructed in the three-dimensional digital model. The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order. When the repulsive force between stacker cranes is detected to be greater than a preset threshold, the task priority is determined based on the urgency of the picking task order and an obstacle avoidance and yielding decision is triggered. According to the obstacle avoidance and yielding decision, the obstacle avoidance trajectory of the stacker crane is calculated, a stacker crane collaborative scheduling instruction without collision risk is generated, and the picking operation is executed according to the stacker crane collaborative scheduling instruction.

[0010] Optionally,

[0011] The steps for constructing a 3D digital model of an automated warehouse using digital twin technology include:

[0012] A three-dimensional digital model of an automated warehouse, including rack units, stacker cranes, and conveyor lines, was constructed using digital twin technology.

[0013] The collected data on warehouse occupancy and stacker crane location are written into a dual buffer, which uses a staggered read / write mechanism to prioritize the data. The collected data is converted into corresponding status information, and the status information is synchronously updated to the three-dimensional digital model according to the urgency priority based on the timestamp.

[0014] The stacker crane position data in the three-dimensional digital model is compared with the actual collected stacker crane position data. The stacker crane position data is corrected using the Kalman filter algorithm, and the stacker crane movement trajectory is smoothed using B-spline curves.

[0015] Optionally,

[0016] In the three-dimensional digital model, the allocation of storage locations based on a deep reinforcement learning algorithm, using the frequency of goods access, the probability of associated purchases, and seasonal demand forecasts as input parameters, and the iterative calculation steps with the optimization objectives of minimizing the picking path and maximizing the centralized storage of associated goods, include:

[0017] The frequency of product access, probability of associated purchase, and seasonal demand prediction results are used as state features input into the state space of the deep reinforcement learning network. The probability of associated purchase is calculated based on historical order data, and the seasonal demand prediction results are obtained by combining the seasonality coefficient and residual term of historical sales data of the product.

[0018] An evaluation network and a target network of the deep reinforcement learning network are constructed. Both the evaluation network and the target network use long short-term memory network units to form hidden layers. The target network updates its parameters through the evaluation network.

[0019] The storage location allocation scheme is used as the action space of the deep reinforcement learning network to construct a reward function. The reward function includes a picking path loss term calculated based on the frequency of goods access and the distance from the storage location to the inlet / outlet, and a storage concentration reward term calculated based on the probability of goods associated purchase and the allocation status of adjacent storage locations.

[0020] A priority experience replay mechanism is used to store and filter state transition samples. Based on the evaluation network and the target network, the time-series difference error is calculated and the network parameters are updated to generate the final cargo location allocation scheme.

[0021] Optionally,

[0022] The steps to construct the reward function include:

[0023] The dynamic distance matrix from each storage location to the inlet / outlet is calculated based on the real-time location of the stacker crane. The product of the dynamic distance matrix and the frequency of goods access is used as the baseline loss value. The obstacle avoidance path length between storage locations is calculated based on the real-time location of the stacker crane. The weighted sum of the obstacle avoidance path length and the baseline loss value is used as the picking path loss item.

[0024] Products with related relationships are divided into multiple association groups using a hierarchical clustering algorithm, and the association strength between products within the association group is calculated. Spatial proximity constraints between products are constructed based on the association strength. The degree of matching between the spatial proximity constraints and the allocation status of adjacent storage locations is used as a storage concentration reward.

[0025] Extract the temporal relationship of product combinations from historical order data, calculate the product association probability under different temporal relationships, determine the deviation between the product association probability and the actual associated purchase probability as the weight adjustment coefficient, and dynamically update the weight coefficient of the storage concentration reward item according to the weight adjustment coefficient.

[0026] A reward function is constructed by combining the picking path loss term with the storage concentration reward term. A reward decay factor is calculated based on the gradient of the reward function. The product of the reward decay factor and the temporal difference error is used as the priority weight of the state transition sample.

[0027] Optionally,

[0028] The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order; when the repulsive force between the stacker cranes is detected to be greater than a preset threshold, the steps of determining the task priority and triggering obstacle avoidance and yielding decision based on the urgency of the picking task order include:

[0029] The urgency of picking tasks is calculated based on order waiting time, number of items in the order, and original order priority.

[0030] A virtual potential field is constructed in the three-dimensional digital model, and the potential energy of the virtual potential field is determined according to the movement speed of the stacker crane and the urgency of the picking task order. The potential energy of the virtual potential field includes a repulsive potential energy component that is proportional to the square of the movement speed of the stacker crane, and an attractive potential energy component that is proportional to the urgency of the picking task order.

[0031] The repulsive force between the stacker cranes is calculated based on the potential energy of the virtual potential field. When the repulsive force is greater than a preset threshold, the urgency of the picking task orders executed by the multiple stacker cranes that have encountered obstacle avoidance conflicts is obtained. The task priority is determined according to the relationship between the urgency levels, and an obstacle avoidance and yielding decision is triggered based on the task priority.

[0032] Optionally,

[0033] The steps for determining the potential energy of the virtual potential field based on the stacker crane's movement speed and the urgency of the picking task order include:

[0034] The motion velocity vector of the stacker crane is decomposed into radial velocity components and tangential velocity components, and the motion state coefficient is determined based on the ratio of the radial velocity component to the tangential velocity component.

[0035] The urgency of the picking task orders is mapped to a spatial attraction coefficient, which increases with the urgency. The order density of local areas is calculated based on the distribution of order target points in the three-dimensional digital model, and the spatial attraction coefficient is dynamically weighted based on the order density.

[0036] An adaptive potential field weight is constructed, which is positively correlated with the motion state coefficient and negatively correlated with the spatial attraction coefficient; when the motion state coefficient is greater than a first threshold, the weight of repulsive potential energy is increased, and when the spatial attraction coefficient is greater than a second threshold, the weight of attractive potential energy is increased.

[0037] The ratio of repulsive and attractive potential energy components is adjusted according to the adaptive potential field weights, and the potential energy magnitude of the virtual potential field is updated in real time.

[0038] Optionally,

[0039] Based on the obstacle avoidance and yielding decision, the steps of calculating the obstacle avoidance trajectory of the stacker crane, generating a stacker crane collaborative scheduling instruction with no collision risk, and executing the picking operation according to the stacker crane collaborative scheduling instruction include:

[0040] The initial obstacle avoidance trajectory is generated by fifth-order polynomial interpolation, and the initial obstacle avoidance trajectory satisfies the constraints of the start and end point positions, velocity, and acceleration.

[0041] A spatiotemporal topology graph is constructed, wherein the nodes of the spatiotemporal topology graph represent the position and state of the stacker crane at different times, and the edges of the spatiotemporal topology graph represent feasible state transitions; the trajectory smoothness and the weighted sum of the target temporal node deviation are calculated as the trajectory cost based on the spatiotemporal topology graph, and the initial obstacle avoidance trajectory is optimized according to the trajectory cost;

[0042] The actual motion state of the stacker crane is collected, and the position deviation between the actual position of the stacker crane and the planned trajectory is calculated. When the position deviation is greater than the fault tolerance threshold, a local optimization interval is constructed based on the current state of the stacker crane and the target time sequence node, and the trajectory is replanned within the local optimization interval.

[0043] Calculate the time window for each stacker crane to reach the target timing node based on the spatiotemporal topology diagram, construct a speed adjustment function based on the time window, and adjust the movement speed of each stacker crane through the speed adjustment function so that multiple stacker cranes pass through the target timing node at staggered times.

[0044] Based on the replanned trajectory and the speed adjustment function, a stacker crane collaborative scheduling instruction is generated; after verifying that the stacker crane collaborative scheduling instruction meets the safety distance constraint, the picking operation is executed.

[0045] Secondly, it provides a digital twin-based automated warehouse location allocation and picking scheduling system, including:

[0046] The first unit is used to construct a three-dimensional digital model of an automated warehouse using digital twin technology;

[0047] The second unit is used to allocate storage locations in the three-dimensional digital model based on a deep reinforcement learning algorithm. It takes the frequency of goods access, the probability of related purchases, and the seasonal demand forecast as input parameters, and performs iterative calculations with the optimization objectives of minimizing the picking path and maximizing the centralized storage of related goods. Based on the iterative calculation results, it generates storage location allocation instructions.

[0048] The third unit is used to execute the goods receiving operation according to the receiving location allocation instruction and to receive picking task orders;

[0049] The fourth unit is used to construct a virtual potential field in the three-dimensional digital model. The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order. When the repulsive force between stacker cranes is detected to be greater than a preset threshold, the task priority is determined based on the urgency of the picking task order and an obstacle avoidance and yielding decision is triggered. According to the obstacle avoidance and yielding decision, the obstacle avoidance trajectory of the stacker crane is calculated, a stacker crane collaborative scheduling instruction without collision risk is generated, and the picking operation is executed according to the stacker crane collaborative scheduling instruction.

[0050] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

[0051] This invention constructs a high-precision 3D digital model of an automated warehouse using digital twin technology, achieving accurate mapping and real-time synchronization between the physical warehouse and the digital model. A deep reinforcement learning algorithm built with a Long Short-Term Memory (LSTM) network is used for dynamic location allocation. This algorithm considers not only the frequency of goods access but also the probability of associated purchases and seasonal demand forecasts, enabling the location allocation scheme to adaptively optimize picking paths and maximize the centralized storage of related goods. Compared to traditional static partitioning strategies, this method can adjust the location allocation strategy in real time according to changes in market demand, significantly improving warehouse space utilization and order picking efficiency. Furthermore, this invention innovatively introduces virtual potential field theory to guide the collaborative scheduling of multiple stacker cranes. By constructing a potential energy model combining movement speed and order urgency, intelligent obstacle avoidance and dynamic adjustment of task priorities among stacker cranes are achieved, effectively solving path conflict problems in multi-machine collaborative operations and improving the responsiveness to urgent orders. Attached Figure Description

[0052] Figure 1 This is a flowchart illustrating the method for allocating and picking in an automated warehouse based on digital twins, as described in an embodiment of the present invention.

[0053] Figure 2 This is a flowchart of stacker crane obstacle avoidance trajectory planning and collaborative scheduling based on spatiotemporal topology graph. Detailed Implementation

[0054] The technical solutions of the present invention will be described below with reference to the accompanying drawings. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0055] Figure 1 This is a flowchart illustrating the method for location allocation and picking scheduling in an automated warehouse based on digital twins, as described in this invention. Figure 1 As shown, the method includes:

[0056] A three-dimensional digital model of the automated warehouse was constructed using digital twin technology;

[0057] In the three-dimensional digital model, storage location allocation is performed based on a deep reinforcement learning algorithm. The frequency of goods access, probability of related purchases, and seasonal demand forecast results are used as input parameters. The optimization objectives are to minimize the picking path and maximize the centralized storage of related goods. Based on the iterative calculation results, storage location allocation instructions are generated.

[0058] Execute the goods receiving operation according to the warehouse location allocation instruction, and receive picking task orders;

[0059] A virtual potential field is constructed in the three-dimensional digital model. The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order. When the repulsive force between stacker cranes is detected to be greater than a preset threshold, the task priority is determined based on the urgency of the picking task order and an obstacle avoidance and yielding decision is triggered. According to the obstacle avoidance and yielding decision, the obstacle avoidance trajectory of the stacker crane is calculated, a stacker crane collaborative scheduling instruction without collision risk is generated, and the picking operation is executed according to the stacker crane collaborative scheduling instruction.

[0060] Optionally,

[0061] The steps for constructing a 3D digital model of an automated warehouse using digital twin technology include:

[0062] A three-dimensional digital model of an automated warehouse, including rack units, stacker cranes, and conveyor lines, was constructed using digital twin technology.

[0063] The collected data on warehouse occupancy and stacker crane location are written into a dual buffer, which uses a staggered read / write mechanism to prioritize the data. The collected data is converted into corresponding status information, and the status information is synchronously updated to the three-dimensional digital model according to the urgency priority based on the timestamp.

[0064] The stacker crane position data in the three-dimensional digital model is compared with the actual collected stacker crane position data. The stacker crane position data is corrected using the Kalman filter algorithm, and the stacker crane motion trajectory is smoothed using B-spline curves, so as to achieve an accurate mapping between the three-dimensional digital model and the actual operating state.

[0065] For example, a 3D digital model of an automated warehouse including racking units, stacker cranes, and conveyor lines is constructed. In this construction process, computer-aided design software is used to create the geometric model of the automated warehouse, and the racking units are parametrically modeled. Each racking unit includes a location number, dimensions, and load-bearing parameters. Taking a large automated warehouse as an example, the racking units are modeled using a 3D mesh structure, with each location having dimensions of 1200mm × 1000mm × 1500mm and a load-bearing capacity of 1500kg. For the stacker cranes, a layered modeling approach is used, including the chassis, lifting mechanism, and picking mechanism, with defined motion parameters such as a maximum walking speed of 5m / s, a maximum lifting speed of 2m / s, and a maximum acceleration of 1.2m / s². 2 The conveyor line model includes a conveyor belt, a turning mechanism, and a sorter. The model takes into account an operating speed of 1.8 m / s and a load capacity of 50 kg / m.

[0066] After completing the geometric model, dynamic behavioral characteristics are added to the digital model. Real-time data is collected through sensor interfaces, including location occupancy data, stacker crane position data, and conveyor line load data. Location occupancy data is collected via infrared or weight sensors, and includes location ID, occupancy status (0 for free, 1 for occupied), and timestamp. Stacker crane position data is acquired via photoelectric encoders and laser rangefinders at a frequency of 100Hz, with an accuracy of ±5mm. Conveyor line load data is collected via pressure sensors and photoelectric switches, with an update frequency of 50Hz. This collected data is written to a dual buffer for processing. The dual buffer design includes two buffers, A and B. When data is written to buffer A, buffer B is available for reading, and then the roles alternate. The size of the dual buffers is set to 8MB, sufficient to store approximately 30 seconds of high-frequency data. The dual buffers use a staggered read / write mechanism to prioritize data, marking and classifying data. High-urgency data (such as stacker crane collision warnings) is prioritized to 5, normal operational data to 3, and statistical data to 1. The buffers are checked every 10ms, and data is processed in priority order.

[0067] The collected raw data needs to be converted into automated warehouse status information. Storage location occupancy data is converted into the loading status of rack units, displayed in a 3D model using different colors (red indicates occupied, green indicates idle). Stacker crane location data is converted into the stacker crane's coordinates and orientation in 3D space, including X-axis position (aisle direction), Y-axis position (height direction), and Z-axis position (depth direction). Conveyor line load data is converted into the flow status of items on the conveyor line, including item ID, current location, and estimated arrival time. The converted status information is synchronously updated to the 3D digital model based on timestamps and priority according to urgency. The update mechanism employs a variable frequency strategy: high-priority information (such as stacker crane position) is updated every 100ms, medium-priority information (such as conveyor line status) every 500ms, and low-priority information (such as storage location occupancy status) every 2000ms. This hierarchical update strategy ensures the real-time nature of critical data while avoiding excessive resource consumption.

[0068] To ensure accurate mapping between the digital model and actual operating conditions, data calibration and trajectory smoothing are performed. The stacker crane position data in the 3D digital model is periodically compared with the actual collected stacker crane position data. During the comparison, the deviation between the model position and the actual position is calculated, and a calibration process is triggered when the deviation exceeds 50mm. The calibration uses a Kalman filter algorithm to correct the stacker crane position data. The Kalman filter algorithm takes the stacker crane's historical position data and current measurement values ​​as input, predicts the position at the next moment, and then corrects the prediction based on the actual measurement values, thereby reducing the impact of random errors. The Kalman filter process sets the observation noise covariance to 0.01 and the process noise covariance to 0.005 to balance tracking accuracy and smoothness. For the stacker crane's motion trajectory, B-spline curves are used for smoothing. B-spline curves construct smooth and continuous curves through control points, effectively eliminating jitter and jumps in the trajectory. In the implementation, a third-order B-spline curve is selected, and the 10 most recent position points are used as control points. The node vectors are uniformly distributed in the interval [0,1]. The generated smooth trajectory can accurately reflect the actual movement state of the stacker crane, while filtering out high-frequency noise. The trajectory data is updated every 50ms to ensure the continuity and accuracy of the visual display.

[0069] This invention improves data processing efficiency and real-time performance through a double buffer and staggered read / write mechanism; it ensures timely synchronization of key data by adopting a priority-based hierarchical update strategy; and it significantly improves the consistency between the digital model and the actual operating state by using the Kalman filter algorithm and B-spline curves to correct and smooth the data.

[0070] Optionally,

[0071] In the three-dimensional digital model, the allocation of storage locations based on a deep reinforcement learning algorithm, using the frequency of goods access, the probability of associated purchases, and seasonal demand forecasts as input parameters, and the iterative calculation steps with the optimization objectives of minimizing the picking path and maximizing the centralized storage of associated goods, include:

[0072] The frequency of product access, probability of associated purchase, and seasonal demand prediction results are used as state features input into the state space of the deep reinforcement learning network. The probability of associated purchase is calculated based on historical order data, and the seasonal demand prediction results are obtained by combining the seasonality coefficient and residual term of historical sales data of the product.

[0073] An evaluation network and a target network of the deep reinforcement learning network are constructed. Both the evaluation network and the target network use long short-term memory network units to form hidden layers. The target network updates its parameters through the evaluation network.

[0074] The storage location allocation scheme is used as the action space of the deep reinforcement learning network. A reward function is constructed to evaluate the quality of the storage location allocation scheme. The reward function includes a picking path loss term calculated based on the frequency of goods access and the distance from the storage location to the inlet / outlet, and a storage concentration reward term calculated based on the probability of goods associated purchase and the allocation status of adjacent storage locations.

[0075] A priority experience replay mechanism is used to store and filter state transition samples. Based on the evaluation network and the target network, the time-series difference error is calculated and the network parameters are updated to generate the final cargo location allocation scheme.

[0076] For example, the state space is first constructed, and the product access frequency, associated purchase probability, and seasonal demand forecast results are input into the deep reinforcement learning network as state features. The product access frequency feature is calculated by statistically analyzing the average daily number of times the product was picked over the past 90 days, and then normalized to keep the value between 0 and 1. For example, if a certain daily necessities were picked 6480 times in the past 90 days, averaging 72 times per day, the normalized value is 0.82. The associated purchase probability is calculated based on historical order data. By analyzing order records from the past 180 days, the co-occurrence frequency of products in the same order is extracted, and the conditional probability between product pairs is calculated. For example, if a certain shampoo and conditioner co-occurred 4280 times in historical orders, and the shampoo appeared alone a total of 7350 times, then the associated purchase probability is 0.58. Seasonal demand forecasts are obtained by decomposing historical sales data of goods into time series, extracting seasonality coefficients and residual terms, and combining this with monthly sales data from the past two years to extract seasonality coefficients. These are then combined with an ARIMA model to predict sales trends for the next three months, resulting in a comprehensive demand forecast. For example, a certain sunscreen might have a seasonality coefficient of 1.85 in summer months, indicating that its demand is significantly higher than the annual average.

[0077] The deep reinforcement learning network architecture comprises an evaluation network and a target network, both with identical structures, using Long Short-Term Memory (LSTM) units for their hidden layers. The evaluation network assesses the value of each action in the current state, while the target network serves as a benchmark to calculate the target Q-value. The evaluation network structure includes an input layer, LSTM hidden layers, and an output layer. The input layer dimension is the number of state features; in this embodiment, it is the number of items multiplied by 3 (corresponding to access frequency, association probability, and seasonal demand, respectively). The hidden layer consists of 128 LSTM units, and the output layer dimension is equal to the action space size. LSTM units can capture the temporal dependencies between item features, making them particularly suitable for handling seasonal variations and associated purchase patterns. The target network structure is the same as the evaluation network, but its parameters are updated at a lower frequency, once every 5000 training iterations. This is achieved through soft parameter updates in the evaluation network, with the update weight factor set to 0.01 to ensure the stability of the learning process. The evaluation network uses the Adam optimizer with a learning rate of 0.0005 and a batch size of 64.

[0078] The location allocation scheme is defined as the action space of the deep reinforcement learning network. In an automated warehouse, each location is uniquely identified by its shelf number, layer number, and column number. For example, the location number for shelf 3, layer 5, column 2 in area A is A-03-05-02. The action space is defined as the mapping relationship for allocating goods to specific locations. For example, action 'a' represents allocating goods i to location j. To reduce the dimensionality of the action space, a hierarchical decision-making method is adopted. First, the area and shelf where the goods are located are determined, and then the specific location is determined, reducing the original action space, which could reach hundreds of thousands of dimensions, to a manageable level. A reward function is constructed to evaluate the merits of the location allocation scheme. The reward function contains two main components: a picking path loss term and a storage concentration reward term. The picking path loss term is calculated based on the frequency of goods access and the distance from the location to the inlet / outlet. The distance is measured using Manhattan distance, which is the sum of the distances traveled along the x, y, and z directions. For example, if a product has a storage frequency of 0.75 and its assigned storage location is 45 meters from the Manhattan distance of the inbound / outbound gate, the corresponding path loss is -33.75 (a negative value indicates a loss). The storage concentration reward is calculated based on the probability of related purchases of products and the allocation status of adjacent storage locations. When a product with high correlation is assigned to an adjacent storage location, a positive reward is given. For example, if the probability of related purchases of products A and B is 0.65, and they are assigned to adjacent storage locations, a reward value of 13.0 will be generated. The total reward function combines these two items through a weighted summation, with a path loss weight of 0.7 and a concentration reward weight of 0.3, balancing the optimization objectives of picking efficiency and storage concentration.

[0079] To improve learning efficiency, this method employs a priority-based experience replay mechanism to store and filter state transition samples. Specifically, a 100,000-sample experience replay buffer is used to store samples in four-tuple form, containing the current state, the action performed, the reward received, and the next state. Each sample is assigned a priority, which is proportional to the temporal difference error; the larger the temporal difference error, the more valuable the knowledge contained in the sample, and the higher its priority. During replay training, samples are sampled according to priority probability, with higher-priority samples having a greater probability of being selected, while still retaining a small probability of random sampling to prevent local optima caused by over-focusing on certain samples. To balance sampling bias, importance sampling weights are introduced to appropriately scale the gradient updates of high-priority samples. Each training iteration samples 64 samples from the replay buffer for batch updates. The network training process uses an ε-greedy strategy to balance exploration and utilization. The initial ε value is set to 0.9, representing a 90% probability of randomly selecting an action, and it linearly decays to 0.1 as training progresses, ensuring that the algorithm can both explore new allocation schemes and utilize learned knowledge. After 50,000 training iterations, the final converged evaluation network parameters are taken and applied to the storage location allocation decision to generate the final storage location allocation scheme. This scheme contains the optimal storage location for each product, output in dictionary format, with the product ID as the key and the storage location number as the value, and is directly passed to the warehouse management to execute the inbound operation.

[0080] This invention effectively captures seasonal variation patterns by introducing an LSTM network to process the temporal relationships of product features; it reduces the complexity of the action space by employing hierarchical decision-making, making the location allocation problem in large-scale automated warehouses solvable; it improves learning efficiency by combining a priority experience playback mechanism; and its multi-objective reward function design takes into account both picking path optimization and centralized storage of related products, realizing intelligent location allocation that dynamically adapts to market demands, significantly improving warehouse space utilization and order picking efficiency.

[0081] Optionally,

[0082] The steps to construct the reward function include:

[0083] The dynamic distance matrix from each storage location to the inlet / outlet is calculated based on the real-time location of the stacker crane. The product of the dynamic distance matrix and the frequency of goods access is used as the baseline loss value. The obstacle avoidance path length between storage locations is calculated based on the real-time location of the stacker crane. The weighted sum of the obstacle avoidance path length and the baseline loss value is used as the picking path loss item.

[0084] Products with related relationships are divided into multiple association groups using a hierarchical clustering algorithm, and the association strength between products within the association group is calculated. Spatial proximity constraints between products are constructed based on the association strength. The degree of matching between the spatial proximity constraints and the allocation status of adjacent storage locations is used as a storage concentration reward.

[0085] Extract the temporal relationship of product combinations from historical order data, calculate the product association probability under different temporal relationships, determine the deviation between the product association probability and the actual associated purchase probability as the weight adjustment coefficient, and dynamically update the weight coefficient of the storage concentration reward item according to the weight adjustment coefficient.

[0086] A reward function is constructed by combining the picking path loss term with the storage concentration reward term. A reward decay factor is calculated based on the gradient of the reward function. The product of the reward decay factor and the temporal difference error is used as the priority weight of the state transition sample.

[0087] For example, the picking path loss term is calculated, and the real-time location information of the stacker cranes is obtained, including coordinate values ​​on the X-axis (shelf length direction), Y-axis (shelf height direction), and Z-axis (aisle width direction). Taking the location data of 5 stacker cranes in an automated warehouse at a certain moment as an example, the coordinates of stacker crane 1 are (25.6, 8.2, 12.3), in meters, indicating that it is located 25.6 meters from the origin in the X-axis direction, 8.2 meters in the Y-axis direction, and 12.3 meters in the Z-axis direction. Based on the real-time location of the stacker cranes and the topology of the automated warehouse, the dynamic distance matrix from each storage location to the inlet / outlet is calculated. This distance matrix considers not only the static geometric distance but also the current location and availability of the stacker crane, dynamically adjusting the optimal picking path. For example, when the static distance of storage location A-05-03-02 from outlet 1 is 45 meters, considering the current positions of stacker cranes 1 and 3, the actual optimal picking distance is 52 meters. Each element of the dynamic distance matrix represents the distance from a specific storage location to a specific inbound / outbound port. For an automated warehouse with 5000 storage locations and 4 inbound / outbound ports, the distance matrix has a dimension of 5000×4. The dynamic distance matrix is ​​multiplied by the product access frequency to obtain the baseline loss value. The product access frequency is calculated by statistically analyzing the number of times each product is picked over the past 30 days; for example, the access frequency for a certain fast-moving consumer good is 0.75 times / day. Considering the obstacle avoidance requirements between stacker cranes during actual picking, the obstacle avoidance path length between storage locations is calculated based on the real-time position of the stacker cranes. The obstacle avoidance path is planned in a 3D grid space using the A* algorithm, with the safe operating distance of the stacker cranes as a constraint. For example, a safe distance of at least 5 meters must be maintained between two stacker cranes. When the planned path results in the stacker crane distance being less than the safe distance, a detour path is generated. Finally, the weighted sum of the obstacle avoidance path length and the baseline loss value is used as the picking path loss term, with weighting coefficients set to 0.7 and 0.3 to balance static path planning and dynamic obstacle avoidance requirements.

[0088] To optimize the storage layout of related products, a hierarchical clustering algorithm is used to divide related products into multiple association groups. This process uses hierarchical clustering, with the probability of associated purchase between products as a similarity measure. A product association matrix is ​​constructed, where each element represents the probability of associated purchase between corresponding product pairs. Taking 10,000 products stored in an automated warehouse as an example, the association matrix is ​​a sparse matrix of 10,000 × 10,000. A bottom-up aggregation strategy is adopted, initially with each product in an independent cluster, gradually merging clusters with high association strength until a preset number of clusters or a threshold for association strength within a cluster is met. In practice, the products are ultimately divided into 350 association groups, each containing an average of 28.6 products. For each association group, the association strength between products within the group is calculated; the association strength is the weighted average of the probability of associated purchase for all products within the group. An example of an association group is the "Baby Care Group," which includes products such as diapers, wipes, and baby powder, with an average association strength of 0.62. Spatial proximity constraints are constructed between products based on their association strength. Products with higher association strength have stricter spatial proximity constraints. The spatial proximity constraint is expressed as the reciprocal of the ideal storage distance between products. For example, two products with an association strength of 0.8 have a spatial proximity constraint value of 0.85, meaning these two products should be assigned to nearby storage locations. The degree of matching between the spatial proximity constraint and the allocation status of adjacent storage locations is used as a storage concentration reward. Products with high association strength are assigned to adjacent or nearby storage locations, receiving a positive reward; conversely, products with low association strength are stored in dispersed locations, receiving a negative reward or a lower positive reward.

[0089] To further improve the accuracy of the reward function, the temporal relationships of product combinations in historical order data are extracted. By analyzing order data from the past 365 days, the order and time intervals of product purchases within orders are identified. For example, the analysis reveals that after a customer purchases a coffee machine, there is an average 68% probability that they will purchase coffee beans within 15 days. Order data is grouped according to time windows, and the probability of product association under different temporal relationships is calculated. Temporal relationships are categorized into four types: immediate association (same order), short-term association (within 7 days), medium-term association (within 30 days), and long-term association (within 90 days). For each type of temporal relationship, the co-occurrence frequency of product pairs within that time window is calculated. For example, the probability of a certain shampoo and conditioner being associated in an immediate association is 0.82, while the probability in a short-term association is 0.65. The calculated temporal association probabilities are compared with the actual observed associated purchasing behavior, and the deviation value is calculated as a weight adjustment coefficient. For example, the predicted association probability of a certain product pair is 0.75, while the actual associated purchase probability is 0.85, with a deviation of 0.1, and the corresponding weight adjustment coefficient is 1.15. The weighting coefficients of the storage concentration reward item are dynamically updated based on the weighting adjustment coefficients, enabling the reward function to adaptively adjust the degree of emphasis placed on different types of associations. The weighting adjustment coefficients are recalculated every 7 days to ensure that the reward function can reflect changes in market demand in a timely manner.

[0090] A complete reward function is constructed by combining the picking path loss term and the storage concentration reward term. The combination method uses linear weighted summation, with initial weights of 0.6 and 0.4, respectively, and the coefficients are dynamically updated based on the weights. The picking path loss term is negative; a smaller value indicates higher picking efficiency. The storage concentration reward term is positive; a larger value indicates a higher degree of concentrated storage of related goods. Taking a certain location allocation scheme evaluation as an example, the calculated picking path loss term is -125.8, the storage concentration reward term is 98.3, and the comprehensive reward value is -33.0. The reward decay factor is calculated based on the gradient of the reward function. Specifically, the rate of change of the current reward value relative to the historical average reward value is calculated and mapped to a range between 0.9 and 1.0 using a soft maximum function. For example, if the current reward value is 15% higher than the historical average, the corresponding reward decay factor is 0.98. The product of the reward decay factor and the temporal difference error is used as the priority weight for state transition samples. The temporal difference error reflects the deviation between the estimated value and the actual value; a larger value indicates that the sample contains more valuable information. For example, if the temporal difference error of a state transition sample is 0.85 and the reward decay factor is 0.96, the corresponding priority weight is 0.816. The priority weight determines the probability that the sample will be selected for network updates during experience replay. The higher the priority weight, the greater the probability of being selected, thereby accelerating the network's learning of important experiences.

[0091] This invention combines a dynamic distance matrix with obstacle avoidance path length to make the picking path loss calculation more consistent with actual operation; it introduces hierarchical clustering and spatial proximity constraints to achieve centralized storage optimization of related goods; a dynamic weight adjustment mechanism based on temporal relationships enables the reward function to adaptively reflect changes in market demand; and a priority strategy combining reward decay factor and temporal difference error improves network learning efficiency and convergence speed, ultimately achieving intelligent location allocation that balances picking efficiency and related storage.

[0092] Optionally,

[0093] The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order; when the repulsive force between the stacker cranes is detected to be greater than a preset threshold, the steps of determining the task priority and triggering obstacle avoidance and yielding decision based on the urgency of the picking task order include:

[0094] The urgency of picking tasks is calculated based on order waiting time, number of items in the order, and original order priority.

[0095] A virtual potential field is constructed in the three-dimensional digital model, and the potential energy of the virtual potential field is determined according to the movement speed of the stacker crane and the urgency of the picking task order. The potential energy of the virtual potential field includes a repulsive potential energy component that is proportional to the square of the movement speed of the stacker crane, and an attractive potential energy component that is proportional to the urgency of the picking task order.

[0096] The repulsive force between the stacker cranes is calculated based on the potential energy of the virtual potential field. When the repulsive force is greater than a preset threshold, the urgency of the picking task orders executed by the multiple stacker cranes that have encountered obstacle avoidance conflicts is obtained. The task priority is determined according to the relationship between the urgency levels, and an obstacle avoidance and yielding decision is triggered based on the task priority.

[0097] For example, in a multi-stacking crane collaborative operation environment, task priority needs to be calculated based on the urgency of picking orders. The urgency of picking orders is determined by three key factors: order wait time, number of items in the order, and original order priority. Order wait time refers to the time interval from when the order is entered to the current moment, measured in minutes. For example, if an order has been waiting for 35 minutes, the wait time will be normalized, mapping 35 minutes to a value of 0.7 between 0 and 1, representing the urgency of the order's wait time. The number of items in the order reflects the complexity and resource requirements of the order; the more items, the longer it takes to complete the picking. For example, an order containing 20 items has a normalized item quantity value of 0.8, indicating a high urgency in terms of size. The original order priority refers to the basic priority assigned when the order is created, usually determined by the order type and customer level. For example, the original priority of an expedited order is 5, the original priority of a regular order is 3, and the original priority of a low-priority order is 1. Normalizing the original priority to between 0 and 1, the normalized value for an expedited order is 1.0. The final urgency level of a picking task order is calculated by weighting these three factors with weights of 0.4, 0.25, and 0.35, respectively, ensuring that waiting time and original priority play a more significant role in the urgency assessment. Taking the order mentioned above as an example, its final urgency level is 0.7×0.4+0.8×0.25+1.0×0.35=0.83, indicating that this order has a high urgency level.

[0098] Constructing a virtual potential field in a 3D digital model is key to achieving intelligent obstacle avoidance for stacker cranes. A virtual potential field is a mathematical model used to describe the interactions between stacker cranes in space. The magnitude of the potential energy of the virtual potential field is determined based on the stacker crane's speed and the urgency of the picking order. The potential energy of the virtual potential field consists of two main components: a repulsive potential energy component and an attractive potential energy component. The repulsive potential energy component is proportional to the square of the stacker crane's speed, meaning that the faster the stacker crane moves, the greater the repulsive potential energy generated, and the easier it is to trigger obstacle avoidance behavior. The repulsive potential energy component is calculated by multiplying the square of the stacker crane's speed by an influence coefficient, set to 0.05. For example, if stacker crane A moves at a speed of 2.5 m / s, its repulsive potential energy component is 0.05 × (2.5 × 2.5) = 0.3125. The attractive potential energy component is proportional to the urgency of the picking order, meaning that the higher the urgency of the order, the higher the priority for the stacker crane to complete the task, and the greater the attractive potential energy generated. The attractive potential energy component is calculated by multiplying the order urgency by an influence coefficient, set to 0.08. For example, if stacker crane B executes an order with an urgency of 0.83, its attractive potential energy component will be 0.08 × 0.83 = 0.0664. The total potential energy of the stacker crane is the difference between the repulsive and attractive potential energy components. The higher the total potential energy, the greater the stacker crane's influence in the virtual potential field. Taking stacker crane A as an example, if its order urgency is 0.75, its total potential energy will be 0.3125 - 0.08 × 0.75 = 0.2525.

[0099] The magnitude of the repulsive force is related to the distance between the stacker cranes and their respective potential energies. The Coulomb force model is used to calculate the repulsive force, which is equal to the product of the potential energies of the two stacker cranes divided by the square of the distance between them, and then multiplied by a scaling factor. The scaling factor is set to 10 to adjust the magnitude of the repulsive force. For example, if the potential energies of stacker crane A and stacker crane B are 0.2525 and 0.3125 respectively, and the distance between them is 8 meters, then the repulsive force between them is 10 × 0.2525 × 0.3125 / (8 × 8) = 0.0123. A preset threshold of 0.01 is set. When the calculated repulsive force exceeds this threshold, the obstacle avoidance decision process is triggered. In practical applications, the threshold selection needs to balance safety and efficiency. A threshold that is too low will lead to frequent obstacle avoidance behaviors, while a threshold that is too high will fail to avoid collision risks in a timely manner. When the detected repulsive force between stacker crane A and stacker crane B is 0.0123, exceeding the preset threshold of 0.01, the urgency level of the picking orders performed by these two stacker cranes is obtained. Assuming that the urgency level of the order executed by stacker crane A is 0.75 and the urgency level of the order executed by stacker crane B is 0.83, the task priorities are determined based on the urgency level, with stacker crane B, which has a higher urgency level, receiving a higher task priority. Based on the task priority, an obstacle avoidance and yielding decision is triggered: stacker crane A, with lower priority, must yield, while stacker crane B, with higher priority, continues along its original path.

[0100] The specific implementation of obstacle avoidance and yielding decisions includes the selection of yielding strategies and path planning. Three yielding strategies are provided: pause and wait, decelerate and yield, and path replanning. The most suitable yielding strategy is selected based on the relative position, direction of movement, and speed difference between the stacker cranes. When two stacker cranes are moving in the same direction and their speed difference is less than 1 m / s, the decelerate yielding strategy is usually chosen. For example, if stacker crane A's speed is 2.5 m / s and stacker crane B's speed is 2.0 m / s, stacker crane A is instructed to decelerate to 1.5 m / s until the distance to stacker crane B increases to a safe range (greater than 15 meters). When two stacker cranes are moving in opposite directions or on intersecting paths, the pause and wait strategy or the path replanning strategy is selected. The pause and wait strategy is suitable for short-term encounter scenarios, such as stacker crane A pausing for 5 seconds before reaching the intersection to wait for stacker crane B to pass. The path replanning strategy is suitable for complex obstacle avoidance scenarios, planning a new movement path for stacker crane A to bypass the working area of ​​stacker crane B. During path replanning, the spatial constraints of the automated warehouse, the mobility of the stacker crane, and the task completion time requirements are considered to generate the optimal obstacle avoidance path. For example, stacker crane A, originally planned to move along the X-axis, is replanned to first rise two layers along the Y-axis, then move along the X-axis, and finally descend back to the target layer, completing the obstacle avoidance process. The stacker crane's operating status is monitored in real time. Once the stacker crane completes the obstacle avoidance maneuver, it automatically resumes normal operation and continues executing the original picking task. If a new collision risk is detected during the obstacle avoidance process, a new round of obstacle avoidance decisions will be immediately triggered to ensure the continuity and safety of the stacker crane's operations.

[0101] To improve the accuracy of obstacle avoidance decisions, a dynamic update mechanism for order urgency has been implemented. As time progresses, the urgency of waiting orders gradually increases, with the urgency of all orders updated every 30 seconds. For example, an order with an urgency of 0.75 will be updated to 0.77 after 30 seconds, ensuring that orders with long waiting times receive higher processing priority. Furthermore, an urgency cap control mechanism is in place to prevent some orders from monopolizing resources due to excessively long waiting times. The maximum urgency of any order is 0.95; exceeding this value will not increase it further, avoiding extreme decision-making. A manual intervention mechanism is also supported. In special circumstances, administrators can manually adjust the urgency of specific orders or directly specify obstacle avoidance strategies to meet specific operational needs.

[0102] This invention achieves an organic combination of physical characteristics and business requirements by quantifying the stacker crane's movement speed and order urgency into a potential field for calculation; the obstacle avoidance mechanism triggered by repulsion can prevent collision risks in a timely manner; the dynamic selection of multi-level obstacle avoidance strategies ensures adaptability in various scenarios; and the dynamic update mechanism of order urgency ensures the reasonable allocation of task priorities, significantly improving the safety and efficiency of collaborative operations of multiple stacker cranes in automated warehouses.

[0103] Optionally,

[0104] The steps for determining the potential energy of the virtual potential field based on the stacker crane's movement speed and the urgency of the picking task order include:

[0105] The velocity vector of the stacker crane is decomposed into radial velocity components and tangential velocity components. The radial velocity component represents the speed at which stacker cranes approach or move away from each other, and the tangential velocity component represents the speed at which stacker cranes cross each other and avoid each other. The motion state coefficient is determined based on the ratio of the radial velocity component to the tangential velocity component.

[0106] The urgency of the picking task orders is mapped to a spatial attraction coefficient, which increases with the urgency. The order density of local areas is calculated based on the distribution of order target points in the three-dimensional digital model, and the spatial attraction coefficient is dynamically weighted based on the order density.

[0107] An adaptive potential field weight is constructed, which is positively correlated with the motion state coefficient and negatively correlated with the spatial attraction coefficient; when the motion state coefficient is greater than a first threshold, the weight of repulsive potential energy is increased, and when the spatial attraction coefficient is greater than a second threshold, the weight of attractive potential energy is increased.

[0108] The ratio of the repulsive potential energy component and the attractive potential energy component is adjusted according to the adaptive potential field weight, and the potential energy of the virtual potential field is updated in real time. The repulsive potential energy component is used to avoid stacker crane collisions, and the attractive potential energy component is used to ensure picking efficiency.

[0109] For example, in determining the potential energy of the virtual potential field, it is necessary to accurately analyze the motion state of the stacker crane. The real-time velocity vector of the stacker crane is obtained, including its magnitude and direction. For any two stacker cranes in the automated warehouse, a local coordinate system is established with their connecting line as the reference. The velocity vector of the stacker crane is decomposed into radial and tangential velocity components. The radial velocity component represents the speed at which the stacker cranes approach or move away from each other, while the tangential velocity component represents the speed at which they cross paths. Taking two stacker cranes in the automated warehouse as an example, the velocity vector of stacker crane A is (1.5, 0.8, 0.3) m / s, and the velocity vector of stacker crane B is (-0.9, 1.2, 0.5) m / s. The direction of the line connecting them is (0.707, 0.707, 0). Through vector projection calculations, the radial velocity component of stacker crane A is 1.62 m / s, and the tangential velocity component is 0.95 m / s; the radial velocity component of stacker crane B is 0.21 m / s, and the tangential velocity component is 1.86 m / s. The motion state coefficient is determined based on the ratio of the radial to tangential velocity components. The motion state coefficient is calculated by comparing the absolute values ​​of the radial and tangential velocity components, then mapping it to a value between 0 and 1 using the sigmoid function. For example, the radial-tangential velocity ratio of stacker crane A is 1.62 / 0.95 = 1.71, corresponding to a motion state coefficient of 0.85; the ratio of the radial-tangential velocity ratio of stacker crane B is 0.21 / 1.86 = 0.11, corresponding to a motion state coefficient of 0.18. A larger motion state coefficient indicates that the stacker crane moves more radially, resulting in a higher collision risk; a smaller coefficient indicates that the stacker crane moves more tangentially, exhibiting better avoidance characteristics.

[0110] The urgency of picking orders is mapped to a spatial attraction coefficient using a piecewise linear function. For urgency levels between 0 and 0.3, the spatial attraction coefficient is 1.2 times the urgency level; for urgency levels between 0.3 and 0.7, it's 1.5 times; and for urgency levels between 0.7 and 1.0, it's 2.0 times. This piecewise mapping ensures that high-urgency orders receive stronger attraction. For example, an order with an urgency level of 0.83 has a spatial attraction coefficient of 0.83 × 2.0 = 1.66, indicating strong spatial attraction. To address the uneven distribution of orders in space, the order density of local areas is calculated based on the distribution of order target points in the 3D digital model. The automated warehouse space is divided into 5m × 5m × 3m 3D grid cells, and the number of orders within each grid is counted to calculate the order density. For example, if a grid cell has 8 order target points and the average number of orders in adjacent grids is 3, then the relative order density of that grid is 8 / 3 = 2.67. The spatial attraction coefficient is dynamically weighted based on order density. The weighting formula is the original spatial attraction coefficient multiplied by the square root of the order density, and then multiplied by an adjustment factor of 0.8. For the example above, the corrected spatial attraction coefficient is 2.17. This dynamic weighting mechanism ensures that the scheduling needs of order-dense areas can be handled reasonably, avoiding efficiency losses caused by local congestion.

[0111] The adaptive potential field weight is positively correlated with the motion state coefficient and negatively correlated with the spatial attraction coefficient. The basic calculation logic is to divide the motion state coefficient by the spatial attraction coefficient and then multiply by a baseline coefficient of 0.5 to obtain the initial adaptive potential field weight. For example, if the motion state coefficient of stacker crane A is 0.85 and the spatial attraction coefficient of the executed order is 2.17, its initial adaptive potential field weight is 0.85 / 2.17×0.5=0.196. To handle extreme cases, upper and lower limits for the adaptive weight are set at 0.9 and 0.1 respectively, ensuring stability in various scenarios. When the motion state coefficient exceeds a first threshold, the weight of the repulsive potential energy is increased. The first threshold is set to 0.75, indicating that the stacker crane mainly moves in the radial direction, with a higher risk of collision. When the motion state coefficient exceeds 0.75, the repulsive potential energy weight increases by twice the excess. For example, if the motion state coefficient of stacker crane A is 0.85, exceeding the threshold of 0.1, the repulsive potential energy weight increases by 0.1 × 2 = 0.2, and the corrected adaptive potential field weight becomes 0.196 × (1 + 0.2) = 0.235. When the spatial attraction coefficient exceeds the second threshold, the weight of the attractive potential energy increases. The second threshold is set to 1.5, indicating a high degree of urgency for the order. When the spatial attraction coefficient exceeds 1.5, the attractive potential energy weight increases by 1.5 times the excess. For example, if the spatial attraction coefficient of an order is 2.17, exceeding the threshold of 0.67, the attractive potential energy weight increases by 0.67 × 1.5 = 1.005, and the corresponding repulsive potential energy weight decreases to 0.235 / (1 + 1.005) = 0.117. This dual-threshold triggering mechanism ensures that the ratio of repulsive and attractive potential energy can be flexibly adjusted according to the actual situation, guaranteeing both safety and efficiency.

[0112] The repulsive potential energy component is calculated using a Gaussian function, which is directly proportional to the square of the stacker crane's speed and inversely proportional to the square of the distance between stacker cranes. The basic form of the repulsive potential energy component is the square of the speed multiplied by a base coefficient of 0.05, then multiplied by the adaptive potential field weight. For example, if stacker crane A's speed is 1.8 m / s and the adaptive potential field weight is 0.117, its repulsive potential energy component is 1.8 × 1.8 × 0.05 × 0.117 = 0.019. The attractive potential energy component is directly proportional to the urgency of the picking task order and inversely proportional to the distance to the target point. The basic form of the attractive potential energy component is the urgency of the order multiplied by a base coefficient of 0.08, then multiplied by (1 - adaptive potential field weight). For example, if the urgency of the order is 0.83 and the adaptive potential field weight is 0.117, its attractive potential energy component is 0.83 × 0.08 × (1 - 0.117) = 0.059. The total potential energy of the virtual potential field is a combination of repulsive and attractive potential energy components, using a potential energy difference model, meaning the total potential energy equals the attractive potential energy component minus the repulsive potential energy component. In the example above, the total potential energy is 0.059 - 0.019 = 0.04, indicating that in the current state, the attractive force is slightly stronger than the repulsive force. The stacker crane will continue moving towards the target point but will remain vigilant towards surrounding stacker cranes. The potential energy calculation of the virtual potential field is updated every 100 milliseconds to ensure that the potential field can reflect the stacker crane's movement status and changes in order urgency in real time.

[0113] This invention accurately describes the relative motion state between stacker cranes through radial and tangential decomposition of velocity vectors; it introduces a dynamic weighting mechanism for order density to solve the problem of uneven spatial distribution of orders; it adopts a dual-threshold adjustment strategy for adaptive potential field weights to achieve a dynamic balance between safety and efficiency; and it provides a clear basis for motion decision-making for stacker cranes through a difference model of repulsive and attractive potential energy, significantly improving the collaborative operation efficiency and safety of multiple stacker cranes in automated warehouses.

[0114] Optionally,

[0115] Based on the obstacle avoidance and yielding decision, the steps of calculating the obstacle avoidance trajectory of the stacker crane, generating a stacker crane collaborative scheduling instruction with no collision risk, and executing the picking operation according to the stacker crane collaborative scheduling instruction include:

[0116] The initial obstacle avoidance trajectory is generated by fifth-order polynomial interpolation, and the initial obstacle avoidance trajectory satisfies the constraints of the start and end point positions, velocity, and acceleration.

[0117] A spatiotemporal topology graph is constructed, wherein the nodes of the spatiotemporal topology graph represent the position and state of the stacker crane at different times, and the edges of the spatiotemporal topology graph represent feasible state transitions; the trajectory smoothness and the weighted sum of the target temporal node deviation are calculated as the trajectory cost based on the spatiotemporal topology graph, and the initial obstacle avoidance trajectory is optimized according to the trajectory cost;

[0118] The actual motion state of the stacker crane is collected, and the position deviation between the actual position of the stacker crane and the planned trajectory is calculated. When the position deviation is greater than the fault tolerance threshold, a local optimization interval is constructed based on the current state of the stacker crane and the target time node. The trajectory is replanned within the local optimization interval so that the replanned trajectory and the original trajectory smoothly transition at the boundary of the interval.

[0119] Calculate the time window for each stacker crane to reach the target timing node based on the spatiotemporal topology diagram, construct a speed adjustment function based on the time window, and adjust the movement speed of each stacker crane through the speed adjustment function so that multiple stacker cranes pass through the target timing node at staggered times.

[0120] Based on the replanned trajectory and the speed adjustment function, a stacker crane collaborative scheduling instruction is generated; after verifying that the stacker crane collaborative scheduling instruction meets the safety distance constraint, the picking operation is executed.

[0121] Combination Figure 2 This paper describes the process of planning and coordinating the obstacle avoidance trajectory of a stacker crane based on a spatiotemporal topology graph. After determining that the stacker crane needs to avoid obstacles, an initial obstacle avoidance trajectory is generated using fifth-order polynomial interpolation. Fifth-order polynomial interpolation can simultaneously satisfy the position, velocity, and acceleration constraints of the starting and ending points, ensuring the continuity and smoothness of the trajectory. First, the starting and ending state parameters of the obstacle avoidance trajectory are determined, including the three-dimensional coordinate position, velocity vector, and acceleration vector. For example, the starting state of stacker crane A is: position (25.3, 8.2, 12.5) meters, velocity (1.2, 0, 0) meters / second, acceleration (0, 0, 0) meters / second squared; the ending state is: position (38.6, 10.5, 12.5) meters, velocity (1.5, 0, 0) meters / second, acceleration (0, 0, 0) meters / second squared. The total trajectory duration is set to 15 seconds, and 150 points are uniformly sampled on the time axis. The corresponding position coordinates of each sampled point are calculated to form the initial obstacle avoidance trajectory. Fifth-order polynomial interpolation ensures that the trajectory meets preset constraints at the start and end points while maintaining a smooth transition throughout the trajectory. For complex 3D obstacle avoidance scenarios, fifth-order polynomial interpolation is performed separately in the X, Y, and Z directions, and then combined to form a complete 3D trajectory. To optimize the velocity and acceleration curves, the velocity and acceleration at the trajectory sampling points are checked for limits to ensure that the velocity does not exceed the stacker crane's maximum speed limit (usually 3 m / s) and the acceleration does not exceed the maximum acceleration limit (usually 1.2 m / s²). If any limits are exceeded, the polynomial interpolation parameters are adjusted, and a trajectory that meets the constraints is regenerated.

[0122] A spatiotemporal topology graph is a graph structure that unifies the representation of space and time. Its nodes represent the position and state of a stacker crane at different times, and edges represent feasible state transitions. First, a basic topology graph is constructed in three-dimensional space, discretizing the automated warehouse space into a grid of 0.5m × 0.5m × 0.5m. Each grid point is connected to its 26 adjacent grid points, forming a three-dimensional spatial connection. Then, the graph is expanded in the time dimension by adding a timestamp to each spatial grid point, with a time resolution of 0.1 seconds. For example, the node (25.5, 8.0, 12.5, 10.2) represents the stacker crane at time 10.2, located at spatial coordinates (25.5, 8.0, 12.5). Directed edges are established between nodes with adjacent timestamps if their spatial positions satisfy the stacker crane's motion constraints. For example, the edge from node (25.5, 8.0, 12.5, 10.2) to node (25.8, 8.0, 12.5, 10.3) is valid because moving 0.3 meters in 0.1 seconds meets the stacker crane's speed constraint. Based on the constructed spatiotemporal topology graph, the cost function of the trajectory is calculated. The cost function consists of two parts: a weighted sum of trajectory smoothness and target temporal node deviation. Trajectory smoothness is evaluated by calculating the change in acceleration between adjacent sampling points; the smaller the change, the smoother the trajectory. Target temporal nodes refer to the critical locations and time points that the stacker crane must pass through during obstacle avoidance, such as avoidance points and re-merging points into the original trajectory. Target temporal node deviation refers to the difference between the actual passage time and the planned time of the stacker crane. The smoothness weight is set to 0.6, and the temporal node deviation weight is set to 0.4, and the total trajectory cost is calculated. The A* algorithm is used to search for the path with the minimum cost in the spatiotemporal topology graph to optimize the initial obstacle avoidance trajectory. For example, the total cost of the initial trajectory was 28.5, which was reduced to 21.3 after optimization, and the trajectory smoothness was improved by about 25%.

[0123] In actual operation, the stacker crane cannot move precisely along the planned trajectory, requiring real-time tracking and correction. The actual motion state of the stacker crane, including position coordinates, velocity, and acceleration, is collected every 100 milliseconds using position sensors. The deviation between the actual position and the planned trajectory position is calculated using Euclidean distance. For example, if the actual position of the stacker crane at a certain moment is (26.8, 8.3, 12.6), and the corresponding position on the planned trajectory is (27.0, 8.2, 12.5), the calculated position deviation is 0.36 meters. A tolerance threshold of 0.3 meters is set; when the position deviation exceeds this threshold, the trajectory replanning process is triggered. Trajectory replanning does not recalculate the entire trajectory but constructs a local optimization interval based on the stacker crane's current state and the target time node. The starting point of the local optimization interval is the stacker crane's current state, and the ending point is the next target time node; the interval length is typically 3-5 seconds. Within the local optimization interval, fifth-order polynomial interpolation is reapplied to generate a new trajectory segment. To ensure trajectory continuity, a smooth transition mechanism was designed to maintain consistency in position, velocity, and acceleration between the replanned trajectory and the original trajectory at the interval boundaries. For example, if the stacker crane deviates from the original trajectory, with the current state as: position (26.8, 8.3, 12.6), velocity (1.3, 0.2, 0.1), and the next target timing node is (31.5, 9.0, 12.5), it should arrive in 4.2 seconds. Within this local interval, the trajectory is replanned to ensure the stacker crane can smoothly return to the planned trajectory within a specified time.

[0124] The time windows for each stacker crane to reach the target time node are calculated based on the spatiotemporal topology diagram. The time window consists of the earliest and latest arrival times, reflecting the scheduling flexibility of the stacker cranes. For example, the time window for stacker crane A to reach node (31.5, 9.0, 12.5) is [15.8, 17.2] seconds, meaning it can arrive as early as 15.8 seconds and no later than 17.2 seconds. When multiple stacker cranes need to pass through the same or adjacent spatial locations, the overlap of their time windows is analyzed. For stacker cranes with overlapping time windows, priority is ranked based on order urgency and current operating status. Higher-priority stacker cranes gain priority access to their time windows, while lower-priority stacker cranes need to adjust their transit times. Based on the time window analysis, a speed adjustment function is constructed to enable multiple stacker cranes to pass through the target time node at staggered times. The speed adjustment function is calculated individually for each stacker crane, dynamically generated based on key location points on its trajectory and time constraints. When a potential intersection between two stacker cranes is detected in time and space, the higher-priority stacker crane is prioritized to proceed as planned, while the lower-priority stacker crane's speed is adjusted. During the adjustment process, the maximum speed limit, maximum acceleration limit, and safe stopping distance of the stacker cranes are considered, and the operation is divided into acceleration, constant speed, and deceleration phases. For example, when stacker crane A and stacker crane B need to pass through the same intersection point sequentially, stacker crane A, with higher priority, is instructed to maintain its planned speed of 1.5 m / s while stacker crane B begins to decelerate 8 meters from the intersection point, decreasing from 1.8 m / s to 1.2 m / s at a rate of 0.3 m / s². After A passes, B delays for 1.5 seconds before passing the intersection point, and then gradually returns to its planned speed at an acceleration of 0.25 m / s². The delay duration is determined by calculating the minimum safe interval time, which depends on the stacker crane's physical dimensions, loading status, and safety margin coefficient. The adjusted speed curve minimizes energy consumption and avoids unnecessary acceleration and deceleration operations while ensuring safety. The speed adjustment calculation is updated every 200 milliseconds to ensure that it can respond to changes in the movement status of multiple stacker cranes in real time.

[0125] Based on the replanned trajectory and speed adjustment function, collaborative scheduling instructions for the stacker cranes are generated. These instructions consist of a triplet of timestamp, position coordinates, and speed value, and are sent to the stacker crane control at a frequency of 10 Hz. Before issuing the instructions, a final safety verification is performed to ensure that all stacker cranes consistently meet safety distance constraints during instruction execution. The safety distance constraint is defined as the minimum permissible distance between any two stacker cranes, which can be set to 3 meters. The complete process of all stacker cranes operating according to the scheduling instructions is simulated, calculating the minimum distance between any two stacker cranes at any given time to verify if it is below the safety threshold. If a potential safety risk is detected, the speed adjustment function is readjusted or the trajectory is locally replanned until the safety requirements are met. After successful verification, the collaborative scheduling instructions are sent to the control units of each stacker crane. The stacker cranes precisely execute motion control according to the instructions, completing obstacle avoidance and yielding, and continuing the picking operation. A real-time monitoring mechanism is also implemented to continuously track the actual operating status of the stacker cranes. When new deviations or potential risks are detected, a new round of trajectory optimization and scheduling updates is immediately triggered, forming a closed-loop control to ensure the continuity and safety of the picking operation.

[0126] This invention employs fifth-order polynomial interpolation to ensure smooth and continuous trajectories; the construction of a spatiotemporal topology map enables multi-machine collaborative problems to be solved within a unified framework; the trajectory replanning mechanism in local optimization intervals improves the real-time response capability to deviations; the speed adjustment strategy based on time windows enables efficient staggered passage of stacker cranes; and complete safety verification ensures the reliable execution of collaborative scheduling instructions, effectively solving path conflicts and scheduling congestion problems in multiple stacker cranes.

[0127] Secondly, it provides a digital twin-based automated warehouse location allocation and picking scheduling system, including:

[0128] The first unit is used to construct a three-dimensional digital model of an automated warehouse using digital twin technology;

[0129] The second unit is used to allocate storage locations in the three-dimensional digital model based on a deep reinforcement learning algorithm. It takes the frequency of goods access, the probability of related purchases, and the seasonal demand forecast as input parameters, and performs iterative calculations with the optimization objectives of minimizing the picking path and maximizing the centralized storage of related goods. Based on the iterative calculation results, it generates storage location allocation instructions.

[0130] The third unit is used to execute the goods receiving operation according to the receiving location allocation instruction and to receive picking task orders;

[0131] The fourth unit is used to construct a virtual potential field in the three-dimensional digital model. The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order. When the repulsive force between stacker cranes is detected to be greater than a preset threshold, the task priority is determined based on the urgency of the picking task order and an obstacle avoidance and yielding decision is triggered. According to the obstacle avoidance and yielding decision, the obstacle avoidance trajectory of the stacker crane is calculated, a stacker crane collaborative scheduling instruction without collision risk is generated, and the picking operation is executed according to the stacker crane collaborative scheduling instruction.

[0132] Thirdly, a computer-readable storage medium is provided, having stored thereon computer program instructions that, when executed by a processor, implement the aforementioned method.

Claims

1. A method for location allocation and picking scheduling in an automated warehouse based on digital twins, characterized in that, include: A three-dimensional digital model of the automated warehouse was constructed using digital twin technology; In the three-dimensional digital model, storage location allocation is performed based on a deep reinforcement learning algorithm. The frequency of goods access, probability of related purchases, and seasonal demand forecast results are used as input parameters. The optimization objectives are to minimize the picking path and maximize the centralized storage of related goods. Based on the iterative calculation results, storage location allocation instructions are generated. Execute the goods receiving operation according to the warehouse location allocation instruction, and receive picking task orders; A virtual potential field is constructed in the three-dimensional digital model. The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order. When the repulsive force between stacker cranes is detected to be greater than a preset threshold, the task priority is determined based on the urgency of the picking task order, and an obstacle avoidance and yielding decision is triggered. Specifically, this includes: calculating the urgency of the picking task order based on the order waiting time, the number of items in the order, and the original priority of the order; constructing a virtual potential field in the three-dimensional digital model, determining the potential energy of the virtual potential field based on the movement speed of the stacker crane and the urgency of the picking task order, decomposing the movement speed vector of the stacker crane into radial velocity components and tangential velocity components, and determining the motion state coefficient based on the ratio of the radial velocity component to the tangential velocity component; mapping the urgency of the picking task order to a spatial attraction coefficient, which increases with the increase of urgency; calculating the order density of a local area based on the distribution of the order target points in the three-dimensional digital model, and dynamically weighting the spatial attraction coefficient based on the order density; An adaptive potential field weight is constructed, which is positively correlated with the motion state coefficient and negatively correlated with the spatial attraction coefficient. When the motion state coefficient is greater than a first threshold, the weight of repulsive potential energy is increased; when the spatial attraction coefficient is greater than a second threshold, the weight of attractive potential energy is increased. The ratio of repulsive potential energy components to attractive potential energy components is adjusted according to the adaptive potential field weight, and the potential energy of the virtual potential field is updated in real time. The potential energy of the virtual potential field includes a repulsive potential energy component that is proportional to the square of the stacker crane's movement speed and an attractive potential energy component that is proportional to the urgency of the picking task order. The repulsive force between the stacker cranes is calculated based on the potential energy of the virtual potential field. When the repulsive force is greater than a preset threshold, the urgency of the picking task orders executed by the multiple stacker cranes that have encountered obstacle avoidance conflicts is obtained. The task priority is determined according to the magnitude of the urgency, and an obstacle avoidance and yielding decision is triggered based on the task priority. Based on the obstacle avoidance and yielding decision, the obstacle avoidance trajectory of the stacker crane is calculated, a stacker crane collaborative scheduling instruction with no collision risk is generated, and the picking operation is executed according to the stacker crane collaborative scheduling instruction.

2. The method according to claim 1, characterized in that, The steps for constructing a 3D digital model of an automated warehouse using digital twin technology include: A three-dimensional digital model of an automated warehouse, including rack units, stacker cranes, and conveyor lines, was constructed using digital twin technology. The collected data on warehouse occupancy and stacker crane location are written into a dual buffer, which uses a staggered read / write mechanism to prioritize the data. The collected data is converted into corresponding status information, and the status information is synchronously updated to the three-dimensional digital model according to the urgency priority based on the timestamp. The stacker crane position data in the three-dimensional digital model is compared with the actual collected stacker crane position data. The stacker crane position data is corrected using the Kalman filter algorithm, and the stacker crane movement trajectory is smoothed using B-spline curves.

3. The method according to claim 1, characterized in that, In the three-dimensional digital model, the allocation of storage locations based on a deep reinforcement learning algorithm, using the frequency of goods access, the probability of associated purchases, and seasonal demand forecasts as input parameters, and the iterative calculation steps with the optimization objectives of minimizing the picking path and maximizing the centralized storage of associated goods, include: The frequency of product access, probability of associated purchase, and seasonal demand prediction results are used as state features input into the state space of the deep reinforcement learning network. The probability of associated purchase is calculated based on historical order data, and the seasonal demand prediction results are obtained by combining the seasonality coefficient and residual term of historical sales data of the product. An evaluation network and a target network of the deep reinforcement learning network are constructed. Both the evaluation network and the target network use long short-term memory network units to form hidden layers. The target network updates its parameters through the evaluation network. The storage location allocation scheme is used as the action space of the deep reinforcement learning network. A reward function is constructed to evaluate the storage location allocation scheme. The reward function includes a picking path loss term calculated based on the frequency of goods access and the distance from the storage location to the inlet / outlet, and a storage concentration reward term calculated based on the probability of goods associated purchase and the allocation status of adjacent storage locations. A priority experience replay mechanism is used to store and filter state transition samples. Based on the evaluation network and the target network, the time-series difference error is calculated and the network parameters are updated to generate the final cargo location allocation scheme.

4. The method according to claim 3, characterized in that, The steps to construct the reward function include: The dynamic distance matrix from each storage location to the inlet / outlet is calculated based on the real-time location of the stacker crane. The product of the dynamic distance matrix and the frequency of goods access is used as the baseline loss value. The obstacle avoidance path length between storage locations is calculated based on the real-time location of the stacker crane. The weighted sum of the obstacle avoidance path length and the baseline loss value is used as the picking path loss item. Products with related relationships are divided into multiple association groups using a hierarchical clustering algorithm, and the association strength between products within the association group is calculated. Spatial proximity constraints between products are constructed based on the association strength. The degree of matching between the spatial proximity constraints and the allocation status of adjacent storage locations is used as a storage concentration reward. Extract the temporal relationship of product combinations from historical order data, calculate the product association probability under different temporal relationships, determine the deviation between the product association probability and the actual associated purchase probability as the weight adjustment coefficient, and dynamically update the weight coefficient of the storage concentration reward item according to the weight adjustment coefficient. A reward function is constructed by combining the picking path loss term with the storage concentration reward term. A reward decay factor is calculated based on the gradient of the reward function. The product of the reward decay factor and the temporal difference error is used as the priority weight of the state transition sample.

5. The method according to claim 1, characterized in that, Based on the obstacle avoidance and yielding decision, the steps of calculating the obstacle avoidance trajectory of the stacker crane, generating a stacker crane collaborative scheduling instruction with no collision risk, and executing the picking operation according to the stacker crane collaborative scheduling instruction include: The initial obstacle avoidance trajectory is generated by fifth-order polynomial interpolation, and the initial obstacle avoidance trajectory satisfies the constraints of the start and end point positions, velocity, and acceleration. A spatiotemporal topology graph is constructed, wherein the nodes of the spatiotemporal topology graph represent the position and state of the stacker crane at different times, and the edges of the spatiotemporal topology graph represent feasible state transitions; the trajectory smoothness and the weighted sum of the target temporal node deviation are calculated as the trajectory cost based on the spatiotemporal topology graph, and the initial obstacle avoidance trajectory is optimized according to the trajectory cost; The actual motion state of the stacker crane is collected, and the position deviation between the actual position of the stacker crane and the planned trajectory is calculated. When the position deviation is greater than the fault tolerance threshold, a local optimization interval is constructed based on the current state of the stacker crane and the target time sequence node, and the trajectory is replanned within the local optimization interval. Calculate the time window for each stacker crane to reach the target timing node based on the spatiotemporal topology diagram, construct a speed adjustment function based on the time window, and adjust the movement speed of each stacker crane through the speed adjustment function so that multiple stacker cranes pass through the target timing node at staggered times. Based on the replanned trajectory and the speed adjustment function, a stacker crane collaborative scheduling instruction is generated; after verifying that the stacker crane collaborative scheduling instruction meets the safety distance constraint, the picking operation is executed.

6. A digital twin-based automated warehouse location allocation and picking scheduling system, used to implement the method of any one of claims 1-5, characterized in that, include: The first unit is used to construct a three-dimensional digital model of an automated warehouse using digital twin technology; The second unit is used to allocate storage locations in the three-dimensional digital model based on a deep reinforcement learning algorithm. It takes the frequency of goods access, the probability of related purchases, and the seasonal demand forecast as input parameters, and performs iterative calculations with the optimization objectives of minimizing the picking path and maximizing the centralized storage of related goods. Based on the iterative calculation results, it generates storage location allocation instructions. The third unit is used to execute the goods receiving operation according to the receiving location allocation instruction and to receive picking task orders; The fourth unit is used to construct a virtual potential field in the three-dimensional digital model. The potential energy of the virtual potential field is related to the movement speed of the stacker crane and the urgency of the picking task order. When the repulsive force between stacker cranes is detected to be greater than a preset threshold, the task priority is determined based on the urgency of the picking task order and an obstacle avoidance and yielding decision is triggered. According to the obstacle avoidance and yielding decision, the obstacle avoidance trajectory of the stacker crane is calculated, a stacker crane collaborative scheduling instruction without collision risk is generated, and the picking operation is executed according to the stacker crane collaborative scheduling instruction.

7. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Intelligent warehouse logistics method and system

    CN114715581A

  • Unmanned road roller cluster minor radius curve path planning method based on artificial potential field technology

    CN120232423A