Dynamic order picking path optimization method and system based on deep reinforcement learning

By building a Markov decision process based on deep reinforcement learning and optimizing the picking path, the decision delay and path suboptimality problems of traditional picking methods in a dynamic order environment are solved, achieving efficient order picking and path planning.

CN120765155AActive Publication Date: 2025-10-10STATE GRID ZHEJIANG ELECTRIC POWER CO LTD

Patent Information

Application Number
CN202511278746.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-09-09
Publication Date
2025-10-10
Estimated Expiration
2045-09-09

AI Technical Summary

Technical Problem

Traditional order picking methods face a dynamically changing order environment and struggle to achieve optimal matching between order combinations and route selection. Furthermore, the system's responsiveness declines in scenarios with high order arrival rates, making it unable to effectively cope with emergencies. This leads to decision delays and suboptimal routes, especially in multi-block layouts and multi-device collaboration scenarios where route conflicts and task allocation issues arise.

Method used

A method based on deep reinforcement learning is adopted. By collecting warehouse channel layout and picker parameters for structured encoding, a Markov decision process is constructed. The deep reinforcement learning neural network model is used to optimize the picking path, and the picking path is adjusted in real time to minimize the average waiting time of orders.

Benefits of technology

It achieves real-time optimization of dynamic picking routes, improves warehouse order picking efficiency, reduces the average waiting time for orders, and enhances the response speed and intelligence level of the logistics system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120765155A_ABST
    Figure CN120765155A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of warehouse logistics management, in particular to a dynamic picking path optimization method and system based on deep reinforcement learning, and the method comprises the steps: obtaining a picker state vector based on warehouse structured coding data analysis; calculating the movement value density of orders in each channel according to the to-be-processed order queue data and the picker state vector; defining a state space according to the picker state vector and the to-be-processed order queue data; based on the state space and the discrete execution action space, modeling a warehouse order selection problem as a Markov decision process; solving a Markov decision process by using a deep reinforcement learning neural network model based on the real-time state and the moving value density of the picker to obtain an optimal order picking path decision; and controlling a picker to execute picking operation according to the optimal picking path decision. According to the method, the Markov decision process is solved in combination with the deep reinforcement learning neural network model, accurate optimization of the order picking path is realized, and the warehouse order picking efficiency is effectively improved.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of warehouse logistics management, and particularly relates to a dynamic picking path optimization method and system based on deep reinforcement learning. BACKGROUND

[0002] In modern warehouse operation, order picking as the core link of warehouse logistics operation, its efficiency directly affects the overall supply chain performance, however, in the time-sensitive scenarios such as power material warehouse, the efficient processing of dynamic orders has become a technical problem, and the traditional order picking technology has exposed significant limitations in dealing with real-time changing order flow.

[0003] The traditional order picking method is mostly based on the assumption of static order set, and uses Traveling Salesman Problem (TSP) based on fixed order set or heuristic algorithm for route optimization, and does not consider the dynamic arrival characteristics of orders when modeling, which makes it difficult to effectively respond to the dynamic changing order environment. When facing the continuous arrival of new orders, repeated full calculation must be carried out, resulting in decision delay and suboptimal path problem. Especially in the high order arrival rate scene, the system response ability decreases sharply, with the development of warehouse automation trend, although the existing dynamic picking research tries to introduce order batching strategy, most of the researches separate the order batching and path planning by using single-stage optimization strategy, which leads to the lack of organic cooperation between the two key decision-making links, neither can the optimal matching of order combination and path selection be realized, nor can the real-time adjustment demand in dynamic scenarios be met. Some researches try to optimize combination through iterative search, but the problem of exponential growth of calculation complexity exists, which makes it difficult to meet the time constraints of real-time decision. In addition, the multi-objective balance mechanism in dynamic environment has not been effectively established, and the existing methods generally use fixed weight strategy to balance the driving distance and order timeliness and other target parameters, which lacks adaptability to environmental changes. Especially in the emergency order insertion, equipment abnormality and other sudden scenarios, the traditional method is difficult to realize the dynamic weight adjustment of target parameters, resulting in insufficient system robustness.

[0004] At the same time, the existing algorithm is mostly based on the simplified model design of single-zone warehouse, and when dealing with practical complex scenarios such as multi-zone layout and multi-device cooperation, there is generally a model mismatch problem. Especially when the path planning of multiple picking devices needs to be coordinated, the traditional method is difficult to effectively solve the path conflict and task allocation problem between devices, which is easy to produce operation bottleneck. Therefore, developing a dynamic adaptive picking path optimization method has become a technical bottleneck that needs to be broken through in the field of warehouse automation. SUMMARY

[0005] To solve the above technical problems, the present application provides a dynamic picking path optimization method and system based on deep reinforcement learning.

[0006] In a first aspect, the present application provides a dynamic picking path optimization method based on deep reinforcement learning, comprising the following steps: Collecting warehouse channel layout parameters and picker physical parameters to form warehouse infrastructure data, and structurally encoding the warehouse infrastructure data to obtain warehouse structured encoding data; Based on the warehouse structured encoding data, the current position of the picker and the remaining load capacity of the picker are analyzed to obtain a picker state vector; Real-time collection of warehouse order queue data, and calculation of the moving value density of orders in each channel according to the order queue data and the picker state vector; Defining a state space according to the picker state vector and the order queue data, and constructing a discrete execution action space for the picker at each time step; Based on the state space and the discrete execution action space, setting the cumulative reward as the optimization goal to minimize the average waiting time of orders, and modeling the warehouse order picking problem as a Markov decision process; Based on the real-time state of the picker and the moving value density, a pre-constructed deep reinforcement learning neural network model is used to solve the Markov decision process to obtain an optimal picking path decision; According to the optimal picking path decision, the picker performs a picking operation.

[0007] In further embodiments, the step of structurally encoding the warehouse infrastructure data to obtain warehouse structured encoding data comprises: According to the pre-collected warehouse physical structure data, a warehouse discretization coordinate system is established with the center of the entrance cross channel as the origin of the coordinate system; According to the warehouse channel layout parameters, all channels in the warehouse and the storage compartments on both sides of each channel are encoded to obtain channel layout parameter encoding data; Mapping the channel layout parameter encoding data to the warehouse discretization coordinate system to obtain warehouse channel storage encoding data; According to the picker physical parameters, the initial position of the picker is fixed to the entrance cross channel coordinate, and the initial unloaded capacity is set to the maximum load capacity; According to the initial position of the picker and the maximum load capacity, an initial state vector of the picker is obtained; The warehouse channel storage encoding data and the initial state vector of the picker are packaged as a structured data packet to form the warehouse structured encoding data.

[0008] In a further embodiment, the step of parsing the current position of the picker and the remaining cargo capacity of the picker based on the warehouse structured coding data to obtain the picker state vector includes: Acquire the real-time position data of the picker, match the real-time position data of the picker with the structured coding data of the warehouse, parse the channel number and storage cell number of the current location of the picker, and obtain the current position of the picker; Obtaining current load data of the picker, and determining the remaining load capacity of the picker based on the maximum load capacity of the picker in the warehouse structured coding data and the current load data; The current position of the picker and the remaining cargo capacity of the picker are combined into a picker state vector.

[0009] In a further embodiment, the pending order queue data includes the coordinates of item storage locations of pending orders in a warehouse; The movement value density is the sum of the ratios of the order quantity of each picking position in each channel and the minimum relative distance between the current position of the picker and the coordinates of the item storage position when the picker moves up or down.

[0010] In a further embodiment, the step of setting a cumulative reward with minimizing the average waiting time of orders as the optimization objective based on the state space and the discrete execution action space, and modeling the warehouse order picking problem as a Markov decision process includes: Determining a state transition rule for the picker to transition from a current state to a next state based on the state space and the discrete execution action space; Based on the state transition rules, the state transition probability of the picker transferring from the current state to the next state after executing the action is calculated, and a state transition probability model is constructed; With minimizing the average waiting time for orders as the optimization goal, a composite cumulative reward function is constructed based on the picker's action execution results, which includes a picking efficiency reward, a travel distance penalty, and an unloading incentive. The picker's action execution results include at least the time required for the picker to complete the picking order, the travel distance, and the order completion rate. Based on the state space, the discrete execution action space and the state transition probability model, the warehouse order picking problem is modeled as a Markov decision process by maximizing the optimization objective of a composite cumulative reward function.

[0011] In a further embodiment, the state transition rule is specifically: Set the minimum time unit required for the picker to move from one storage location to another, and detect the current position of the picker and perform actions; When it is detected that the picker is at a non-warehouse base location and performs a stationary action, the picker is controlled to remain stationary at the current location for the minimum time unit, and when a new order is detected to arrive during the period when the picker remains stationary, the state is triggered to transition to the next state; When it is detected that the picker is at the warehouse base position and performs the unloading action, the picker is controlled to unload all the items it carries and the remaining loading capacity of the picker is reset to the maximum loading capacity; When the picker is detected to be performing a movement action at the warehouse base position, if the picker is in the rear cross aisle or front cross aisle, the left and right movement actions are allowed; During the picker's movement, the order status changes and picker status changes are monitored in real time. If any of the following conditions are detected: the picker reaches the cross channel and cannot continue to move in the original direction of movement when performing the movement, a new order arrival condition is detected during the picker's movement, and at least one item is picked up during the picker's movement, the state is triggered to transfer to the next state.

[0012] In a further embodiment, the compound cumulative reward function is specifically: Where, is the compound cumulative reward at time step t; is the picking efficiency reward coefficient; The maximum load capacity that a picker can carry in a single pick; is the remaining cargo capacity of the picker at time step t; is the unloading incentive adjustment coefficient; is the number of storage locations moved by the picker in time step t; is the number of items picked up by the picker in time step t; is the action performed by the picker at time step t, Indicates that the picker performs unloading action.

[0013] In a further embodiment, the deep reinforcement learning neural network model includes an input layer, a two-branch feature extraction module, a feature fusion module, a multi-layer feature compression module and an output layer connected in sequence; The dual-branch feature extraction module includes an order status feature extraction unit and a picker status feature extraction unit connected in parallel; both the order status feature extraction unit and the picker status feature extraction unit are composed of a fully connected layer and a linear rectifier activation function; The multi-layer feature compression module includes at least three levels of fully connected layers connected in series.

[0014] In a further embodiment, the step of solving a Markov decision process by using a pre-constructed deep reinforcement learning neural network model based on the picker real-time state and the mobile value density to obtain an optimal picking path decision comprises: combining the mobile value density of each channel with the order real-time state into an order state vector; the dimension of the order state vector is twice the total number of channels; performing nonlinear transformation on the picker real-time state by the picker state feature extraction unit to extract a picker high-dimensional feature vector; performing nonlinear transformation on the order state vector by the order state feature extraction unit to extract an order distribution feature vector; concatenating the picker high-dimensional feature vector and the order distribution feature vector to generate a state fusion feature vector; performing hierarchical feature conversion on the state fusion feature vector by a multi-layer feature compression module to obtain a decision-oriented high-order feature representation; mapping the decision-oriented high-order feature representation to the action space dimension, and calculating the action expected cumulative reward value of each action taken under the current state by a fully connected layer; selecting the action with the highest action expected cumulative reward value as the optimal action of the picker under the current state, and generating an optimal picking path decision according to the optimal action.

[0015] In a second aspect, the present application provides a dynamic picking path optimization system based on deep reinforcement learning, which comprises: a data acquisition module for collecting warehouse channel layout parameters and picker physical parameters, forming warehouse basic structure data, and structurally encoding the warehouse basic structure data to obtain warehouse structured encoding data; a state analysis module for analyzing the current position of the picker and the remaining cargo capacity of the picker based on the warehouse structured encoding data to obtain a picker state vector; a mobile analysis module for collecting real-time warehouse order queue data, and calculating the mobile value density of orders in each channel according to the order queue data and the picker state vector; a space construction module for defining a state space according to the picker state vector and the order queue data, and constructing a discrete execution action space of the picker at each time step; a picking modeling module for modeling the warehouse order picking problem as a Markov decision process based on the state space and the discrete execution action space, and setting the cumulative reward with the optimization goal of minimizing the average waiting time of orders; A path decision module is used to solve the Markov decision process based on the real-time status of the picker and the movement value density using a pre-built deep reinforcement learning neural network model to obtain the optimal picking path decision; The picking execution module is used to control the picker to perform the picking operation according to the optimal picking path decision.

[0016] The present invention provides a dynamic picking path optimization method and system based on deep reinforcement learning. The method collects warehouse channel layout parameters and picker physical parameters to form warehouse infrastructure data, and performs structured encoding on the warehouse infrastructure data to obtain warehouse structured encoding data; analyzes the current position of the picker and the remaining cargo capacity of the picker based on the warehouse structured encoding data to obtain the picker state vector; collects the warehouse's pending order queue data in real time, and calculates the mobile value density of orders in each channel based on the pending order queue data and the picker state vector; defines a state space based on the picker state vector and the pending order queue data, and constructs a discrete execution action space for the picker at each time step; based on the state space and the discrete execution action space, sets a cumulative reward with minimizing the average waiting time of orders as the optimization goal, and models the warehouse order picking problem as a Markov decision process; based on the real-time state of the picker and the mobile value density, uses a pre-constructed deep reinforcement learning neural network model to solve the Markov decision process to obtain the optimal picking path decision; and controls the picker to perform the picking operation based on the optimal picking path decision. Compared with the existing technology, this method collects a variety of warehouse data and structures them into codes, and dynamically adjusts the picking path through deep reinforcement learning, thereby realizing real-time optimization of the dynamic picking path, effectively improving the efficiency of warehouse order picking, reducing the average waiting time of orders, and enhancing the response speed and intelligence level of the warehouse logistics system. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 This is a flow chart of a dynamic picking path optimization method based on deep reinforcement learning provided by an embodiment of the present invention; Figure 2 This is an example diagram of a rectangular warehouse layout provided by an embodiment of the present invention; Figure 3 This is a schematic diagram of the deep reinforcement learning neural network model architecture provided by an embodiment of the present invention; Figure 4 This is a block diagram of a dynamic picking path optimization system based on deep reinforcement learning provided by an embodiment of the present invention.

[0018] Explanation of the accompanying drawings: 101, data acquisition module; 102, state analysis module; 103, movement analysis module; 104, space construction module; 105, picking modeling module; 106, path decision module; 107, picking execution module. DETAILED DESCRIPTION

[0019] The following describes the embodiments of the present invention in detail with reference to the accompanying drawings. The embodiments are provided for illustrative purposes only and are not to be construed as limiting the present invention. The accompanying drawings are provided for reference and illustration only and do not constitute a limitation on the scope of protection of the present invention. Many changes may be made to the present invention without departing from the spirit and scope of the present invention.

[0020] Figure 1 This is a flow chart of a dynamic picking path optimization method based on deep reinforcement learning provided by an embodiment of the present invention. The embodiment of the present invention provides a dynamic picking path optimization method based on deep reinforcement learning, such as Figure 1 As shown, the method includes the following steps: S1. Collect warehouse channel layout parameters and picker physical parameters to form warehouse infrastructure data, and perform structured coding on the warehouse infrastructure data to obtain warehouse structured coding data.

[0021] In some embodiments, the step of performing structured encoding on the warehouse infrastructure data to obtain warehouse structured encoded data includes: Based on the pre-collected warehouse physical structure data, a warehouse discretized coordinate system is established with the center of the entrance cross channel as the origin of the coordinate system; Encode all aisles in the warehouse and the storage cells on both sides of each aisle according to the warehouse aisle layout parameters to obtain aisle layout parameter encoding data; Mapping the channel layout parameter encoding data to the warehouse discretized coordinate system to obtain warehouse channel storage encoding data; Fixing the initial position of the picker to the entrance cross-channel coordinates according to the physical parameters of the picker, and setting the initial unloaded capacity to the maximum loaded capacity; Obtaining an initial state vector of the picker according to the initial position of the picker and the maximum cargo capacity; The warehouse channel storage coding data and the picker initial state vector are encapsulated into a structured data packet to form warehouse structured coding data.

[0022] The embodiment mainly aims at the path planning problem of a single picker when performing an order picking task in a rectangular warehouse, which is composed of a single block, has multiple parallel channels inside, and the parallel channels are connected with each other through two cross channels. The inventory storage locations are distributed on both sides of each channel. The picker needs to collect the requested items from the specified storage locations. When the picker moves in the channel, the picker can access the storage locations on both sides of the channel. To simplify the problem, the embodiment assumes that the horizontal movement distance in the channel can be ignored. When the picker needs to move between different channels, the picker must use one of the two cross channels. These cross channels do not set storage locations and only serve as paths for transferring between channels. It should be noted that the picker can only enter or leave a channel through one of the cross channels. In addition, it is assumed that the size and priority of all to-be-picked items are the same. The picker starts from the starting point of the warehouse to collect the assigned items and complies with the capacity limit during transportation, and finally delivers the items back to the warehouse. Figure 2 A rectangular warehouse layout example diagram is shown, which provides two alternative paths that the picker can choose to collect the next assigned order. The path distance is represented by The main goal of the embodiment is to minimize the average waiting time of the order and thus improve the throughput of the warehouse operation.

[0023] The embodiment is applicable to warehouse logistics centers with rectangular layout, especially to small and medium-sized warehouses or automated warehouse systems that use single picker operation mode. Specifically, typical application scenarios include e-commerce warehouse order picking scenarios, retail distribution center replenishment scenarios, and medical or precision parts warehouse scenarios. For e-commerce warehouse order picking scenarios, the picker needs to efficiently plan the path to complete multiple order picking in batches under capacity constraints to minimize order waiting time. For retail distribution center replenishment scenarios, the picker needs to pick goods from the storage area to the sorting area to support order picking modes such as goods-to-person or person-to-goods, thereby optimizing the path and efficiency of a single picking task. For medical or precision parts warehouse scenarios, the precision requirement for item storage locations is extremely high in the medical or precision parts warehouse field. Structured modeling is used to ensure that the picker strictly follows the channel rules for movement and avoids the generation of invalid paths, thereby ensuring the accuracy and efficiency of warehouse operations.

[0024] In the process of warehouse structured coding, this embodiment first collects data to obtain the basic structural data of the warehouse. The basic structural data of the warehouse includes warehouse channel layout parameters and picker physical parameters. Specifically, the warehouse channel layout parameters include the number of parallel picking channels and the number of storage cells on both sides of each channel. These parameters together define the physical layout of the warehouse; the picker physical parameters include its initial position, which is usually set to the base coordinates (0, 0) that coincide with the entrance cross channel, and the maximum capacity, that is, the maximum number of items that the picker can carry in a single picking. By collecting these parameters, this embodiment can construct the basic structural data of the warehouse. After completing the data collection, this embodiment performs structured coding on the basic structural data of the warehouse. In this process, this embodiment first establishes discrete coordinate axes in the horizontal and vertical directions based on the physical structure data of the warehouse, with the center of the entrance cross channel as the origin of the discrete coordinate system according to the pre-set discretization accuracy, thereby establishing a warehouse discretized coordinate system, wherein the physical structure data of the warehouse can be Including but not limited to basic information such as the size and position of each part of the warehouse (such as channels, storage cells, etc.), then according to the warehouse channel layout parameters, all channels in the warehouse and the storage cells on both sides of each channel are encoded to obtain channel layout parameter encoding data, and these channel layout parameter encoding data are mapped to the warehouse discretized coordinate system according to the actual position of the channels and storage cells corresponding to each encoding in the warehouse, thereby generating warehouse channel storage encoding data. For example, this embodiment can determine the coordinate range of the channel in the coordinate system according to the direction and length of the channel; for the storage cells on both sides of the channel, their specific coordinate positions in the coordinate system are determined according to their relative positions with the channel and their order on the channel; for a certain horizontal channel, its coordinate range is from the starting coordinate to the ending coordinate in the horizontal direction, and is a fixed value in the vertical direction; for a certain storage cell on the channel, its specific coordinate offset in the horizontal direction is determined according to its encoding order, thereby obtaining its accurate coordinates in the coordinate system, thereby determining the position information of each channel and storage cell in the warehouse discretized coordinate system.

[0025] Meanwhile, according to the physical parameters of the sorter, the initial position of the sorter is fixed as the coordinate of the entrance intersection passage, that is, the initial coordinate position of the sorter in the warehouse discretization coordinate system is the coordinate corresponding to the center of the entrance intersection passage, and the initial unloaded cargo capacity of the sorter is set as the maximum loaded cargo capacity. In this embodiment, the initial position coordinate (the coordinate values in the horizontal and vertical directions) of the sorter and the initial loaded cargo capacity value are combined into a vector, and then the initial state vector of the sorter is obtained. Finally, the warehouse passage storage encoding data and the initial state vector of the sorter are packaged, the warehouse passage storage encoding data is taken as a data set which contains the position information of each passage and storage grid in the warehouse discretization coordinate system, and the initial state vector of the sorter is taken as a separate data element. The two parts are combined into a structured data packet, which is organized and stored according to a predefined data structure format, for example, the warehouse passage storage encoding data and the initial state vector of the sorter are packaged in a specific data format (such as JSON or XML), to form warehouse structured encoding data, so as to ensure the integrity and readability of the data. In the data logic processing stage, the warehouse passage and storage position are mapped to the discrete coordinates of the state space, for example, the left i-th storage grid of passage n corresponds to the vertical coordinate (2n-1), and the right storage grid corresponds to 2n, while the horizontal coordinate is used to distinguish the order picking passage from the intersection passage. The order storage position is represented as a two-tuple (passage number, vertical index), which is convenient for calculating the relative distance between the sorter and the target (for example, when moving upwards, the distance is the difference between the current vertical index and the target index). In this embodiment, the physical structure of the warehouse and the state information of the sorter can be represented in a unified and standardized form through structured encoding, which provides strong support for the intelligentization and automation of the warehouse management system.

[0026] S2. Analyzing the current position of the sorter and the remaining loaded cargo capacity of the sorter based on the warehouse structured encoding data to obtain a sorter state vector.

[0027] In some embodiments, the step of analyzing the current position of the sorter and the remaining loaded cargo capacity of the sorter based on the warehouse structured encoding data to obtain a sorter state vector comprises: obtaining real-time position data of the sorter, matching the real-time position data of the sorter with the warehouse structured encoding data, analyzing the passage number and storage grid number where the sorter is currently located to obtain the current position of the sorter; obtaining current load weight data of the sorter, and determining the remaining loaded cargo capacity of the sorter according to the maximum loaded cargo capacity of the sorter in the warehouse structured encoding data and the current load weight data; combining the current position of the sorter and the remaining loaded cargo capacity of the sorter into a sorter state vector.

[0028] Specifically, this embodiment collects the position information of the picker in the warehouse in real time through a positioning system pre-deployed in the warehouse (such as a positioning device based on ultra-wideband technology or radio frequency identification technology, etc.), and obtains the real-time position data of the picker. This data can accurately reflect the specific position coordinates of the picker in the warehouse space, and then matches the collected real-time position data of the picker with the pre-obtained warehouse structured coding data. The warehouse structured coding data contains the position information of all channels in the warehouse and the storage cells on both sides of each channel in the warehouse discrete coordinate system, and the channels and storage cells are uniquely coded. In the matching process, the real-time position coordinates of the picker are matched. Compare with the coordinate range of each channel and storage cell in the warehouse structured coding data. When the real-time position coordinates of the picker fall within the coordinate range of a certain channel, determine that the channel is the channel where the picker is currently located, and obtain the corresponding channel number, so as to further judge the relative position relationship between the picker position and the storage cells on both sides of the channel. If the picker position is close to a storage cell on one side and is within the coordinate range of the storage cell, determine that the storage cell is the storage cell where the picker is currently located, and obtain the corresponding storage cell serial number. In this way, the channel number and storage cell serial number of the picker's current location can be obtained, and the channel number and storage cell serial number of the picker's current location are used as the picker's current position.

[0029] At the same time, this embodiment uses a weighing sensor installed on the picker to collect the weight information of the goods currently carried by the picker in real time, obtain the current load data of the picker, and extract the maximum cargo capacity information of the picker from the warehouse structured coding data. The maximum cargo capacity of the picker is subtracted from the current load data, and the difference obtained is the remaining cargo capacity of the picker. This embodiment combines the obtained current position of the picker (including the channel number and storage cell number) and the remaining cargo capacity of the picker in a certain order. For example, this embodiment can first arrange the channel number, storage cell number, and remaining cargo capacity in sequence to form an ordered data combination. This data combination is the picker state vector, which comprehensively reflects the current status information of the picker in the warehouse.

[0030] S3. Collect the warehouse's pending order queue data in real time, and calculate the movement value density of orders in each channel based on the pending order queue data and the picker state vector.

[0031] Specifically, during the warehouse order processing process, this embodiment collects pending order data in real time and sorts the pending order data according to the order generation time to obtain pending order queue data, which contains the storage location coordinates of the items of the pending orders in the warehouse. At the same time, this embodiment combines the current status of the picker (including position and remaining capacity) and dynamically calculates the mobile value density of orders in each channel based on the sorting of the pending order queue data by order generation time. and , where n represents the nth channel, , N represents the total number of picking channels in the warehouse. In this embodiment, each channel n is composed of and To describe its order status, so as to dynamically update the status value of the order in each channel. The order status value can be expressed as , where, specifically, represents the movement value density of the order when the picker moves upward in the first channel at time t, represents the movement value density of the order when the picker moves downward in the second channel at time t, Represents the picker in ( ) The moving value density of the order when it moves upward at time t in the channel, This means that the picker is in ( ) The movement value density of the order when moving downward at time t in the channel. These values ​​are calculated by dividing the value of each order in the channel by the distance of the order from the picker based on the moving direction, and then summing up all orders. It should be noted that the distance between the picker and the order refers to the minimum number of moves required for the picker to reach the location of the order, in units of the number of storage locations. This calculation method comprehensively considers the distribution density of orders and the movement cost of the picker, provides a quantitative basis for path optimization, and can more accurately simulate decision-making in the picking process, ensuring the efficiency and real-time performance of the picking strategy. In this embodiment, for each picking position i in channel n, the number of orders at that position is counted. , and calculate its minimum relative distance to the current position of the picker or ,final, and They represent the weighted sum of all orders in the corresponding moving direction of the picker. The moving value density is defined as the sum of the ratios of the number of orders at each picking location in each channel and the minimum relative distance between the current position of the picker and the coordinates of the item storage location when the picker moves up or down. The specific calculation formula is: Where, For the picker in ( ) The moving value density of orders when moving upward within the channel at time t; is the number of pending orders at the i-th picking location in the n-th aisle; is the minimum relative distance (in terms of storage locations) that the picker needs to reach the picking position i when moving upwards; L is the number of picking locations in each aisle; For the picker in ( ) The moving value density of orders when moving downward within the channel at time t; is the minimum relative distance for the picker to reach the picking position i when moving downward.

[0032] S4. Define a state space based on the picker state vector and the pending order queue data, and construct a discrete execution action space of the picker at each time step.

[0033] S5. Based on the state space and the discrete execution action space, a cumulative reward is set with minimizing the average waiting time of orders as the optimization goal, and the warehouse order picking problem is modeled as a Markov decision process.

[0034] In some embodiments, based on the state space and the discrete execution action space, the steps of modeling the warehouse order picking problem as a Markov decision process with minimizing the average order waiting time as the optimization objective include: Based on the state space and discrete execution action space, determine the state transition rule for the picker to transfer from the current state to the next state; Based on the state transition rules, the state transition probability of the picker transferring from the current state to the next state after executing the action is calculated, and a state transition probability model is constructed; With minimizing the average waiting time for orders as the optimization goal, a composite cumulative reward function is constructed based on the picker's action execution results, which includes a picking efficiency reward, a travel distance penalty, and an unloading incentive. The picker's action execution results include at least the time required for the picker to complete the picking order, the travel distance, and the order completion rate. Based on the state space, the discrete execution action space and the state transition probability model, the warehouse order picking problem is modeled as a Markov decision process by maximizing the optimization objective of a composite cumulative reward function.

[0035] Specifically, this embodiment proposes a modeling method based on a finite Markov decision process (MDP) for the real-time warehouse order picking problem. The real-time warehouse order picking problem is modeled as a finite Markov decision process, fully utilizing its dynamic and sequential decision-making characteristics. The structured description of the problem is as follows: State space S: the state at time step t This includes the orders that currently need to be picked, the location of these orders in the warehouse, and the current location and status of the picker; Discrete execution action space A: the action selected at each time step t This can include moving in a specific direction (such as left, right, up, or down), remaining stationary, and releasing items at the base; State transition probability model: executing actions Will lead to a new state The generation of To status The transition is determined by the action taken and the current state; Immediate reward function R: Immediate reward function Assigning numerical values ​​based on the effectiveness of actions taken. Rewards can be based on factors such as the time required to complete a picking order, the distance traveled, and the success of order completion. Strategy obtained by the algorithm :Strategy A strategy for selecting actions based on the current state is defined, and the optimization goal is to find an optimal strategy that maximizes the cumulative reward.

[0036] in, represents the system state at time step t; represents the action taken at time step t.

[0037] The goal of the algorithm optimization is to determine the optimal strategy to maximize the expected cumulative reward within a limited time range. Once a strategy is selected, the warehousing system will evolve in discrete time steps. The specific process of each step is: observe the current state ; Strategy-based Select Action ; Execute the action to generate a new state Received reward In this embodiment, the Markov decision process is used to mathematically model the state transition process of the warehouse system. The Markov decision process provides a structured framework for modeling decision processes that are affected by both random factors and the actions of decision makers. The Markov decision process consists of a five-tuple (S, A, P, R, ) modeling, where S represents the system state set; A represents a finite set of actions; P represents the state transition probability; R represents the immediate reward function; represents the discount factor used to balance current rewards and future rewards. The following is the specific Markov decision implementation process for the problem: In this embodiment, the system state of time step t is expressed as ,in, is the state of the picker; The status of existing orders in the warehouse. This embodiment further decomposes the status into: Picker Status Contains four components 、 、 and , specifically: Indicates the horizontal position of the picker, that is, the position of the picker relative to the cross channel. Its value can be -1, 0 or 1, where -1 means the picker is in the rear cross channel, 1 means the picker is in the front cross channel, and 0 means the picker is in a picking channel.

[0038] and Both represent the vertical position, that is, the position of the picker relative to the picking channel. When the picker is located in the nth channel, and The two values ​​are (2n-1) and (2n), and each picking channel is represented by two consecutive indices (rather than a single index) because Defined as a vector of size 2N, this maintains consistency between the two components of the state and facilitates the interpretation of the neural network.

[0039] Indicates the remaining loading capacity of the picker. If , it means that the picker is full and can no longer pick up additional items; if It means that the picker is empty, and K is the maximum loading capacity of the picker.

[0040] when The exact position of the picker within a picking lane is implicitly encoded in the order status when In this embodiment, this state representation method ensures consistency between various parts of the system state, which is conducive to the use of neural network processing. Through this state definition and transition process modeling, this embodiment can effectively simulate decision-making in the picking process, thereby optimizing the picking strategy. The number of picking locations and the number of orders at each location jointly affect the formulation of the picking strategy, and the distance moved up or down to a certain picking location is directly related to the movement cost required for the picker to approach the order. It can more accurately evaluate the expected effects of different actions, thereby optimizing the decision-making in the picking process. In some embodiments, the state transition rule is specifically as follows: Set the minimum time unit required for the picker to move from one storage location to another, and detect the current position of the picker and perform actions; When it is detected that the picker is at a non-warehouse base location and performs a stationary action, the picker is controlled to remain stationary at the current location for the minimum time unit, and when a new order is detected to arrive during the period when the picker remains stationary, the state is triggered to transition to the next state; When it is detected that the picker is at the warehouse base position and performs the unloading action, the picker is controlled to unload all the items it carries and the remaining loading capacity of the picker is reset to the maximum loading capacity; When the picker is detected to be performing a movement action at the warehouse base position, if the picker is in the rear cross aisle or front cross aisle, the left and right movement actions are allowed; During the picker's movement, the order status changes and picker status changes are monitored in real time. If any of the following conditions are detected: the picker reaches the cross channel and cannot continue to move in the original direction of movement when performing the movement, a new order arrival condition is detected during the picker's movement, and at least one item is picked up during the picker's movement, the state is triggered to transfer to the next state.

[0041] Specifically, in this embodiment, in a given state Next, action Determines the picker's decision and action Specifically, =0 means the picker unloads the item when it is at the base, otherwise it remains stationary; and The actions correspond to moving right and moving left respectively, but are only allowed to be executed in the cross channel. It should be noted that in this embodiment, the actions are limited to feasible options. For example, when the picker is located in the picking channel When the vehicle is in motion, it is prohibited to move left or right; and Corresponding to upward movement and downward movement respectively, in addition, when the picker is in the rear cross channel When the vehicle is in the cross passage ahead, it is prohibited to move upwards. When the state is transferred, the system is prohibited from moving downward. This constraint mechanism effectively avoids invalid operations and ensures the actual executability of all generated strategies. During the state transfer process, the system evolves from one state to another according to the actions taken by the picker and the randomness of the order arrival. The minimum time unit required for a picker to move between adjacent storage locations if executed at a location other than the warehouse base =0, the picker stays at the current position Time, now and Same, unless a new order arrives; if executed at the warehouse base location =0, the picker will unload all items, resulting in the picker state becomes (1, 2n-1, 2n, K); and The action is only performed in the cross channel, and the time required is proportional to the distance between the channels. Valid when , the vertical coordinate changes with a step size of 2, When in action, the picker status The vertical position part of the step size is increased by 2 to obtain the picker state at time (t+1) ; When in action, the picker status The vertical position of the part is reduced by step 2 to obtain the state of the picker at time (t+1) .

[0042] when and When the action is executed, if the picker reaches the cross channel and cannot move forward, and a new order arrives and changes the status of the existing order in the warehouse Or the picker collects at least one item, which will lead to the generation of a new state. Assuming that the picker collects all items in the current storage location before transitioning to the new state, execute and When the action is taken, whether there is a new order arriving or not, the status of the existing order in the warehouse The relative distance changes caused by the movement of the picker will change. For the new picker state , whose horizontal position may change, but the vertical position remains the same, and the remaining capacity depends on the number of items collected during the action.

[0043] For the Markov decision process reward function, at time step t, the composite cumulative reward function is specifically defined as: Where, is the compound cumulative reward at time step t; is the picking efficiency reward coefficient; The maximum load capacity that a picker can carry in a single pick; is the remaining cargo capacity of the picker at time step t; is the unloading incentive adjustment coefficient, and its value range is between 0 and 1; is the number of storage locations moved by the picker in time step t; is the number of items picked up by the picker in time step t; is the action performed by the picker at time step t, Indicates that the picker performs unloading action.

[0044] When the picker executes at the base =0 action, the number of items is A smaller unloading incentive adjustment coefficient means that picking items is more valuable than unloading at the base. The picker will tend to maximize capacity utilization before going to the base, thereby minimizing the travel distance; a larger unloading incentive adjustment coefficient means that the two actions are of equal value. The picker will tend to unload items as soon as possible to minimize waiting time and delivery time at the same time. This reward mechanism optimizes the picking strategy by adjusting the unloading incentive adjustment coefficient, so that the picker can decide whether to prioritize maximizing the number of items carried or returning to the base as soon as possible to unload according to the specific situation, thereby finding the best balance between efficiency and response speed. It should be noted that the goal of this embodiment is to train a strategy to maximize the discounted cumulative reward by establishing a deep neural network model. Specifically, this embodiment adopts an improved Q learning algorithm, the objective function is the discounted cumulative reward, and the discounted cumulative reward calculation formula is: Where, is the cumulative discounted return at time step t, i.e., the discounted cumulative reward; is the discount factor. In this embodiment, , the discount factor is used to balance the importance of immediate rewards and future rewards, and the smaller The value indicates that the current action is more important than the future action. Value emphasizes future rewards; is the discount factor for the next k steps; is the future time step ( ) moment; k is the count of future steps, that is, the increment of the number of steps relative to the current moment t; To infinity.

[0045] This embodiment uses a deep neural network to approximate the optimal Q function: Where, is the optimal action value function, that is, the maximum expected reward that can be obtained by taking action a in state s. State s may include information such as the picker location and order distribution. Reflects the specific state-action pair The expected maximum value of the discounted cumulative reward that can be obtained by executing subsequent actions according to the optimal strategy; is the action taken in state s. In the warehouse order picking problem, the action can be moving right, moving left, moving up, moving down, or unloading items at the base; A strategy for selecting an action based on the current state; is the expectation operator, which is used to find the mean of a random variable (such as order arrivals); is the cumulative discounted return at time step t, i.e., the discounted cumulative reward; is the system state at time t, where time t indicates the temporal nature of the state.

[0046] The input of the deep neural network model is a state-action pair, and the output is an estimated expected reward. During the training process, the network parameters are continuously adjusted through gradient descent, and the optimal policy is finally derived: Where, is the optimal strategy for selecting actions based on the current state, that is, the strategy that maximizes the expected discounted cumulative reward among all possible strategies; is the parameter extraction operator, which is used to find the The action a that obtains the maximum value.

[0047] This embodiment adjusts the Q-learning mechanism to construct an optimal strategy that maximizes the reward. However, due to the complexity of the environment, the function cannot be directly accessed. However, since the neural network is a universal function approximator, a neural network can be constructed and trained to approximate it. In this way, the Q-learning method is used to approximate the optimal Q function through the neural network to formulate a strategy. This can solve complex environmental decision-making problems and find an action strategy that can maximize long-term rewards. This relies on the powerful fitting ability of the neural network to deal with various uncertainties in the environment. In summary, this embodiment uses the generalization ability of the neural network to effectively solve the dimensionality curse problem of traditional Q-learning in high-dimensional state space.

[0048] S6. Based on the real-time status of the picker and the mobile value density, a pre-built deep reinforcement learning neural network model is used to solve the Markov decision process to obtain the optimal picking path decision.

[0049] S7. Control the picker to perform picking operations according to the optimal picking path decision.

[0050] In some embodiments, the deep reinforcement learning neural network model includes an input layer, a two-branch feature extraction module, a feature fusion module, a multi-layer feature compression module, and an output layer connected in sequence; the two-branch feature extraction module includes an order status feature extraction unit and a picker status feature extraction unit connected in parallel; the order status feature extraction unit and the picker status feature extraction unit are both composed of a fully connected layer and a linear rectifier activation function; the multi-layer feature compression module includes at least three fully connected layers connected in series. In this embodiment, based on the real-time state of the picker and the mobile value density, a pre-built deep reinforcement learning neural network model is used to solve the Markov decision process to obtain the optimal picking path decision step, including: Combine the mobile value density of each channel and the real-time order status into an order status vector; the dimension of the order status vector is twice the total number of channels; The real-time state of the picker is subjected to nonlinear transformation by the picker state feature extraction unit to extract a high-dimensional feature vector of the picker; Performing a nonlinear transformation on the order status vector through an order status feature extraction unit to extract an order distribution feature vector; Perform channel splicing on the picker high-dimensional feature vector and the order distribution feature vector to generate a state fusion feature vector; The state fusion feature vector is subjected to hierarchical feature conversion through a multi-layer feature compression module to obtain a decision-oriented high-order feature representation; Mapping the decision-oriented high-level feature representation to the action space dimension, and calculating the expected cumulative reward value of each action taken in the current state through a fully connected layer; The action with the highest expected cumulative reward value is selected as the optimal action of the picker in the current state, and the optimal picking path decision is generated based on the optimal action.

[0051] Specifically, the deep reinforcement learning neural network model proposed in this embodiment is used to process the optimal action prediction of the picker in a dynamic warehouse environment. The deep reinforcement learning neural network model is used as a Q network to approximate the action value function The deep reinforcement learning neural network model contains multiple fully connected layers, which process the picker status and order status respectively, and then combine the two to make the final action prediction. Figure 3 This is a schematic diagram of the deep reinforcement learning neural network model architecture. The specific processing flow is: the input of the deep reinforcement learning neural network model is divided into the picker state and order status Two parts, where the picker state is a four-dimensional vector represented by ;Order status is a vector of size 2N, where N represents the total number of picking channels in the warehouse. This structure enables the network to process information from the picker status and order status separately, receive this information through different input layers, and integrate it in subsequent network layers to more accurately predict the optimal action under a given state, effectively utilizing deep learning capabilities to solve complex decision-making problems and achieve efficient order picking in a dynamic warehouse environment.

[0052] In the picker state feature extraction unit, the picker state input is processed by the first fully connected layer, which uses the ReLU (Rectified Linear Unit) activation function to convert the four-dimensional picker state into a high-dimensional feature vector of the picker. The feature space dimension of the high-dimensional feature vector of the picker is , this transformation method can capture the key features of the picker state, and the picker state feature extraction unit weight matrix dimension is At the same time, in the order status feature extraction unit, the order status is input into the second fully connected layer for processing. The second fully connected layer also uses the ReLU activation function to convert the 2N-dimensional order status vector into an order distribution feature vector. The feature space dimension of the order distribution feature vector is , the second fully connected layer can extract key features from the warehouse order status, and the order status feature extraction unit weight matrix dimension is ,Then, the feature fusion module concatenates the picker high-dimensional feature vector and the order distribution feature vector into a ,size of ( )-dimensional state fusion feature vector, integrating information from the two states to form a comprehensive representation of the system state.

[0053] Next, this embodiment inputs the spliced ​​state fusion feature vector into the multi-layer feature compression module for hierarchical feature conversion. The multi-layer feature compression module includes a multi-level fully connected layer. In this embodiment, for the convenience of description, this embodiment sets the multi-layer feature compression module to include a three-level fully connected layer, which are the third fully connected layer, the fourth fully connected layer and the fifth fully connected layer connected in sequence. The third fully connected layer includes units, and uses the ReLU activation function to transform the state fusion feature vector; the fourth fully connected layer reduces the dimension of the feature vector output by the third fully connected layer to units, and the fourth fully connected layer also uses the ReLU activation function; the fifth fully connected layer further reduces the dimension of the feature vector output by the fourth fully connected layer to units, and the fifth fully connected layer also uses the ReLU activation function, where 、 and are the feature layer dimensions (i.e., the number of neurons) of different fully connected layers in the multi-layer feature compression module, i.e., the intermediate dimensions of the state fusion feature vector gradually reduced by the three-level fully connected layer. is the weight matrix dimension of the third fully connected layer; is the weight matrix dimension of the fourth fully connected layer; is the weight matrix dimension of the fifth fully connected layer; is the dimension of the output layer weight matrix; is the number of neurons in the third fully connected layer, i.e., the initial processing dimension of the state fusion feature vector; is the number of neurons in the fourth fully connected layer, i.e., the intermediate feature compression dimension; is the number of neurons in the fifth fully connected layer, i.e., the final decision-oriented high-order feature representation dimension; The size of the action space, that is, the number of different actions that the picker can take during the warehouse order picking process, and the prediction of the output layer corresponding to different actions, determine the dimension of the output layer and the type of information ultimately output by the neural network.

[0054] Through the above steps, the network gradually refines and compresses the information extracted from the picker status and order status, providing a more compact and high-order feature representation for the final action prediction. Multi-level processing helps capture complex patterns in the data and supports more accurate decision making. Finally, the output layer converts The decision-oriented high-order feature representation of the dimensional model is mapped to the action space, the Q value of each possible action is predicted, and the action with the highest Q value is selected as the optimal action of the picker. The Q value is the expected cumulative reward value of the action. This process realizes the selection from state to optimal action through a deep neural network, ensuring that the action with the maximum expected reward is taken under a given state, effectively guiding the operation of the picker in the warehouse, improving the efficiency and accuracy of order picking in a dynamic warehouse environment, capturing the complex interaction between the picker and warehouse orders, providing efficient decision support for the picker, and improving overall operational efficiency and performance.

[0055] It should be noted that in the training process of the deep reinforcement learning neural network model, this embodiment adopts a dual network architecture of policy network and target network for training. The policy network is used for real-time decision-making. The policy network selects actions based on the epsilon-greedy strategy. The epsilon-greedy strategy balances the exploration of the action space and the utilization of the learned value; the target network serves as a stable reference, and its parameters are regularly updated from the learning parameters of the policy network. In order to enhance the learning process, this embodiment introduces an experience replay mechanism. The experience replay mechanism serves as a repository to save the experience of the intelligent agent in the form of state transitions. During the training process, the deep reinforcement learning intelligent agent will learn from small batches of samples randomly sampled and converted from the experience replay, which enables past experience to be effectively reused and stabilizes the learning process.

[0056] Training is performed over Ne cycles, each containing Ns steps, where Ne is the number of training cycles and Ns is the number of steps per cycle. In each step, the agent first observes the current state. Next, this embodiment uses the epsilon-greedy method and the policy network to select an action and obtain a corresponding reward when transitioning to the next state. The current state, action, next state, and reward together constitute a transition, which is stored in the experience replay as the experience replay capacity. Subsequently, this embodiment randomly selects mini-batches of transitions from the experience replay, and the policy network is trained based on these experiences. If the number of transitions in the experience replay is not enough to form a mini-batch, the training step is temporarily skipped. The training procedure for each mini-batch is a multi-step process: first, the state, next state, action, and reward within the batch are stacked for parallel processing. Then, the policy network is set to training mode and the target network is put into evaluation mode. This setup allows the policy network to calculate the state-action value while using the target network to calculate the expected state-action value. The difference between the two is used to calculate the loss. This embodiment uses the Huber Loss function (Huber Loss). Function) to increase robustness, the calculated loss is back-propagated through the policy network to promote its optimization. To ensure stability, the target network is updated less frequently, specifically once every Nupdate steps. The update follows the soft update rule and uses the target network update coefficient , the new target network weight is updated as the weighted average of the policy network weight and the existing target network weight. This gradual update mechanism can prevent large fluctuations and promote a smoother learning process. Among them, Nupdate is the update interval of the target network; is the target network update coefficient, which is used to softly update the target network weights. This embodiment effectively solves the stability problem in deep reinforcement learning by separating the policy network and the target network and combining the experience replay mechanism. It can be applied to practical application scenarios with continuous state space, such as dynamic warehouse order picking.

[0057] In summary, the embodiment proposes a dynamic picking path optimization method based on deep reinforcement learning (DRL), which can solve the dynamic order picking problem in a single-block warehouse layout served by autonomous picking equipment. Not only does it provide a solution to the current problem, but it also lays the foundation for future research and can explore more complex warehouse environments. Specifically, the deep reinforcement learning framework can be extended to multi-block layouts and scenarios involving multiple coordinated picking equipment. Through the prediction and adaptation capabilities of deep reinforcement learning, warehouse operations can balance efficiency and flexibility, quickly respond to dynamic changes in customer demand. For example, in terms of multi-block warehouse layout, existing models can be tested and optimized in larger and more complex warehouse structures, further improving the operational efficiency of large-scale warehouses. In terms of multi-device collaboration, research can be conducted on how to achieve effective collaboration between multiple autonomous picking equipment to further improve picking efficiency and reduce bottlenecks.

[0058] The embodiment of the present application provides a dynamic picking path optimization method based on deep reinforcement learning. The method collects warehouse channel layout parameters and picker physical parameters to form warehouse basic structure data, and structurally encodes the warehouse basic structure data to obtain warehouse structured encoding data. The current position of the picker and the remaining cargo capacity of the picker are analyzed based on the warehouse structured encoding data to obtain a picker state vector. The warehouse order queue data is collected in real time, and the moving value density of the orders in each channel is calculated based on the order queue data and the picker state vector. The state space is defined based on the picker state vector and the order queue data, and the discrete execution action space of the picker at each time step is constructed. Based on the state space and the discrete execution action space, the cumulative reward with the optimization goal of minimizing the average waiting time of the orders is set, and the warehouse order picking problem is modeled as a Markov decision process. Based on the real-time state of the picker and the moving value density, the Markov decision process is solved by using a pre-constructed deep reinforcement learning neural network model to obtain an optimal picking path decision. The picker performs the picking operation according to the optimal picking path decision. Compared with the prior art, the method collects and structurally encodes various warehouse data, dynamically adjusts the picking path through deep reinforcement learning, realizes real-time optimization of the dynamic picking path, effectively improves the efficiency of warehouse order picking, reduces the average waiting time of the orders, and enhances the response speed and intelligent level of the warehouse logistics system.

[0059] It should be noted that the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0060] In one embodiment, as Figure 4 As shown, the embodiment of the present application provides a dynamic picking path optimization system based on deep reinforcement learning, which comprises: The data acquisition module 101 is configured to collect the warehouse channel layout parameters and the picker physical parameters, form warehouse infrastructure data, and structure code the warehouse infrastructure data to obtain warehouse structured code data. The state analysis module 102 is configured to analyze the current position of the picker and the remaining load capacity of the picker based on the warehouse structured code data to obtain a picker state vector. The movement analysis module 103 is configured to collect the order queue data of the warehouse in real time, and calculate the movement value density of the orders in each channel according to the order queue data and the picker state vector. The space construction module 104 is configured to define a state space according to the picker state vector and the order queue data, and construct a discrete execution action space of the picker at each time step. The picking modeling module 105 is configured to set a cumulative reward with the optimization goal of minimizing the average waiting time of the orders based on the state space and the discrete execution action space, and model the warehouse order picking problem as a Markov decision process. The path decision module 106 is configured to solve the Markov decision process by using a pre-constructed deep reinforcement learning neural network model based on the real-time state of the picker and the movement value density to obtain an optimal picking path decision. The picking execution module 107 is configured to control the picker to perform picking operations according to the optimal picking path decision.

[0061] The specific limitations of the dynamic picking path optimization system based on deep reinforcement learning can refer to the limitations of the dynamic picking path optimization method based on deep reinforcement learning described above, which will not be repeated here. Those skilled in the art can realize that the various modules and steps described in combination with the embodiments disclosed in the present application can be realized in hardware, software or a combination of both. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. The skilled person can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.

[0062] The embodiment of the application provides a dynamic picking path optimization system based on deep reinforcement learning, the system collects warehouse channel layout parameters and picker physical parameters through a data acquisition module, forms warehouse basic structure data, and carries out structured coding on the warehouse basic structure data to obtain warehouse structured coding data; a state analysis module analyzes the current position of a picker and the residual load capacity of the picker based on the warehouse structured coding data to obtain a picker state vector; a movement analysis module collects order queue data to be processed of the warehouse in real time, and calculates the movement value density of orders in each channel according to the order queue data to be processed and the picker state vector; a space construction module defines a state space according to the picker state vector and the order queue data to be processed, and constructs a discrete execution action space of the picker at each time step; a picking modeling module sets a cumulative reward with an optimization target of minimizing the average order waiting time based on the state space and the discrete execution action space, models the warehouse order picking problem as a Markov decision process; a path decision module solves the Markov decision process by using a pre-constructed deep reinforcement learning neural network model based on the real-time state of the picker and the movement value density to obtain an optimal picking path decision; and a picking execution module controls the picker to perform a picking operation according to the optimal picking path decision. Compared with the prior art, the system collects various data of the warehouse and carries out structured coding, dynamically adjusts the picking path through deep reinforcement learning, realizes real-time optimization of the dynamic picking path, effectively improves the efficiency of warehouse order picking, reduces the average order waiting time, and enhances the response speed and intelligent level of the warehouse logistics system.

[0063] The above-described embodiments only express several preferred embodiments of the application, and the description is relatively specific and detailed, but it should not be understood as limiting the scope of the patent. It should be noted that for ordinary skilled in the art, several improvements and replacements can be made without departing from the technical principles of the application, and these improvements and replacements should be regarded as the protection scope of the application. Therefore, the protection scope of the patent of the application should be subject to the protection scope of the claims.

Claims

1. A dynamic picking path optimization method based on deep reinforcement learning, characterized in that: The following steps are involved: Collecting warehouse aisle layout parameters and picker physical parameters to form warehouse infrastructure data, and performing structured coding on the warehouse infrastructure data to obtain warehouse structured coding data; Analyzing the current position of the picker and the remaining cargo capacity of the picker based on the warehouse structured coding data to obtain a picker state vector; Collecting the warehouse's pending order queue data in real time, and calculating the movement value density of orders in each channel based on the pending order queue data and the picker state vector; A state space is defined according to the picker state vector and the pending order queue data, and a discrete execution action space of the picker at each time step is constructed; Based on the state space and the discrete execution action space, a cumulative reward is set with minimizing the average waiting time of orders as the optimization goal, and the warehouse order picking problem is modeled as a Markov decision process; Based on the real-time status of the picker and the mobile value density, a pre-built deep reinforcement learning neural network model is used to solve the Markov decision process to obtain the optimal picking path decision; The picker is controlled to perform the picking operation according to the optimal picking path decision.

2. A dynamic picking path optimization method based on deep reinforcement learning according to claim 1, characterized in that: The step of performing structured coding on the warehouse infrastructure data to obtain warehouse structured coding data includes: Based on the pre-collected warehouse physical structure data, a warehouse discretized coordinate system is established with the center of the entrance cross channel as the origin of the coordinate system; Encode all aisles in the warehouse and the storage cells on both sides of each aisle according to the warehouse aisle layout parameters to obtain aisle layout parameter encoding data; Mapping the channel layout parameter encoding data to the warehouse discretized coordinate system to obtain warehouse channel storage encoding data; Fixing the initial position of the picker to the entrance cross-channel coordinates according to the physical parameters of the picker, and setting the initial unloaded capacity to the maximum loaded capacity; Obtaining an initial state vector of the picker according to the initial position of the picker and the maximum cargo capacity; The warehouse channel storage coding data and the picker initial state vector are encapsulated into a structured data packet to form warehouse structured coding data.

3. The dynamic picking path optimization method based on deep reinforcement learning according to claim 1, characterized in that: The step of analyzing the current position of the picker and the remaining cargo capacity of the picker based on the warehouse structured coding data to obtain the picker state vector includes: Acquire the real-time position data of the picker, match the real-time position data of the picker with the structured coding data of the warehouse, parse the channel number and storage cell number of the current location of the picker, and obtain the current position of the picker; Obtaining current load data of the picker, and determining the remaining load capacity of the picker based on the maximum load capacity of the picker in the warehouse structured coding data and the current load data; The current position of the picker and the remaining cargo capacity of the picker are combined into a picker state vector.

4. The method for dynamic picking path optimization based on deep reinforcement learning according to claim 1, characterized in that: The pending order queue data includes the storage location coordinates of items in the warehouse for pending orders; The movement value density is the sum of the ratios of the order quantity of each picking position in each channel and the minimum relative distance between the current position of the picker and the coordinates of the item storage position when the picker moves up or down.

5. The method for dynamic picking path optimization based on deep reinforcement learning according to claim 1, characterized in that: The steps of setting a cumulative reward with minimizing the average waiting time of orders as an optimization goal based on the state space and the discrete execution action space, and modeling the warehouse order picking problem as a Markov decision process include: Determining a state transition rule for the picker to transition from a current state to a next state based on the state space and the discrete execution action space; Based on the state transition rules, the state transition probability of the picker transferring from the current state to the next state after executing the action is calculated, and a state transition probability model is constructed; With minimizing the average waiting time for orders as the optimization goal, a composite cumulative reward function is constructed based on the picker's action execution results, which includes a picking efficiency reward, a travel distance penalty, and an unloading incentive. The picker's action execution results include at least the time required for the picker to complete the picking order, the travel distance, and the order completion rate. Based on the state space, the discrete execution action space and the state transition probability model, the warehouse order picking problem is modeled as a Markov decision process by maximizing the optimization objective of a composite cumulative reward function.

6. A dynamic picking path optimization method based on deep reinforcement learning as claimed in claim 5, characterized in that: The state transition rules are specifically as follows: Set the minimum time unit required for the picker to move from one storage location to another, and detect the current position of the picker and perform actions; When it is detected that the picker is at a non-warehouse base location and performs a stationary action, the picker is controlled to remain stationary at the current location for the minimum time unit, and when a new order is detected to arrive during the period when the picker remains stationary, the state is triggered to transition to the next state; When it is detected that the picker is at the warehouse base position and performs the unloading action, the picker is controlled to unload all the items it carries and the remaining loading capacity of the picker is reset to the maximum loading capacity; When the picker is detected to be performing a movement action at the warehouse base position, if the picker is in the rear cross aisle or front cross aisle, the left and right movement actions are allowed; During the picker's movement, the order status changes and picker status changes are monitored in real time. If any of the following conditions are detected: the picker reaches the cross channel and cannot continue to move in the original direction of movement when performing the movement, a new order arrival condition is detected during the picker's movement, and at least one item is picked up during the picker's movement, the state is triggered to transfer to the next state.

7. The method for dynamic picking path optimization based on deep reinforcement learning according to claim 5, characterized in that: The composite cumulative reward function is specifically: Where, is the compound cumulative reward at time step t; is the picking efficiency reward coefficient; The maximum load capacity that a picker can carry in a single pick; is the remaining cargo capacity of the picker at time step t; is the unloading incentive adjustment coefficient; is the number of storage locations moved by the picker in time step t; is the number of items picked up by the picker in time step t; is the action performed by the picker at time step t, Indicates that the picker performs unloading action.

8. The method for dynamic picking path optimization based on deep reinforcement learning according to claim 1, characterized in that: The deep reinforcement learning neural network model includes an input layer, a dual-branch feature extraction module, a feature fusion module, a multi-layer feature compression module and an output layer connected in sequence; The dual-branch feature extraction module includes an order status feature extraction unit and a picker status feature extraction unit connected in parallel; both the order status feature extraction unit and the picker status feature extraction unit are composed of a fully connected layer and a linear rectifier activation function; The multi-layer feature compression module includes at least three levels of fully connected layers connected in series.

9. The method for dynamic picking path optimization based on deep reinforcement learning according to claim 8, characterized in that: The steps of solving the Markov decision process based on the real-time state of the picker and the mobile value density using a pre-built deep reinforcement learning neural network model to obtain the optimal picking path decision include: Combine the mobile value density of each channel and the real-time order status into an order status vector; the dimension of the order status vector is twice the total number of channels; The real-time state of the picker is subjected to nonlinear transformation by the picker state feature extraction unit to extract a high-dimensional feature vector of the picker; Performing a nonlinear transformation on the order status vector through the order status feature extraction unit to extract an order distribution feature vector; Perform channel splicing on the picker high-dimensional feature vector and the order distribution feature vector to generate a state fusion feature vector; The state fusion feature vector is subjected to hierarchical feature conversion through a multi-layer feature compression module to obtain a decision-oriented high-order feature representation; Mapping the decision-oriented high-level feature representation to the action space dimension, and calculating the expected cumulative reward value of each action taken in the current state through a fully connected layer; The action with the highest expected cumulative reward value is selected as the optimal action of the picker in the current state, and the optimal picking path decision is generated based on the optimal action.

10. A dynamic picking path optimization system based on deep reinforcement learning, characterized in that: The system comprises: A data acquisition module is used to collect warehouse channel layout parameters and picker physical parameters to form warehouse infrastructure data, and to perform structured coding on the warehouse infrastructure data to obtain warehouse structured coding data; a state parsing module, configured to parse the current position of the picker and the remaining cargo capacity of the picker based on the warehouse structured coding data to obtain a state vector of the picker; A movement analysis module, configured to collect data on the warehouse's pending order queues in real time, and calculate the movement value density of orders in each channel based on the pending order queue data and the picker state vector; A space construction module, configured to define a state space according to the picker state vector and the pending order queue data, and to construct a discrete execution action space of the picker at each time step; A picking modeling module is configured to set a cumulative reward with minimizing the average waiting time of orders as an optimization goal based on the state space and the discrete execution action space, and model the warehouse order picking problem as a Markov decision process; A path decision module is used to solve the Markov decision process based on the real-time status of the picker and the movement value density using a pre-built deep reinforcement learning neural network model to obtain the optimal picking path decision; The picking execution module is used to control the picker to perform the picking operation according to the optimal picking path decision.

Citation Information

Patent Citations

  • Multi-AGV path planning method and system based on deep reinforcement learning

    CN119105508A

  • Warehouse inventory management and optimization method based on artificial intelligence

    CN120494697A

  • Warehouse storage and distribution wave optimization method and system

    CN120579928A

  • System and method for ride order dispatching

    US20200193834A1

  • Method for using reinforcement learning to optimize order fulfillment

    WO2024028839A1

Cited By

  • Slot allocation strategy optimization method and system based on reinforcement learning

    CN121257873A

  • Intelligent supply chain logistics warehouse storage location distribution method and system and computer readable storage medium

    CN122175512A