An optimization method for the storage location allocation of a three-dimensional warehouse for non - font shelves

By building a reinforced learning model of the DQN framework, the problem of improper distribution of cargo spaces in non-font-shaped shelf three-dimensional warehouses is solved, the efficiency of entry and exit is optimized, and dynamic planning and experience utilization is improved.

CN115860239BActive Publication Date: 2025-07-04KEDA INTELLIGENT IOT TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211594440.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-12-13
Publication Date
2025-07-04
Estimated Expiration
2042-12-13

AI Technical Summary

Technical Problem

The existing technology lacks effective cargo space allocation optimization methods in non-font-shaped shelf three-dimensional warehouses, resulting in low efficiency in inlet and exit, unable to perform dynamic planning, and there is a deviation between the logistics scenario modeling and the actual situation.

Method used

Build a reinforcement learning model based on the DQN framework, define agents, state space, action space and reward returns, train the model through the Poisson arrival process, use the backpropagation algorithm to update parameters, and optimize cargo space allocation.

Benefits of technology

The inlet and exit efficiency of non-font-shaped shelf three-dimensional warehouses is improved, repeated calculations are reduced, experience utilization is improved, and dynamic cargo space allocation is optimized.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115860239B_ABST
    Figure CN115860239B_ABST
Patent Text Reader

Abstract

The present invention discloses a three-dimensional warehouse storage location allocation optimization method for non-shaped shelves, including constructing a reinforcement learning model based on the DQN framework, and defining the agent, state space, action space, reward return, and its optimization goal in the reinforcement learning model; initializing all parameter values and policies of the reinforcement learning model, batch generating inbound and outbound tasks, and training the reinforcement learning model through the inbound and outbound tasks with Poisson arrival; using the backpropagation algorithm to derive the policy gradient, calculating the gradient descent to update the DQN network parameters, obtaining the trained reinforcement learning model, and applying it to the three-dimensional warehouse for intelligent storage location optimization. The present invention obtains the optimal allocation scheme by training the model with the inbound and outbound tasks of goods, and solves the problem of storage location allocation in the non-shaped shelf warehouse in dynamic storage and large-scale logistics scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent warehousing industry, and particularly relates to an optimization method for the storage location allocation of a three-dimensional warehouse for non-shaped shelves. Background Art

[0002] Automated three-dimensional warehouses are widely used in industrial warehousing due to their characteristics of low floor area, high throughput efficiency, and intelligent integrated control. Among them, three-dimensional warehouses with non-shaped shelves are the most common in the logistics industry. The non-shaped shelf is a type of shelf designed with single rows on both outer sides and double rows inside. Among the factors affecting the access operation efficiency of non-shaped shelves, the main one is the storage location allocation problem. Due to uncontrollable external factors and the characteristics of non-shaped shelves themselves, if the storage location is not properly allocated when the goods are stored in the warehouse, it will increase the distance and operation time of inbound and outbound operations, and reduce the inbound and outbound efficiency and the benefits of logistics enterprises. At present, the research on intelligent three-dimensional warehousing in China started relatively late, and the dynamic storage for large-scale logistics still develops slowly, especially lacking the optimization control research applicable to the storage location allocation of non-shaped shelves.

[0003] The deficiencies of the existing technology are as follows. There are currently proposed storage location allocation methods with set rules, such as optimizing the placement position of the storage location according to the ratio of the turnover rate of the goods. Some researchers establish a fitness function with the inbound and outbound frequency of the goods and the shelf stability as the optimization objectives, and optimize it through a heuristic algorithm. However, the existing optimization methods for the storage location allocation of three-dimensional warehouses have certain limitations: the existing methods consider few realistic factors, and there are deviations between the logistics scenario modeling and the actual situation; it is necessary to artificially design the objective function and constraint conditions, and the requirements for prior knowledge are relatively high; it can only consider the optimal situation of the warehouse at the current time point and cannot perform dynamic planning for inbound and outbound operations. Summary of the Invention

[0004] The purpose of the present invention is to overcome the deficiencies of the existing technology. To achieve the above purpose, an optimization method for the storage location allocation of a three-dimensional warehouse for non-shaped shelves is adopted to solve the problems raised in the above background art.

[0005] An optimization method for the storage location allocation of a three-dimensional warehouse for non-shaped shelves, the specific steps include:

[0006] Step S1, construct a reinforcement learning model based on the DQN framework, and define the intelligent agent, state space, action space, reward return, and its optimization objective in the reinforcement learning model;

[0007] Step S2, initialize all parameter values and policies of the reinforcement learning model, and randomly generate inbound and outbound tasks;

[0008] Step S3: Input the inbound and outbound tasks based on the Poisson arrival process into the reinforcement learning model to train the constructed reinforcement learning model;

[0009] Step S4: Use the backpropagation algorithm to derive the policy gradient and calculate the gradient descent to update the DQN network parameters;

[0010] Step S5: Repeat the above steps S3 and S4 to obtain the trained reinforcement learning model and apply it to the stereoscopic warehouse for intelligent optimization of storage locations.

[0011] As a further solution of the present invention: The specific steps of the said step S1 include:

[0012] S11: Definition of state space:

[0013] Obtain the information of the storage locations in the stereoscopic warehouse, including the information of the goods stored in the storage locations, the information of the stacker, and the executable task information. At the same time, encode each storage location, and the encoding is expressed as:

[0014] S = (P 1, P 2, P 3, … t, D, E) ∈ [1, …, t];

[0015] Among them, P i represents the type of goods, t represents the number of storage locations, i represents the unique encoding corresponding to each storage location, D represents the information of the stacker, and E represents the current executable task information;

[0016] S12: Definition of action space:

[0017] Set the action space by adopting different inbound and outbound rules. At the same time, pre-calculate the distance between each storage location and the inbound and outbound points and store it in the distance matrix. The action space is set with four types of actions, including selecting the storage cell with the closest moving distance for outbound, selecting the empty storage cell with the shortest distance for inbound, selecting the empty storage cell near the middle position of the bottom layer for inbound, and having no executable inbound and outbound tasks and waiting for new tasks to arrive;

[0018] S13: Definition of reward return:

[0019] Obtain the feedback given by the reward return to the agent's action for the state and use this feedback to guide the agent's learning. Set the optimization goal to maximize the obtained reward return;

[0020] The said reward return is set to normalize the task completion time and use its opposite number as the reward return.

[0021] As a further solution of the present invention: The specific steps of the said step S2 include:

[0022] Randomly generate storage bins and goods, and randomly select n storage bin positions;

[0023] Generate the storage bin status and the stacker status. Among them, the storage bin status indicates the presence or absence of goods in the storage bin, and the initial goods type is 0, indicating not selected.

[0024] As a further solution of the present invention: The specific steps of the step S3 include:

[0025] First, scan to obtain the currently arrived tasks and determine whether they can be executed;

[0026] Adopt an inbound task or an outbound task. Based on the DQN network and the greedy strategy, select the empty storage bin with the shortest distance for inbound or the empty storage bin near the middle position of the bottom layer for inbound;

[0027] At the same time, continuously update whether there are new arrived tasks during the execution of the task. If not, wait; if so, execute the new task;

[0028] Obtain the reward return r and the new state s', and store them in the preset experience pool in the DQN network.

[0029] As a further solution of the present invention: The specific steps of the step S4 include:

[0030] S41. Priority experience extraction:

[0031] First, sample a batch of data from the experience pool and extract the experience in a probabilistic manner. Then the actual extraction probability of each experience is:

[0032]

[0033] Among them, j =|δ t +|, j is the number of experiences in the experience pool, δ t is the TD deviation. Set the non-uniform sampling probability p j to be proportional to the TD deviation δ t , and perform normalization processing on p j to obtain the actual sampling probability P(j) of each experience;

[0034] By appropriately adjusting the learning rate α to eliminate the deviation, the expression is:

[0035] α←α·(np t ) -β

[0036] Among them, n is the number of experiences participating in the sampling, β∈(0,1];

[0037] S42. Value network update:

[0038] Introduce a target network, estimate the TD target value from the target network for updating the DQN network. Update the target network by using the new DQN network. Update the DQN network according to the sampled experience (s t , a, r t , s t+1 ), and the specific formula is as follows:

[0039] TD target value:

[0040] TD deviation: δ t = Q(s t , a t ; w) - y t

[0041] Gradient descent:

[0042] where s t represents the state at time t, a represents the action selection, a t represents the action selection at time t, Q(s t , a t ; w) represents the value estimation of the target network for the state-action selection at the current time t, r t is the immediate reward obtained from the current action selection at time t, γ ∈ (0, 1] represents the decay of the value estimation of future states, s t+1 represents the next state, represents the value estimation of the maximum state-action selection at time t + 1, α represents the learning rate of the algorithm update, and w represents the gradient descent.

[0043] As a further solution of the present invention: The specific step S5 is to train the reinforcement learning model by setting the number of training and test rounds and the number of tasks included in each round.

[0044] Compared with the prior art, the present invention has the following technical effects:

[0045] By adopting the above technical solution, by constructing a reinforcement learning model based on the DQN framework, then improving the training mechanism, adjusting the learning parameters, simplifying the action space, considering the characteristics of non-character-shaped shelves and adding actions for inbound at the middle position near the bottom layer, and pre-computing the distance between each storage location and the inbound / outbound point and storing it in the distance matrix, the repetitive and redundant calculations are reduced.

[0046] Using the method of updating with experience replay, sampling is performed using the experience replay pool and used for parameter update. By setting random extraction from the experience pool, the correlation of experiences can be cut off, and each experience can be learned multiple times, improving the utilization rate of experiences.

[0047] Finally, a target network is introduced to estimate the TD target value from the target network, which is then used to update the DQN network. After a period of time, the target network is updated using the new DQN network. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] The following is a detailed description of the specific embodiments of the present invention with reference to the accompanying drawings:

[0049] Figure 1 It is a schematic diagram of the steps of the allocation optimization method for the disclosed embodiments of the present application;

[0050] Figure 2 It is a flowchart of the reinforcement learning model for the disclosed embodiments of the present application. SPECIFIC EMBODIMENTS

[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0052] Please refer to Figure 1 and Figure 2 , in the embodiments of the present invention, a three-dimensional warehouse location allocation optimization method for a non-shaped shelf specifically includes the following steps:

[0053] Step S1: Construct a reinforcement learning model based on the DQN framework, and define the agent, state space, action space, reward return, and its optimization goal in the reinforcement learning model. The specific steps include:

[0054] S11: Definition of the state space S:

[0055] Obtain the information of the three-dimensional warehouse locations, including the information of the goods stored in the locations, the stacker information, and the executable task information. At the same time, each location is encoded, and the encoding is represented as:

[0056] S = (P 1, P 2, P 3, … t, D, E) ∈ [1, …, t];

[0057] Among them, P i represents the type of goods, t represents the number of locations, i represents the unique code corresponding to each location, D represents the stacker information, and E represents the current executable task information;

[0058] S12: Definition of the action space:

[0059] When setting the action space, different inbound and outbound rules are used instead of specific storage locations to set the action space. At the same time, by pre-computing the distances between each storage location and the inbound and outbound points and storing them in a distance matrix, the computing time of the program is effectively reduced. Four types of actions are set in the action space, including selecting the storage bin with the closest moving distance for outbound, selecting the empty storage bin with the shortest distance for inbound, selecting the empty storage bin near the middle position of the bottom layer for inbound, and having no executable inbound or outbound task and waiting for a new task to arrive;

[0060] S13. Reward definition:

[0061] Obtain the feedback given by the state of the reward to the agent's action, and use this feedback to guide the agent's learning. The set optimization goal is to maximize the obtained reward;

[0062] The said reward is set as normalizing the task completion time and using its opposite number as the reward;

[0063] Step S2. Initialize all parameter values and policies of the reinforcement learning model, and randomly generate inbound and outbound tasks. The specific steps include:

[0064] Step S21. Randomly generate storage bins and goods, and randomly select n storage bin positions;

[0065] Step S22. Generate the storage bin state and the stacker state. Among them, the storage bin state indicates the presence or absence of goods in the storage bin. The initial goods type is 0, indicating not selected. The initial action of the stacker is set to 5 to distinguish the first operation.

[0066] Step S3. Input the said inbound and outbound tasks based on the Poisson arrival process into the reinforcement learning model, and train the constructed reinforcement learning model. The specific steps include:

[0067] First, scan to obtain the currently arrived tasks and judge whether they can be executed;

[0068] Take the inbound task or the outbound task. Based on the DQN network and the greedy strategy, select the empty storage bin with the shortest distance for inbound or the empty storage bin near the middle position of the bottom layer for inbound;

[0069] At the same time, update in real time whether there are new arrived tasks during the task execution. If not, wait. If so, execute the new task;

[0070] Obtain the reward r and the new state s', and store the obtained parameters (s, a, r, s') into the preset experience pool in the DQN network. The experience pool is the basic design of the DQN algorithm. The improvement in this embodiment lies in the priority extraction of the experience pool and Double DQN.

[0071] Step S4: Use the backpropagation algorithm to derive the policy gradient and calculate the gradient descent to update the DQN network parameters. The specific steps are as follows:

[0072] S41: Prioritized experience sampling:

[0073] First, sample a batch of data from the experience pool. To prevent the network from overfitting, sample the experience in a probabilistic manner. Then the actual sampling probability of each experience is:

[0074]

[0075] where, j = |δ t +|, δ t is the TD error. Set the non-uniform sampling probability p j to be proportional to the TD error δ t . Normalize p j to obtain the actual sampling probability P(k) of each experience. ∈ is a very small value to prevent the probability of an experience with a TD error of 0 from being 0 when sampled;

[0076] Since experiences are sampled with different probabilities, there is a bias in the DQN prediction;

[0077] Adjust the learning rate α accordingly to eliminate the bias. The expression is:

[0078] α ← α · (np t ) -β

[0079] where n is the number of experiences participating in the sampling, and β ∈ (0, 1].

[0080] S42: Value network update:

[0081] Introduce a target network to estimate the TD target value from the target network for updating the DQN network. Update the DQN network by using the new DQN network to update the target network. According to the sampled experience (s t , a, r t , s t+1 ), update the DQN network. The specific formula is as follows:

[0082] TD target value:

[0083] TD error: δ t = Q(s t , a t ; w) - y t

[0084] Gradient descent:

[0085] Among them, s t represents the state at time t, a represents the action selection, and a t represents the action selection at time t. Q(s t , a t ; w) represents the value estimation of the target network for the state-action selection at the current time t. r t is the immediate reward obtained from the current action selection at time t. γ ∈ (0, 1] represents the attenuation of the value estimation of future states. s t+1 represents the next state, represents the value estimation of the maximum state-action selection at time t + 1. α represents the learning rate of the algorithm update, and w represents gradient descent.

[0086] Step S5: Repeat the above steps S3 and S4 to obtain a trained reinforcement learning model and apply it to the stereoscopic warehouse for storage location optimization. The specific implementation method is to set the specific number of training and test rounds, as well as the number of tasks included in each round, and then train the reinforcement learning model.

[0087] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents and should be included within the protection scope of the present invention.

Claims

1. A method for optimizing the allocation of storage locations in a three-dimensional warehouse for non - shaped shelves, characterized in that, The specific steps include: Step S1: Construct a reinforcement learning model based on the DQN framework, and define the agent, state space, action space, reward, and its optimization goal in the reinforcement learning model; Step S2: Initialize all parameter values and inbound and outbound strategies of the reinforcement learning model, and randomly generate inbound and outbound tasks; Step S3: Input the inbound and outbound tasks based on the Poisson arrival process into the reinforcement learning model, and train the constructed reinforcement learning model; Step S4: Use the backpropagation algorithm to derive the policy gradient, calculate the gradient descent to update the DQN network parameters, and its specific steps include: S41: Prioritized experience sampling: First, sample a batch of data from the experience pool, and extract experiences in a probabilistic manner. Then the actual extraction probability of each experience is: where p j = |δ t + ∈|, j is the number of experiences in the experience pool, δ t is the TD error, set the non-uniform sampling probability p j to be proportional to the TD error δ t , and normalize p j to obtain the actual sampling probability P(j) of each experience; By appropriately adjusting the learning rate α to eliminate bias, the expression is: α ← α·(np t ) -β where n is the number of experiences participating in the sampling, and β ∈ (0, 1]; S42: Value network update: Introduce a target network, estimate the TD target value from the target network for updating the DQN network, update the target network by using the new DQN network, and update the DQN network according to the sampled experience (s t , a, r t , s t+1 ). The specific formula is as follows: TD target value: TD deviation: δ t = Q(s t , a t ; w) - y t Gradient descent: Among them, s t represents the state at time t, a represents the action selection, and a t represents the action selection at time t. Q(s t , a t ; w) represents the value estimation of the target network for the state-action selection at the current time t. r t is the immediate reward obtained by the current action selection at time t. γ ∈ (0, 1] represents the attenuation of the value estimation of future states. s t+1 represents the next state, represents the value estimation of the maximum state-action selection at time t + 1. α represents the learning rate of the algorithm update, and w represents gradient descent; Step S5: Repeat the above steps S3 and S4 to obtain a trained reinforcement learning model, and apply it to the automated storage and retrieval system for intelligent optimization of storage locations.

2. The three-dimensional warehouse location allocation optimization method for non-shaped shelves according to claim 1, wherein The specific steps of the said step S1 include: S11: State space definition: Obtain the information of the storage locations in the automated storage and retrieval system, including the information of the goods stored in the storage locations, the information of the stacker crane, and the information of the executable tasks. At the same time, encode each storage location, and the encoding is expressed as: S = (P 1, P 2, P 3, …P t, D, E) i ∈ [1, …, t]; Among them, P i represents the type of goods, t represents the number of storage locations, i represents the unique code corresponding to each storage location, D represents the information of the stacker, and E represents the information of the currently executable task; S12: Action space definition: Set the action space by using different inbound and outbound rules. At the same time, pre-calculate the distance between each storage location and the inbound and outbound points, and store it in the distance matrix. Four types of actions are set in the action space, including selecting the storage cell with the shortest moving distance for outbound, selecting the empty storage cell with the shortest distance for inbound, selecting the empty storage cell near the middle position of the bottom layer for inbound, and having no executable inbound and outbound tasks and waiting for new tasks to arrive; S13: Reward definition: Obtain the reward to determine the feedback given by the state to the agent's action, and use this feedback to guide the agent's learning. Set the optimization goal to maximize the obtained reward; The said reward is set as normalizing the task completion time and using its negative value as the reward.

3. The three-dimensional warehouse location allocation optimization method for a non-shaped shelf according to claim 1, wherein, The specific steps of the said step S2 include: Randomly generate storage cells and goods, and randomly select n storage cell positions; Generate the storage cell state and the stacker crane state. Among them, the storage cell state indicates the presence or absence of goods in the storage cell, and the initial goods type is 0, indicating not selected.

4. The optimized method for allocating storage locations in a three-dimensional warehouse for non-shaped shelves according to claim 1, characterized in that, The specific steps of the said step S3 include: First, scan to obtain the currently arrived tasks and judge whether they can be executed; Take the inbound task or the outbound task, and based on the DQN network and the greedy strategy, select the empty storage cell with the shortest distance for inbound or the empty storage cell near the middle position of the bottom layer for inbound; At the same time, update in real time whether there are new tasks arriving during the task execution. If not, wait; if so, execute the new tasks; Obtain the reward r and the new state s', and store them in the experience pool preset in the DQN network.

5. The optimized method for allocating storage locations in a three-dimensional warehouse for non-shaped shelves according to claim 1, wherein The specific content of the said step S5 is to train the reinforcement learning model by setting the number of training and testing rounds, and the number of tasks included in each round.

Citation Information

Patent Citations

  • Storage location allocation method and device

    CN107341629A

  • Dynamic management method and system for centralized distribution positions

    CN115222224A