Intelligent logistics digital warehouse management method based on deep learning
By applying deep learning and reinforcement learning technologies in warehouse management, dynamically optimizing SKU storage layout and picking paths, the inefficiency problem of traditional warehouse management models when SKU changes and order demand surges is solved, and more efficient picking and order processing is achieved.
Patent Information
- Application Number
- CN202510205295.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-24
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
When SKU changes, order demand surges or warehouse layout adjustments, it is difficult to adapt quickly, resulting in unreasonable picking paths, low picking efficiency, insufficient dynamic adjustment capabilities, and affecting the overall order processing efficiency.
Using a digital warehouse management method for intelligent logistics based on deep learning, SKU storage optimization strategies and picking paths are calculated through deep learning models (such as Transformer and GNN), combined with reinforcement learning to train picking robots, dynamically adjust SKU storage layout and picking paths, and optimize task scheduling and allocation.
Adaptive adjustment of SKU dynamic changes has been realized, the intelligence of warehousing operations has been improved, the picking path has been shortened, the picking efficiency and order fulfillment capabilities have been improved, and the warehouse operation cost has been reduced.
Smart Images

Figure CN119990985A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of warehouse management, and in particular to a smart logistics digital warehouse management method based on deep learning. Background Art
[0002] With the rapid development of smart logistics, warehouse management has gradually evolved from relying on manual operations to digitalization and automation; modern warehouses generally use warehouse management systems (WMS) and automated sorting equipment to improve inventory management speed.
[0003] With the expansion of business scale and the continuous change of market demand, the current management model efficiency is gradually insufficient in the e-commerce warehouse and fast-moving consumer goods logistics environment where the inventory unit SKU changes frequently. The order sorting system in the e-commerce warehouse mainly relies on the rule engine for optimization, such as order classification, ABC sorting strategy, etc. This method is difficult to adapt to SKU changes, including new product listings and fluctuations in short-term hot-selling products, resulting in the inability to adjust the warehouse storage layout and picking path in time; after the SKU is changed, the operation is still performed according to the established storage rules, resulting in unreasonable picking paths, and pickers or automatic sorting robots frequently travel back and forth between different areas, increasing invalid operation time;
[0004] At the same time, the dynamic order pool optimization capability is insufficient, and the storage location adjustment of high-frequency order goods is delayed, further exacerbating the decline in warehouse operation efficiency. When order demand surges or SKUs undergo large-scale adjustments, the traditional static rule-based optimization method has a delayed response and is unable to meet the needs of efficient fulfillment, affecting the overall order processing efficiency. Therefore, a smart logistics digital warehouse management method based on deep learning is urgently needed to solve such problems. Summary of the invention
[0005] In view of the above existing problems, the present invention is proposed.
[0006] The present invention provides a smart logistics digital warehouse management method based on deep learning to solve the problems that traditional order sorting relies on fixed rules, has poor adaptability to SKU changes, lagging path optimization, insufficient dynamic adjustment capabilities, and insufficient picking efficiency.
[0007] In order to solve the above technical problems, the present invention provides the following technical solutions:
[0008] The embodiment of the present invention provides a smart logistics digital warehouse management method based on deep learning, which includes:
[0009] Step S1, obtaining order data, parsing the SKU information, quantity requirements and priority in the order data, classifying them in combination with historical order data, and generating SKU related data;
[0010] Step S2, based on the SKU associated data generated in step S1, the deep learning model Transformer is used to calculate the SKU storage optimization strategy, the SKU storage location is adjusted through the attention mechanism, and the storage layout is generated;
[0011] Step S3, based on the optimized storage layout generated in step S2, using reinforcement learning to train the picking robot to learn the optimal picking path;
[0012] Combine GNN to calculate the optimal allocation of AGV and manual picking tasks, and generate picking paths and task scheduling plans;
[0013] Step S4, executing AGV and manual picking operations based on the picking path and task scheduling plan generated in step S3;
[0014] After picking is completed, the SKU picking time and path optimization results are recorded, and the data is passed to step S5 as feedback data;
[0015] Step S5, based on the feedback data, using self-supervised learning to dynamically adjust the warehouse management strategy, automatically optimize the SKU storage layout, picking path and task scheduling, and apply it to the subsequent order sorting process.
[0016] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, wherein: in step S1, the node feature learning algorithm GraphSAGE is used to calculate the SKU correlation, generate a SKU association matrix, and optimize the high-frequency SKU combination based on the SKU correlation to generate optimized SKU association data.
[0017] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, the steps of obtaining order data, parsing the SKU information, quantity requirements and priority in the order data, and classifying in combination with the historical order data are as follows:
[0018] Parse the order data and define the order data set as D:
[0019] D={d1,d2,…,d n},
[0020] Where D represents the order data set, d i Represents a single order data, n represents the total number of orders,
[0021] The order data structure is d i :
[0022] d i =(SKU i ,q i ,pi ),
[0023] Among them, SKU i represents the stock unit, q i Indicates the SKU demand quantity, p i Indicates order priority and calculates SKU demand weight. The calculation formula is:
[0024]
[0025] Among them, w i represents the demand weight of SKU, α is the demand weighting coefficient, β is the priority weighting coefficient, max(q) represents the maximum demand among all orders, and max(p) represents the highest priority among all orders.
[0026] Classify the historical order data and define the historical order data set as H:
[0027] H={h1,h2,…,h m},
[0028] Among them, H represents the historical order dataset, h i represents a single historical order data, m represents the total number of historical orders,
[0029] Classify SKU orders, the classification formula is:
[0030] C k ={SKU i ∣d i ∈H k},
[0031] Among them, C k represents SKU category k, H k Represents a subset of historical orders with similar SKU demand characteristics.
[0032] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, the step of generating SKU associated data is as follows:
[0033] Graph SAGE is used to calculate the SKU relevance and construct the SKU relationship graph G:
[0034] G=(V,E),
[0035] Among them, G represents the SKU relationship graph, V represents the SKU node set, and E represents the edge set between SKUs.
[0036] The SKU relevance edge weight is e ij :
[0037]
[0038] Among them, e ij represents the correlation weight between SKUi and SKUj, f(SKU i ,SKU j ) represents the co-occurrence frequency of SKUi and SKUj in historical orders, ∑ j f(SKU i ,SKU j ) represents the sum of the co-occurrence frequencies of SKUi and all SKUs,
[0039] GraphSAGE iteratively updates the SKU relevance, and the update formula is:
[0040]
[0041] in, represents the embedding vector of SKUi after the kth iteration, σ is a nonlinear activation function, W k represents the weight matrix of the kth layer, AGG (k) is the aggregation function at the kth level, represents the embedding vector of SKUj after the k-1th iteration, represents the set of adjacent nodes of SKUi,
[0042] Generate SKU association matrix M:
[0043]
[0044] Among them, M is the final SKU association matrix, K is the maximum number of iterations of Graph SAGE, and N is the total number of SKUs.
[0045] Optimize high-frequency SKU combinations based on correlation. The optimization adjustment formula is:
[0046]
[0047] Among them, S * is the optimized high-frequency SKU combination, S is the optional high-frequency SKU combination, e ij represents the correlation weight between SKUs, and T is the maximum capacity limit of the SKU combination.
[0048] As a preferred solution of the deep learning-based smart logistics digital warehouse management method described in the present invention, in which: in step S2, the storage layout is optimized in combination with the graph neural network GNN, the SKU storage area is dynamically adjusted to place the high-frequency SKU in the optimal storage position, and the optimized storage layout is generated.
[0049] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, the step of using the deep learning model Transformer to calculate the SKU storage optimization strategy and adjusting the SKU storage location through the attention mechanism is as follows:
[0050] The SKU associated data M generated in step S1 is used as the input of Transformer. Let the SKU storage state sequence be X:
[0051] X={x1,x2,…,x N},
[0052] Among them, X represents the SKU storage status sequence, x i Represents the position vector of SKUi in the current storage layout, N is the total number of SKUs,
[0053] Calculate the SKU position code and use the position coding mechanism to enable Transformer to capture the relative relationship of SKUs in the storage space. The formula is:
[0054]
[0055] Among them, PE (i,2j) is the 2j-th dimension component of the SKUi position vector, PE (i,2j+1) is the 2j+1th dimension component of the SKUi position vector, d is the SKU position encoding dimension, i is the SKU index, j is the dimension index,
[0056] A multi-head self-attention mechanism is used to calculate the storage optimization weights between SKUs. The calculation formula is:
[0057]
[0058] Among them, A is the SKU storage optimization weight matrix, Q = XW Q is the query matrix, K = XW K is the bond matrix, W Q ,W K is the trainable parameter matrix, d k is the attention dimension,
[0059] Calculate the SKU storage location adjustment vector using the formula:
[0060] Z=AV,
[0061] Where Z is the SKU position adjustment vector, V = XW V is the value matrix, W V is the trainable parameter matrix,
[0062] Generate storage layout L′, L′=argmax L(Z),
[0063] Among them, L′ is the optimized SKU storage layout, and L is all possible SKU storage solutions.
[0064] As a preferred solution of the deep learning-based smart logistics digital warehouse management method described in the present invention, the step of optimizing the storage layout by combining the graph neural network GNN is as follows:
[0065] Set the SKU storage graph and build a graph model for SKU storage optimization:
[0066] G s =(V s ,E s ),
[0067] Among them, G s is the SKU storage layout diagram, V s is the SKU storage node set, E s Optimize the edge set for storage between SKUs,
[0068] Calculate the weights between SKU storage areas using the following formula:
[0069]
[0070] in, The storage optimization weight between SKUi and SKUj, f s (x i ,x j ) represents the similarity between SKUi and SKUj in the storage area,
[0071] GNN iteratively updates the SKU position, and the update formula is:
[0072]
[0073] in, is the storage optimization vector of SKUi after the kth iteration, Optimize the weight matrix for layer k storage, AGG (k) is the k-th level aggregation function,
[0074] Generate optimized storage layout L * ,
[0075]
[0076] Among them, L * This is the final optimized SKU storage layout.
[0077] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, in step S3, the step of generating a picking path and a task scheduling plan is as follows:
[0078] Reinforcement learning is used to train the picking robot, which performs SKU picking tasks in the warehouse. The path optimization goal is defined as minimizing the total picking time T.
[0079] Set the warehouse environment to E:
[0080] E=(S,A,P,R),
[0081] Among them, E represents the warehouse environment, S is the state space, which represents the positions of different SKUs and the states of the robot, A is the action space, which represents the possible movement directions of the robot, P is the state transition probability, and R is the reward function.
[0082] Set the state to mean that the robot is in state s at time step t t :
[0083] s t =(x t ,y t ,v t ,SKU t ),
[0084] Among them, s t Indicates the current state of the robot, x t ,y t is the coordinate position of the robot in the warehouse, v t is the current speed of the robot, SKU t is the current target SKU,
[0085] Set the reward function R t ,
[0086] R t =-(λ1d t +λ2T t +λ3C t ),
[0087] Among them, R t is the reward of the current time step, d t is the current moving distance, T t is the cumulative picking time, C t is the number of path conflicts, λ1,λ2,λ3 are weight coefficients,
[0088] The reinforcement learning Qlearning training strategy is adopted, and the process formula is:
[0089]
[0090] Among them, Q(s t ,a t ) is the Q value, indicating the current state s t Select action a t The expected return, α is the learning rate, γ is the discount factor, a is the possible action at the current time step, belonging to the action space A,
[0091] Indicates that in state s t+1 Select the maximum Q value of the optimal action;
[0092] The step of generating a picking path and a task scheduling plan also includes:
[0093] Combine GNN to calculate the optimal allocation of AGV and manual picking tasks:
[0094] Construct task allocation graph G p :
[0095] G p =(V p ,E p ),
[0096] Among them, G p is the picking task allocation graph, V p is a set of task nodes, including AGV tasks and manual tasks, E p is the association edge between tasks,
[0097] Calculate the task priority using the following formula:
[0098]
[0099] Among them, P i is the priority of task i, w i is the task weight, d i is the task execution distance, ∈ is a smoothing term to prevent the denominator from being zero,
[0100] The GNN calculation task allocation strategy is adopted, and the calculation formula is:
[0101]
[0102] in, is the embedding vector of task i after the kth iteration, Assign a weight matrix to the k-th layer task,
[0103] Generate optimized picking path R * , expressed as:
[0104]
[0105] Among them, R * is the optimized picking path, where R is the optional picking path, and d i is the total moving distance of path i, T i is the total picking time of path i.
[0106] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, the steps of executing AGV and manual picking operations based on the picking path and task scheduling plan generated in step S3 are as follows:
[0107] Based on the optimized picking path R calculated in step S3 * , perform AGV and manual picking operations,
[0108] Set the picking task execution status to T pick ={t1,t2,…,t N},
[0109] Among them, T pick is the SKU picking task set, t i For the picking task of SKUi,
[0110] Calculate the AGV and manual task allocation ratio using the following formula:
[0111]
[0112] Among them, r AGV is the AGV task ratio, r Human is the proportion of manual picking tasks, N AGV The number of SKUs that the AGV is responsible for, N Human The number of SKUs that the person is responsible for;
[0113] Calculate the picking completion time:
[0114] For SKUi, the picking time is defined as T i :
[0115] T i =d i / v i +t proc,i ,
[0116] Among them, T i is the picking completion time of SKUi, d i is the distance from the storage location of SKUi to the picking point, v i is the picking speed of the corresponding execution subject, t proc,i is the picking processing time of SKUi,
[0117] Calculate the total time T of the picking task total, the formula is:
[0118] T total =max(T i ),
[0119] Among them, T total The total completion time for all SKU picking tasks;
[0120] In step S4, after picking is completed, the SKU picking time and path optimization results are recorded, and the picking path deviation is calculated. The calculation formula is:
[0121]
[0122] Where, Δd i is the deviation between the actual travel distance of SKUi and the theoretical optimal distance, is the actual travel distance of SKUi, The optimal path distance calculated for SKUi,
[0123] Calculate the actual picking time deviation using the formula:
[0124]
[0125] Where, ΔT i is the deviation between the actual picking time of SKUi and the theoretical optimal time, The actual picking time for SKUi. The optimal picking time calculated for SKUi,
[0126] Define the feedback data set as F:
[0127] F={(Δd i ,ΔT i )|i∈T pick},
[0128] Where F is the feedback data set, (Δd i ,ΔT i ) is the picking deviation data of SKUi;
[0129] The steps of dynamically adjusting the warehouse management strategy based on the feedback data using self-supervised learning, automatically optimizing the SKU storage layout, picking path and task scheduling, and applying them to the subsequent order sorting process are as follows:
[0130] Using self-supervised learning to dynamically optimize SKU storage layout:
[0131] In the feedback data set recorded in step S4, the features F required for storage layout optimization are extracted s :
[0132] Fs ={(Δd i ,ΔT i ,L i )|i∈T pick},
[0133] Among them, F s Optimize the feedback dataset for storage layout, Δd i is the deviation between the actual travel distance of SKUi and the optimal path distance, ΔT i is the deviation between the actual picking time of SKUi and the optimal time, L i is the storage location of SKUi,
[0134] Calculate the storage layout optimization loss and define the storage layout loss function as L store :
[0135]
[0136] Among them, L store Optimize the loss for storage layout, α s ,β s Optimize weight coefficients for storage,
[0137] Self-supervised learning is used to adjust the SKU storage location. The adjustment formula is:
[0138]
[0139] in, The updated storage location for SKUi. is the old storage location of SKUi, η s is the learning rate.
[0140] As a preferred solution of the deep learning-based intelligent logistics digital warehouse management method described in the present invention, in step S5, self-supervised learning is used to optimize the picking path, specifically:
[0141] Extract picking path optimization feedback data F p :
[0142] F p ={(Δd i ,ΔT i ,R i )|i∈T pick},
[0143] Among them, F p Feedback dataset for picking route optimization, R i is the picking path corresponding to SKUi,
[0144] Calculate the picking path optimization loss, the calculation formula is:
[0145]
[0146] Among them, L path Optimize the loss for picking paths, α p ,β p is the path optimization weight coefficient,
[0147] Self-supervised learning is used to optimize the picking path. The formula is:
[0148]
[0149] in, Updated picking path for SKUi, is the old picking path of SKUi, η p is the learning rate;
[0150] Optimizing task scheduling using self-supervised learning:
[0151] Define the task scheduling feedback data as F t :
[0152] F t ={(Δd i ,ΔT i ,P i )|i∈T pick},
[0153] Among them, F t Optimizing feedback dataset for task scheduling, P i is the task scheduling strategy corresponding to SKUi,
[0154] Calculate the task scheduling optimization loss, the calculation formula is:
[0155]
[0156] Among them, L task Optimize the loss for task scheduling, α t ,β t Optimize weight coefficients for task scheduling,
[0157] Self-supervised learning is used to optimize task scheduling. The optimization formula is:
[0158]
[0159] in, Updated task scheduling strategy for SKUi, is the old task scheduling strategy of SKUi, η t is the learning rate,
[0160] The optimized storage layout, picking path and task scheduling strategy are applied to subsequent order processing.
[0161] The beneficial effects of the present invention are as follows: the present invention calculates the correlation of SKUs through Graph SAGE, constructs a SKU association matrix, and optimizes the SKU combination based on the SKU co-occurrence frequency to improve the SKU storage efficiency; uses Transformer to calculate the SKU storage optimization strategy, analyzes the storage relationship between SKUs through the attention mechanism, and combines GNN to optimize the storage layout, so that high-frequency SKUs can be placed in the optimal storage area and the picking path length can be reduced; in the picking task scheduling link, reinforcement learning is used to train the picking robot, which autonomously learns the optimal picking path in the warehouse environment, combines GNN to calculate the optimal allocation of AGV and manual tasks, balances the collaborative operation of automated equipment and manual operations, and effectively reduces path conflicts and invalid movements.
[0162] After the task is completed, the present invention records the picking time and path optimization results, and uses the feedback data for self-supervised learning to optimize the warehouse management strategy, so that SKU storage, picking paths and task scheduling can be dynamically adjusted to adapt to changes in market demand.
[0163] In summary, the present invention realizes adaptive adjustment to the dynamic changes of SKUs, improves the intelligence level of warehousing operations, makes the storage layout more compact, the picking path shorter, and the scheduling and allocation more reasonable, thereby significantly improving the picking efficiency and order fulfillment capabilities, while reducing warehouse operating costs and improving the intelligence level of the overall logistics system. BRIEF DESCRIPTION OF THE DRAWINGS
[0164] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without paying creative work.
[0165] Figure 1 This is a flow chart of the deep learning-based intelligent logistics digital warehouse management method of the present invention. DETAILED DESCRIPTION
[0166] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific implementation methods of the present invention are described in detail below in conjunction with the accompanying drawings.
[0167] In the following description, many specific details are set forth to facilitate a full understanding of the present invention, but the present invention may also be implemented in other ways different from those described herein, and those skilled in the art may make similar generalizations without violating the connotation of the present invention. Therefore, the present invention is not limited to the specific embodiments disclosed below.
[0168] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The term "in one embodiment" that appears in different places in this specification does not necessarily refer to the same embodiment, nor does it refer to a separate or selective embodiment that is mutually exclusive with other embodiments.
[0169] Example 1, reference Figure 1 , this embodiment provides a smart logistics digital warehouse management method based on deep learning, including the following steps:
[0170] Step S1, obtaining order data, parsing the SKU information, quantity requirements and priority in the order data, classifying them in combination with historical order data, and generating SKU related data;
[0171] In step S1, the node feature learning algorithm GraphSAGE is used to calculate the SKU correlation, generate a SKU correlation matrix, and optimize the high-frequency SKU combination based on the SKU correlation to generate optimized SKU correlation data;
[0172] The steps to obtain order data, parse the SKU information, quantity requirements and priority in the order data, and classify it based on historical order data are as follows:
[0173] Parse the order data and define the order data set as D:
[0174] D={d1,d2,…,d n},
[0175] Where D represents the order data set, d i Represents a single order data, n represents the total number of orders,
[0176] The order data structure is d i :
[0177] d i =(SKU i ,q i ,p i ),
[0178] Among them, SKU i represents the stock unit, q i Indicates the SKU demand quantity, p i Indicates the order priority.
[0179] Calculate the SKU demand weight using the following formula:
[0180]
[0181] Among them, w i represents the demand weight of SKU, α is the demand weighting coefficient, β is the priority weighting coefficient, max(q) represents the maximum demand among all orders, and max(p) represents the highest priority among all orders.
[0182] Classify the historical order data and define the historical order data set as H:
[0183] H={h1,h2,…,h m},
[0184] Among them, H represents the historical order dataset, h i represents a single historical order data, m represents the total number of historical orders,
[0185] Classify SKU orders, the classification formula is:
[0186] C k ={SKU i ∣d i ∈H k},
[0187] Among them, C k represents SKU category k, H k represents a subset of historical orders with similar SKU demand characteristics,
[0188] The steps to generate SKU associated data are:
[0189] Graph SAGE is used to calculate the SKU relevance and construct the SKU relationship graph G:
[0190] G=(V,E),
[0191] Among them, G represents the SKU relationship graph, V represents the SKU node set, and E represents the edge set between SKUs.
[0192] The SKU relevance edge weight is e ij :
[0193]
[0194] Among them, e ij represents the correlation weight between SKUi and SKUj, f(SKU i ,SKU j ) represents the co-occurrence frequency of SKUi and SKUj in historical orders, Σj f(SKU i ,SKU j ) represents the sum of the co-occurrence frequencies of SKUi and all SKUs,
[0195] GraphSAGE iteratively updates the SKU relevance, and the update formula is:
[0196]
[0197] in, represents the embedding vector of SKUi after the kth iteration, σ is a nonlinear activation function, W k represents the weight matrix of the kth layer, AGG (k) is the aggregation function at the kth level, represents the embedding vector of SKUj after the k-1th iteration, represents the set of adjacent nodes of SKUi,
[0198] Generate SKU association matrix M:
[0199]
[0200] Among them, M is the final SKU association matrix, K is the maximum number of iterations of Graph SAGE, and N is the total number of SKUs.
[0201] Optimize high-frequency SKU combinations based on correlation. The optimization adjustment formula is:
[0202]
[0203] Among them, S * is the optimized high-frequency SKU combination, S is the optional high-frequency SKU combination, e ij represents the correlation weight between SKUs, T is the maximum capacity limit of SKU combination,
[0204] Specifically, in step S1, the order data is parsed to extract SKU information, demand quantity and priority, and classified in combination with historical order data to optimize SKU storage management and picking scheduling; a SKU relationship graph is constructed, and the correlation between SKUs is calculated through the Graph SAGE algorithm to generate a SKU association matrix. By optimizing high-frequency SKU combinations, the SKU combination storage efficiency is improved, and storage space fragmentation and SKU picking time are reduced;
[0205] Step S2, based on the SKU associated data generated in step S1, the deep learning model Transformer is used to calculate the SKU storage optimization strategy, the SKU storage location is adjusted through the attention mechanism, and the storage layout is generated;
[0206] In step S2, the storage layout is optimized by combining the graph neural network GNN, dynamically adjusting the SKU storage area so that the high-frequency SKU is in the optimal storage position, and generating the optimized storage layout;
[0207] The deep learning model Transformer is used to calculate the SKU storage optimization strategy. The steps of adjusting the SKU storage location through the attention mechanism are as follows:
[0208] The SKU associated data M generated in step S1 is used as the input of Transformer. Let the SKU storage state sequence be X:
[0209] X={x1,x2,…,x N},
[0210] Among them, X represents the SKU storage status sequence, x i Represents the position vector of SKUi in the current storage layout, N is the total number of SKUs,
[0211] Calculate the SKU position code and use the position coding mechanism to enable Transformer to capture the relative relationship of SKUs in the storage space. The formula is:
[0212]
[0213] Among them, PE (i,2j) is the 2j-th dimension component of the SKUi position vector, PE (i,2j+1) is the 2j+1th dimension component of the SKUi position vector, d is the SKU position encoding dimension, i is the SKU index, j is the dimension index,
[0214] A multi-head self-attention mechanism is used to calculate the storage optimization weights between SKUs. The calculation formula is:
[0215]
[0216] Among them, A is the SKU storage optimization weight matrix, Q = XW Q is the query matrix, K = XW K is the bond matrix, W Q ,W K is the trainable parameter matrix, d k is the attention dimension,
[0217] Calculate the SKU storage location adjustment vector using the formula:
[0218] Z=AV,
[0219] Where Z is the SKU position adjustment vector, V = XW V is the value matrix, W V is the trainable parameter matrix,
[0220] Generate storage layout L′, L′=argmax L (Z),
[0221] Among them, L′ is the optimized SKU storage layout, and L is all possible SKU storage solutions;
[0222] The steps for optimizing storage layout by combining graph neural network GNN are:
[0223] Set the SKU storage graph and build a graph model for SKU storage optimization:
[0224] G s =(V s ,E s ),
[0225] Among them, G s is the SKU storage layout diagram, V s is the SKU storage node set, E s Optimize the edge set for storage between SKUs,
[0226] Calculate the weights between SKU storage areas using the following formula:
[0227]
[0228] in, The storage optimization weight between SKUi and SKUj, f s (x i ,x j ) represents the similarity between SKUi and SKUj in the storage area,
[0229] GNN iteratively updates the SKU position, and the update formula is:
[0230]
[0231] in, is the storage optimization vector of SKUi after the kth iteration, Optimize the weight matrix for layer k storage, AGG (k) is the k-th level aggregation function,
[0232] Generate optimized storage layout L * ,
[0233]
[0234] Among them, L * The final optimized SKU storage layout;
[0235] Specifically, Transformer is used here to calculate the SKU storage optimization strategy, and the SKU storage location is adjusted based on the attention mechanism to improve the picking efficiency; GNN is combined to optimize the storage layout so that high-frequency SKUs are in the optimal storage area, reducing the SKU access time and the picking path length;
[0236] Here, the SKU storage area is dynamically adjusted to match the SKU association with the storage layout, improve the continuity of the SKU picking path, and reduce the movement cost within the warehouse;
[0237] Step S3, based on the optimized storage layout generated in step S2, using reinforcement learning to train the picking robot to learn the optimal picking path;
[0238] Combine GNN to calculate the optimal allocation of AGV and manual picking tasks, and generate picking paths and task scheduling plans;
[0239] In step S3, the steps of generating the picking path and task scheduling plan are as follows:
[0240] Reinforcement learning is used to train the picking robot, which performs SKU picking tasks in the warehouse. The path optimization goal is defined as minimizing the total picking time T.
[0241] Set the warehouse environment to E:
[0242] E=(S,A,P,R),
[0243] Among them, E represents the warehouse environment, S is the state space, which represents the positions of different SKUs and the states of the robot, A is the action space, which represents the possible movement directions of the robot, P is the state transition probability, and R is the reward function.
[0244] Set the state to mean that the robot is in state s at time step t t :
[0245] s t =(x t ,y t ,v t ,SKU t ),
[0246] Among them, s t Indicates the current state of the robot, x t ,y t is the coordinate position of the robot in the warehouse, v t is the current speed of the robot, SKU t is the current target SKU,
[0247] Set the reward function R t ,
[0248] Rt =-(λ1d t +λ2T t +λ3C t ),
[0249] Among them, R t is the reward of the current time step, d t is the current moving distance, T t is the cumulative picking time, C t is the number of path conflicts, λ1,λ2,λ3 are weight coefficients,
[0250] The reinforcement learning Qlearning training strategy is adopted, and the process formula is:
[0251]
[0252] Among them, Q(s t ,a t ) is the Q value, indicating the current state s t Select action a t The expected return, α is the learning rate, γ is the discount factor, a is the possible action at the current time step, belonging to the action space A,
[0253] Indicates that in state s t+1 Select the maximum Q value of the optimal action;
[0254] The steps of generating picking routes and task scheduling plans also include:
[0255] Combine GNN to calculate the optimal allocation of AGV and manual picking tasks:
[0256] Construct task allocation graph G p :
[0257] G p =(V p ,E p ),
[0258] Among them, G p is the picking task allocation graph, V p is a set of task nodes, including AGV tasks and manual tasks, E p is the association edge between tasks,
[0259] Calculate the task priority using the following formula:
[0260]
[0261] Among them, P i is the priority of task i, w i is the task weight, d iis the task execution distance, ∈ is a smoothing term to prevent the denominator from being zero,
[0262] The GNN calculation task allocation strategy is adopted, and the calculation formula is:
[0263]
[0264] in, is the embedding vector of task i after the kth iteration, Assign a weight matrix to the k-th layer task,
[0265] Generate optimized picking path R * , expressed as:
[0266]
[0267] Among them, R * is the optimized picking path, where R is the optional picking path, and d i is the total moving distance of path i, T i is the total picking time for path i;
[0268] Specifically, reinforcement learning is used here to train the picking robot to learn the optimal picking path to minimize the SKU picking time. At the same time, GNN is combined to calculate the optimal allocation of AGV and manual picking tasks, balancing the picking efficiency and the rationality of human-machine collaboration; making the execution of picking tasks more intelligent and efficient;
[0269] Step S4, executing AGV and manual picking operations based on the picking path and task scheduling plan generated in step S3;
[0270] After picking is completed, the SKU picking time and path optimization results are recorded, and the data is passed to step S5 as feedback data;
[0271] Based on the picking path and task scheduling plan generated in step S3, the steps for executing AGV and manual picking operations are:
[0272] Based on the optimized picking path R calculated in step S3 * , perform AGV and manual picking operations,
[0273] Set the picking task execution status to T pick ={t1,t2,…,t N},
[0274] Among them, T pick is the SKU picking task set, t i For the picking task of SKUi,
[0275] Calculate the AGV and manual task allocation ratio using the following formula:
[0276]
[0277] Among them, r AGV is the AGV task ratio, r Human is the proportion of manual picking tasks, N AGV The number of SKUs that the AGV is responsible for, N Human The number of SKUs that the person is responsible for;
[0278] Calculate the picking completion time:
[0279] For SKUi, the picking time is defined as T i :
[0280] T i =d i / v i +t proc,i ,
[0281] Among them, T i is the picking completion time of SKUi, d i is the distance from the storage location of SKUi to the picking point, v i is the picking speed of the corresponding execution subject, t proc,i is the picking processing time of SKUi,
[0282] Calculate the total time T of the picking task total , the formula is:
[0283] T total =max(T i ),
[0284] Among them, T total The total completion time for all SKU picking tasks;
[0285] In step S4, after picking is completed, the SKU picking time and path optimization results are recorded, and the picking path deviation is calculated. The calculation formula is:
[0286]
[0287] Where, Δd i is the deviation between the actual travel distance of SKUi and the theoretical optimal distance, is the actual travel distance of SKUi, The optimal path distance calculated for SKUi,
[0288] Calculate the actual picking time deviation using the formula:
[0289]
[0290] Where, ΔTi is the deviation between the actual picking time of SKUi and the theoretical optimal time, The actual picking time for SKUi. The optimal picking time calculated for SKUi,
[0291] Define the feedback data set as F:
[0292] F={(Δd i ,ΔT i )|i∈T pick},
[0293] Where F is the feedback data set, (Δd i ,ΔT i ) is the picking deviation data of SKUi;
[0294] Based on the feedback data, the warehouse management strategy is dynamically adjusted using self-supervised learning to automatically optimize the SKU storage layout, picking path and task scheduling, and then applied to the subsequent order sorting process.
[0295] Using self-supervised learning to dynamically optimize SKU storage layout:
[0296] In the feedback data set recorded in step S4, the features F required for storage layout optimization are extracted s :
[0297] F s ={(Δd i ,ΔT i ,L i )|i∈T pick},
[0298] Among them, F s Optimize the feedback dataset for storage layout, Δd i is the deviation between the actual travel distance of SKUi and the optimal path distance, ΔT i is the deviation between the actual picking time of SKUi and the optimal time, L i is the storage location of SKUi,
[0299] Calculate the storage layout optimization loss and define the storage layout loss function as L store :
[0300]
[0301] Among them, L store Optimize the loss for storage layout, α s ,β s Optimize weight coefficients for storage,
[0302] Self-supervised learning is used to adjust the SKU storage location. The adjustment formula is:
[0303]
[0304] in, The updated storage location for SKUi. is the old storage location of SKUi, η s is the learning rate;
[0305] Specifically, based on the picking path and task scheduling plan generated in step S3, AGV and manual picking operations are executed, and the SKU picking time and path optimization results are recorded, the deviation between the actual picking path and the optimal path is calculated, and the time consumption of the picking task is analyzed, so as to obtain the deviation information that occurs during the actual execution process;
[0306] Step S5, based on the feedback data, using self-supervised learning to dynamically adjust the warehouse management strategy, automatically optimize the SKU storage layout, picking path and task scheduling, and apply it to the subsequent order sorting process;
[0307] In step S5, self-supervised learning is used to optimize the picking path. Specifically:
[0308] Extract picking path optimization feedback data F p :
[0309] F p ={(Δd i ,ΔT i ,R i )|i∈T pick},
[0310] Among them, F p Feedback dataset for picking route optimization, R i is the picking path corresponding to SKUi,
[0311] Calculate the picking path optimization loss, the calculation formula is:
[0312]
[0313] Among them, L path Optimize the loss for picking paths, α p ,β p is the path optimization weight coefficient,
[0314] Self-supervised learning is used to optimize the picking path. The formula is:
[0315]
[0316] in, Updated picking path for SKUi, is the old picking path of SKUi, ηp is the learning rate;
[0317] Optimizing task scheduling using self-supervised learning:
[0318] Define the task scheduling feedback data as F t :
[0319] F t ={(Δd i ,ΔT i ,P i )|i∈T pick},
[0320] Among them, F t Optimizing feedback dataset for task scheduling, P i is the task scheduling strategy corresponding to SKUi,
[0321] Calculate the task scheduling optimization loss, the calculation formula is:
[0322]
[0323] Among them, L task Optimize the loss for task scheduling, α t ,β t Optimize weight coefficients for task scheduling,
[0324] Self-supervised learning is used to optimize task scheduling. The optimization formula is:
[0325]
[0326] in, Updated task scheduling strategy for SKUi, is the old task scheduling strategy of SKUi, η t is the learning rate,
[0327] The optimized storage layout, picking path and task scheduling strategy are applied to subsequent order processing;
[0328] Specifically, the feedback data generated in step S4 is used here to dynamically optimize the SKU storage layout, picking path and task scheduling strategy using self-supervised learning; the warehouse optimization strategy can be dynamically adjusted according to real-time data, continuously improving warehouse operation efficiency and reducing logistics costs.
[0329] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A smart logistics digital warehouse management method based on deep learning, characterized by: include, Step S1, obtaining order data, parsing the SKU information, quantity requirements and priority in the order data, classifying them in combination with historical order data, and generating SKU related data; Step S2, based on the SKU associated data generated in step S1, the deep learning model Transformer is used to calculate the SKU storage optimization strategy, the SKU storage location is adjusted through the attention mechanism, and the storage layout is generated; Step S3, based on the optimized storage layout generated in step S2, using reinforcement learning to train the picking robot to learn the optimal picking path; Combine GNN to calculate the optimal allocation of AGV and manual picking tasks, and generate picking paths and task scheduling plans; Step S4, executing AGV and manual picking operations based on the picking path and task scheduling plan generated in step S3; After picking is completed, the SKU picking time and path optimization results are recorded, and the data is passed to step S5 as feedback data; Step S5, based on the feedback data, using self-supervised learning to dynamically adjust the warehouse management strategy, automatically optimize the SKU storage layout, picking path and task scheduling, and apply it to the subsequent order sorting process.
2. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 1, characterized in that: In step S1, the node feature learning algorithm GraphSAGE is used to calculate the SKU correlation, generate a SKU correlation matrix, and optimize the high-frequency SKU combination based on the SKU correlation to generate optimized SKU correlation data.
3. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 2, characterized in that: The steps of obtaining order data, parsing the SKU information, quantity requirements and priority in the order data, and classifying in combination with historical order data are as follows: Parse the order data and define the order data set as D: D={d1,d2,…,d n }, Where D represents the order data set, d i Represents a single order data, n represents the total number of orders, The order data structure is d i : d i =(SKU i ,q i ,p i ), Among them, SKU i represents the stock unit, q i Indicates the SKU demand quantity, p i Indicates the order priority. Calculate the SKU demand weight using the following formula: Among them, w i represents the demand weight of SKU, α is the demand weighting coefficient, β is the priority weighting coefficient, max(q) represents the maximum demand among all orders, and max(p) represents the highest priority among all orders. Classify the historical order data and define the historical order data set as H: H={h1,h2,…,h m }, Among them, H represents the historical order dataset, h i represents a single historical order data, m represents the total number of historical orders, Classify SKU orders, the classification formula is: C k ={SKU i ∣d i ∈H k }, Among them, C k represents SKU category k, H k Represents a subset of historical orders with similar SKU demand characteristics.
4. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 3, characterized in that: The steps of generating SKU associated data are: Graph SAGE is used to calculate the SKU relevance and construct the SKU relationship graph G: G=(V,E), Among them, G represents the SKU relationship graph, V represents the SKU node set, and E represents the edge set between SKUs. The SKU relevance edge weight is e ij : Among them, e ij represents the correlation weight between SKUi and SKUj, f(SKU i ,SKU j ) represents the co-occurrence frequency of SKUi and SKUj in historical orders, Σ j f(SKU i ,SKU j ) represents the sum of the co-occurrence frequencies of SKUi and all SKUs, GraphSAGE iteratively updates the SKU relevance, and the update formula is: in, represents the embedding vector of SKUi after the kth iteration, σ is a nonlinear activation function, W k represents the weight matrix of the kth layer, AGG (k) is the aggregation function at the kth level, represents the embedding vector of SKUj after the k-1th iteration, represents the set of adjacent nodes of SKUi, Generate SKU association matrix M: Among them, M is the final SKU association matrix, K is the maximum number of iterations of Graph SAGE, and N is the total number of SKUs. Optimize high-frequency SKU combinations based on correlation. The optimization adjustment formula is: Among them, S * is the optimized high-frequency SKU combination, S is the optional high-frequency SKU combination, e ij represents the correlation weight between SKUs, and T is the maximum capacity limit of the SKU combination.
5. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 4, characterized in that: In step S2, the storage layout is optimized in combination with the graph neural network GNN, the SKU storage area is dynamically adjusted to place high-frequency SKUs in the optimal storage location, and an optimized storage layout is generated.
6. A smart logistics digital warehouse management method based on deep learning as claimed in claim 5, characterized in that: The steps of using the deep learning model Transformer to calculate the SKU storage optimization strategy and adjusting the SKU storage location through the attention mechanism are as follows: The SKU associated data M generated in step S1 is used as the input of Transformer. Let the SKU storage state sequence be X: X={x1,x2,…,x N }, Among them, X represents the SKU storage status sequence, x i Represents the position vector of SKUi in the current storage layout, N is the total number of SKUs, Calculate the SKU position code and use the position coding mechanism to enable Transformer to capture the relative relationship of SKUs in the storage space. The formula is: Among them, PE (i,2j) is the 2j-th dimension component of the SKUi position vector, PE (i,2j+1 ) is the 2j+1th dimension component of the SKUi position vector, d is the SKU position encoding dimension, i is the SKU index, j is the dimension index, A multi-head self-attention mechanism is used to calculate the storage optimization weights between SKUs. The calculation formula is: Among them, A is the SKU storage optimization weight matrix, Q = XW Q is the query matrix, K = XW K is the bond matrix, W Q ,W K is the trainable parameter matrix, d k is the attention dimension, Calculate the SKU storage location adjustment vector using the formula: Z=AV, Where Z is the SKU position adjustment vector, V = XW V is the value matrix, W V is the trainable parameter matrix, Generate storage layout L′, L′=argmax L (Z), Among them, L′ is the optimized SKU storage layout, and L is all possible SKU storage solutions.
7. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 6, characterized in that: The steps of optimizing storage layout by combining the graph neural network GNN are: Set the SKU storage graph and build a graph model for SKU storage optimization: G s =(V s ,E s ), Among them, G s is the SKU storage layout diagram, V s is the SKU storage node set, E s Optimize the edge set for storage between SKUs, Calculate the weights between SKU storage areas using the following formula: in, The storage optimization weight between SKUi and SKUj, f s (x i ,x j ) represents the similarity between SKUi and SKUj in the storage area, GNN iteratively updates the SKU position, and the update formula is: in, is the storage optimization vector of SKUi after the kth iteration, Optimize the weight matrix for layer k, AGG (k) is the k-th level aggregation function, Generate optimized storage layout L * , Among them, L * This is the final optimized SKU storage layout.
8. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 7, characterized in that: In step S3, the step of generating a picking path and a task scheduling plan is as follows: Reinforcement learning is used to train the picking robot, which performs SKU picking tasks in the warehouse. The path optimization goal is defined as minimizing the total picking time T. Set the warehouse environment to E: E=(S,A,P,R), Among them, E represents the warehouse environment, S is the state space, which represents the positions of different SKUs and the states of the robot, A is the action space, which represents the possible movement directions of the robot, P is the state transition probability, and R is the reward function. Set the state to mean that the robot is in state s at time step t t : with t =(x t ,y r ,in t ,SKU t ), Among them, s t Indicates the current state of the robot, x t ,y t is the coordinate position of the robot in the warehouse, v t is the current speed of the robot, SKU t is the current target SKU, Set the reward function R t , R t =-(λ1d t +λ2T t +λ3C t ), Among them, R t is the reward of the current time step, d t is the current moving distance, T t is the cumulative picking time, C t is the number of path conflicts, λ1,λ2,λ3 are weight coefficients, The reinforcement learning Qlearning training strategy is adopted, and the process formula is: Among them, Q(s t ,a t ) is the Q value, indicating the current state s t Select action a t The expected return, α is the learning rate, γ is the discount factor, a is the possible action at the current time step, belonging to the action space A, Indicates that in state s t+1 Select the maximum Q value of the optimal action; The step of generating a picking path and a task scheduling plan also includes: Combine GNN to calculate the optimal allocation of AGV and manual picking tasks: Construct task allocation graph G p : G p =(V p ,E p ), Among them, G p is the picking task allocation graph, V p is a set of task nodes, including AGV tasks and manual tasks, E p is the association edge between tasks, Calculate the task priority using the following formula: Among them, P i is the priority of task i, w i is the task weight, d i is the task execution distance, ∈ is a smoothing term to prevent the denominator from being zero, The GNN calculation task allocation strategy is adopted, and the calculation formula is: in, is the embedding vector of task i after the kth iteration, Assign a weight matrix to the k-th layer task, Generate optimized picking path R * , expressed as: Among them, R * is the optimized picking path, where R is the optional picking path, and d i is the total moving distance of path i, T i is the total picking time of path i.
9. A smart logistics digital warehouse management method based on deep learning as claimed in claim 8, characterized in that: The steps of executing AGV and manual picking operations based on the picking path and task scheduling plan generated in step S3 are: Based on the optimized picking path R calculated in step S3 * , perform AGV and manual picking operations, Set the picking task execution status to T pick ={t1,t2,…,t N }, Among them, T pick is the SKU picking task set, t i For the picking task of SKUi, Calculate the AGV and manual task allocation ratio using the following formula: Among them, r AGV is the AGV task ratio, r Human is the proportion of manual picking tasks, N AGV The number of SKUs that the AGV is responsible for, N Human The number of SKUs that the person is responsible for; Calculate the picking completion time: For SKUi, the picking time is defined as T i : T i =d i / v i +t proc,i , Among them, T i is the picking completion time of SKUi, d i is the distance from the storage location of SKUi to the picking point, v i is the picking speed of the corresponding execution subject, t proc,i is the picking processing time of SKUi, Calculate the total time T of the picking task total , the formula is: T total =max(T i ), Among them, T total The total completion time for all SKU picking tasks; In step S4, after picking is completed, the SKU picking time and path optimization results are recorded, and the picking path deviation is calculated. The calculation formula is: Where, Δd i is the deviation between the actual travel distance of SKUi and the theoretical optimal distance, is the actual travel distance of SKUi, The optimal path distance calculated for SKUi, Calculate the actual picking time deviation using the formula: Where, ΔT i is the deviation between the actual picking time of SKUi and the theoretical optimal time, The actual picking time for SKUi. The optimal picking time calculated for SKUi, Define the feedback data set as F: F={(Δd i ,ΔT i )∣i∈T pick }, Where F is the feedback data set, (Δd i ,ΔT i ) is the picking deviation data of SKUi; The steps of dynamically adjusting the warehouse management strategy based on the feedback data using self-supervised learning, automatically optimizing the SKU storage layout, picking path and task scheduling, and applying them to the subsequent order sorting process are: Using self-supervised learning to dynamically optimize SKU storage layout: In the feedback data set recorded in step S4, the features F required for storage layout optimization are extracted s : F s ={(Δd i ,ΔT i ,L i )∣i∈T pick }, Among them, F s Optimize the feedback dataset for storage layout, Δd i is the deviation between the actual travel distance of SKUi and the optimal path distance, ΔT i is the deviation between the actual picking time of SKUi and the optimal time, L i is the storage location of SKUi, Calculate the storage layout optimization loss and define the storage layout loss function as L store : Among them, L store Optimize the loss for storage layout, α s ,β d Optimize weight coefficients for storage, Self-supervised learning is used to adjust the SKU storage location. The adjustment formula is: in, The updated storage location for SKUi. is the old storage location of SKUi, η s is the learning rate.
10. The method for managing a smart logistics digital warehouse based on deep learning as claimed in claim 9, characterized in that: In step S5, self-supervised learning is used to optimize the picking path. Specifically: Extract picking path optimization feedback data F p : F p ={(Δd i ,ΔT i ,R i )∣i∈T pick }, Among them, F p Feedback dataset for picking route optimization, R i is the picking path corresponding to SKUi, Calculate the picking path optimization loss, the calculation formula is: Among them, L path Optimize the loss for picking paths, α p ,β p is the path optimization weight coefficient, Self-supervised learning is used to optimize the picking path. The formula is: in, Updated picking path for SKUi, is the old picking path of SKUi, η p is the learning rate; Optimizing task scheduling using self-supervised learning: Define the task scheduling feedback data as F t : F t ={(Δd i ,ΔT i ,P i )∣i∈T pick }, Among them, F t Optimizing feedback dataset for task scheduling, P i is the task scheduling strategy corresponding to SKUi, Calculate the task scheduling optimization loss, the calculation formula is: Among them, L task Optimize the loss for task scheduling, α t ,β t Optimize weight coefficients for task scheduling, Self-supervised learning is used to optimize task scheduling. The optimization formula is: in, Updated task scheduling strategy for SKUi, is the old task scheduling strategy of SKUi, η t is the learning rate, The optimized storage layout, picking path and task scheduling strategy are applied to subsequent order processing.
Citation Information
Cited By
Supply chain warehouse management system and method based on AI digital intelligence
CN120374018A
Vertical lifting type sorting warehouse and control method thereof
CN120841063A
Slot allocation strategy optimization method and system based on reinforcement learning
CN121257873A
In-warehouse sorting operation system simulation optimization method, equipment and medium
CN121580877A
Warehouse management control system and method based on reinforcement learning
CN122312038A