Warehouse goods allocation dynamic optimization method based on AI
By acquiring multidimensional data to calculate the correlation between goods and the expected picking time, a reinforcement learning model is constructed to dynamically optimize the storage location. This solves the problem that traditional warehouse storage location management methods cannot adapt to market changes, and achieves efficient, adaptive, and globally optimal layout of the warehousing system.
Patent Information
- Application Number
- CN202511598458.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-04
- Publication Date
- 2026-01-13
AI Technical Summary
Traditional warehouse location management methods cannot adapt to dynamic changes in market demand, resulting in excessively long picking paths, low operational efficiency, difficulty in effectively utilizing the correlation between goods, increased labor and time costs, and limitations on the responsiveness and flexibility of the warehousing system.
By acquiring multidimensional data, calculating the correlation values between goods and the expected picking time, a reinforcement learning model is constructed. Using the time-sensitive location-goods matching score as the basis for decision-making, a dynamic location optimization plan is output. Combined with the reward function, the handling decision is guided to achieve global and long-term intelligent decision-making.
It enables intelligent adjustment of cargo layout based on real-time and future business dynamics, significantly shortens order picking paths, improves warehousing operation efficiency and timeliness, avoids frequent ineffective handling caused by short-sighted decisions, and ensures that the warehouse layout remains in an optimal state in a constantly changing environment.
Smart Images

Figure CN121329281A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of intelligent warehousing and logistics management, and particularly relates to a warehouse storage space dynamic optimization method based on AI. BACKGROUND
[0002] Warehouse management is a core link in supply chain management, and the efficiency of storage space management directly affects the storage density, order picking speed and overall operating cost of the warehouse. Traditional warehouse storage space management methods mostly adopt static or semi-static strategies, such as dividing goods into ABC categories according to their historical outbound frequency and allocating them to fixed areas. Although this method is simple and easy to implement, its disadvantages are also very obvious: they cannot adapt to dynamic changes in market demand, such as seasonal fluctuations, promotional activities or sudden orders. When business data changes, fixed storage space allocation strategies can lead to long picking paths, low operational efficiency, and difficulty in effectively utilizing the correlation between goods, thereby increasing unnecessary labor and time costs and limiting the overall response capability and flexibility of the warehouse system. Therefore, how to intelligently and prospectively adjust the layout of goods according to real-time and future business dynamics to achieve global optimization of warehouse space and operational efficiency is a technical problem that needs to be solved in the current warehouse automation and intelligentization field. SUMMARY
[0003] In view of the above problems existing in the prior art, the present application aims to provide a warehouse storage space dynamic optimization method based on AI, comprising the following steps: Step S1, obtaining multi-dimensional data, the multi-dimensional data including goods information, warehouse storage space information, and historical business data and future business data.
[0004] Step S2, based on the multi-dimensional data, calculating the goods correlation value between any two goods, and determining the expected picking time of each kind of goods according to the future business data.
[0005] Step S3, for each combination of a target warehouse storage space and a to-be-allocated good in the warehouse, calculating a time-sensitive storage space-goods matching degree score at the current time, the score being used to quantify the degree of appropriateness of placing the to-be-allocated good in the target warehouse storage space at the current time; Step S4, constructing and using a reinforcement learning model, taking the time-sensitive storage space-goods matching degree score as a key basis for decision-making of the reinforcement learning model, and outputting a dynamic storage space optimization scheme.
[0006] Further, the time-sensitive storage space-goods matching degree score is calculated according to the following formula, wherein represents the current time the goods Time-sensitive location-goods matching score for placement in the target storage location; Indicates the index of the target storage location. Indicates the index of goods to be allocated. Indicates goods The future basic outbound frequency; This represents the collection of goods already stored within a predefined physical neighborhood of the target storage location. Represents a set One of the goods indexes. Indicates goods With goods The correlation value between goods Indicates the target storage location To the preset warehouse operation baseline point The physical distance; Represent the natural logarithm function; Indicates with goods Related, depending on the current moment Dynamically changing time urgency factor.
[0007] Furthermore, the time urgency factor is calculated according to the following formula: In the formula: Indicates the expected picking time for the goods; A preset urgency sensitivity coefficient; is a very small positive constant used to prevent the denominator from being zero; exp is the natural exponential function.
[0008] Furthermore, the calculation of the future basic outbound frequency specifically includes: statistically analyzing the historical outbound frequency of goods based on the historical business data, and applying a time decay factor to the historical outbound frequency; extracting the future outbound demand of goods from the future business data; and fusing the historical outbound frequency processed by the time decay factor with the future outbound demand to obtain the future basic outbound frequency.
[0009] Furthermore, the calculation of the cargo correlation value includes the following steps: Step S101: The outbound order records extracted from the historical business data are formatted into a transaction dataset, where each outbound order corresponds to a transaction, and the transaction contains all the items of the goods in the order.
[0010] Step S102: Based on a preset minimum support threshold, analyze the transaction dataset to identify all combinations of goods with a frequency not lower than the minimum support threshold, as frequent itemsets.
[0011] Step S103: Based on the frequent itemsets and the preset minimum confidence threshold, generate association rules and filter out rules with a confidence level not lower than the minimum confidence threshold to form a strong association rule set.
[0012] Step S104: For any two goods k and j, according to the set of strong association rules, multiply the support of their co-occurrence by the larger of the confidence scores from goods k to goods j and from goods j to goods k, and the resulting value is the goods association value.
[0013] Furthermore, the reinforcement learning model uses a reward function R to guide its learning process, which is specifically defined as: ,in The change in the sum of the time-sensitive location-cargo matching scores of all occupied storage locations in the warehouse after a handling instruction is executed.
[0014] A physical handling cost is determined by the weight of the goods in the handling instruction, the handling distance, and the type of equipment required to perform the handling.
[0015] The reward is a high-urgency task reward whose value is proportional to the reduction in the expected picking time of the goods brought about by the handling instruction, and the reward is calculated only when the time urgency factor of the goods being handled is higher than a preset threshold.
[0016] Furthermore, the reinforcement learning model selects the optimal transport instruction by calculating an action value function Q. The process is as follows: the current state and a candidate transport instruction are jointly encoded into an input vector; the input vector is fed into a pre-trained deep neural network; the deep neural network outputs a scalar value, which is the action value function Q of the candidate transport instruction, representing the long-term reward prediction of executing the instruction in the current state.
[0017] Furthermore, the current state input of the reinforcement learning model includes a state vector containing all the following information: the current time t; the occupancy status and physical attributes of all storage locations, including size and maximum load capacity; the real-time inventory, static attributes, and corresponding expected picking time of all goods; and the pre-calculated matrix of the correlation values between all goods.
[0018] Further, before outputting the dynamic storage space optimization scheme, a multi-dimensional verification step is further included, each carrying instruction in the dynamic storage space optimization scheme is verified to ensure that the size specification, maximum load capacity, and environmental requirements of the target storage space of the carrying instruction, such as temperature zone or humidity zone, are compatible with the self attributes and storage requirements of the goods to be moved.
[0019] Further, the method is configured to be triggered to be executed, including continuously monitoring a preset trigger condition, the trigger condition being receiving new future business data or reaching a preset time interval; when the trigger condition is met, the method is automatically executed, the time-sensitive storage space-goods matching degree score of all goods is recalculated, and a new round of generation of the dynamic storage space optimization scheme is started to ensure that the warehouse layout is continuously and adaptively adjusted to business changes.
[0020] Compared with the prior art, the beneficial effects of the present application are: (1) the present application can accurately and prospectively quantitatively evaluate the suitability of storage space allocation by constructing a multi-dimensional dynamic matching model integrating the basic frequency of goods, the correlation degree of adjacent goods, the physical distance, and the time urgency. Unlike traditional methods that rely on single historical data or static rules, the present method can dynamically integrate the expected picking time in future business data into the decision-making process, so that the storage space adjustment not only bases on historical laws, but also actively responds to the upcoming picking task, realizes the change from "passive response" to "active prediction", significantly shortens the order picking path, and improves the overall efficiency and timeliness of warehouse operation; (2) the present application innovatively introduces a reinforcement learning model and designs a compound reward function considering matching degree gain, physical carrying cost, and high urgency task reward, thereby realizing global and long-term intelligent decision-making. The model can overcome the greedy strategy that only focuses on single movement optimization, balance immediate income and future cost by learning long-term return in the whole warehouse state, and output an overall optimization scheme considering carrying economy. This decision-making mechanism based on global state and long-term value makes the storage space optimization scheme more scientific and stable, can effectively avoid frequent invalid carrying caused by short-sighted decision-making, and ensures that the warehouse layout maintains a long-term, dynamic, and adaptive optimal state in the continuously changing business environment. BRIEF DESCRIPTION OF DRAWINGS
[0021] Figure 1 An exemplary step flowchart of the dynamic optimization method of the present application.
[0022] Figure 2 An exemplary step flowchart of the calculation of the goods correlation degree value of the present application. DETAILED DESCRIPTION
[0023] The present application will be further described below in conjunction with specific embodiments.
[0024] As Figure 1 shown, an AI-based warehouse storage space dynamic optimization method provided by the embodiment includes the following steps: Step S1, obtaining multi-dimensional data, the multi-dimensional data including goods information, warehouse storage space information, and historical business data and future business data; in an embodiment, the S1 step is the cornerstone of the entire storage space optimization method, aiming to provide comprehensive and accurate data support for subsequent intelligent analysis and decision-making. The system will integrate with warehouse management systems, enterprise resource planning systems, manufacturing execution systems MES, sensor networks, and market forecasting systems, etc. to obtain these multi-dimensional data. The goods information covers the detailed attributes of all goods in the inventory, such as the SKU code, category, volume, weight, storage requirement temperature zone, fragility, and hazardous property of the goods. The warehouse storage space information includes the detailed parameters of all storage units in the warehouse, such as the number, size, maximum load capacity, belonging temperature and humidity zone, and whether it is an automated picking space of each storage space. The historical business data refers to all operation data related to goods storage and picking in the past period of time, such as historical warehouse entry and exit records, picking order data, goods residence time in each storage space, historical picking path and duration, etc. The future business data is a prediction of future warehouse operation, such as expected sales of goods in the future period, promotion plan, new product warehouse entry plan, seasonal demand fluctuation, and expected picking order volume, etc. These data are collected through data collection interfaces to the central data platform of the system, providing input for AI model training and reasoning.
[0025] As Figure 2 shown, an exemplary step flowchart for calculating the goods correlation value of the embodiment includes the following steps: Step S101, formatting the warehouse order records extracted from the historical business data into transaction data sets, wherein each warehouse order corresponds to a transaction containing all goods items in the order; in an embodiment, the S101 step is to prepare data for association rule mining. The system will extract all warehouse order records from the historical business data obtained in S1. Each warehouse order, regardless of the number of goods it contains, will be formatted into an independent "transaction" record. This transaction record contains a list of all goods items in the order. For example, if an order buys goods A, B, and C, the order corresponds to a transaction {A, B, C}. This process converts the original order data into a "transaction data set" suitable for association rule mining.
[0026] Step S102, based on the preset minimum support threshold, analyze the transaction dataset to mine all the goods combinations whose occurrence frequency is not lower than the minimum support threshold, as the frequent item sets; in an embodiment, S102 step is to identify the combinations of goods that often appear together. The system takes the transaction dataset generated by S101 as input, and analyzes it using association rule mining algorithms such as Apriori algorithm or FP-growth algorithm. In the analysis process, the system will be filtered according to the preset "minimum support threshold". The minimum support represents the frequency of a goods combination appearing in the total transaction dataset. For example, if the minimum support is set to 1%, only the goods combination that appears in at least 1% of the orders will be identified as "frequent item set". These frequent item sets represent the goods sets that are often operated together in sales or picking.
[0027] Step S103, based on the frequent item sets and the preset minimum confidence threshold, generate association rules, and filter out the rules whose confidence is not lower than the minimum confidence threshold, to form a strong association rule set; in an embodiment, S103 step is to find the strong association relationship between goods. The system takes the frequent item sets identified by S102 as input, and generates various possible "association rules". An association rule is usually expressed as "if A is purchased, then B is likely to be purchased", that is, A->B. For each generated association rule, the system will calculate its "confidence", which measures the probability of purchasing the rule's post-item goods in the case of purchasing the rule's pre-item goods. The system will be filtered according to the preset "minimum confidence threshold". For example, if the minimum confidence is set to 70%, only the rules whose confidence reaches or exceeds 70% will be included in the "strong association rule set". These strong association rules clearly indicate which goods have a strong tendency to be picked together.
[0028] Step S104, for any two goods k and j, the support of their co-occurrence is multiplied by the larger one of the confidence from good k to good j and the confidence from good j to good k according to the strong association rule set, and the obtained value is the goods correlation value. In an embodiment, the S3 step is the key for the system to evaluate each potential goods placement scheme. The system will exhaustively or intelligently screen out each “goods-bin” combination for all available target storage bins in the warehouse and all goods to be allocated into the warehouse. For each such combination, the system will calculate a “time-sensitive goods-bin matching score” at the current time. This score is a quantitative indicator that takes into account multiple factors and dynamically changes over time, and is used to measure the appropriateness of placing a particular good in a particular bin at a particular time. For example, if a bin is close to the picking path and the goods to be allocated are about to be picked, their matching score will be high. This score is one of the core inputs for the reinforcement learning model to make decisions.
[0029] Step S2, based on multi-dimensional data, calculate the goods correlation value between any two goods, and determine the expected picking time of each kind of goods according to future business data; in an embodiment, the calculation of the goods correlation value can reveal which goods are often picked together or stored together. For example, in an e-commerce warehouse, if chips and beverages often appear in the same order, their correlation is high, which suggests that they should be placed in adjacent bins to improve picking efficiency. The calculation of the goods correlation value is usually realized by data mining algorithms.
[0030] Step S3, for each combination of target storage bins in the warehouse and goods to be allocated, a time-sensitive goods-bin matching score is calculated at the current time, and the score is used to quantify the appropriateness of placing the goods to be allocated in the target storage bin at the time.
[0031] In an embodiment, the S3 step is the key for the system to evaluate each potential goods placement scheme. The system will exhaustively or intelligently screen out each “goods-bin” combination for all available target storage bins in the warehouse and all goods to be allocated into the warehouse. For each such combination, the system will calculate a “time-sensitive goods-bin matching score” at the current time. This score is a quantitative indicator that takes into account multiple factors and dynamically changes over time, and is used to measure the appropriateness of placing a particular good in a particular bin at a particular time. For example, if a bin is close to the picking path and the goods to be allocated are about to be picked, their matching score will be high. This score is one of the core inputs for the reinforcement learning model to make decisions.
[0032] In this embodiment, the time-sensitive goods-bin matching score is calculated according to the following formula, ,in Indicates the current time , to transport goods Time-sensitive location-goods matching score for placement in the target storage location; Indicates the index of the target storage location. Indicates the index of goods to be allocated. Indicates goods The future basic outbound frequency; This represents the collection of goods already stored within a predefined physical neighborhood of the target storage location. Represents a set One of the goods indexes. Indicates goods With goods The correlation value between goods Indicates the target storage location To the preset warehouse operation baseline point The physical distance; Represent the natural logarithm function; Indicates with goods Related, depending on the current moment A dynamically changing time urgency factor. In one embodiment, the formula takes into account the future outbound frequency of goods, their relevance to surrounding goods, the accessibility of the storage location, and the time urgency of picking the goods themselves, to provide a comprehensive and dynamic matching assessment.
[0033] Step S4: Construct and use a reinforcement learning model, taking the time-sensitive location-cargo matching score as the key basis for the reinforcement learning model to make decisions, and output a dynamic location optimization scheme.
[0034] The time urgency factor is calculated using the following formula: In the formula: This indicates the expected picking time for the goods, usually at a future time. As a preset urgency sensitivity coefficient, in this embodiment, the effect of time urgency on the system can be adjusted by configuring this parameter. The degree of impact; is a very small positive constant used to prevent the denominator from being zero; exp is the natural exponential function.
[0035] In this embodiment, the calculation of the future basic outbound frequency includes: statistically analyzing the historical outbound frequency of goods based on historical business data, and applying a time decay factor to process the historical outbound frequency; extracting the future outbound demand of goods from future business data; and fusing the historical outbound frequency processed by the time decay factor with the future outbound demand to obtain the future basic outbound frequency. In one embodiment, the calculation of the future basic outbound frequency is to more accurately predict the future movement of goods. The system will statistically analyze the historical outbound frequency of each type of goods based on the historical business data in S1, such as the average number of outbound shipments per week or month in the past. To reflect timeliness, the system will apply a time decay factor to the historical outbound frequency, for example, giving higher weight to recent outbound data and lower weight to older data. At the same time, the system will also extract the future outbound demand of goods from the future business data in S1, which may come from sales forecasts, promotional plans, etc. Finally, the historical outbound frequency processed by the time decay factor will be fused with the future outbound demand, for example, by weighted averaging, to obtain the "future basic outbound frequency" of the goods. This frequency is a quantitative indicator of the probability that the goods will be picked soon.
[0036] Reinforcement learning models use a reward function R to guide their learning process, which is defined in detail as follows: R is the immediate reward the reinforcement learning model receives after executing a transport instruction. The model's goal is to maximize the long-term cumulative reward. This refers to the change in the sum of time-sensitive location-goods matching scores for all occupied storage locations in the warehouse after a handling instruction is executed. When a handling instruction is executed to move goods from an old location to a new location, the system will recalculate the sum of matching scores for all occupied storage locations and the goods on them. This is the difference between the new total and the old total. This item represents the positive contribution of this handling operation to the overall optimization of the warehouse. If the total overall matching score increases after the handling, it indicates a good optimization effect, and the reward is positive; otherwise, it is negative.
[0037] This represents the physical handling cost, determined by the weight of the goods in the handling instruction, the handling distance, and the type of equipment required to perform the handling. It is a negative reward, representing the resource consumption required for the actual handling operation, and the model attempts to minimize this cost. Its value depends on the weight of the goods in the handling instruction, the handling distance, and the type of equipment required to perform the handling. For example, moving heavy objects, long-distance handling, or requiring large equipment such as forklifts will have higher costs and a more negative reward.
[0038] This is a high-urgency task reward, its value being proportional to the reduction in the expected picking time of the goods brought about by the handling instruction, and it is only calculated if the time urgency factor of the handled goods is higher than a preset threshold. This is a positive incentive designed to encourage the model to prioritize urgent tasks. Its value is proportional to the reduction in the expected picking time of the goods brought about by the handling instruction. For example, if a piece of goods is expected to be picked within 2 hours, but becomes pickable within 10 minutes after handling, the reward will be high. This reward is only calculated if the time urgency factor of the handled goods is higher than a preset threshold, ensuring that the system only focuses on truly urgent goods.
[0039] Through this reward function, the reinforcement learning model is trained to learn how to make handling decisions, which can optimize the overall warehouse layout, respond to emergency tasks, and minimize unnecessary handling costs.
[0040] In this embodiment, the reinforcement learning model selects the optimal transport instruction by calculating an action value function Q. The process is as follows: the current state and a candidate transport instruction are jointly encoded into an input vector; the input vector is fed into a pre-trained deep neural network; the deep neural network outputs a scalar value, which is the action value function Q of the candidate transport instruction, representing the long-term reward prediction of executing the instruction in the current state.
[0041] The current state input of the reinforcement learning model is a state vector containing all the following information: the current time t; the occupancy status and physical attributes of all storage locations, including size and maximum load capacity; the real-time inventory, static attributes, and corresponding expected picking time of all goods; and the pre-calculated matrix of goods correlation values among all goods.
[0042] Before outputting the dynamic storage location optimization plan, a multi-dimensional verification step is also included to verify each handling instruction in the dynamic storage location optimization plan, to ensure that the size and specifications of the target storage location of the handling instruction, the maximum load-bearing capacity, and the environmental requirements of the temperature or humidity zone are compatible with the properties and storage requirements of the goods to be moved.
[0043] The method is configured to be executed on a trigger basis, including continuously monitoring preset trigger conditions, which are the receipt of new future business data or the arrival of a preset time interval; when the trigger conditions are met, the method is executed automatically, recalculating the time-sensitive location-goods matching score of all goods, and initiating the generation of a new round of dynamic location optimization schemes to ensure that the warehouse layout makes continuous adaptive adjustments to business changes.
[0044] In one embodiment, the method is configured for triggered execution to ensure that the warehouse layout continuously adapts to changes in business operations. The system continuously monitors preset trigger conditions through a scheduling service. These trigger conditions can be passive, such as when the system receives new future business data, like new bulk inbound plans or urgent orders, or when existing forecast data changes significantly. They can also be active, such as reaching a preset time interval, like performing a global optimization every morning or a local optimization every hour. Once a trigger condition is met, the system automatically executes the location optimization method. This includes recalculating the time-sensitive location-goods matching score for all goods, as the original scores may no longer be accurate due to changes in time and business data. Subsequently, the system initiates a new round of dynamic location optimization scheme generation, re-evaluating and outputting the latest optimal location allocation strategy through a reinforcement learning model. This triggered, adaptive mechanism ensures that the warehouse layout is always in an optimal state, flexibly responding to constantly changing business needs and environmental conditions.
[0045] The above description is merely an example and illustration of the present invention. Those skilled in the art can make various modifications or additions to the specific embodiments described, or use similar methods to replace them, as long as they do not deviate from the scope defined by the invention, and all such modifications and additions should fall within the protection scope of the present invention.
Claims
1. A method for dynamic optimization of warehouse storage locations based on AI, characterized in that, Includes the following steps: Step S1: Obtain multidimensional data, which includes cargo information, warehouse location information, historical business data, and future business data; Step S2: Based on the multidimensional data, calculate the correlation value between any two types of goods, and determine the expected picking time for each type of goods based on the future business data. Step S3: For each combination of target storage location and goods to be allocated in the warehouse, calculate a time-sensitive location-goods matching score at the current moment. The score is used to quantify the suitability of placing the goods to be allocated in the target storage location at that moment. Step S4: Construct and use a reinforcement learning model, taking the time-sensitive location-cargo matching score as the key basis for the reinforcement learning model to make decisions, and output a dynamic location optimization scheme.
2. The AI-based dynamic optimization method for warehouse storage locations according to claim 1, characterized in that: The time-sensitive location-cargo matching score is calculated using the following formula. ,in Indicates the current time , to transport goods Time-sensitive location-goods matching score for placement in the target storage location; Indicates the index of the target storage location. Indicates the index of goods to be allocated. Indicates goods The future basic outbound frequency; This represents the collection of goods already stored within a predefined physical neighborhood of the target storage location. Represents a set One of the goods indexes. Indicates goods With goods The correlation value between goods Indicates the target storage location To the preset warehouse operation baseline point The physical distance; Represent the natural logarithm function; Indicates with goods Related, depending on the current moment Dynamically changing time urgency factor.
3. The AI-based dynamic optimization method for warehouse storage locations according to claim 2, characterized in that: The time urgency factor is calculated according to the following formula: In the formula: Indicates the expected picking time for the goods; A preset urgency sensitivity coefficient; is a very small positive constant used to prevent the denominator from being zero; exp is the natural exponential function.
4. The AI-based dynamic optimization method for warehouse storage locations according to claim 2, characterized in that: The calculation of the future basic outbound frequency includes: statistically analyzing the historical outbound frequency of goods based on the historical business data, and applying a time decay factor to the historical outbound frequency. Extract the future outbound demand of goods from the future business data; merge the historical outbound frequency after time decay factor processing with the future outbound demand to obtain the future basic outbound frequency.
5. The AI-based dynamic optimization method for warehouse storage locations according to claim 2, characterized in that: The calculation of the cargo correlation value includes the following steps: Step S101: The outbound order records extracted from the historical business data are formatted into a transaction dataset, wherein each outbound order corresponds to a transaction, and the transaction contains all the items of the goods in the order; Step S102: Based on a preset minimum support threshold, analyze the transaction dataset to discover all combinations of goods with a frequency not lower than the minimum support threshold, as frequent itemsets. Step S103: Based on the frequent itemsets and the preset minimum confidence threshold, generate association rules and filter out rules with a confidence level not lower than the minimum confidence threshold to form a strong association rule set. Step S104: For any two goods k and j, according to the set of strong association rules, multiply the support of their co-occurrence by the larger of the confidence scores from goods k to goods j and from goods j to goods k, and the resulting value is the goods association value.
6. The AI-based dynamic optimization method for warehouse storage locations according to claim 1, characterized in that: The reinforcement learning model uses a reward function R to guide its learning process, which is specifically defined as follows: ,in The change in the sum of the time-sensitive location-cargo matching scores of all occupied storage locations in the warehouse after a handling instruction is executed; The cost of physical handling is determined by the weight of the goods in the handling instruction, the handling distance, and the type of equipment required to perform the handling. The reward is a high-urgency task reward whose value is proportional to the reduction in the expected picking time of the goods brought about by the handling instruction, and the reward is calculated only when the time urgency factor of the goods being handled is higher than a preset threshold.
7. The AI-based dynamic optimization method for warehouse storage locations according to claim 1, characterized in that: The reinforcement learning model selects the optimal transport instruction by calculating an action value function Q. The process is as follows: the current state and a candidate transport instruction are jointly encoded into an input vector; the input vector is fed into a pre-trained deep neural network; the deep neural network outputs a scalar value, which is the action value function Q of the candidate transport instruction, representing the long-term reward prediction of executing the instruction in the current state.
8. The AI-based dynamic optimization method for warehouse storage locations according to claim 1, characterized in that: The current state input of the reinforcement learning model is a state vector containing all the following information: the current time t; the occupancy status and physical attributes of all storage locations, including size and maximum load capacity; the real-time inventory, static attributes, and corresponding expected picking time of all goods; and the pre-calculated matrix of the correlation values between all goods.
9. The AI-based dynamic optimization method for warehouse storage locations according to claim 1, characterized in that: Before outputting the dynamic storage location optimization scheme, a multi-dimensional verification step is also included to verify each handling instruction in the dynamic storage location optimization scheme to ensure that the size specifications, maximum load-bearing capacity, and environmental requirements of the target storage location of the handling instruction are compatible with the properties and storage requirements of the goods to be moved.
10. The AI-based dynamic optimization method for warehouse storage locations according to claim 1, characterized in that: The method is configured to be executed on a trigger basis, including continuously monitoring preset trigger conditions, wherein the trigger conditions are receiving new future business data or reaching a preset time interval; Once the triggering condition is met, the method is executed automatically, recalculating the time-sensitive location-cargo matching score for all goods, and initiating the generation of a new round of dynamic location optimization schemes to ensure that the warehouse layout continuously adapts to changes in business operations.
Citation Information
Cited By
Warehouse-in and warehouse-out management method and system for stored goods
CN121961434A
Warehouse storage space dynamic allocation optimization algorithm and system based on deep reinforcement learning
CN122492069A