Supply chain multi-level storage intelligent scheduling and collaboration method and system

Through long-term short-term memory networks and multi-agent reinforcement learning models, combined with edge computing and rule knowledge base, the problems of low prediction accuracy and low synergy efficiency in traditional warehousing management are solved, inventory balance and rapid response are achieved, and the stability and efficiency of the supply chain are improved.

CN120258693AActive Publication Date: 2025-07-04SHANDONG XINDA IOT APPL TECH CO LTD

Patent Information

Application Number
CN202510732817.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-07-04
Estimated Expiration
2045-06-04

AI Technical Summary

Technical Problem

In the traditional supply chain multi-level warehousing management, there are low demand forecasting accuracy, unbalanced inventory allocation and low coordination efficiency, which makes it difficult to deal with emergencies and lack of effective inter-node coordination mechanisms, resulting in high operating costs and reduced customer satisfaction.

Method used

Long-term memory network is used to predict demand, build a multi-agent reinforcement learning model, use each warehouse node as an independent agent, generate scheduling instructions through edge computing and rule knowledge base, and conduct emergency collaborative requests and local negotiations in case of inventory abnormalities or resource conflicts.

Benefits of technology

It improves the accuracy of demand forecasting, achieves improved inventory equality, improved turnover efficiency and reduced operating costs, enhances the resilience and stability of the supply chain, ensures the executability and business compliance of decisions, shortens response time, and reduces the computing pressure and communication overhead of the central node.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258693A_ABST
    Figure CN120258693A_ABST
Patent Text Reader

Abstract

The invention provides a supply chain multistage warehousing intelligent scheduling and collaboration method and system, and relates to the technical field of intelligent warehousing, and the method comprises the steps: carrying out the hierarchical modeling prediction of the short-term and long-term demands of each stage of warehouse through employing a long-short-term memory network based on historical order data, and obtaining the inventory demand; setting each warehouse node as an independent intelligent agent, and generating an inventory allocation strategy through strategy iteration optimization among the intelligent agents; converting the inventory allocation strategy into a scheduling instruction containing an allocation object, an allocation quantity and allocation time based on a rule knowledge base; and in the edge computing unit of each warehouse node, receiving a scheduling instruction and performing sorting execution, when inventory abnormity or resource conflict is detected, initiating an emergency cooperation request to an adjacent warehouse node, returning response information by the warehouse node receiving the emergency cooperation request according to the own resource state, and performing scheduling according to the response information. And determining an emergency processing scheme through local negotiation between the nodes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to intelligent warehousing technology, and particularly to an intelligent scheduling and coordination method and system for multi-level warehousing in the supply chain. Background Art

[0002] Traditional multi-level warehousing management in the supply chain mainly relies on manual experience and static rules for inventory allocation, and it is difficult to adapt to the complex and changeable market demands. Such methods generally have problems such as low demand forecasting accuracy, unbalanced inventory allocation, and low coordination efficiency, resulting in an increase in the overall supply chain operation cost and a decrease in customer satisfaction.

[0003] Existing technologies usually adopt a centralized decision-making model to handle inventory management, with high computational complexity and slow response speed, and it is difficult to respond to emergencies in real time. At the same time, due to the lack of an effective inter-node coordination mechanism, each warehouse often acts independently, resulting in low overall resource utilization rate of the system and the coexistence of inventory backlog and shortage.

[0004] With the development of Internet of Things and artificial intelligence technologies, intelligent warehousing management has become possible, but existing solutions still have problems such as complex model training, disconnection between decision execution, and weak emergency handling ability. Especially in case of emergencies, the system lacks a flexible response mechanism, and it is difficult to achieve intelligent coordination between warehouse nodes and optimal allocation of resources, seriously affecting the stability and resilience of the supply chain. Summary of the Invention

[0005] In view of the deficiencies of the existing technologies, the present invention provides an intelligent scheduling and coordination method and system for multi-level warehousing in the supply chain, which can solve the problems in the existing technologies.

[0006] The present invention provides an intelligent scheduling and coordination method for multi-level warehousing in the supply chain, including: Obtaining real-time inventory data and historical order data of warehouses at all levels in the multi-level warehousing network; Based on the historical order data, using a long short-term memory network to hierarchically model and predict the short-term and long-term demands of warehouses at all levels to obtain inventory demand quantities; Constructing a multi-agent reinforcement learning model, setting each warehouse node as an independent agent, where the state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand quantity, and available resource quantity of each warehouse node, and the action space includes the inventory allocation quantity and allocation timing, and generating an inventory allocation strategy through policy iteration optimization among agents; Converting the inventory allocation strategy into a scheduling instruction including the allocation object, allocation quantity, and allocation time based on a rule knowledge base; In the edge computing units of each warehouse node, the scheduling instructions are received and sorted for execution. When inventory anomalies or resource conflicts are detected, an emergency collaboration request is initiated to adjacent warehouse nodes. The warehouse nodes that receive the emergency collaboration request return response information based on their own resource status. Based on the response information, an emergency handling plan is determined through local negotiation between nodes.

[0007] Optionally, The steps of hierarchically modeling and predicting the short-term and long-term demands of each level of warehouse using a long short-term memory network based on the historical order data to obtain the inventory demand quantity include: Dividing the historical order data into daily order sequences and weekly order sequences; Constructing an adaptive hierarchical prediction network, including a short-term prediction layer and a long-term prediction layer. The short-term prediction layer receives the daily order sequence and dynamically adjusts the prediction time window according to the inventory goods category and seasonal characteristics. The long-term prediction layer receives the weekly order sequence and captures long-cycle patterns through a memory enhancement module. The output error gradient of the short-term prediction layer is passed to the parameter update process of the long-term prediction layer, and the trend characteristics of the long-term prediction layer are embedded into the input end of the short-term prediction layer; Training the adaptive hierarchical prediction network using a multi-scale loss function. The multi-scale loss function includes a short-term prediction loss term, a long-term prediction loss term, and a regularization term, and dynamically weights the loss terms based on the time-scale characteristics of each prediction layer; Integrating the prediction results to obtain an initial prediction value; calibrating the initial prediction value based on the warehouse capacity constraint, and generating an inventory demand quantity prediction result with a confidence interval in combination with the prediction standard deviation.

[0008] Optionally, Constructing a multi-agent reinforcement learning model, setting each warehouse node as an independent agent. The state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand quantity, and available resource quantity of each warehouse node. The action space includes the inventory allocation quantity and the allocation timing. The steps of generating an inventory allocation strategy through policy iteration among agents include: Setting each warehouse node as an independent agent, constructing the state vector of the independent agent. The state vector includes real-time inventory data, inventory demand quantity, and available resource quantity; constructing an adjacency matrix based on the logistics channels between each warehouse node. The adjacency matrix is used to represent the connection relationship between each warehouse node; Constructing the action space of the independent agent. The action space includes an allocation quantity matrix and an allocation timing vector. The allocation quantity matrix is used to represent the inventory allocation quantity between warehouse nodes, and the allocation timing vector is used to represent the allocation execution time; Construct a multi-objective reward function, which includes a local reward function and a global reward function. The local reward function calculates the reward values of each independent agent based on inventory balance, turnover efficiency, and scheduling cost, and the global reward function calculates the global collaborative reward value based on the inventory level difference of adjacent warehouse nodes; Use the Proximal Policy Optimization (PPO) algorithm to train the independent agents. During the training process, construct a hierarchical experience buffer to store state transition samples. The hierarchical experience buffer includes a local buffer for storing the local experience of the independent agents and a global buffer for storing global state information; use a parameter sharing mechanism to train the policy network for the same type of independent agents; Based on the trained policy network, generate an inventory allocation strategy according to the real-time state of each warehouse node.

[0009] Optionally, The steps of constructing a multi-objective reward function, which includes a local reward function and a global reward function, and the local reward function calculates the reward values of each independent agent based on inventory balance, turnover efficiency, and scheduling cost, and the global reward function calculates the global collaborative reward value based on the inventory level difference of adjacent warehouse nodes, are as follows: Classify the items based on the historical order data of each warehouse node and determine the target inventory level; calculate the standardized inventory deviation of each warehouse node based on the target inventory level, and determine the inventory balance degree reward value as the product of the standardized inventory deviation and the time decay factor; Calculate the daily turnover rate, weekly turnover rate, and monthly turnover rate based on the outbound volume, beginning inventory, and ending inventory of each warehouse node respectively, and determine the turnover efficiency reward value as the weighted sum with a preset weight coefficient; Determine the scheduling cost reward value based on the weighted sum of storage cost, transportation cost, and operation cost; Calculate the ratio of the inventory level difference between adjacent warehouses to construct a difference matrix, calculate the distance weight coefficient based on the logistics distance between adjacent warehouses, and determine the global collaborative reward value as the product of the distance weight coefficient, the difference matrix, the warehouse level collaboration coefficient, and the congestion penalty factor; Calculate the performance improvement rate of the inventory balance degree reward value, turnover efficiency reward value, scheduling cost reward value, and global collaborative reward value within a preset evaluation period; construct a multiple linear regression model based on the performance improvement rate to calculate the correlation coefficient matrix, calculate the contribution degree of each reward value according to the correlation coefficient matrix, multiply the difference between the contribution degree and the current weight by the learning rate to obtain the weight adjustment amount, and update the weight coefficients of each reward value under the weight adjustment constraint conditions to obtain the multi-objective reward function.

[0010] Optionally, The steps of training independent agents using the Proximal Policy Optimization (PPO) algorithm and constructing a hierarchical experience buffer to store state transition samples during the training process include: The hierarchical experience buffer includes a local experience buffer for storing local state transition samples of each independent agent, a collaborative experience buffer for storing interaction samples of adjacent agents, and a global experience buffer for storing global state samples; Calculate the importance score of the state transition sample. The importance score is determined based on the reward value, state transition frequency, state influence range, and policy difference degree. Deposit the state transition sample into the corresponding-level experience buffer according to the importance score; Calculate the sampling probability of the sample based on the sample timeliness and the number of policy updates. Select training samples from the hierarchical experience buffer according to the sampling probability; Update the policy network parameters of a single agent based on the samples in the local experience buffer, update the policy network parameters of the adjacent agent group based on the samples in the collaborative experience buffer, and update the policy network parameters of all agents based on the samples in the global experience buffer; Construct a knowledge sharing weight matrix and calculate the element values of the knowledge sharing weight matrix based on the historical decision-making data and network topology of the agents; Divide the objective function of the Proximal Policy Optimization algorithm into an immediate optimization term, a periodic optimization term, and a long-term optimization term. Calculate the weight coefficients of each optimization term based on the knowledge sharing weight matrix and system state information, and optimize and update the policy network parameters.

[0011] Optionally, The steps of converting the inventory allocation policy into a scheduling instruction including the allocation object, allocation quantity, and allocation time based on the rule knowledge base include: Construct a rule knowledge base including physical constraint rules, business constraint rules, and priority rules, and parse the inventory allocation policy into an initial scheduling plan including warehouse node pairs, allocation quantities, and allocation time windows; Based on the rule knowledge base, perform constraint verification on the initial scheduling plan, and construct a rule matching degree scoring function. The rule matching degree scoring function calculates the matching degree score based on the degree of physical constraint satisfaction, business goal achievement, and resource utilization rate, and automatically corrects the scheduling plan that fails the constraint verification or has a matching degree score lower than the threshold; Based on the scheduling delay rate, resource conflict rate, and inventory volatility in the historical scheduling execution data, calculate the execution risk probability of the corrected scheduling plan in different time windows, generate an execution priority considering risk aversion, and convert the sorted scheduling plan into a scheduling instruction including the allocation object, allocation quantity, and allocation time.

[0012] Optionally, The warehouse node that receives the emergency collaboration request returns response information according to its own resource status. The steps of determining the emergency handling plan through local negotiation among nodes according to the response information include: Calculate the logistics correlation based on the historical logistics frequency and current logistics intensity of the warehouse node, calculate the resource response ability based on the ratio of the required resource quantity of the request to its own available resource quantity, and generate response information including response time, response cost, and resource status according to the logistics correlation and resource response ability; Obtain the task completion rate and resource stability of the warehouse nodes participating in the negotiation, calculate the node credibility in combination with the response information, generate an influence weight based on the number of hops between nodes, and the influence weight decreases exponentially with the increase of the number of hops, and determine the initial negotiation weight according to the node credibility and influence weight; Receive the negotiation plans of neighboring warehouse nodes, calculate the weighted average value of the plans based on the initial negotiation weight, combine the weighted average value with the current optimal plan to form an updated local plan, and calculate the variance value of the updated local plan within the negotiation group as the consistency index; When the update amplitude of the plan is continuously lower than the preset amplitude threshold, adjust the weights of cost, time, and resource consumption in the negotiation plan according to the historical execution effect, and generate a random perturbation value positively correlated with the stagnation duration to restart the negotiation; generate an emergency handling plan when the consistency index reaches the preset condition; Monitor the heartbeat information and response delay of the warehouse nodes participating in the negotiation. When the heartbeat information is interrupted or the response delay exceeds the preset value, obtain the uncompleted negotiation tasks of the nodes with abnormal heartbeat, calculate the task reception suitability based on the resource status, response time, and logistics distance of the remaining warehouse nodes, and reallocate the uncompleted negotiation tasks under the condition of meeting the resource constraints.

[0013] In the second aspect of the embodiments of the present invention, a multi-level warehouse intelligent scheduling and collaboration system for the supply chain is provided, including: The first unit is used to obtain the real-time inventory data and historical order data of each level of warehouses in the multi-level warehouse network; The second unit is used to hierarchically model and predict the short-term and long-term demands of each level of warehouses based on the historical order data by using a long short-term memory network to obtain the inventory demand; The third unit is used to construct a multi-agent reinforcement learning model, set each warehouse node as an independent agent. The state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand, and available resource quantity of each warehouse node, and the action space includes the inventory allocation quantity and allocation timing. Generate an inventory allocation strategy through policy iteration optimization among agents; A fourth unit, configured to convert the inventory deployment strategy into a scheduling instruction including deployment objects, deployment quantities, and deployment times based on a rule knowledge base; A fifth unit, configured to receive the scheduling instruction in the edge computing units of each warehouse node and execute it in sequence. When inventory anomalies or resource conflicts are detected, an emergency collaboration request is sent to adjacent warehouse nodes. The warehouse node that receives the emergency collaboration request returns response information according to its own resource status. According to the response information, an emergency handling plan is determined through local negotiation between nodes.

[0014] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the foregoing method is implemented.

[0015] The present invention realizes hierarchical prediction of demand through a long short-term memory network, significantly improves the accuracy of demand prediction, and effectively reduces inventory fluctuations and out-of-stock risks. The multi-agent reinforcement learning model enables each warehouse node to autonomously learn the optimal deployment strategy, realizes the improvement of inventory balance, the improvement of turnover efficiency, and the reduction of operating costs, and optimizes the overall supply chain resource allocation efficiency.

[0016] The policy conversion mechanism based on the rule knowledge base of the present invention ensures the executability and business compliance of decisions, and significantly reduces the execution anomaly rate. The edge computing architecture realizes the local processing of decisions, significantly shortens the response time, improves the system throughput, reduces the computing pressure and communication overhead of the central node, and makes the system more efficient and scalable.

[0017] The distributed emergency collaboration mechanism of the present invention enables the system to have self-healing capabilities in the face of abnormal situations, quickly forms an emergency handling plan through local negotiation between nodes, shortens the abnormal handling time, improves the abnormal resolution rate, significantly enhances the resilience and stability of the supply chain system, effectively responds to market fluctuations and supply interruption risks, and provides a more reliable supply chain guarantee for enterprises. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 It is a schematic flowchart of the multi-level warehouse intelligent scheduling and collaboration method for the supply chain according to the embodiments of the present invention; Figure 2 It is a diagram showing the influence of the gradual iteration of each technical component of the present invention on the system performance. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The technical solutions in the embodiments of the present invention will be described below with reference to the accompanying drawings in the embodiments of the present invention. These specific embodiments may be combined with each other. For the same or similar concepts or processes, they may not be repeated in some embodiments.

[0020] Figure 1This is a schematic flowchart of the intelligent scheduling and coordination method for multi-level warehousing in the supply chain of the present invention. As Figure 1 shown, the method includes: Obtain the real-time inventory data and historical order data of each level of warehouse in the multi-level warehousing network; Based on the historical order data, use a long short-term memory network to hierarchically model and predict the short-term and long-term demands of each level of warehouse to obtain the inventory demand; Construct a multi-agent reinforcement learning model, set each warehouse node as an independent agent. The state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand, and available resource quantity of each warehouse node. The action space includes the inventory allocation quantity and allocation timing. Generate an inventory allocation strategy through policy iteration optimization among agents; Convert the inventory allocation strategy into a scheduling instruction including the allocation object, allocation quantity, and allocation time based on the rule knowledge base; In the edge computing unit of each warehouse node, receive the scheduling instruction and execute it in sequence. When inventory anomalies or resource conflicts are detected, initiate an emergency collaboration request to adjacent warehouse nodes. The warehouse node that receives the emergency collaboration request returns response information according to its own resource status. According to the response information, determine an emergency handling plan through local negotiation among nodes.

[0021] Optionally, Based on the historical order data, the steps of using a long short-term memory network to hierarchically model and predict the short-term and long-term demands of each level of warehouse to obtain the inventory demand include: Divide the historical order data into daily order sequences and weekly order sequences; Construct an adaptive hierarchical prediction network, including a short-term prediction layer and a long-term prediction layer. The short-term prediction layer receives the daily order sequence and dynamically adjusts the prediction time window according to the inventory goods category and seasonal characteristics. The long-term prediction layer receives the weekly order sequence and captures long-cycle patterns through a memory enhancement module; the output error gradient of the short-term prediction layer is transmitted to the parameter update process of the long-term prediction layer, and the trend characteristics of the long-term prediction layer are embedded into the input end of the short-term prediction layer; Train the adaptive hierarchical prediction network using a multi-scale loss function. The multi-scale loss function includes a short-term prediction loss term, a long-term prediction loss term, and a regularization term, and dynamically weights the loss terms based on the time-scale characteristics of each prediction layer; Integrate the prediction results to obtain an initial prediction value; calibrate the initial prediction value based on the warehousing capacity constraint, and generate an inventory demand prediction result with a confidence interval in combination with the prediction standard deviation.

[0022] In this embodiment, historical order data of warehouses at all levels in the multi-level warehousing network is obtained. These data usually include fields such as order date, order quantity, goods category, customer information, etc. The obtained historical order data is preprocessed, including data cleaning, outlier handling, and data standardization. During the data cleaning process, missing values are detected and processed. For example, for missing order data on a certain day, the average value of adjacent dates can be used for filling. Outlier handling is to identify and adjust those data points that significantly deviate from the normal range, such as sudden surges in orders caused by promotional activities or holidays. Data standardization is to convert data of different magnitudes to the same scale. The preprocessed historical order data is divided into daily order sequences and weekly order sequences. The daily order sequence records the order quantity for each day and is mainly used for short-term demand forecasting; while the weekly order sequence summarizes the order quantity for each week and is used for long-term trend forecasting. For example, for the food category in a certain warehouse, the daily sequence records the daily order volume such as "Date 1: 150 pieces, Date 2: 165 pieces...", and the weekly sequence records "Week 1: 1050 pieces, Week 2: 1120 pieces...".

[0023] An adaptive hierarchical prediction network is constructed. This network includes a short-term prediction layer and a long-term prediction layer, and there is an information interaction mechanism between the two layers. The short-term prediction layer adopts a long short-term memory network structure with an attention mechanism, mainly processes the daily order sequence, and captures short-term fluctuation characteristics; the long-term prediction layer adopts a bidirectional long short-term memory network with a memory enhancement module, processes the weekly order sequence, and captures long-term trend characteristics. The short-term prediction layer dynamically adjusts the prediction time window according to the inventory goods category and seasonal characteristics. For example, for clothing products with strong seasonality, the prediction time window is set to 7 - 14 days; while for daily necessities with stable demand, the prediction time window is set to 3 - 7 days. In specific implementation, volatility indicators of historical data of each product category, such as the coefficient of variation, are analyzed, and then the appropriate prediction window size is dynamically determined according to these indicators. The smaller the time window, the higher the prediction accuracy but the shorter the covered time range; the larger the window, the further into the future can be predicted, but the accuracy is relatively reduced.

[0024] The memory enhancement module is a structure containing multiple memory units, and each memory unit stores historical pattern information for a specific time span. For example, for products with obvious seasonality, the memory unit will pay special attention to the sales pattern in the same period last year; for products affected by the economic cycle, the memory unit will capture trend changes over a longer time span. This design enables the model to "remember" and utilize long-term historical patterns, improving the prediction ability for future long-term trends.

[0025] A two-way information flow mechanism is established between the two prediction layers. The output error gradient of the short-term prediction layer is transmitted to the parameter update process of the long-term prediction layer, enabling the long-term prediction to be adjusted considering the results of the short-term prediction; at the same time, the trend features of the long-term prediction layer are embedded into the input end of the short-term prediction layer, enabling the short-term prediction to be fine-tuned based on the long-term trend. This design enables the two prediction layers to complement and correct each other, forming a more accurate prediction result.

[0026] During the network training stage, a multi-scale loss function is used to train the adaptive hierarchical prediction network. This loss function includes a short-term prediction loss term, a long-term prediction loss term, and a regularization term. The short-term prediction loss term mainly measures the difference between the daily prediction result and the actual value, the long-term prediction loss term measures the difference between the weekly prediction result and the actual value, and the regularization term is used to prevent the model from overfitting. The loss terms are dynamically weighted based on the time-scale characteristics of each prediction layer. For example, for products with high volatility, the weight of the short-term prediction loss term will be higher; for products with obvious trends, the weight of the long-term prediction loss term will be higher.

[0027] After training is completed, the prediction results of the short-term prediction layer and the long-term prediction layer are integrated to obtain the initial prediction value. The integration method uses weighted average, and the weights are dynamically adjusted according to the historical prediction accuracy. For example, if the short-term prediction layer performs better recently, its weight will increase accordingly; conversely, if the long-term prediction layer captures important seasonal trends, its weight will increase. This dynamic adjustment mechanism enables the integrated result to take into account both short-term accuracy and long-term trends. Finally, the initial prediction value is calibrated based on the warehouse capacity constraint. If the predicted demand exceeds the maximum storage capacity of the warehouse, the prediction value will be adjusted according to the historical overflow handling experience to ensure that the prediction result is feasible in actual operation. At the same time, the uncertainty during the prediction process is combined, and the prediction standard deviation is calculated to generate an inventory demand prediction result with a confidence interval. For example, a result such as "the predicted demand for the next 7 days is 1000 ± 50 pieces" will be given, where ±50 pieces represents the range of prediction uncertainty, helping inventory managers formulate more reasonable inventory strategies.

[0028] The present invention realizes the accurate prediction of multi-level warehousing demand by constructing an adaptive hierarchical prediction network, which has significant advantages compared with traditional methods; the two-way information interaction mechanism between the short-term and long-term prediction layers effectively integrates the features of different time scales, the dynamically adjusted prediction time window adapts to the demand characteristics of different commodities, and the training method of the multi-scale loss function improves the generalization ability of the model. The prediction calibration based on warehousing constraints and the generation of confidence intervals make the prediction results more practical and reliable, providing accurate data support for inventory management decisions, effectively reducing inventory costs, and improving the supply chain response speed and customer satisfaction.

[0029] Optionally, Construct a multi-agent reinforcement learning model, set each warehouse node as an independent agent, the state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand, and available resources of each warehouse node, the action space includes the inventory allocation quantity and the allocation timing, and the steps of generating an inventory allocation strategy through policy iteration among agents include: Set each warehouse node as an independent agent, construct the state vector of the independent agent, and the state vector includes real-time inventory data, inventory demand, and available resources; construct an adjacency matrix based on the logistics channels between each warehouse node, and the adjacency matrix is used to represent the connection relationship between each warehouse node; Construct the action space of the independent agent, the action space includes an allocation quantity matrix and an allocation timing vector, the allocation quantity matrix is used to represent the inventory allocation quantity between warehouse nodes, and the allocation timing vector is used to represent the allocation execution time; Construct a multi-objective reward function, the multi-objective reward function includes a local reward function and a global reward function, the local reward function calculates the reward value of each independent agent based on inventory balance, turnover efficiency, and scheduling cost, and the global reward function calculates the global collaborative reward value based on the inventory level difference between adjacent warehouse nodes; Use the proximal policy optimization algorithm to train the independent agent, construct a hierarchical experience buffer to store state transition samples during the training process, the hierarchical experience buffer includes a local buffer for storing the local experience of the independent agent and a global buffer for storing global state information; use a parameter sharing mechanism to train the policy network for the same type of independent agent; Based on the trained policy network, generate an inventory allocation strategy according to the real-time state of each warehouse node.

[0030] In this embodiment, each warehouse node in the multi-level warehousing network is set as an independent agent. In practical applications, for example, the supply chain network of a large e-commerce enterprise includes 5 central warehouses, 25 regional warehouses, and 100 front-end warehouses, and each warehouse node is set as an independent agent. Construct a state vector for each agent, and the state vector includes three types of key information: real-time inventory data, inventory demand, and available resources. The real-time inventory data records the current inventory levels of various commodities, such as "Commodity A: 500 pieces, Commodity B: 320 pieces"; the inventory demand is the demand forecast value for different future time periods obtained through the aforementioned prediction method, such as "Demand for Commodity A in the next 3 days: 150 pieces, Demand for Commodity A in the next 7 days: 380 pieces"; the available resources include the remaining storage space in the warehouse, the number of schedulable vehicles, the number of available loading and unloading workers, etc., such as "Remaining storage space: 2000 cubic meters, Schedulable vehicles: 15".

[0031] Construct an adjacency matrix based on the logistics channels between each warehouse node. If there is a direct logistics channel between two warehouse nodes, the corresponding matrix element value is 1; otherwise, it is 0. For example, a 130×130 adjacency matrix (corresponding to 130 warehouse nodes) is constructed. Among them, the central warehouse is connected to all regional warehouses, the regional warehouses are connected to the front-end warehouses within their coverage, and the front-end warehouses in different regions are usually not directly connected. This adjacency matrix plays three key roles in the subsequent intelligent agent training and decision-making process: it limits the communication range between intelligent agents, enabling each intelligent agent to exchange information only with directly connected nodes; it determines the boundary conditions for executable deployment operations, and the intelligent agent can only initiate deployment to adjacent nodes; it serves as the weight basis when calculating the global reward function, and the inventory balance degree between adjacent nodes contributes more to the global reward. This constraint mechanism based on the actual logistics network topology structure ensures that the generated deployment strategy is physically realizable.

[0032] The deployment quantity matrix describes the quantity of various commodities that this warehouse deploys to other adjacent warehouses. For example, the deployment quantity matrix of a regional warehouse includes information such as "deploy 50 pieces of commodity A to front-end warehouse 1, deploy 30 pieces of commodity B to front-end warehouse 2". The deployment timing vector describes the time points for executing these deployment operations, such as "deployment 1 execution time: T + 1 day, deployment 2 execution time: T + 3 days". To discretize the action space, the deployment quantity is divided into several gears by percentage (such as 0%, 20%, 40%, 60%, 80%, 100%), and the deployment timing is also divided into several time periods (such as immediately, within 1 day, within 3 days, within 7 days). This design makes the action space tractable while retaining sufficient flexibility.

[0033] Construct a multi-objective reward function, including a local reward function and a global reward function. Use the Proximal Policy Optimization algorithm to train independent intelligent agents. The core idea of this algorithm is to ensure training stability by restricting the policy update step size. To improve training efficiency and generalization performance, a parameter sharing mechanism is adopted for the policy network training of the same type of independent intelligent agents. Specifically, warehouse nodes at the same level (such as all regional warehouses) share the policy network parameters, so that the experience of each intelligent agent can be used to update the shared policy, greatly increasing the training sample size. For example, in a certain actual deployment, 25 regional warehouses share a set of policy network parameters, enabling the experience of each regional warehouse to be utilized by other regional warehouses, accelerating training convergence and improving the robustness of the policy.

[0034] During the training process, a phased training strategy is adopted. First, each agent conducts local training based on the samples in the local buffer to learn the basic inventory allocation strategy. Then, collaborative training is carried out based on the samples in the global buffer to learn the allocation strategy considering the overall benefits. In actual training, pre-training is first performed with historical data, then fine-tuning is carried out in a simulated environment, and finally small-scale testing is conducted in the actual environment and the application scope is gradually expanded.

[0035] Input the real-time inventory data, demand forecast, and resource status of each warehouse node into the trained policy network. The network outputs the probability distribution of the allocation quantity and the allocation timing, and selects the action with the highest probability as the final allocation decision. For example, the policy network gives the decision "Allocate 200 pieces of commodity X from the central warehouse to regional warehouse A, and the execution time is 8 am tomorrow". These decisions are combined to form the inventory allocation strategy for the entire warehousing network, guiding the execution of the actual inventory scheduling.

[0036] The present invention realizes intelligent inventory allocation in the warehousing network by constructing a multi-agent reinforcement learning model. Each warehouse is set as an independent agent to have the ability of distributed decision-making. The design of the multi-objective reward function takes into account both local optimization and global coordination. The proximal policy optimization algorithm and the hierarchical experience buffer improve the training efficiency and policy quality; the parameter sharing mechanism enhances the generalization ability of the model, enabling it to adapt to different warehouse scenarios; through the policy iteration optimization among agents, an optimal allocation strategy that can balance the inventory distribution, improve the turnover efficiency, and reduce the operation cost can be generated, providing an effective solution for the intelligent management of multi-level warehousing networks.

[0037] Optionally, Construct a multi-objective reward function, the multi-objective reward function includes a local reward function and a global reward function. The steps of calculating the reward value of each independent agent by the local reward function based on the inventory balance degree, turnover efficiency, and scheduling cost, and calculating the global coordination reward value by the global reward function based on the inventory level difference of adjacent warehouse nodes include: Classify the items based on the historical order data of each warehouse node and determine the target inventory level; calculate the standardized inventory deviation of each warehouse node based on the target inventory level, and determine the inventory balance degree reward value as the product of the standardized inventory deviation and the time decay factor; Calculate the daily turnover rate, weekly turnover rate, and monthly turnover rate respectively based on the outbound volume, beginning inventory, and ending inventory of each warehouse node, and determine the turnover efficiency reward value as the weighted sum with the preset weight coefficient; Determine the scheduling cost reward value based on the weighted sum of the storage cost, transportation cost, and operation cost; Calculate the ratio of inventory level differences between adjacent warehouses to construct a difference matrix, calculate the distance weight coefficient based on the logistics distance between adjacent warehouses, and determine the global collaborative reward value as the product of the distance weight coefficient, the difference matrix, the warehouse level collaborative coefficient, and the congestion penalty factor; Calculate the performance improvement rate of the inventory balance reward value, turnover efficiency reward value, scheduling cost reward value, and global collaborative reward value within a preset evaluation period; construct a multiple linear regression model based on the performance improvement rate to calculate the correlation coefficient matrix, calculate the contribution degree of each reward value according to the correlation coefficient matrix, multiply the difference between the contribution degree and the current weight by the learning rate to obtain the weight adjustment amount, and update the weight coefficient of each reward value under the weight adjustment constraint conditions to obtain a multi-objective reward function.

[0038] In this embodiment, the items are classified and the target inventory level is determined based on the historical order data of each warehouse node. Item classification is usually carried out according to the sales speed and demand volatility. For example, the goods are divided into category A (high-frequency fast-selling products), category B (medium-frequency regular products), and category C (low-frequency long-tail products). For a regional warehouse, the shampoo with a daily average sales volume exceeding 100 pieces and a volatility coefficient less than 0.3 is classified as category A goods, the skin care products with a daily average sales volume between 20 - 100 pieces and a volatility coefficient between 0.3 - 0.6 are classified as category B goods, and the special seasonal products with a daily average sales volume less than 20 pieces or a volatility coefficient greater than 0.6 are classified as category C goods. For different categories of goods, different methods are used to determine the target inventory level. Category A goods usually adopt the method of adding safety inventory to the demand forecast value. For example, the 7-day demand forecast of a certain category A good is 700 pieces, the volatility standard deviation is 50 pieces, and its target inventory is set at 850 pieces (forecast demand plus 3 times the standard deviation); category B goods adopt the method of adding buffer inventory to the historical average consumption; category C goods adopt the method of combining the minimum order quantity and periodic replenishment.

[0039] Calculate the standardized inventory deviation of each warehouse node based on the target inventory level. For each product, subtract the target inventory from the actual inventory and then divide by the target inventory to obtain the relative deviation value. For example, if the actual inventory of a certain type A product in a warehouse is 900 pieces and the target inventory is 850 pieces, the relative deviation is (900 - 850) / 850 = 0.059, indicating a slightly excessive inventory. Take the absolute value of the deviation values of all products and then calculate the weighted average by sales amount to obtain the overall standardized inventory deviation of the warehouse. The smaller this deviation value is, the closer the actual inventory is to the target level and the more accurate the inventory management is. Multiply the standardized inventory deviation by the time decay factor to determine the inventory balance reward value. The time decay factor is a value that decreases over time. For example, it can be set as the negative t-th power of e (t is the number of days passed after inventory adjustment). The purpose of this design is to encourage the inventory to quickly reach the balanced state and maintain the balance for a long time. If the inventory balance improves just after adjustment but becomes unbalanced again after a few days, the reward value will decrease due to the time decay factor. The final inventory balance reward value is equal to 1 minus the decayed standardized inventory deviation, so that the smaller the deviation, the greater the reward.

[0040] The turnover rate is calculated as the outbound quantity in a specific period divided by the average inventory. The average inventory is equal to the sum of the beginning inventory and the ending inventory, and then divided by 2. For example, if a warehouse ships out 600 pieces of goods in a day, with a beginning inventory of 1200 pieces and an ending inventory of 1000 pieces, the daily turnover rate is 600 / ((1200 + 1000) / 2) = 0.55, indicating that 55% of the inventory is turned over on average every day. The calculation principles of the weekly turnover rate and the monthly turnover rate are the same, except for the different statistical periods. The turnover rates in different periods reflect the inventory flow efficiency on different time scales, and combined, they can comprehensively evaluate the turnover situation. For example, the weights can be set as 0.2 for the daily turnover rate, 0.3 for the weekly turnover rate, and 0.5 for the monthly turnover rate, indicating a higher emphasis on the long-term turnover effect. If the daily, weekly, and monthly turnover rates of a certain warehouse are 0.55, 3.5, and 14.0 respectively, the weighted turnover efficiency is 0.55×0.2 + 3.5×0.3 + 14.0×0.5 = 8.15. Normalize this value (such as dividing by the expected maximum value) to obtain the turnover efficiency reward value between 0 and 1.

[0041] The scheduling cost reward value is determined based on the weighted sum of storage costs, transportation costs, and operation costs. Storage costs include expenses such as warehouse rent, equipment depreciation, and energy consumption allocated to each commodity; transportation costs include vehicle usage fees, fuel costs, labor costs, etc.; operation costs include operation expenses such as picking, packaging, and loading and unloading. For example, a certain scheduling operation involves 500 commodities, with storage costs of 1,000 yuan, transportation costs of 3,000 yuan, and operation costs of 2,000 yuan, and the total cost is 6,000 yuan. Compare this cost with the benchmark cost (such as the historical average cost), calculate the cost reduction ratio, and then convert this ratio into a scheduling cost reward value between 0 and 1. The lower the cost, the higher the reward value.

[0042] Calculate the ratio of the inventory level difference between adjacent warehouses to construct a difference matrix. For each pair of adjacent warehouses, calculate the ratio of the inventory level difference between them, that is, the absolute value of the difference between the normalized inventory levels (actual inventory divided by target inventory) of the two warehouses. For example, if the normalized inventory level of warehouse A is 1.1 (inventory slightly exceeds the standard), and the adjacent warehouse B is 0.85 (inventory slightly falls short), then the difference ratio between them is |1.1 - 0.85| = 0.25. The difference ratios of all pairs of adjacent warehouses form a difference matrix. Calculate the distance weight coefficient based on the logistics distance between adjacent warehouses. The closer the pair of warehouses, the higher the importance of their inventory balance coordination, and the larger the weight coefficient. For example, the weight coefficient can be set as the benchmark value (such as 1.0) divided by the normalized distance, where the normalized distance is the actual distance divided by the average warehouse spacing in the network. If the distance between two warehouses is 50 kilometers and the average network spacing is 100 kilometers, then the normalized distance is 0.5, and the weight coefficient is 1.0 / 0.5 = 2.0. Determine the global coordination reward value as the product of the distance weight coefficient, the difference matrix, the warehouse level coordination coefficient, and the congestion penalty factor. The warehouse level coordination coefficient reflects the importance of coordination between different levels of warehouses. For example, the coordination coefficient between the central warehouse and the regional warehouse is higher than that between the regional warehouse and the front-end warehouse. The congestion penalty factor reduces the reward value when the logistics channel is congested to avoid generating scheduling strategies that exacerbate congestion. The final global coordination reward value is obtained by weighted summing the coordination rewards for all pairs of adjacent warehouses and then normalizing to a value between 0 and 1.

[0043] Calculate the performance improvement rates of the inventory balance reward value, turnover efficiency reward value, scheduling cost reward value, and global collaboration reward value within a preset evaluation period (e.g., every 30 days). The performance improvement rate is the percentage change in the average reward value of the current evaluation period compared to the previous evaluation period. For example, if the average inventory balance reward in the current period is 0.75 and in the previous period is 0.7, the improvement rate is (0.75 - 0.7) / 0.7 = 7.14%. Based on these performance improvement rates, construct a multiple linear regression model with the overall performance improvement as the dependent variable and the improvement rates of each individual reward as independent variables, and calculate the correlation coefficient matrix through regression analysis. The correlation coefficient reflects the contribution degree of each reward item to the overall performance. According to the correlation coefficient matrix, calculate the contribution degree of each reward value, that is, the absolute value of each coefficient divided by the sum of the absolute values of all coefficients. Multiply the difference between the contribution degree and the current weight by the learning rate to obtain the weight adjustment amount. For example, if the contribution degree of the inventory balance reward is 0.4, the current weight is 0.3, and the learning rate is set to 0.2, the adjustment amount is (0.4 - 0.3) × 0.2 = 0.02. Under the weight adjustment constraints (such as the sum of all weights is 1 and each individual weight is not less than 0.1), update the weight coefficients of each reward value to form an adaptive multi-objective reward function.

[0044] The present invention realizes the effective guidance of multi-agent warehouse scheduling by constructing a multi-objective reward function including local rewards and global rewards. The inventory balance reward ensures that the inventory level is close to the optimal target, the turnover efficiency reward promotes the efficient flow of resources, the scheduling cost reward reduces the operating cost, and the global collaboration reward optimizes the overall network collaboration effect. The adaptive weight adjustment mechanism based on the performance improvement rate enables the reward function to be dynamically optimized according to the actual effect, better adapting to different warehouse environments and business stages. This multi-dimensional and adaptive reward mechanism provides a clear learning direction for the agents, effectively improving the inventory management quality and the overall efficiency of the supply chain.

[0045] Optionally, When using the proximal policy optimization algorithm to train independent agents, the steps of constructing a hierarchical experience buffer to store state transition samples during the training process include: The hierarchical experience buffer includes a local experience buffer for storing local state transition samples of each independent agent, a collaborative experience buffer for storing interaction samples of adjacent agents, and a global experience buffer for storing global state samples; Calculate the importance score of the state transition sample, which is determined based on the reward value, state transition frequency, state influence range, and policy difference degree, and store the state transition sample into the corresponding-level experience buffer according to the importance score; Calculate the sampling probability of the sample based on the sample timeliness and the number of policy updates, and select training samples from the hierarchical experience buffer according to the sampling probability; Update the policy network parameters of a single agent based on samples in the local experience buffer, update the policy network parameters of the adjacent agent group based on samples in the collaborative experience buffer, and update the policy network parameters of all agents based on samples in the global experience buffer; Construct a knowledge sharing weight matrix and calculate the element values of the knowledge sharing weight matrix based on the historical decision-making data of the agents and the network topology; Divide the objective function of the proximal policy optimization algorithm into an immediate optimization term, a periodic optimization term, and a long-term optimization term, calculate the weight coefficients of each optimization term based on the knowledge sharing weight matrix and the system state information, and optimize and update the policy network parameters.

[0046] In this embodiment, the hierarchical experience buffer includes three levels of buffers: a local experience buffer, a collaborative experience buffer, and a global experience buffer. The local experience buffer is used to store the local state transition samples of each independent agent, mainly including the state, action, reward, and next state information of the agent. For example, for a regional warehouse agent, its local experience buffer stores samples such as "State: current inventory level 75%, demand forecast 120 pieces per day, 8 available vehicles; Action: allocate 30 pieces of commodity X to the upstream warehouse A; Reward: local reward value 0.65; Next state: inventory level 70%, demand forecast 118 pieces per day, 7 available vehicles".

[0047] The collaborative experience buffer is used to store the interaction samples between adjacent agents, mainly including the collaborative state, joint action, and shared reward information involving multiple adjacent agents. For example, the collaborative samples between a regional warehouse and the three upstream warehouses it covers include information such as "Collaborative state: 20% overstock in the regional warehouse, 15% out-of-stock in upstream warehouse A, 10% out-of-stock in upstream warehouse B, normal inventory in upstream warehouse C; Joint action: the regional warehouse allocates 15% of its inventory to A and 10% to B; Shared reward: global collaborative reward 0.75".

[0048] The global experience buffer is used to store the global state samples of the entire warehousing network, mainly including hierarchical state information, global performance indicators, and reward data. For example, the global samples include information such as "Global state: total inventory level 85%, inter-regional inventory difference coefficient 0.15, network congestion degree 0.25; Performance: order fulfillment rate 92%, average delivery time 1.8 days; Reward: global reward value 0.82". This hierarchical design enables learning and optimizing strategies from multiple levels from local to global.

[0049] Calculate the importance score of each state transition sample to determine which layer of buffer to store the sample in. The importance score is calculated based on four key metrics: reward value, state transition frequency, state influence range, and policy difference degree. The reward value refers to the size of the reward obtained by the sample. The higher the reward, the more important the sample. The state transition frequency refers to the frequency of this type of state transition in history. The lower the frequency of the state transition, the rarer it is, and thus the more important it is. The state influence range refers to the number of agents affected by this state transition. The wider the influence range, the higher the importance. The policy difference degree refers to the degree of difference between the action of this sample and the action generated by the current policy. The greater the difference, the more new information the sample contains, and thus the more important it is. The importance score is obtained by weighted summation of these four metrics. For example, for a sample with a reward value of 0.9 (after normalization), a state transition frequency of 0.1 (low occurrence frequency), a state influence range of 0.7 (affecting multiple agents), and a policy difference degree of 0.6 (significant difference from the current policy), with the weights of each metric being 0.3, 0.2, 0.3, and 0.2 respectively, the importance score of this sample is 0.9×0.3 + 0.1×0.2 + 0.7×0.3 + 0.6×0.2 = 0.67. Store the sample in the corresponding level of the experience buffer according to the importance score. Usually, two thresholds T1 and T2 are set (e.g., T1 = 0.3, T2 = 0.7). Samples with an importance score lower than T1 are stored in the local experience buffer, samples with scores between T1 and T2 are stored in the collaborative experience buffer, and samples with scores higher than T2 are stored in the global experience buffer. In this way, samples with higher global significance will be stored in higher-level buffers for learning by a wider range of agents.

[0050] During the training process, the sample timeliness refers to the freshness of the sample, usually represented by the reciprocal of the sample age (i.e., the number of training rounds elapsed since the sample was stored in the buffer). For example, the timeliness of a sample stored 10 rounds ago is 1 / 10 = 0.1, while the timeliness of the latest stored sample is 1. The number of policy updates refers to the number of times the parameters of the policy network have been updated since the sample was generated. The more updates there are, the less representative the sample is under the current policy, and the lower the sampling probability should be. Use the function of sample timeliness and the number of policy updates as the sampling probability. For example, it can be set that the sampling probability is equal to the timeliness multiplied by e to the power of negative λ times the number of policy updates, where λ is a coefficient controlling the decay rate (e.g., λ = 0.05). For a sample with a timeliness of 0.5 and having experienced 10 policy updates, its sampling probability is 0.5×e (-0.05×10) = 0.5×0.61 = 0.31. Select training samples from the hierarchical experience buffer according to the calculated sampling probability to ensure that new samples and representative samples have a higher chance of being selected.

[0051] Update the policy network parameters of agents within the corresponding range based on samples from different-level buffers respectively. Based on samples from the local experience buffer, only update the policy network parameters of a single agent; based on samples from the collaborative experience buffer, update the policy network parameters of the group of adjacent agents (i.e., a set of interconnected agents); based on samples from the global experience buffer, update the policy network parameters of all agents. For example, in a certain training iteration, select samples from the local buffer of regional warehouse A and only update the policy network of the agent in this warehouse; select samples from the collaborative buffer containing regional warehouse A and its covered group of front warehouses and update the policy networks of all agents in this group of warehouses; select samples from the global buffer and update the policy networks of all agents in the entire network. This hierarchical update mechanism not only ensures local optimization but also promotes global collaboration.

[0052] Construct a knowledge sharing weight matrix and calculate the element values of this matrix based on the historical decision data and network topology of the agents. The knowledge sharing weight matrix is an N×N matrix (N is the number of agents), and the matrix element (i, j) represents the weight of agent i transmitting knowledge to agent j. The weight value is determined based on two factors: historical decision similarity and network topology distance. Historical decision similarity is calculated by comparing the decisions made by two agents in similar states. The more similar the decisions, the higher the similarity. For example, if two regional warehouses both tend to allocate inventory to front warehouses in the state of overstock, their decision similarity is high; if one tends to allocate while the other tends to maintain, the similarity is low. Network topology distance refers to the number of connection hops between two agents in the warehousing network. The fewer the hops, the closer the distance and the higher the weight.

[0053] Use the function of historical decision similarity and network topology distance as the weight value. For example, it can be set that the weight value is equal to the decision similarity multiplied by e to the power of negative d times the topology distance, where d is a coefficient controlling the influence of distance (such as d = 0.5). For two agents with a decision similarity of 0.8 and a topology distance of 2, the weight value is 0.8×e (-0.5×2) =0.8×0.37 = 0.30. This design enables stronger knowledge sharing between agents with similar decision patterns and close network positions.

[0054] The objective function of the Proximal Policy Optimization (PPO) algorithm is divided into three optimization terms: the immediate optimization term, the periodic optimization term, and the long-term optimization term. The immediate optimization term focuses on maximizing the immediate reward at the current time step. The periodic optimization term focuses on the cumulative reward in the medium term. The long-term optimization term focuses on long-term stability and performance. The knowledge influence index of each agent is extracted from the knowledge sharing weight matrix, and the calculation method is to sum the element values in the corresponding row of the matrix for this agent and then normalize. The higher the influence of an agent, the greater the weight of its long-term optimization term. At the same time, the current state features are analyzed, including the inventory pressure index (the degree of difference between the current inventory and the target inventory), the demand volatility index (the coefficient of variation of the recent demand forecast), and the resource saturation (the ratio of available resources to required resources). When the inventory pressure is high, the weight of the immediate optimization term increases; when the demand is stable and resources are sufficient, the weight of the periodic optimization term increases; when operating in a critical business cycle (such as the end of a quarter or the annual planning period), the weight of the long-term optimization term increases. These state-based weight adjustment factors are combined with the knowledge sharing-based weight adjustment factors to generate the final weight coefficients of the three optimization terms. For example, for a warehouse agent with a knowledge influence of 0.6, a current inventory pressure of 0.8, a demand volatility of 0.3, and a resource saturation of 0.5, the calculated weight of the immediate optimization term is 0.5, the weight of the periodic optimization term is 0.3, and the weight of the long-term optimization term is 0.2. The Proximal Policy Optimization algorithm serves as the core algorithm for agent training here. Its key feature is to maintain the amplitude of policy change within a controllable range when updating the policy. The policy gradient is calculated using the state transition samples sampled from the hierarchical experience buffer, and the difference between the new and old policies is restricted to avoid an overly large update step. For example, in inventory allocation decisions, it is ensured that the updated allocation policy changes moderately compared to the original policy, preventing extreme situations of over-allocation or under-allocation. This stable policy learning method, combined with the hierarchical experience buffer and the knowledge sharing mechanism, enables each agent to gradually improve its inventory allocation policy while maintaining decision stability.

[0055] By constructing a hierarchical experience buffer and optimizing the policy update mechanism, the present invention significantly improves the training efficiency and policy quality of multi-agent reinforcement learning; the design of the hierarchical experience buffer enables simultaneous attention to local optimization and global cooperation, the importance score calculation mechanism ensures that key samples are fully utilized, and the sampling strategy based on timeliness and the number of policy updates improves the representativeness of training samples. This solution innovatively divides the update process into two dimensions: "which agents to update" and "how to update these parameters"; the former solves the update scope problem through the hierarchical buffer, and the latter solves the update quality problem through the knowledge sharing matrix and multi-time scale optimization; the knowledge sharing weight matrix promotes the experience transfer between agents, and the division of the three-layer optimization objectives enables good performance on different time scales. The combined action of these innovative mechanisms provides an efficient and reliable learning method for intelligent scheduling in multi-level warehousing networks, greatly improving the intelligence level and operation efficiency of inventory management.

[0056] Optionally, The step of converting the inventory allocation policy into a scheduling instruction including the allocation object, the allocation quantity, and the allocation time based on the rule knowledge base includes: Construct a rule knowledge base including physical constraint rules, business constraint rules, and priority rules, and parse the inventory allocation policy into an initial scheduling plan including warehouse node pairs, allocation quantities, and allocation time windows; Based on the rule knowledge base, perform constraint verification on the initial scheduling plan, and construct a rule matching degree scoring function. The rule matching degree scoring function calculates the matching degree score based on the degree of physical constraint satisfaction, the achievement degree of business objectives, and the resource utilization rate, and automatically corrects the scheduling plan that fails the constraint verification or has a matching degree score lower than the threshold; Based on the scheduling delay rate, resource conflict rate, and inventory volatility in the historical scheduling execution data, calculate the execution risk probability of the corrected scheduling plan in different time windows, generate an execution priority considering risk aversion, and convert the sorted scheduling plan into a scheduling instruction including the allocation object, the allocation quantity, and the allocation time.

[0057] In this embodiment, a rule knowledge base including physical constraint rules, business constraint rules, and priority rules is constructed. Physical constraint rules mainly describe the hard limit conditions of logistics, such as "the maximum load of a single truck is 3 tons", "Class A dangerous goods cannot be mixed with food", "the transportation time of cold chain goods shall not exceed 4 hours", etc. These rules are usually related to physical facilities, transportation tools, and commodity characteristics, and are hard constraints that cannot be violated. Business constraint rules mainly describe the management requirements and business process regulations of enterprise operations, such as "the allocation of high-value goods must be approved by the supervisor", "the inventory level of promotional goods shall not be lower than 120% of the predicted demand", "the allocated quantity must be an integer multiple of the standard packaging unit", etc. These rules are usually related to enterprise management processes, business objectives, and customer service standards, and are important conditions for ensuring the normal operation of the business. Priority rules mainly describe the execution priority order of different allocation tasks, such as "stockout risk allocation takes precedence over inventory balance allocation", "allocations related to Class A customer orders take precedence over Class B customers", "allocations of goods approaching the expiration date take precedence over new product allocations", etc. These rules are usually related to business importance, customer level, and timeliness requirements, and are the basis for reasonably arranging tasks under limited resources. These rules are stored in a structured manner. For example, for the physical constraint rule "the maximum load of a single truck is 3 tons", it can be stored as "rule type: physical constraint; rule object: transportation tool; rule condition: truck model XYZ; constraint content: maximum load; constraint value: 3 tons; constraint severity: high". This structured storage method facilitates the rapid retrieval and application of relevant rules.

[0058] Parse the inventory allocation strategy into an initial scheduling plan including warehouse node pairs, allocation quantity, and allocation time window. For example, the strategy indicates that "Warehouse A allocates 20% of the inventory of Product X to Warehouse B within the next 3 days". Convert these abstract strategy descriptions into specific scheduling plans. For example, receive the strategy output from the reinforcement learning model: "The central warehouse allocates 25% of the inventory of Commodity Category A to Regional Warehouse 3, and the allocation time window is within the next 48 hours". Query the current inventory data and find that the inventory of Commodity Category A in the central warehouse is 8000 pieces, and 25% of it is 2000 pieces. Therefore, the initial scheduling plan is parsed as "allocation object: central warehouse → regional warehouse 3; allocated commodity: category A; allocation quantity: 2000 pieces; allocation time window: within the next 48 hours".

[0059] Conducting constraint verification on the initial scheduling plan includes three aspects: physical constraint verification, business constraint verification, and resource conflict verification. Physical constraint verification checks whether the scheduling plan violates physical limitations. For example, it checks whether the total weight of 2,000 items exceeds the load capacity of the available transportation vehicles. Business constraint verification checks whether the scheduling plan complies with business regulations. For example, it checks whether the remaining inventory in the central warehouse after allocation meets the minimum safety inventory requirements. Resource conflict verification checks whether the scheduling plan has resource conflicts with the already scheduled tasks. For example, it checks whether the required transportation vehicles are occupied by other tasks during the specified time window.

[0060] Continuing with the above case, the verification found that: the total weight of 2,000 items of product category A is 5 tons, exceeding the load capacity of a single standard truck of 3 tons; after allocation, the remaining inventory in the central warehouse is 6,000 items, higher than the safety inventory line of 4,000 items, meeting the business requirements; however, within the next 48 hours, the number of trucks already scheduled for use makes the available transportation capacity insufficient to complete the transportation of 5 tons of goods at one time. Therefore, this initial scheduling plan fails the physical constraint verification and resource conflict verification.

[0061] Construct a rule matching degree scoring function to calculate the matching degree score based on the degree of physical constraint satisfaction, the achievement degree of business goals, and resource utilization rate. The degree of physical constraint satisfaction measures the degree of compliance of the scheduling plan with physical constraints. For example, it can be calculated as the inverse ratio of "actual weight / maximum load capacity"; the achievement degree of business goals measures the contribution of the scheduling plan to business goals. For example, it can be calculated as the improvement amplitude of inventory balance after the execution of allocation; the resource utilization rate measures the utilization efficiency of the scheduling plan for resources. For example, it can be calculated as the ratio of "actual loading volume / maximum loading volume".

[0062] For a scheduling plan that meets both requirements (i.e., passes the constraint verification and the matching degree score is higher than the threshold), it is considered that the plan has met the feasibility conditions and does not need to be automatically corrected and can directly enter the next step of processing. For a scheduling plan that fails the constraint verification or the matching degree score is lower than the threshold (such as 0.75), automatic correction is performed. The correction strategies include batch allocation, adjustment of the time window, and replacement of resources, etc. Batch allocation splits a large-scale allocation into multiple small batches to make each batch comply with physical constraints; adjusting the time window arranges the allocation tasks in a time period with fewer resource conflicts; replacing resources means finding alternative transportation tools or routes to complete the allocation tasks. In the above case, the allocation of 2,000 items is automatically split into two batches, with 1,000 items (2.5 tons) in each batch, meeting the single-vehicle load limit; and the first batch is arranged to be completed within 24 hours, and the second batch is arranged to be completed within 24 - 48 hours to adapt to the available transportation capacity. The corrected scheduling plan becomes "Allocation object: Central warehouse → Regional warehouse 3; Allocation goods: Category A; Allocation quantity: 1,000 items in the first batch, 1,000 items in the second batch; Allocation time: within 24 hours for the first batch, within 24 - 48 hours for the second batch".

[0063] Based on historical scheduling execution data, calculate the execution risk probability of the revised scheduling plan in different time windows. The historical data includes indicators such as scheduling delay rate, resource conflict rate, and inventory volatility. The scheduling delay rate reflects the deviation between the actual completion time and the planned time of past similar scheduling tasks; the resource conflict rate reflects the frequency of resource contention in the same time window in the past; the inventory volatility reflects the difference between the actual change and the expected change of the past inventory level.

[0064] Extract historical records similar to the current scheduling plan from the historical database, such as extracting deployment tasks of similar scale on the route of "central warehouse → regional warehouse 3". Analysis shows that within the 24-hour time window, the historical scheduling delay rate of this route is 15%, the resource conflict rate is 10%, and the inventory volatility is 8%; while within the 24-48-hour time window, the historical scheduling delay rate is 25%, the resource conflict rate is 18%, and the inventory volatility is 12%. This indicates that the execution risk of the second batch is significantly higher than that of the first batch.

[0065] Generate an execution priority considering risk aversion, and convert the sorted scheduling plan into scheduling instructions. The execution priority is determined comprehensively based on task importance, timeliness requirements, and execution risk. Task importance is determined by the degree of business impact, timeliness requirements are determined by the deployment time limit, and execution risk is determined by the aforementioned risk probability. Tasks with higher priorities are allocated resources and executed first.

[0066] Convert the sorted scheduling plan into scheduling instructions including specific execution details. The scheduling instructions include: deployment objects (source warehouse and target warehouse), deployed goods (specific to the SKU level), deployment quantity (accurate to the number of pieces), deployment time (accurate to hours and minutes), resources used (specific vehicle numbers or personnel arrangements), and operation requirements (such as handling precautions for special goods), etc.

[0067] In the above case, the finally generated scheduling instructions are: "Instruction number: DS2023051001; Deployment object: Central warehouse CD001 → Regional warehouse RD003; Deployed goods: Class A SKU10015 - 10025; Deployment quantity: 1000 pieces; Deployment time: 2023 - 05 - 10 08:00; Resources used: Logistics vehicle TK0058, operators OP108, OP109; Operation requirements: Standard loading and unloading process B12" and "Instruction number: DS2023051002; Deployment object: Central warehouse CD001 → Regional warehouse RD003; Deployed goods: Class A SKU10015 - 10025; Deployment quantity: 1000 pieces; Deployment time: 2023 - 05 - 11 10:00; Resources used: Logistics vehicle TK0062, operators OP112, OP115; Operation requirements: Standard loading and unloading process B12".

[0068] The present invention realizes the efficient transformation from abstract strategies to specific executable instructions by constructing a rule knowledge base and a multi - stage transformation process; the construction of the rule knowledge base ensures that the scheduling instructions comply with physical constraints and business requirements, the constraint verification and automatic correction of the initial scheduling plan improve the feasibility of the plan, and the risk assessment and priority sorting based on historical data enhance the reliability of scheduling execution. This rule - based strategy transformation mechanism effectively solves the "last mile" problem between AI decision - making and actual execution, significantly improves the implementation effect of inventory deployment strategies, reduces the execution exception rate, enhances the overall operation efficiency of the warehousing network, and provides key technical support for enterprises to achieve intelligent inventory management.

[0069] Optionally, The steps for the warehouse node that receives the emergency collaboration request to return response information according to its own resource status and determine the emergency handling plan through local negotiation among nodes include: Calculate the logistics correlation based on the historical logistics frequency and current logistics intensity of the warehouse node, calculate the resource response ability based on the ratio of the required resource quantity of the request to the available resource quantity of itself, and generate response information including response time, response cost, and resource status according to the logistics correlation and resource response ability; Obtain the task completion rate and resource stability of the warehouse nodes participating in the negotiation, calculate the node credibility in combination with the response information, generate an influence weight based on the number of hops between nodes, and the influence weight decreases exponentially with the increase of the number of hops, and determine the initial negotiation weight according to the node credibility and influence weight; Receive the negotiation plans of neighboring warehouse nodes, calculate the weighted average value of the plans based on the initial negotiation weight, combine the weighted average value with the current optimal plan to form an updated local plan, and calculate the variance value of the updated local plan within the negotiation group as the consistency index; When the update amplitude of the plan is continuously lower than the preset amplitude threshold, the cost, time and resource consumption weights in the negotiation plan are adjusted according to the historical execution effect, and a random disturbance value positively correlated with the duration of the stagnation is generated to restart the negotiation; when the consistency index reaches the preset condition, an emergency treatment plan is generated; Monitor the heartbeat information and response delay of the warehouse nodes participating in the negotiation. When it is detected that the heartbeat information is interrupted or the response delay exceeds the preset value, obtain the unfinished negotiation tasks of the abnormal heartbeat node, calculate the task reception fitness based on the resource status, response time and logistics distance of the remaining warehouse nodes, and reallocate the unfinished negotiation tasks according to the task reception fitness while satisfying the resource constraints.

[0070] In this embodiment, a scheduling execution module is deployed in the edge computing unit of each warehouse node, which receives the aforementioned generated scheduling instructions and executes them according to priority. When an inventory anomaly or resource conflict is detected during the execution process, the module automatically initiates an emergency coordination request to the adjacent warehouse node. Inventory anomalies include a variety of situations, such as actual inventory lower than the record (inventory loss), inventory quality problems (such as damaged goods), and a sudden large number of orders leading to insufficient inventory. Resource conflicts include situations where the scheduling cannot be executed as originally planned, such as transportation tool failure, personnel absence, warehouse equipment failure, etc. For example, when a forward warehouse is performing a distribution task, it is found that the actual inventory of Class A goods is 30% less than the record, resulting in the inability to meet the distribution needs of the day. At this time, an emergency coordination request is automatically initiated to the adjacent regional warehouse and other forward warehouses, and the request content includes "missing commodity type: Class A; missing quantity: 150 pieces; demand urgency: high; response time required: within 2 hours".

[0071] The warehouse node that receives the emergency coordination request evaluates its response capability based on two key indicators: logistics relevance and resource responsiveness. Logistics relevance is calculated based on historical logistics frequency and current logistics intensity. Historical logistics frequency refers to the number of logistics interactions between two warehouse nodes in the past period of time (such as 30 days), and current logistics intensity refers to the number of logistics tasks currently being executed between the two nodes. For example, if there were 25 logistics interactions between two warehouse nodes in the past 30 days and there are currently 3 logistics tasks being executed, the logistics relevance can be calculated as "historical frequency / reference period+current intensity", that is, 25 / 30+3=3.83. The higher the logistics relevance, the closer the collaboration between the two nodes and the higher the coordination efficiency. Resource responsiveness is calculated based on the ratio of the amount of resources required for the request to the amount of available resources. Resources include inventory resources, transportation resources, and human resources. For example, if the request requires 150 pieces of Class A goods, and the node currently has 200 pieces of Class A goods available, the inventory resource response ratio is 200 / 150=1.33, indicating that the request can be fully met; if the request requires 2 transport vehicles, and the node currently has only 1 available, the transport resource response ratio is 1 / 2=0.5, indicating that it can only be partially met. The minimum value of the overall resource response capacity is taken as the comprehensive resource response ratio, reflecting the "barrel principle".

[0072] Based on the logistics correlation and resource response capability, the receiving node generates response information including response time, response cost and resource status. Response time refers to the time required to complete the requested task, including preparation time and logistics time; response cost includes resource allocation cost, transportation cost and opportunity cost; resource status describes in detail the quantity and quality of various resources that can be provided. For example, a regional warehouse returns a response message to the aforementioned out-of-stock forward warehouse: "Available Class A goods: 120 pieces; response time: 90 minutes; response cost: 2,000 yuan; resource status: 1 transport vehicle available, 2 operators available".

[0073] The node initiating the request obtains the task completion rate and resource stability of each responding node participating in the negotiation, and calculates the node credibility based on the response information. The task completion rate is the proportion of collaborative tasks successfully completed by the node in the past period of time, and the resource stability reflects the consistency of the node's resource supply. For example, if a responding node has a task completion rate of 92% in the past and a resource stability score of 0.85 (full score 1), combined with the completion time and resource status promised in its response information, its node credibility can be calculated as 0.92×0.85×(committed resources / requested resources)=0.92×0.85×(120 / 150)=0.626. The higher the node credibility, the greater the reference value of its response information.

[0074] Generate influence weights based on the number of hops between nodes. The influence weights decrease exponentially as the number of hops increases. The number of hops refers to the length of the shortest path between two nodes in the network. Nodes directly connected have a hop count of 1, nodes connected through one intermediate node have a hop count of 2, and so on. The influence weight can be calculated by "the base weight multiplied by the attenuation coefficient to the power of the number of hops". For example, if the base weight is 1 and the attenuation coefficient is 0.7, then the influence weight of a node with a hop count of 1 is 1×0.7 1 =0.7, and the influence weight of a node with a hop count of 2 is 1×0.7 2 =0.49. This design enables nodes with a shorter physical distance to have a greater influence in the negotiation, which conforms to the actual characteristics of the logistics network. The initial negotiation weight is the product of the node credibility and the influence weight. For example, if the credibility of a certain node is 0.626 and the influence weight is 0.7, then its initial negotiation weight is 0.626×0.7 = 0.438. Normalize the initial negotiation weights of all participating nodes so that the sum of all weights is 1, which serves as the final negotiation weight.

[0075] During the negotiation process, each node receives the negotiation plans of its neighboring nodes and calculates the weighted average of the plans based on the negotiation weights. For example, a certain node receives the plans of three neighboring nodes, which are "quantity of goods allocation: 100 pieces, 80 pieces, 120 pieces" respectively, and the corresponding negotiation weights are "0.5, 0.3, 0.2", then the weighted average is 100×0.5 + 80×0.3 + 120×0.2 = 98 pieces. The node combines this weighted average with the current optimal plan (such as the best locally feasible plan) to form an updated local plan. The combination method can be linear interpolation. For example, "updated plan = current optimal plan × (1 - learning rate) + weighted average × learning rate", where the learning rate is a parameter between 0 and 1 that controls the aggressiveness of the plan update.

[0076] Calculate the variance value of the updated local plans within the negotiation group as the consistency index. The smaller the variance value, the closer the plans of each node are, and the more consistent the negotiation is. For example, if the updated plans of three nodes are "quantity of goods allocation: 95 pieces, 98 pieces, 102 pieces" respectively, then the variance value is small, indicating that the negotiation is close to being consistent; if the plan differences are large, such as "80 pieces, 100 pieces, 120 pieces", then the variance value is large, indicating that the negotiation needs to continue.

[0077] When the update amplitude of the plan is continuously lower than the preset amplitude threshold, it indicates that the negotiation has fallen into a local stagnation state. For example, if in 5 consecutive rounds of negotiation, the update amplitude of the plan is less than 2% each time, while the preset threshold is 3%, then the negotiation is considered to have stagnated. At this time, adjust the weights of cost, time, and resource consumption in the negotiation plan according to the historical execution effect. For example, if historical data shows that reducing the cost weight and increasing the time weight can effectively break the stagnation in similar situations, then the weights will be adjusted accordingly. At the same time, generate a random perturbation value that is positively correlated with the duration of the stagnation to restart the negotiation. The perturbation value can be designed as "basic perturbation amplitude × (1 + number of stagnant rounds / 10)", so that the longer the stagnation time, the greater the perturbation, increasing the probability of jumping out of the local optimum.

[0078] When the consistency index reaches the preset condition, generate the final emergency handling plan. The preset condition can be "the variance value is less than the threshold and the change amplitude in 3 consecutive rounds is less than 1%", indicating that the nodes have reached a stable and consistent plan. The final plan details the specific tasks of each participating node, including the type and quantity of resources provided, the execution time, and the responsible person, etc. For example, the final plan is "Regional Warehouse A provides 100 pieces of Class A goods, and Forward Warehouse B provides 50 pieces of Class A goods, which are delivered to the requesting node before 14:00 and 14:30 respectively, and are executed by their respective logistics teams."

[0079] Real-time monitor the heartbeat information and response delay of the warehouse nodes participating in the negotiation. The heartbeat information is the status notification regularly sent by the node, indicating that the node is running normally; the response delay is the time interval from when the node receives the message to when it returns the response. When it is detected that the heartbeat information of a certain node is interrupted (such as not receiving the heartbeat 3 times in a row) or the response delay exceeds the preset value (such as 3 times the normal value), it is judged that the node has a failure or a network interruption. Obtain the unfinished negotiation tasks of the node with abnormal heartbeat, and calculate the task reception suitability based on the resource status, response time, and logistics distance of the remaining normal nodes. The task reception suitability is an ability index to measure the node's ability to undertake the tasks of the failed node, and the calculation method is the weighted average of "resource sufficiency × time responsiveness × distance suitability". For example, if the resource sufficiency of a certain node is 0.8 (resources are basically sufficient), the time responsiveness is 0.9 (able to respond in time), and the distance suitability is 0.7 (the distance is relatively close but not the closest), then its task reception suitability is 0.8×0.4 + 0.9×0.4 + 0.7×0.2 = 0.82 (assuming the weights are 0.4, 0.4, and 0.2 respectively).

[0080] Reallocate the uncompleted negotiation tasks according to the task reception fitness under the premise of meeting the resource constraints. The resource constraints ensure that the task reallocation will not cause excessive strain on the resources of the receiving nodes. The allocation principle is to give priority to nodes with high fitness and sufficient resources. If necessary, a large task can be split among multiple nodes to complete together. For example, if the failed node was originally planned to provide 100 items, the task can be reallocated as "Node C provides 60 items and Node D provides 40 items" to ensure the reliable execution of the overall emergency plan.

[0081] Figure 2 This is the impact diagram of the system performance after the gradual iteration of each technical component of the present invention. With the addition of key technical components, the negotiation convergence rounds have been significantly reduced from 25 rounds in the basic version to 3 rounds in the final version, improving the negotiation efficiency. At the same time, the resource utilization rate has increased from 65.2% to 85.1%, and the solution consistency has increased from 69.0% to 87.4%. Especially after adding the stagnation detection and heartbeat monitoring mechanisms, the system performance has a significant leap, the convergence rounds have been greatly reduced, and at the same time, a high resource utilization rate and solution consistency have been maintained. This shows that the technical solution of the present invention can effectively balance resource utilization and solution quality while improving the decision-making efficiency. The present invention effectively solves the problems of sudden inventory anomalies and resource conflicts in the multi-level warehousing network by constructing a distributed emergency collaboration mechanism. The local response ability based on edge computing enables the system to maintain operation when the central system is unavailable. The comprehensive evaluation of logistics correlation and resource response ability ensures the scientific nature of collaborative decision-making. The dynamic calculation of node credibility and influence weight improves the negotiation efficiency and solution quality. The automatic detection and restart mechanism for negotiation stagnation enhances the system's ability to find the optimal solution, and the heartbeat monitoring and task reallocation mechanisms improve the system's resilience in the face of node failures. This distributed collaboration architecture significantly reduces the exception handling time, improves the resource utilization efficiency, enhances the overall resilience of the supply chain, and provides strong technical support for enterprises to cope with the complex and changeable market environment.

[0082] In the second aspect of the embodiments of the present invention, a multi-level warehousing intelligent scheduling and collaboration system for the supply chain is provided, including: A first unit for obtaining real-time inventory data and historical order data of each level of warehouse in the multi-level warehousing network; A second unit for hierarchically modeling and predicting the short-term and long-term demands of each level of warehouse based on the historical order data by using a long short-term memory network to obtain the inventory demand; A third unit for constructing a multi-agent reinforcement learning model, setting each warehouse node as an independent agent. The state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand, and available resource quantity of each warehouse node, and the action space includes the inventory allocation quantity and allocation timing. An inventory allocation strategy is generated through policy iteration optimization among the agents; A fourth unit, configured to convert the inventory deployment strategy into a scheduling instruction including deployment objects, deployment quantities, and deployment times based on a rule knowledge base; A fifth unit, configured to receive the scheduling instruction in edge computing units of each warehouse node, sort and execute the scheduling instruction, and when detecting inventory anomalies or resource conflicts, initiate an emergency collaboration request to adjacent warehouse nodes. The warehouse node that receives the emergency collaboration request returns response information according to its own resource status, and determines an emergency handling plan through local negotiation between nodes according to the response information.

[0083] In a third aspect of the embodiments of the present invention, a computer-readable storage medium is provided, on which computer program instructions are stored. When the computer program instructions are executed by a processor, the foregoing method is implemented.

Claims

1. A multi-level warehouse intelligent scheduling and collaboration method for a supply chain, characterized in that Including: Obtain the real-time inventory data and historical order data of warehouses at all levels in a multi-level warehousing network; Based on the historical order data, use a long short-term memory network to hierarchically model and predict the short-term and long-term demands of warehouses at all levels to obtain the inventory demand; Construct a multi-agent reinforcement learning model, set each warehouse node as an independent agent, the state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand, and available resource volume of each warehouse node, and the action space includes the inventory allocation quantity and allocation timing. Generate an inventory allocation strategy through policy iteration optimization among agents; Convert the inventory allocation strategy into a scheduling instruction including the allocation object, allocation quantity, and allocation time based on a rule knowledge base; In the edge computing unit of each warehouse node, receive the scheduling instruction and execute it in sequence. When inventory anomalies or resource conflicts are detected, initiate an emergency collaboration request to adjacent warehouse nodes. The warehouse node that receives the emergency collaboration request returns response information according to its own resource status. According to the response information, determine an emergency handling plan through local negotiation between nodes.

2. The method according to claim 1, characterized in that, Based on the historical order data, the steps of using a long short-term memory network to hierarchically model and predict the short-term and long-term demands of warehouses at all levels to obtain the inventory demand include: Divide the historical order data into daily order sequences and weekly order sequences; Construct an adaptive hierarchical prediction network, including a short-term prediction layer and a long-term prediction layer. The short-term prediction layer receives the daily order sequence and dynamically adjusts the prediction time window according to the inventory goods category and seasonal characteristics. The long-term prediction layer receives the weekly order sequence and captures long-cycle patterns through a memory enhancement module; the output error gradient of the short-term prediction layer is transmitted to the parameter update process of the long-term prediction layer, and the trend characteristics of the long-term prediction layer are embedded into the input end of the short-term prediction layer; Train the adaptive hierarchical prediction network using a multi-scale loss function. The multi-scale loss function includes a short-term prediction loss term, a long-term prediction loss term, and a regularization term, and dynamically weights the loss terms based on the time-scale characteristics of each prediction layer; Integrate the prediction results to obtain an initial prediction value; calibrate the initial prediction value based on the warehousing capacity constraint, and generate an inventory demand prediction result with a confidence interval in combination with the prediction standard deviation.

3. The method according to claim 1, wherein The steps of constructing a multi-agent reinforcement learning model, setting each warehouse node as an independent agent, and generating an inventory allocation strategy through policy iteration optimization among agents include: Set each warehouse node as an independent agent, construct the state vector of the independent agent, and the state vector includes real-time inventory data, inventory demand, and available resource volume; construct an adjacency matrix based on the logistics channels between each warehouse node, and the adjacency matrix is used to represent the connection relationship between each warehouse node; Construct the action space of the independent agent, and the action space includes an allocation quantity matrix and an allocation timing vector. The allocation quantity matrix is used to represent the inventory allocation quantity between warehouse nodes, and the allocation timing vector is used to represent the allocation execution time; Construct a multi-objective reward function, which includes a local reward function and a global reward function. The local reward function calculates the reward value of each independent agent based on inventory balance, turnover efficiency, and scheduling cost, and the global reward function calculates the global collaborative reward value based on the inventory level difference of adjacent warehouse nodes; Use the Proximal Policy Optimization (PPO) algorithm to train the independent agents. During the training process, construct a hierarchical experience buffer to store state transition samples. The hierarchical experience buffer includes a local buffer for storing the local experience of independent agents and a global buffer for storing global state information; use a parameter sharing mechanism to train the policy network for the same type of independent agents; Based on the trained policy network, generate an inventory allocation strategy according to the real-time state of each warehouse node.

4. The method according to claim 3, characterized in that The steps for constructing the multi-objective reward function include: Classify items based on the historical order data of each warehouse node and determine the target inventory level; calculate the standardized inventory deviation of each warehouse node based on the target inventory level, and determine the inventory balance reward value as the product of the standardized inventory deviation and the time decay factor; Calculate the daily turnover rate, weekly turnover rate, and monthly turnover rate based on the outbound volume, initial inventory, and final inventory of each warehouse node respectively, and determine the turnover efficiency reward value as the weighted sum with a preset weight coefficient; Determine the scheduling cost reward value based on the weighted sum of storage cost, transportation cost, and operation cost; Calculate the ratio of inventory level differences between adjacent warehouses to construct a difference matrix, calculate the distance weight coefficient based on the logistics distance between adjacent warehouses, and determine the global collaborative reward value as the product of the distance weight coefficient, the difference matrix, the warehouse level collaboration coefficient, and the congestion penalty factor; Calculate the performance improvement rate of the inventory balance reward value, turnover efficiency reward value, scheduling cost reward value, and global collaborative reward value within a preset evaluation period; construct a multiple linear regression model based on the performance improvement rate to calculate the correlation coefficient matrix, calculate the contribution degree of each reward value according to the correlation coefficient matrix, multiply the difference between the contribution degree and the current weight by the learning rate to obtain the weight adjustment amount, and update the weight coefficient of each reward value under the weight adjustment constraint condition to obtain the multi-objective reward function.

5. The method according to claim 3, wherein The steps for using the Proximal Policy Optimization (PPO) algorithm to train the independent agents and constructing a hierarchical experience buffer to store state transition samples during the training process include: The hierarchical experience buffer includes a local experience buffer for storing the local state transition samples of each independent agent, a collaborative experience buffer for storing the interaction samples of adjacent agents, and a global experience buffer for storing global state samples; Calculate the importance score of the state transition sample, which is determined based on the reward value, state transition frequency, state influence range, and policy difference degree. Store the state transition sample in the corresponding level of the experience buffer according to the importance score; Calculate the sampling probability of the sample based on the sample timeliness and the number of policy updates, and select training samples from the hierarchical experience buffer according to the sampling probability; Update the policy network parameters of a single agent based on samples in the local experience buffer, update the policy network parameters of the adjacent agent group based on samples in the collaborative experience buffer, and update the policy network parameters of all agents based on samples in the global experience buffer; Construct a knowledge sharing weight matrix and calculate the element values of the knowledge sharing weight matrix based on the historical decision-making data and network topology of the agents; Divide the objective function of the proximal policy optimization algorithm into an immediate optimization term, a periodic optimization term, and a long-term optimization term, calculate the weight coefficients of each optimization term based on the knowledge sharing weight matrix and system state information, and optimize and update the policy network parameters.

6. The method according to claim 1, characterized in that, The steps of converting the inventory allocation policy into a scheduling instruction including the allocation object, allocation quantity, and allocation time based on the rule knowledge base are as follows: Construct a rule knowledge base including physical constraint rules, business constraint rules, and priority rules, and parse the inventory allocation policy into an initial scheduling plan including warehouse node pairs, allocation quantities, and allocation time windows; Based on the rule knowledge base, perform constraint verification on the initial scheduling plan, and construct a rule matching degree scoring function. The rule matching degree scoring function calculates the matching degree score based on the degree of satisfaction of physical constraints, the achievement of business goals, and resource utilization rate, and automatically corrects the scheduling plan that fails the constraint verification or has a matching degree score lower than the threshold; Based on the scheduling delay rate, resource conflict rate, and inventory volatility in the historical scheduling execution data, calculate the execution risk probability of the corrected scheduling plan in different time windows, generate an execution priority considering risk aversion, and convert the sorted scheduling plan into a scheduling instruction including the allocation object, allocation quantity, and allocation time.

7. The method according to claim 1, characterized in that, The steps for the warehouse node that receives the emergency collaboration request to return response information based on its own resource status and determine the emergency handling plan through local negotiation among nodes according to the response information are as follows: Calculate the logistics correlation based on the historical logistics frequency and current logistics intensity of the warehouse node, calculate the resource response ability based on the ratio of the required resources of the request to the available resources of itself, and generate response information including response time, response cost, and resource status according to the logistics correlation and resource response ability; Obtain the task completion rate and resource stability of the warehouse nodes participating in the negotiation, calculate the node credibility in combination with the response information, generate an influence weight based on the number of hops between nodes, and the influence weight decreases exponentially with the increase of the number of hops, and determine the initial negotiation weight according to the node credibility and influence weight; Receive the negotiation plans of the neighboring warehouse nodes, calculate the weighted average value of the plans based on the initial negotiation weight, combine the weighted average value with the current optimal plan to form an updated local plan, and calculate the variance value of the updated local plan in the negotiation group as a consistency index; When the update amplitude of the plan is continuously lower than the preset amplitude threshold, adjust the weights of cost, time, and resource consumption in the negotiation plan according to the historical execution effect, and generate a random perturbation value positively correlated with the stagnation duration to restart the negotiation; generate an emergency handling plan when the consistency index reaches the preset condition; Monitor the heartbeat information and response latency of the warehouse nodes participating in the negotiation. When the interruption of the heartbeat information or the response latency exceeds the preset value is detected, obtain the unfinished negotiation tasks of the nodes with abnormal heartbeats, calculate the task reception suitability based on the resource status, response time, and logistics distance of the remaining warehouse nodes, and reallocate the unfinished negotiation tasks under the condition of meeting the resource constraints according to the task reception suitability.

8. A multi-level warehouse intelligent scheduling and collaboration system for a supply chain, which is used to implement the method described in any one of the foregoing claims 1-7, and is characterized in that, Including: The first unit is used to obtain the real-time inventory data and historical order data of each level of warehouses in the multi-level warehousing network; The second unit is used to hierarchically model and predict the short-term and long-term demands of each level of warehouses based on the historical order data by using a long short-term memory network to obtain the inventory demand; The third unit is used to construct a multi-agent reinforcement learning model, set each warehouse node as an independent agent, the state space of the multi-agent reinforcement learning model includes the real-time inventory data, inventory demand, and available resource volume of each warehouse node, the action space includes the inventory allocation quantity and allocation timing, and generate an inventory allocation strategy through policy iteration optimization among the agents; The fourth unit is used to convert the inventory allocation strategy into a scheduling instruction including the allocation object, allocation quantity, and allocation time based on the rule knowledge base; The fifth unit is used to receive and execute the scheduling instruction in the edge computing unit of each warehouse node in order. When inventory anomalies or resource conflicts are detected, an emergency collaboration request is sent to adjacent warehouse nodes. The warehouse node that receives the emergency collaboration request returns response information according to its own resource status, and an emergency handling plan is determined through local negotiation between the nodes according to the response information.

9. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by a processor, the method described in any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • A distributed inventory scheduling system and an improved method

    CN109255481A

  • Logistics order analysis method, apparatus and device, and storage medium

    CN117952397A

  • Intelligent logistics supply chain digital management system based on data analysis

    CN119005838A

  • Multi-supply chain scheduling method and system based on global Critic multi-agent algorithm

    CN119090223A

  • Coal mining intelligent scheduling method based on artificial intelligence platform

    CN119090243A

Cited By

  • Multi-intelligent-vehicle material collaborative distribution method and system

    CN120494447A

  • Multi-Intelligent Vehicle Collaborative Delivery Method and System

    CN120494447B

  • Green building material intelligent scheduling method and system based on reinforcement learning

    CN120822785A

  • Supply chain material inventory control method and system

    CN120931211A

  • AI agent emergency order insertion dynamic decision production scheduling method, medium and system

    CN121010187A