A data-driven perishable multi-level inventory level optimization method and system
By using deep learning and deep reinforcement learning algorithms, a multi-level inventory optimization model for perishable goods is constructed. This model solves the problem of dynamic adjustment of demand forecasting and multi-level inventory management in the perishable goods supply chain, achieving efficient multi-level logistics scheduling and precise delivery of perishable goods, and reducing inventory waste and spoilage.
Patent Information
- Application Number
- CN202511165173.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-20
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-08-20
AI Technical Summary
In existing technologies for multi-level inventory management of perishable goods, demand forecasting relies on static models and lacks dynamic adjustment capabilities. This makes it difficult to adapt to seasonal fluctuations and unforeseen events in the perishable goods supply chain. Furthermore, insufficient multi-level inventory optimization strategies lead to inventory waste and spoilage problems.
By employing deep learning and deep reinforcement learning algorithms, combined with historical sales data and external feature data, a demand prediction network and a multi-level transshipment optimization model are constructed. Through deep reinforcement learning, multi-level logistics scheduling is carried out to optimize transshipment volume and delivery frequency, thereby reducing the spoilage and waste of perishable goods.
It enables high-precision forecasting of demand for perishable goods and dynamic optimization of multi-level transfer volumes, improving the flexibility and efficiency of the supply chain and reducing waste caused by spoilage of perishable goods due to transportation and warehousing.
Smart Images

Figure CN120655215B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of deep reinforcement learning, and in particular to a data-driven perishable multi-level inventory level optimization method and system. BACKGROUND
[0002] Perishable products (such as fresh food, flowers, medicines, etc.) have the characteristics of short shelf life and high demand volatility, and their supply chain management faces significant challenges. In a multi-level distribution system, how to dynamically grasp the demand of retail terminals and adjust the inventory level and transfer strategy accordingly is the key to affecting the quality of perishable products and reducing losses. In recent years, with the development of artificial intelligence and data-driven methods, advanced algorithms such as deep learning and reinforcement learning have been introduced into the field of supply chain management, especially in the multi-level inventory management scenario. How to achieve end-to-end demand prediction and transfer optimization has gradually become a hot topic of interest in the industry.
[0003] The existing invention scheme mainly adopts two types of method paths:
[0004] One is a kind of engineering material multi-level inventory control method and system provided by the patent with the announcement number CN111292044B. This scheme improves the rationality of the allocation plan by constructing multi-level inventory levels and simulation mechanisms, but still has several limitations. First, its dependence on demand prediction is relatively static, only relying on monthly plans or historical average to calculate daily demand, without introducing complex variables (such as weather, market activities) for dynamic modeling, resulting in insufficient prediction sensitivity and responsiveness. Second, the generation of the allocation plan is mainly based on inventory upper and lower limit logic and reasonable inventory levels set by experience, lacking a system learning and automatic optimization mechanism, making it difficult to adapt to supply environments with large seasonal demand fluctuations or frequent unexpected events. In addition, although this method introduces simulation, the decision feedback mechanism is not closed-loop, and the simulation results are not used to continue driving parameter self-updating after result determination, making it difficult for the overall strategy to solve the multi-period multi-level inventory problem. The above scheme cannot adapt to the inventory optimization prediction of perishable products.
[0005] Another type is provided by a patent with publication number CN113592309B, a multi-level inventory quota formulation method based on data driving, which emphasizes data driving and model iteration, and can improve the adaptability of inventory quota, but is still weak in overall supply chain coordination optimization. The core idea focuses on "quota setting", that is, determining the upper and lower limits of inventory in each warehouse, and does not involve more complex inventory turnover strategies or dynamic replenishment path optimization, especially in the horizontal coordination between multi-level warehouse nodes. In addition, although this method uses neural networks and other technologies to improve prediction ability, it is mainly suitable for materials with stable periodic demand, and has poor adaptability for perishable goods or fast-moving consumer goods with complex demand structure or sudden characteristics. At the same time, its algorithm strategy is still based on pre-set static relationships (such as service level and safety stock correspondence table), which limits the flexibility and expansion capability of the model under multi-objective conditions.
[0006] Therefore, it is urgent to provide a data-driven perishable multi-level inventory level optimization method and system, which faces the perishable supply chain scene, provides an efficient dynamic decision-making model to solve the multi-level inventory level optimization problem. SUMMARY
[0007] In view of the problems existing in the prior art, the present application provides a data-driven perishable multi-level inventory level optimization method and system, which faces the perishable supply chain scene, uses deep learning and deep reinforcement learning algorithm, realizes demand prediction based on real-time sales data and external feature data, optimizes multi-period transfer quantity for multi-warehouse and multi-retail point supply chain structure, and finally realizes multi-level logistics scheduling from supply point to warehouse and then to retail point, optimizes distribution frequency and batch combination.
[0008] The technical scheme of the present application is as follows:
[0009] A data-driven perishable multi-level inventory level optimization method, comprising the following steps:
[0010] S1, constructing a neural network, that is, a demand prediction network; for a plurality of retail nodes, collecting historical sales data and external feature data of the perishable goods, and inputting the neural network, training the neural network;
[0011] S2, collecting transportation environment information of the perishable goods; the transportation environment information includes time series order quantity and transfer quantity; the time series order quantity represents the predicted order demand quantity of the jth retail node for the perishable goods output by the neural network on the tth day; the transfer quantity represents the transfer quantity of any one warehouse node to any one retail node for the perishable goods;
[0012] S3, constructing and initializing a deep reinforcement learning network, including a training network with the same structure and a target network ; initializing an experience replay pool;
[0013] S4, defining a state and an action based on the transportation environment information; the state represents all the transit quantities on the t-th day; the action represents the transit quantities dispatched on the t-th day; based on the transportation environment information, a reward function is constructed, which is used to generate a reward value for the state and the action selected by it ;
[0014] It can be understood that the transit quantity in the state refers to the quantity of goods that have been dispatched but have not yet arrived at the destination and are located on the way; for example, on the first day, 50 units of goods are dispatched from warehouse 1 to retail store 1, which need 3 days to arrive, so on the second day, the 50 units still need 2 days to arrive, that is, 50 is the transit quantity in the state.
[0015] The transit quantity dispatched on the same day refers to the change quantity, which refers to the quantity of goods dispatched from the warehouse node to the sales node on a specified date, that is, the newly added quantity. For example, on the second day, 30 units of goods are dispatched from warehouse 1 to retail store 1, so 30 is the transit quantity dispatched on the same day.
[0016] S5, iteratively training the training network;
[0017] The training is: selecting an action according to the state , completing state update , generating state transition , and storing the state transition as a sample in the experience replay pool; randomly selecting a plurality of samples from the experience replay pool, updating the network parameters in combination with the target function ; periodically synchronizing to ;
[0018] S6, outputting an optimal transit quantity sequence from the target network.
[0019] In deep reinforcement learning, the training network includes an action space, and the action space has a plurality of actions; during the iterative training process, for a state, a most suitable action is selected to make the current state reach the next state.
[0020] The deep reinforcement learning network is used for estimating Q values corresponding to a current state and a selected action. The network for training includes an input layer, a hidden layer and an output layer; the input layer receives encoded features of the state and the action; the hidden layer adopts a multi-layer fully connected network, and includes an activation function ReLU; the output layer outputs Q values of all possible actions, i.e. expected values of selecting each action in the current state. Parameters are synchronized from the training network to the target network every M steps (iteration times).
[0021] Perishable goods have the characteristics of continuous corruption over time, and strict requirements are required for transportation and storage time to ensure usability. The data-driven multi-level inventory level optimization method for perishable goods provided by the application realizes two major parts in the process, one is to predict the unknown demand of each retail point for perishable goods with high precision based on the multi-level supply chain network structure, the attributes of perishable goods, historical sales data and external feature data through deep learning, and the other is to optimize the multi-level transfer quantity from the warehouse end to the retail end through deep reinforcement learning method, maximize the transfer efficiency and accurate distribution of perishable goods, and reduce the phenomenon of perishable goods corruption and waste caused by transportation reasons, too much distribution / warehouse overstock, perishable goods characteristics and the like.
[0022] As a further optimization of the above scheme, the transportation environment information includes supply chain structure information, perishable goods information and transportation data;
[0023] The supply chain structure information includes the number I of warehouse nodes and the number J of retail nodes;
[0024] The perishable goods information includes the deterioration rate of the perishable goods , the transfer quantity, the unit deterioration cost and the timing order quantity , ;
[0025] The transportation data includes the lead time ; the lead time indicates the transportation time required from the i-th warehouse node to the j-th retail node, .
[0026] As a further optimization of the above scheme, the state is represented as ;
[0027] wherein, ; indicates the in-transit transfer quantity between the i-th warehouse node and the j-th retail node with a lead time of days and an interval of days; L is the maximum value of ;
[0028] The action is represented as ;
[0029] wherein, ; represents the day, the i-th warehouse node to the j-th retail node issued the transfer of the perishable goods quantity;
[0030] when i = 1, ; when i > 1, ;
[0031] wherein, is an indicator function, when represents 1, otherwise 0; represents the remaining coefficient of the perishable goods when the lead time is ;
[0032] Before the iterative training, set the initial state , at this time, . Based on the initial state, the next state is obtained in turn.
[0033] As a further optimization of the above scheme, the reward function is represented as:
[0034] ;
[0035] wherein, is the unit selling price of the perishable goods, represents the fixed transportation cost from the i-th warehouse node to the j-th retail node j; represents the unit variable transportation cost from the i-th warehouse node to the j-th retail node j; is an indicator function, when represents 1, otherwise 0; represents the remaining coefficient of the perishable goods when the lead time is .
[0036] As a further optimization of the above scheme, the state update is represented as:
[0037] when , , k = k-1;
[0038] when , ;
[0039] wherein, is an indicator function, when represents 1, otherwise 0; the updated as The parameters.
[0040] Assuming L=10 and k=5, the update is represented as follows: .
[0041] As a further optimization of the above scheme, based on a greedy strategy, the state is... Select Action ,Right now:
[0042] ;
[0043] Where ε represents the preset exploration rate; |A| represents the number of actions; Indicates finding The largest 'a'; further, Where e represents the natural constant; n represents the number of iterations; This represents the exploration temperature value, which is typically initialized to a very small value, such as 0.00005. As the number of iterations increases, the exploration rate gradually approaches 0.
[0044] For the sample, calculate the target Q value, i.e.:
[0045] + ;
[0046] The objective function is expressed as:
[0047] ;
[0048] Where γ is a preset discount factor; N is the number of samples drawn; The target Q value corresponding to the b-th sample; Indicates finding The largest The network parameter ω is updated by minimizing L(ω).
[0049] As a further optimization of the above scheme, the optimal transshipment volume sequence is the transshipment volume prediction for the target date T and each day prior to it; the optimal transshipment volume sequence is expressed as follows: ,Right now:
[0050] .
[0051] As a further optimization of the above scheme, the construction and training of the neural network are as follows:
[0052] S1-1, divide the data set into a training set and a test set according to a preset ratio; specifically, for the sample quantity of the data set, 8:2 is respectively allocated to the training set and the test set. The data set includes one-to-one historical sales data and external feature data, that is, historical sales data and external feature data on the same date; a minimum error loss value is set ;
[0053] S1-2, confirm the input layer, the hidden layer and the output layer, and adopt a feedforward ReLU network; although other activation functions such as Sigmoid and Tanh can be used, the ReLU network has good theoretical properties in terms of mathematical simplicity, for example, the universal approximation property; wherein the hidden layer includes K layers, each layer has e neurons; let , then
[0054] ;
[0055] wherein, represents a feature vector input into the input layer; represents an output vector before activation at the kth layer of neural network; represents an output vector after activation at the kth layer; represents a ReLU activation mode; and respectively represent a weight vector and a bias vector of the kth layer of hidden layer; represents a parameter set of the neural network, that is ; and respectively represent a weight vector and a bias vector of the output layer; that is, for an input feature vector x, the neural network outputs a predicted order demand;
[0056] S1-3, back propagation, according to the chain rule, the gradient is descended along the reverse direction of the neural network, and the loss function is minimized; the loss function is represented as:
[0057] ;
[0058] wherein, represents a mean square error, and N is the sample quantity in the test set; represents ; b and h represent a unit backorder cost and a unit overstock cost of one retail node; is a cumulative distribution function of , that is, a predicted output obtained by the neural network after inputting a feature vector; and respectively represent the external feature data and the historical sales data corresponding to the i-th sample; specifically, the historical sales data includes serialized sales records in the form of (historical date, retail point, sales volume).
[0059] S1-4, convergence determination and effect verification; only when convergence is determined; the test set is verified using the converged neural network. That is, the values corresponding to the feature vectors (external feature data) of the test set are taken as inputs, and it is determined whether the output values of the neural network model are consistent with the corresponding demand values (historical sales data). If the overall accuracy meets the requirements (greater than or equal to a preset accuracy ), it indicates that the generalization ability is strong, otherwise it indicates overfitting, and the neural network needs to be adjusted, for example, increasing the training set, reducing the network structure dimension size, introducing regularization in the loss function, and introducing a random dropout mechanism in the parameter network.
[0060] As a further optimization of the above scheme, one of the external feature data is composed of multiple observable external features, represented in the form of a vector; the external features include the geographic location of the retail node, the passenger flow of the retail node within a specified time period, the weather data of the geographic location of the retail node, and the date type within a specified time period.
[0061] According to the geographic location, the radiation range of sales can be calculated; according to the passenger flow, the total amount of potential consumers can be calculated / predicted; according to the weather data, the influence of weather conditions on sales can be reflected; the date type can be divided into weekdays, weekends, holidays, etc., and the influence of the date type on sales is explored by the neural network.
[0062] The application also provides a data-driven perishable multi-level inventory level optimization system, which applies a data-driven perishable multi-level inventory level optimization method as described above;
[0063] The system includes a historical data module, a demand prediction module, a transfer data module, a transfer optimization module, and a storage data module.
[0064] The historical data module and the demand prediction module jointly perform S1; the historical data module is used to collect and organize the historical sales data and the external feature data; and the demand prediction module is used to output the time series order quantity.
[0065] The transfer data module and the transfer optimization module jointly perform S2-S6; wherein the transfer data module is used to collect transportation environment information about the perishable goods and initialize parameters of the deep reinforcement learning network; and the transfer optimization module is used to train the deep reinforcement learning network and output the optimal transfer quantity sequence.
[0066] The storage data module is used for storing data generated when S1-S6 are performed.
[0067] Compared with the prior art, the present application has the following beneficial effects:
[0068] The present application constructs a dynamic transfer strategy through deep learning and deep reinforcement learning method. Specifically, through a deep learning algorithm, a neural network combines historical sales data and future observable external feature data, and the model can accurately describe the distribution of uncertain demand; then, the predicted demand, perishable product attributes, supply chain structure and other resource information are sorted out and input into the "transfer optimization module" to realize multi-level logistics scheduling from the supply point to the warehouse and then to the retail point, and to optimize the distribution frequency and batch combination. Finally, the transfer efficiency and accurate distribution of perishable products are maximized, and the phenomenon of perishable product spoilage and waste caused by transportation, excessive distribution / storage, and perishable product characteristics is reduced. BRIEF DESCRIPTION OF DRAWINGS
[0069] Figure 1 is a structural schematic diagram of a neural network provided by an embodiment of the present application;
[0070] Figure 2 is a data interaction schematic diagram of a data-driven multi-level inventory level optimization method for perishable products provided by an embodiment of the present application;
[0071] Figure 3 is a module schematic diagram of a data-driven multi-level inventory level optimization system for perishable products provided by an embodiment of the present application. DETAILED DESCRIPTION
[0072] In order to make the purpose, technical scheme and advantages of the present application clearer, the technical scheme in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor fall within the scope of protection of the present application.
[0073] As shown in Figure 1 , Figure 2 The present embodiment provides a data-driven multi-level inventory level optimization method for perishable products, comprising the following steps:
[0074] S1, a neural network, i.e. a demand prediction network, is constructed; for a plurality of retail nodes, historical sales data and external feature data about perishable products are collected and input into the neural network, and the neural network is trained.
[0075] An external feature data is composed of a plurality of observable external features, represented in the form of a vector; in the embodiment, the external features include the geographical location of the retail node, the foot traffic of the retail node in a specified time period, the weather data of the geographical location where the retail node is located, the date type in the specified time period, i.e., the external feature data = {geographical location, foot traffic, weather data, date type}. The radiation range of sales can be calculated according to the geographical location; the total amount of potential consumers can be calculated / predicted according to the foot traffic; the influence of weather conditions on sales can be reflected according to the weather data; the date type can be divided into weekdays, weekends, holidays, etc., and the influence of the date type on sales is explored by the neural network based on this.
[0076] Specifically, the historical sales data includes serialized sales records in the form of (historical date, retail point, sales volume).
[0077] Specifically, in the embodiment, the construction and training of the neural network are as follows:
[0078] S1-1, divide the data set into a training set and a test set according to a preset proportion; specifically, for the sample quantity of the data set, divide the training set and the test set according to 8:2 respectively. The data set includes one-to-one historical sales data and external feature data, i.e., the historical sales data and the external feature data on the same date; set the minimum error loss value .
[0079] S1-2, confirm the input layer, the hidden layer and the output layer, and use the feedforward ReLU network; although other activation functions such as Sigmoid and Tanh can be used, the ReLU network has good theoretical properties in terms of mathematical simplicity, for example, the universal approximation property; wherein the hidden layer includes K layers, each layer has e neurons; let , then
[0080] ;
[0081] wherein, represents the feature vector input into the input layer; represents the output vector before activation at the kth layer of the neural network; represents the output vector after activation at the kth layer; represents the ReLU activation mode; and represent the weight vector and the bias vector of the kth hidden layer respectively; represents the parameter set of the neural network, i.e., ; and represent the weight vector and the bias vector of the output layer respectively; , which means that for an input feature vector x, the neural network outputs the predicted order demand.
[0082] S1-3, back propagation, gradient descent along the neural network in the opposite direction according to the chain rule, minimizing the loss function; the loss function is represented as:
[0083] ;
[0084] wherein, represents the mean square error, and N is the number of samples in the test set; represents ; b and h represent the unit backorder cost and the unit overstock cost of a retail node; is the cumulative distribution function of , that is, the predicted output obtained by the neural network after inputting the feature vector; and represent the external feature data and the historical sales data corresponding to the i-th sample, respectively.
[0085] S1-4, convergence judgment and effect verification; only when , it is judged to be converged; the converged neural network is used to verify the test set. That is, the values corresponding to the feature vectors (external feature data) of the test set are taken as inputs, and it is calculated whether the output values of the neural network model are consistent with the corresponding demand values (historical sales data). If the overall accuracy meets the requirements (greater than or equal to the preset accuracy ), it means that the generalization ability is strong, otherwise it means overfitting, and the neural network needs to be adjusted, for example, increasing the training set, reducing the network structure dimension size, introducing regularization in the loss function, and introducing a random dropout mechanism in the parameter network.
[0086] S2, collect the transportation environment information of perishable goods. In this embodiment, the transportation environment information includes supply chain structure information, perishable goods information and transportation data.
[0087] The supply chain structure information includes the number I of warehouse nodes and the number J of retail nodes.
[0088] The perishable goods information includes the perishable rate , the unit perishable cost , the transfer quantity and the time series order quantity , ; the time series order quantity represents the predicted order demand of the j-th retail node for perishable goods output by the neural network on the t-th day. The transfer quantity represents the transfer quantity of any one warehouse node to any one retail node for perishable goods.
[0089] The transportation data includes the lead time ; the lead time This represents the transportation time required from the i-th warehouse node to the j-th retail node. .
[0090] S3. Construct and initialize a deep reinforcement learning network, including a training network with the same structure. and target network Initialize the experience replay pool. A deep reinforcement learning network is used to estimate the Q-values corresponding to the current state and the chosen action. The training network consists of an input layer, hidden layers, and an output layer; the input layer receives the encoded features of the state and action; the hidden layers employ a multi-layer fully connected network, including the ReLU activation function; the output layer outputs the Q-values of all possible actions, i.e., the expected value of choosing each action in the current state.
[0091] S4. Define status based on transportation environment information. and actions ;state This represents the total amount of transshipments in transit on day t; the action represents the amount of transshipments dispatched on day t.
[0092] Understandably, the "in-transit" transit volume refers to the status quantity, that is, the quantity of goods that have been dispatched but have not yet arrived at their destination and are en route; for example, on the first day, warehouse 1 dispatches 50 units of goods to retail store 1, which takes 3 days to arrive, then on the second day, these 50 units will take another 2 days to arrive, that is, 50 is the "in-transit" transit volume.
[0093] The "same-day shipment" volume refers to the change in quantity, which is the number of goods shipped from the warehouse node to the sales node on a specified date, i.e., the new quantity. For example, if warehouse 1 continues to ship 30 units of goods to retail store 1 on the next day, then 30 is the "same-day shipment" volume.
[0094] Specifically, status Represented as ;
[0095] in, ; Indicates the lead time is Day, the interval between the i-th warehouse node and the j-th retail node Daily transit volume; L is The maximum value.
[0096] action Represented as ;
[0097] in, ; Indicates the first The amount of perishable goods transferred from the i-th warehouse node to the j-th retail node per day.
[0098] when i = 1, when i > 1, ;
[0099] wherein, is an indicator function, when denotes 1, otherwise 0; denotes the lead time is , the remaining coefficient of perishable goods.
[0100] The reward function is constructed based on the transportation environment information, and the reward function is used to generate a reward value for a state and a selected action of the state ; in the embodiment, the reward function is represented as:
[0101] ;
[0102] wherein, is the unit selling price of perishable goods, denotes the fixed transportation cost from the i-th warehouse node to the j-th retail node j; denotes the unit variable transportation cost from the i-th warehouse node to the j-th retail node j; is an indicator function, when denotes 1, otherwise 0; denotes the lead time is , the remaining coefficient of perishable goods. denotes that the parameter needs to be passed in for calculating the reward value .
[0103] S5, iteratively train the training network; before the iterative training, set an initial state , at this time, all . Based on the initial state, the next state is obtained in turn. In deep reinforcement learning, the training network contains an action space, and the action space has multiple actions; during the iterative training process, for a state, a most suitable action is selected, so that the current state reaches the next state.
[0104] The training is: according to the state , an action is selected, the state update is completed, the state transition is generated, and the state transition is stored in the experience replay pool as a sample;
[0105] Specifically, in the embodiment, according to the greedy strategy, an action is selected for the state , that is:
[0106] ;
[0107] wherein ε represents a preset exploration rate. ; wherein e represents a natural constant; n represents the number of iterations; represents an exploration temperature value, initialized as 0.00005. As the number of iterations increases, the exploration rate gradually approaches 0. |A| represents the number of actions; represents finding the a that maximizes .
[0108] In the present embodiment, the state update is represented as:
[0109] When , , k = k - 1; here it means that the change in the amount of transit transportation is calculated first, and then the remaining transportation time is reduced by one day. For example, assuming that L = 10, k = 5, then the update is represented as .
[0110] When , .
[0111] wherein is an indicator function, which represents 1 when represents 0; the updated is used as a parameter of .
[0112] A number of samples are randomly selected from the experience replay pool, and the network parameters are updated in combination with the target function ; specifically, for each sample, the target Q value is calculated, i.e.:
[0113] + ;
[0114] The target function is represented as:
[0115] ;
[0116] wherein γ is a preset discount factor; N is the number of extracted samples; corresponding to the bth sample; represents finding the that maximizes ; by minimizing L(ω), the network parameters ω are updated. Every M days (M times of iteration), the is synchronized to ;
[0117] S6, outputting the optimal transit volume sequence from the target network. In the present embodiment, the optimal transit volume sequence is the transit volume prediction for the target date T and each day before; the optimal transit volume sequence is represented as i.e.
[0118] .
[0119] Perishable goods have the characteristics of continuous corruption with time growth, in order to ensure the use, there are strict requirements for transportation and storage time. The application provides a data-driven perishable multi-level inventory level optimization method, which is divided into two parts in the process, one is based on multi-level supply chain network structure, perishable properties, historical sales data and external characteristic data, through deep learning, the unknown demand of each retail point for perishable goods is high-precision predicted, and the other is through deep reinforcement learning method, the multi-level transfer quantity optimization from warehouse end to retail end is carried out, the transfer efficiency and accurate distribution of perishable goods are maximized, and the phenomenon of perishable goods corruption and waste caused by transportation reasons, too much distribution / storage overstock, perishable characteristics and the like is reduced.
[0120] As Figure 3 shown, the embodiment also provides a data-driven perishable multi-level inventory level optimization system, which applies the data-driven perishable multi-level inventory level optimization method as described above;
[0121] The system comprises a historical data module, a demand prediction module, a transfer data module, a transfer optimization module and a storage data module.
[0122] The historical data module and the demand prediction module jointly execute S1; the historical data module is used for collecting and arranging historical sales data and external characteristic data; and the demand prediction module is used for outputting time series order quantity.
[0123] The transfer data module and the transfer optimization module jointly execute S2-S6; wherein the transfer data module is used for collecting transportation environment information about perishable goods, and parameter initialization of a deep reinforcement learning network; and the transfer optimization module is used for training the deep reinforcement learning network, and outputting an optimal transfer quantity sequence.
[0124] The storage data module is used for storing data generated when S1-S6 are executed.
[0125] According to the disclosure and teaching of the above description, those skilled in the art of the present application can also make changes and modifications to the above embodiments. Therefore, the present application is not limited to the specific embodiments disclosed and described above, and some modifications and changes of the present application should also fall within the protection scope of the claims of the present application. In addition, although some specific terms are used in the specification, these terms are only for convenience of explanation and do not constitute any limitation on the present application.
Claims
1. A data-driven method for optimizing multi-level inventory of perishable goods, characterized in that, Including the following steps: S1. Construct a neural network; For multiple retail nodes, collect historical sales data and external feature data of perishable goods respectively, and input them into the neural network to train the neural network; One external feature data consists of multiple observable external features, represented in the form of a vector; The external features include the geographical location of the retail node, the customer flow of the retail node within a specified time period, the weather data of the geographical location of the retail node, and the date type within the specified time period; S2. Collect information on the transportation environment of perishable goods; Transportation environment information includes time-series order volume and transshipment volume; time-series order volume This represents the predicted ordering demand for perishable goods at the j-th retail node output by the neural network on day t; the transfer volume is the number of perishable goods transferred from any warehouse node to any retail node. S3. Construct and initialize the deep reinforcement learning network, including training the network. Target network And the experience replay pool; S4. Define status based on transportation environment information. and actions ;state This represents the total amount of goods being transported on day t; action This represents the transshipment volume dispatched on day t; a reward function is constructed based on transportation environment information, which is used to generate reward values for the state and its selected action. ; S5. Iteratively train the training network; Training is based on the state. Select Action Completed status update Generate state transition The state transition is stored as a sample in the experience replay pool; Multiple samples are randomly selected from the experience replay pool, and the network parameters are updated in conjunction with the objective function. Regularly Synchronize to network parameters ; S6. Output the optimal transport volume sequence from the target network.
2. The data-driven method for optimizing multi-level inventory of perishable goods according to claim 1, characterized in that, The transportation environment information includes supply chain structure information, perishable goods information, and transportation data. The supply chain structure information includes the number of warehouse nodes (I) and the number of retail nodes (J); The perishable information includes the spoilage rate of the perishable product. The aforementioned transfer volume and unit spoilage cost and the time-series order quantity , ; The transportation data includes lead time. The aforementioned lead time This represents the transportation time required from the i-th warehouse node to the j-th retail node. .
3. The data-driven method for optimizing multi-level inventory of perishable goods according to claim 2, characterized in that, The state Represented as ; in, ; The lead time is indicated as Day, the interval between the i-th warehouse node and the j-th retail node Daily transit volume; L is The maximum value; The action Represented as ; in, ; Indicates the first The number of perishable goods transferred from the i-th warehouse node to the j-th retail node on a given day; When i=1 When i > 1, ; in, For indicator functions, when It represents 1 if it is not 1, otherwise it represents 0; Indicates the lead time is Within, the residual coefficient of the perishable goods; Before the iterative training, an initial state is set. ,at this time, .
4. The data-driven multi-level inventory level optimization method for perishable goods according to claim 3, characterized in that, The reward function is expressed as follows: ; in, The unit price of the perishable product. This represents the fixed transportation cost from the i-th warehouse node to the j-th retail node j; This represents the unit variable transportation cost from the i-th warehouse node to the j-th retail node j; For indicator functions, when It represents 1 if it is not 1, otherwise it represents 0; Indicates the lead time is At that time, the residual coefficient of the perishable product.
5. The data-driven method for optimizing multi-level inventory of perishable goods according to claim 3, characterized in that, The state update is represented as follows: when hour, , k=k-1; when hour, ; in, For indicator functions, when Represents 1 if it is 1, otherwise represents 0; Updated As The parameters.
6. The data-driven method for optimizing multi-level inventory of perishable goods according to claim 1, characterized in that, According to the greedy strategy, the state is... Select Action ,Right now: ; Where ε represents the preset exploration rate; |A| represents the number of actions; Indicates finding The largest 'a'; For the sample, calculate the target Q value, i.e.: + ; The objective function is expressed as: ; Where γ is a preset discount factor; N is the number of samples drawn; The target Q value corresponding to the b-th sample; Indicates finding The largest The network parameter ω is updated by minimizing L(ω).
7. The data-driven method for optimizing multi-level inventory of perishable goods according to claim 3, characterized in that, The optimal transshipment volume sequence is the transshipment volume prediction for the target date T and each day prior to it; the optimal transshipment volume sequence is expressed as follows: ,Right now: 。 8. The data-driven method for optimizing multi-level inventory of perishable goods according to claim 1, characterized in that, The construction and training of the neural network are as follows: S1-1. Divide the dataset into a training set and a test set according to a preset ratio; the dataset includes one-to-one historical sales data and external feature data; set a minimum error loss value. ; S1-2, Confirm the input layer, hidden layer, and output layer, using a feedforward ReLU network; the hidden layer consists of K layers; Let Then we have: ; in, This represents the feature vector input to the input layer; This represents the output vector before activation at the k-th layer of the neural network; This represents the output vector after activation at the k-th layer; Indicates the ReLU activation mode; and These represent the weight vector and bias vector of the k-th hidden layer, respectively. This represents the set of parameters of a neural network, i.e. ; and These represent the weight vector and bias vector of the output layer, respectively. S1-3. Backpropagation: According to the chain rule, gradient descent is performed in the reverse direction of the neural network to minimize the loss function; the loss function is expressed as: ; in, This represents the mean squared error, where N is the number of samples in the test set. express b and h represent the unit out-of-stock cost and unit backlog cost of one of the retail nodes; for The cumulative distribution function; and These represent the external feature data and the historical sales data corresponding to the i-th sample, respectively. S1-4. Convergence determination and effect verification; only when... When convergence is achieved, the converged neural network is used to verify the test set.
9. A data-driven multi-level inventory level optimization system for perishable goods, characterized in that, The method for optimizing the multi-level inventory of perishable goods based on any one of claims 1 to 8 is applied. The system includes a historical data module, a demand forecasting module, a transit data module, a transit optimization module, and a data storage module; The historical data module and the demand forecasting module jointly execute S1; the historical data module is used to collect and organize the historical sales data and the external characteristic data; the demand forecasting module is used to output the time-series order quantity; The transfer data module and the transfer optimization module jointly execute S2-S6; wherein, the transfer data module is used to collect transportation environment information based on the perishable goods, and initialize the parameters of the deep reinforcement learning network; the transfer optimization module is used to train the deep reinforcement learning network and output the optimal transfer volume sequence; The data storage module is used to store the data generated during the execution of S1-S6.
Citation Information
Patent Citations
A method and system for multi-level inventory control of engineering materials
CN111292044B
A data-driven multi-level inventory quota formulation method
CN113592309B
Drug inventory dynamic optimization method based on deep learning
CN118710194A
Supply chain distribution network multi-level inventory cost optimization method
CN119313121A