A supply chain resource control optimization method based on deep reinforcement learning

By constructing a multi-stage supply chain environment and using a soft actor-critic agent with deep reinforcement learning to dynamically adjust inventory strategies, the complexity problem in the collaborative control optimization of supply chain resources is solved, and the efficiency and flexibility of inventory management are improved.

CN119784292BActive Publication Date: 2025-10-17GUANGDONG UNIV OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411900442.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-23
Publication Date
2025-10-17
Estimated Expiration
2044-12-23

AI Technical Summary

Technical Problem

Existing technologies face the problems of many uncertain factors, high decision-making complexity, and increasing complexity of supply chain structure in the collaborative control optimization of supply chain resources. Traditional methods are difficult to achieve the optimization of overall operations.

Method used

A supply chain resource control optimization method based on deep reinforcement learning is adopted to construct a multi-stage supply chain environment. A soft actor-critic agent is used for dynamic inventory management. Decisions are made through state feature dense networks, action feature dense networks, soft actor networks and soft critic networks to dynamically adjust inventory strategies.

Benefits of technology

Effectively reduce inventory backlogs and out-of-stock risks, improve supply chain response speed and flexibility, achieve significant improvements in overall operational efficiency, and enhance the strategy generalization ability and robustness of intelligent agents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119784292B_ABST
    Figure CN119784292B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of supply chains and discloses a supply chain resource control optimization method based on deep reinforcement learning, which comprises the following steps: constructing a multi-stage supply chain environment; constructing a soft actor-critic intelligent agent with action decomposition; training the soft actor-critic intelligent agent using the multi-stage supply chain environment to obtain a trained soft actor-critic intelligent agent; and deploying the trained soft actor-critic intelligent agent to a target supply chain to output optimal action decisions. The application proposes a deep reinforcement learning intelligent agent based on problems such as demand fluctuation, complex structure, high-dimensional state and action space, and complex reward functions. The intelligent agent can predict future market trends by analyzing historical production order demands and out-of-stock records through the use of a feature extraction module and an action decomposition mechanism. Based on real-time supply chain inventory data, the intelligent agent can dynamically adjust inventory management strategies to effectively deal with various problems in supply chain resource collaborative control optimization, effectively reduce inventory accumulation and out-of-stock risks, improve the response speed and flexibility of the supply chain, and ultimately significantly improve the overall operating efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of supply chain resource control optimization, and particularly relates to a supply chain resource control optimization method based on deep reinforcement learning. BACKGROUND

[0002] According to data provided by the China Federation of Logistics and Purchasing, the ratio of total social logistics cost to GDP in China in the first half of 2024 was 14.2%, while the logistics cost accounted for about 7% of GDP in the United States. Compared with developed countries, the proportion of logistics cost to GDP in China is higher, which indicates that there is still room for improvement in supply chain efficiency compared with developed countries. Supply chain resource coordination control is one of the key technologies for enterprises to gain market competitive advantage, which focuses on exploring the most efficient product production, transportation and distribution path to ensure accurate market demand within the planned framework. This technology covers every link from raw materials to final products, forming a complete process chain.

[0003] The prior art discloses a multi-period inventory optimization management method based on a supply chain. Downstream node enterprises adopt a (s, Q) production strategy to control production and inventory. At the same time, a dynamic inventory model corresponding to the upstream node enterprises is established, and the upstream and downstream node enterprises work together to minimize the total inventory cost on the entire supply chain. This method is different from the case where each node enterprise only pursues the lowest inventory cost in a non-cooperative manner. This method can also derive the optimal production preparation period, optimal production period and optimal production capacity on the entire supply chain within a production cycle. In the prior art, the (s, Q) strategy is used to control production and inventory. In this strategy, when the inventory level drops to the reorder point , a procurement action is triggered to replenish a fixed amount of inventory. These strategies are widely used in practice because of their simplicity. However, there are many difficulties in the process of supply chain resource coordination control optimization. First, the supply chain network is full of uncertainty factors such as demand fluctuations and supply disruptions, which greatly increase the complexity of decision-making. Second, global economic integration has made the supply chain structure more complex, requiring more frequent interaction and collaboration between manufacturers and retailers to achieve overall operational optimization. The above difficulties make traditional rule-based methods perform poorly. SUMMARY

[0004] The purpose of the present application is to overcome the problems existing in the prior art, provide a supply chain resource control optimization method based on deep reinforcement learning, which can dynamically adjust the inventory management strategy and effectively cope with various problems in supply chain resource collaborative control optimization; not only reduces the inventory backlog and the risk of stockout, but also improves the response speed and flexibility of the supply chain, and finally realizes the significant improvement of the overall operation efficiency.

[0005] In order to achieve the above purpose, the present application provides a supply chain resource control optimization method based on deep reinforcement learning, which comprises the following steps:

[0006] Constructing a multi-stage supply chain environment;

[0007] Constructing a soft actor-critic intelligent agent with action decomposition;

[0008] Training the soft actor-critic intelligent agent using the multi-stage supply chain environment to obtain a trained soft actor-critic intelligent agent, the soft actor-critic intelligent agent with action decomposition comprises a state feature dense network, an action feature dense network, a soft actor network, two soft critic main networks and two soft critic target networks, the state feature dense network is used to extract the features corresponding to the state in the multi-stage supply chain environment, denoted as the first feature; the soft actor network is used to decide the action according to the first feature; the action feature dense network is used to extract the features of the decision action according to the action, denoted as the second feature; the soft critic main network is used to calculate the Q value of the current time t according to the first feature and the second feature ; the soft critic target network is used to calculate the Q value of the next time t+1 , the Q value represents the expected value of the future cumulative reward that can be obtained after selecting a certain action in the current state;

[0009] Deploying the trained soft actor-critic intelligent agent to the target supply chain to output the best action decision.

[0010] Further, the multi-stage supply chain environment comprises at least one raw material supplier, at least one manufacturer and a plurality of retailers, and at least one of the state space , demand function, action space , state transition , reward function is designed, the state space is the set of all possible states in the multi-stage supply chain environment; the demand function is the demand change of each retailer with randomness and seasonality; the action space is the set of all possible actions in the multi-stage supply chain environment; the state transition For an action made by the agent in the state of the multi-stage supply chain environment at the current time, the multi-stage supply chain environment follows the state transition update the environment to give the state at the next time; the reward function For an action made by the agent in the state of the multi-stage supply chain environment at the current time, the environment should give the reward at the current time according to the reward function.

[0011] Further, the state space The possible states at the current time t Specifically represented as follows:

[0012]

[0013]

[0014]

[0015]

[0016] wherein, represents the state at the current time t; represents the inventory level vector of the manufacturer and each retailer at the current time t; represents the backorder level vector of each retailer that does not meet the customer demand at the current time t; represents the demand level vector of each retailer at the previous three times t-3, , and so on. represents the inventory level of the manufacturer at the current time t; represents the inventory level of the retailer i at the current time t; represents the backorder level of the retailer i that does not meet the demand at the current time t; represents the customer demand level of the retailer i at the current time t.

[0017] Further, the demand function is specifically as follows:

[0018]

[0019]

[0020] wherein, represents the demand level of the retailer i at the current time t; represents the maximum demand level of each retailer; represents the constant of pi; represents the offset; represents the maximum number of steps in the supply chain environment; represents a random perturbation term, which follows a uniform distribution .

[0021] Further, the action space possible actions at the current time t is specified as follows:

[0022]

[0023] where, represents the action at the current time t; represents the quantity of products manufactured by the manufacturer at the current time t; represents the quantity of products shipped by the manufacturer to the retailer i at the current time t.

[0024] Further, the state transition is deterministic in the multi-stage supply chain environment, and is realized according to the material balance constraint is specified as follows:

[0025]

[0026] The state transition updates the inventory level vector of the manufacturer and the retailers at the current time t , the backorder level vector of the retailers that are not satisfied at the current time t is specified as follows:

[0027]

[0028]

[0029]

[0030]

[0031]

[0032] where, represents the maximum inventory capacity of the manufacturer; represents the maximum inventory capacity of the retailer i.

[0033] Further, the multi-stage supply chain environment follows a reward function gives the reward at the current time is specified as follows

[0034]

[0035] where, ​represents the final reward given by the environment; represents the selling price of the product; represents the selling quantity of the product for retailer i; represents the production cost of the product; represents the warehousing cost of the product for manufacturer and retailer i; represents the penalty cost when the demand of the product is not satisfied by the retailer; represents the transportation cost for each truck to deliver the product to the retailer; represents the maximum load of each truck; represents the upward rounding operation on the numerical value;

[0036] Specifically, represents the total selling revenue of the product at the current time, wherein is a parameter preset for the supply chain environment, is calculated according to the selling situation of each retailer, and the specific calculation formula is as follows:

[0037]

[0038] represents the total production cost of the product at the current time, wherein is a parameter preset for the supply chain environment, is part of the action ;

[0039] represents the total warehousing cost of the product at the current time, wherein is a parameter preset for the supply chain environment, is part of the state ;

[0040] represents the total penalty cost of the product for which the demand is not satisfied at the current time, wherein is a parameter preset for the supply chain environment, is part of the state ;

[0041] represents the total transportation cost of the product to each retailer at the current time, wherein and are parameters preset for the supply chain environment, is part of the action .

[0042] Further, the specific formula for calculating the expected value of the future cumulative reward is as follows:

[0043]

[0044]

[0045] in, represents the agent’s strategy, represents the total discounted reward in the future, Represents the discount factor.

[0046] Furthermore, the soft actor network is used to decide an action based on the first feature, specifically including: inputting the first feature into the soft actor network to obtain a first action, inputting the first action into an intermediate layer of the soft actor network to obtain a second action, splicing the first action and the second action, and outputting a decision action.

[0047] Furthermore, the multi-stage supply chain environment is used to train the soft actor-critic agent to obtain a trained soft actor-critic agent, specifically comprising:

[0048] (1) Randomly initialize the parameters of the state feature dense network, the parameters of the action feature dense network, the parameters of the soft actor network, and the parameters of the two soft critic main networks, and copy the parameters of the two soft critic main networks to the parameters of the two soft critic target networks, and use the experience replay pool Capacity Set to 0 to initialize the adaptive temperature coefficient ;

[0049] (2) Collect training samples, the soft actor-critic agent will be the state of the current supply chain environment Input to Strategy Select the corresponding action , using the action Interact with the multi-stage supply chain environment and observe the immediate rewards returned by the supply chain environment and the state at the next moment , the four-tuple sample Store in experience replay pool and put the experience back into the pool Capacity Set to , determine whether the number of steps in the multi-stage supply chain environment reaches the set maximum number of steps If it is reached, the supply chain environment is reset, otherwise the above sampling process is repeated until the experience replay pool Capacity Reach the number of samples required for training;

[0050] (3) Using the Experience Replay Pool The samples in update the parameters of the state feature dense network, the parameters of the action feature dense network, the parameters of the soft actor network, and the parameters of the two soft critic main networks;

[0051] (4) At fixed time intervals, use the parameters of the two soft critic main networks described in step (3) and Update the parameters of the two soft critic target networks and , the specific update method is as follows:

[0052]

[0053]

[0054] in, Indicates the soft update parameter, which is used to control the update amplitude;

[0055] (5) Determine whether the number of training times reaches the set number. If so, stop training. Otherwise, go to step (2) to obtain the trained soft actor-critic agent.

[0056] Compared with the prior art, the present invention has the following beneficial effects:

[0057] This paper addresses the challenges in the field of collaborative control and optimization of supply chain resources, such as demand fluctuations, complex structures, high-dimensional state and action spaces, and complex reward functions. It proposes a deep reinforcement learning-based intelligent agent. Using a feature extraction module and an action decomposition mechanism, this intelligent agent is able to predict future market trends by analyzing historical production order demand and out-of-stock records. It can also dynamically adjust inventory management strategies based on real-time supply chain inventory data, effectively addressing various issues in the collaborative control and optimization of supply chain resources. This not only reduces inventory backlogs and out-of-stock risks, but also improves the responsiveness and flexibility of the supply chain, ultimately significantly improving overall operational efficiency.

[0058] We also construct a reasonable multi-stage supply chain environment that can provide the agent with sufficient information to guide it to make wise decisions. We also introduce appropriate perturbations, which not only enhances the agent's strategy generalization ability and enables it to adapt to new environments, but also improves the agent's overall robustness and adaptability. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a flow chart of a supply chain resource control optimization method based on deep reinforcement learning according to Example 1 of the present invention;

[0060] Figure 2 This is the multi-stage supply chain environment of Example 1 of the present invention;

[0061] Figure 3 is a supply chain environment dynamics flowchart of embodiment 1 of the present application;

[0062] Figure 4 is an overall structure diagram of an action decomposition soft actor-critic network of embodiment 1 of the present application;

[0063] Figure 5 is a feature dense network structure diagram of embodiment 1 of the present application;

[0064] Figure 6 is a soft actor network structure diagram of embodiment 1 of the present application;

[0065] Figure 7 is a soft critic network structure diagram of embodiment 1 of the present application. DETAILED DESCRIPTION

[0066] The specific embodiments of the present application are described in further detail below in conjunction with the accompanying drawings and examples. The following examples are used to illustrate the present application, but are not used to limit the scope of the present application.

[0067] Embodiment 1

[0068] As shown in Figure 1 , a supply chain resource control optimization method based on deep reinforcement learning according to the preferred embodiment of the present application comprises:

[0069] S1: Construct a multi-stage supply chain environment;

[0070] In a feasible embodiment, as Figures 2-3 , the multi-stage supply chain environment for agent learning comprises a raw material supplier, a manufacturer, and multiple retailers. The manufacturer needs to order raw materials from the raw material supplier, and the manufacturer processes the raw materials into products and then transports the products to multiple retailers. Customers will purchase products at retailers, and customers will generate purchase demand for products at retailers. At each time of the supply chain environment, the agent needs to decide how many products the manufacturer should produce and how many products should be shipped to each retailer to maximize the satisfaction of customer demand and maximize profits.

[0071] The above-constructed supply chain environment conforms to the Markov decision process. Specifically, the constructed supply chain environment also designs a state space , a demand function, an action space , a state transition , a reward function , and a discount factor, wherein the state space refers to all possible states The set of possible states at the current time t The specific expressions are as follows:

[0072]

[0073]

[0074]

[0075]

[0076] in, Represents the state at the current time t; represents the inventory level vector of the manufacturer and each retailer at the current time t; represents the out-of-stock level vector of each retailer that fails to meet customer demand at the current time t; represents the demand level vector of each retailer at the first three moments t-3, , And so on; represents the manufacturer’s inventory level at the current time t; represents the inventory level of retailer i at the current time t; represents the out-of-stock level of unmet demand of retailer i at the current time t; represents the customer demand level of retailer i at the current time t.

[0077] Design a demand function with stochastic and seasonal characteristics for each retailer, which is expressed as follows:

[0078]

[0079]

[0080] in, represents the demand level of retailer i at the current time t; Indicates the maximum demand level for each retailer; represents the circumference constant of pi; Indicates the offset; Indicates the maximum number of steps in the supply chain environment; represents the random disturbance term, which conforms to Uniform distribution The designed demand function can approximately simulate the random and seasonal customer demand in the real supply chain environment. The setting makes it so that each time the supply chain environment is reset, the customer demand level will not be exactly the same, further increasing the randomness of the environment to better simulate the real environment.

[0081] Action Space denotes the set of all possible actions in the supply chain environment at the current time t is specified as follows:

[0082]

[0083] where, denotes the action at the current time t; denotes the quantity of products manufactured by the manufacturer at the current time t; denotes the quantity of products shipped by the manufacturer to the retailer i at the current time t.

[0084] The state transition denotes the state of the supply chain environment at the current time t For the action made by the agent , the environment should follow the state transition to update the environment and give the state at the next time In the supply chain environment constructed in the present application, the state transition probability is deterministic, and the state transition is realized according to the material balance constraint which is specified as follows:

[0085]

[0086] The state transition mainly updates the inventory level vector of the manufacturer and the retailer at the current time t the backorder level vector of the retailer that does not meet the demand at the current time t which is specified as follows:

[0087]

[0088]

[0089]

[0090]

[0091]

[0092] where, denotes the maximum inventory capacity of the manufacturer; denotes the maximum inventory capacity of the retailer i.

[0093] The reward function denotes the state of the supply chain environment at the current time t For the action made by the agent , the environment should follow the reward function to give the reward at the current time . Specifically, it is expressed as follows:

[0094]

[0095] wherein, represents the final reward given by the environment; represents the selling price of the product; represents the selling quantity of the product of retailer i; represents the production cost of the product; represents the warehousing cost of the product of manufacturer and retailer i; represents the penalty cost when the demand of the retailer is not satisfied; represents the transportation cost of each truck to the retailer; represents the maximum load of each truck; The symbol represents the upward rounding operation on the numerical value.

[0096] The first term represents the total selling revenue of the product at the current time, wherein is a parameter preset for the supply chain environment, will be calculated according to the selling situation of each retailer, and the specific calculation formula is as follows:

[0097]

[0098] The second term represents the total production cost of the product at the current time, wherein is a parameter preset for the supply chain environment, is part of the action ;

[0099] The third term represents the total warehousing cost of the product at the current time, wherein is a parameter preset for the supply chain environment, is part of the state ;

[0100] The fourth term represents the total penalty cost of the product that is not satisfied at the current time, wherein is a parameter preset for the supply chain environment, is part of the state ;

[0101] The fifth term represents the total transportation cost of the product to each retailer at the current time, wherein and are parameters preset for the supply chain environment, is part of the action .

[0102] Discount Factor It is a measure of the importance of future rewards relative to immediate rewards in the supply chain environment. A real number in the interval. Discount factor The core idea is that future rewards are not as valuable as immediate rewards because the realization of future rewards is uncertain and delayed. The agent is based on the state of the supply chain environment at the current time t. , decide the action that is appropriate for the current state , which determines the quantity of products that the manufacturer should produce at the current time t , and the quantity of the product that should be shipped to each retailer The agent uses the action it has decided Interact with the supply chain environment.

[0103] like Figure 3 As shown, at time t, the supply chain environment receives the action decided by the agent Finally, the dynamic flow chart of the supply chain environment consists of the following sequence of events:

[0104] (1) Supply chain environment based on the actions of the agent decision , extracting the actions that determine the manufacturer's production of products , that is, at time t, the manufacturer produces products.

[0105] (2) Supply chain environment based on the actions of the agent , extract the quantity of products shipped by the decision manufacturer to each retailer , that is, deliver to retailer 1 at time t Products shipped to Retailer 2 products, and so on.

[0106] (3) In the supply chain environment, after retailers receive products shipped by manufacturers, they meet the customer's demand for products. If the inventory level of each current retailer If all the demand can be met, there will be no shortage level , otherwise calculate the out-of-stock level .

[0107] (4) In the supply chain environment, after retailers meet customer needs, they begin to calculate their profits at that moment. Total Profit Total sales revenue of the product Subtract the total production cost of producing the product , minus the total warehousing costs of storing the product , minus the total penalty cost of unmet demand products , minus the total transportation cost of transporting products to each retailer .

[0108] (5) The supply chain environment updates the state of the environment according to the state transition above.

[0109] S2: Construct an action-decomposed soft actor-critic agent;

[0110] In a feasible embodiment, the action-decomposed soft actor-critic agent includes a state feature dense network, an action feature dense network, a soft actor network, two soft critic master networks, and two soft critic target networks. The structure of the action-decomposed soft actor-critic agent is shown in Figure 4 The state feature dense network is used to extract the features corresponding to the state in the multi-stage supply chain environment, denoted as the first feature. The soft actor network is used to decide the action according to the first feature. The structure of the soft actor network is shown in Figure 5 The action feature dense network is used to extract the features of the decision action according to the action. The structure of the state (action) feature dense network is shown in Figure 5 denoted as the second feature. The soft critic master network is used to calculate the Q value of the current time t according to the first feature and the second feature The soft critic target network is used to calculate the Q value of the next time t+1 The structure of the soft critic network is shown in Figure 6 The Q value represents the expected value of the future cumulative reward that can be obtained after selecting a certain action in the current state. The specific formula is as follows:

[0111]

[0112]

[0113] wherein, represents the policy of the agent, represents the total discounted reward in the future, represents the discount factor.

[0114] The above policy is the core of the action-decomposed soft actor-critic algorithm agent, which is composed of a state feature dense network and a soft actor network. The working principle of the policy is as follows:

[0115] The current state ​The input state feature dense network extracts features and inputs the features into the soft actor network. The soft actor network will first output a part of the final action, then input this part of the action into the middle layer of the network for calculation, and then output the remaining part of the final action. Finally, the two are spliced ​​together to form the action of the final decision of the intelligent agent. .

[0116] S3: training the soft actor-critic agent using the multi-stage supply chain environment to obtain a trained soft actor-critic agent;

[0117] In a feasible embodiment, step S3 specifically includes:

[0118] S3.1: Initialize the network and experience replay pool, and set the parameters of the state feature dense network mentioned above , parameters of the action feature dense network , parameters of the soft actor network , parameters of the two soft critic main networks and Randomly initialize and copy the parameters of the two soft critic main networks and Give the parameters of the two soft critic target networks and . Replay experience to the pool Capacity Set to 0 to initialize the adaptive temperature coefficient .

[0119] S3.2: Environmental sampling and sample storage, the agent observes the current state of the supply chain environment , which is input into the agent's policy Select an action , use actions Interact with the supply chain environment and observe the immediate rewards returned by the supply chain environment and the state at the next moment . The four-tuple sample Store in experience replay pool and put the experience back into the pool Capacity Set to If the supply chain environment reaches the maximum number of steps set , then reset the supply chain environment. Repeat Step 2 until the experience replay pool Capacity When the number of samples required for training is reached, jump to Step 3.

[0120] S3.3: Feature Dense Network, Soft Actor Network and Soft Critic Main Network Parameter Update:

[0121] S3.3.1: Randomly sample a batch of quadruple samples from the experience replay pool ; ;

[0122] S3.3.2: Extract the features of the next time state using the state feature dense network, calculate the optimal action that can be taken at the next time using the soft actor network , and the policy entropy of the action ; Extract the features of the optimal action that can be taken at the next time using the action feature dense network, calculate the two Q values of the state-action pair at the next time using the two soft critic target networks and . Select the smaller Q value from the two Q values , calculate the target value according to the smaller Q value , the policy entropy of the optimal action that can be taken at the next time , and the reward at the current time , the specific formula is as follows:

[0123]

[0124] S3.3.3: Extract the features of the current time state using the state feature dense network , extract the features of the current time action using the action feature dense network , and calculate the two Q values of the state-action pair at the current time using the two soft critic main networks and ;

[0125] S3.3.4: Set the loss function of the soft critic main network, construct the loss function of the soft critic main network for the two Q values and and the target value , the specific formula is as follows:

[0126]

[0127] Apply gradient descent method to the above loss function to update the state feature dense network parameters , action feature dense network parameters , and two soft critic main network parameters and ;

[0128] S3.3.5: Extract the features of the current time state using the state feature dense network ​​​characteristics of the current state and action, and the policy entropy of the best action that can be taken at the current time and the policy entropy characteristics of the current state and action, and the policy entropy of the best action that can be taken at the current time characteristics of the current state and action, and the policy entropy of the best action that can be taken at the current time two Q values and the smaller one ;

[0129] S3.3.6: Set the loss function of the soft actor network, and the smaller Q value obtained in Step 3.5 and the policy entropy of the best action that can be taken at the current time The loss function of the soft actor network is constructed, and the specific formula is as follows:

[0130]

[0131] The gradient descent method is applied to the above loss function to update the parameters of the soft actor network ;

[0132] S3.3.7: Set the adaptive temperature coefficient loss function according to the policy entropy of the best action that can be taken at the current time The loss function of the adaptive temperature coefficient is constructed, and the specific formula is as follows:

[0133]

[0134] wherein, is the target policy entropy, which is a hyperparameter of the algorithm. The gradient descent method is applied to the above loss function to update the adaptive temperature coefficient .

[0135] S3.4: Soft critic target network parameter update:

[0136] Every fixed time interval, the parameters of the two soft critic main networks and are used to update the parameters of the two soft critic target networks and , and the specific update method is as follows:

[0137]

[0138]

[0139] wherein, indicates the soft update parameter, which is used to control the magnitude of the update.

[0140] S3.5: Training cycle and end:

[0141] If the training round reaches the set number of times, the training ends, otherwise jump to S3.2.

[0142] S4: Deploy the trained soft actor-critic intelligent agent to the target supply chain, and output the optimal action decision.

[0143] In a feasible embodiment, after the training is completed, the soft actor-critic algorithm intelligent agent of action decomposition will obtain an optimal strategy . The strategy accepts the state output and outputs the optimal action under the current state. The trained intelligent agent is deployed to the real supply chain environment, and the intelligent agent receives the state provided by the real supply chain environment in real time, and uses the strategy to calculate the optimal action decision, and according to the action decision, guide the real supply chain to carry out the operation of producing and transporting products.

[0144] This embodiment proposes a deep reinforcement learning intelligent agent based on the challenges in the field of supply chain resource collaborative control optimization, such as demand fluctuations, complex structure, high-dimensional state and action space, and complex reward function. The intelligent agent can predict future market trends by analyzing historical production order demand and out-of-stock records through the use of feature extraction module and action decomposition mechanism, and can dynamically adjust inventory management strategy based on real-time supply chain inventory data, effectively dealing with various problems in supply chain resource collaborative control optimization, not only reducing inventory backlog and out-of-stock risk, but also improving the response speed and flexibility of the supply chain, and ultimately achieving significant improvement in overall operating efficiency.

[0145] A reasonable multi-stage supply chain environment is also constructed, which can provide sufficient information for the intelligent agent to make wise decisions, and appropriate disturbances are also introduced, which not only enhances the strategy generalization ability of the intelligent agent, enabling it to adapt to new environments, but also improves the overall robustness and adaptability of the intelligent agent.

[0146] Embodiment 2

[0147] The application also provides a computer readable storage medium having a computer program stored thereon, wherein the computer storage program is executed by a processor to realize the supply chain resource control optimization method based on deep reinforcement learning according to embodiment 1.

[0148] To sum up, the embodiment of the present application provides a supply chain resource control optimization method and storage medium based on deep reinforcement learning, which proposes a deep reinforcement learning agent based on the challenges in the field of supply chain resource collaborative control optimization, such as demand fluctuation, complex structure, high-dimensional state and action space, and complex reward function. The agent can predict future market trends by analyzing historical production order demand and out-of-stock records through the use of a feature extraction module and an action decomposition mechanism, and can dynamically adjust the inventory management strategy based on real-time supply chain inventory data, effectively dealing with various problems in supply chain resource collaborative control optimization, not only reducing inventory backlog and out-of-stock risk, but also improving the response speed and flexibility of the supply chain, ultimately achieving significant improvement in overall operating efficiency. A reasonable multi-stage supply chain environment is also constructed, which can provide sufficient information for the agent to guide it to make wise decisions, and appropriate disturbances are introduced, which not only enhance the strategy generalization ability of the agent, enabling it to adapt to new environments, but also improve the overall robustness and adaptability of the agent.

[0149] The above only describes the preferred embodiments of the present application, and it should be pointed out that for ordinary skilled persons in the technical field, several improvements and replacements can be made without departing from the technical principles of the present application, and these improvements and replacements should also be considered as the protection scope of the present application.

Claims

1. A supply chain resource control optimization method based on deep reinforcement learning, characterized in that: The method comprises the following steps: Building a multi-stage supply chain environment; Building a soft actor-critic agent for action decomposition; The multi-stage supply chain environment is used to train the soft actor-critic agent to obtain a trained soft actor-critic agent. The action-decomposed soft actor-critic agent includes a state feature dense network, an action feature dense network, a soft actor network, two soft critic main networks and two soft critic target networks. The state feature dense network is used to extract the features corresponding to the states in the multi-stage supply chain environment, which are recorded as the first features; the soft actor network is used to decide actions based on the first features; the action feature dense network is used to extract the features of the decision-making actions based on the actions, which are recorded as the second features; the soft critic main network is used to calculate the Q value of the current time t based on the first features and the second features ; The soft critic target network is used to calculate the Q value at the next time t+1 , the Q value represents the expected value of the future cumulative reward that can be obtained after selecting a certain action in the current state; The trained soft actor-critic agent is deployed to the target supply chain to output the optimal action decision.

2. A supply chain resource control optimization method based on deep reinforcement learning according to claim 1, characterized in that: The multi-stage supply chain environment includes at least one raw material supplier, at least one manufacturer and multiple retailers, and at least a state space is designed. , demand function, action space , state transfer , reward function One of the state space is the set of all possible states in the multi-stage supply chain environment; the demand function is the demand change of each retailer with random and seasonal characteristics; the action space is the set of all possible actions in the multi-stage supply chain environment; the state transition The action taken by the agent in the current state of the supply chain environment, the multi-stage supply chain environment follows the state transition Update the environment and give the state of the next moment; the reward function Under the current state of the multi-stage supply chain environment, for the actions taken by the agent, the environment should give the current reward according to the reward function.

3. A supply chain resource control optimization method based on deep reinforcement learning according to claim 2, characterized in that: The state space Possible states at the current time t The specific expressions are as follows: in, Represents the state at the current time t; represents the inventory level vector of the manufacturer and each retailer at the current time t; represents the out-of-stock level vector of each retailer that fails to meet customer demand at the current time t; represents the demand level vector of each retailer at the first three moments t-3, , And so on; represents the manufacturer’s inventory level at the current time t; represents the inventory level of retailer i at the current time t; represents the out-of-stock level of unmet demand of retailer i at the current time t; represents the customer demand level of retailer i at the current time t.

4. A supply chain resource control optimization method based on deep reinforcement learning according to claim 2, characterized in that: The demand function is as follows: in, represents the demand level of retailer i at the current time t; Indicates the maximum demand level for each retailer; represents the circumference constant of pi; Indicates the offset; Indicates the maximum number of steps in the supply chain environment; represents the random disturbance term, which conforms to Uniform distribution .

5. The supply chain resource control optimization method based on deep reinforcement learning according to claim 2 is characterized in that: The action space Possible actions at the current time t The specific expressions are as follows: in, represents the action at the current time t; represents the number of products produced by the manufacturer at the current time t; represents the quantity of products shipped by the manufacturer to the retailer i at the current time t.

6. A supply chain resource control optimization method based on deep reinforcement learning according to claim 2, characterized in that: The state transfer It is deterministic in the multi-stage supply chain environment and realizes state transition according to material balance constraints. , specifically expressed as follows: State Transfer The inventory level vectors of manufacturers and retailers at the current time t are , the out-of-stock level vector of the retailer’s unmet demand at the current time t Update, specifically as follows: in, Indicates the manufacturer's maximum inventory capacity; represents the maximum inventory capacity of retailer i.

7. The supply chain resource control optimization method based on deep reinforcement learning according to claim 2 is characterized in that: The multi-stage supply chain environment follows the reward function Give the current reward , specifically expressed as follows in, Represents the final reward given by the environment; Indicates the selling price of the product; represents the quantity of product sold by retailer i; Indicates the production cost of the product; represents the storage cost of product i for the manufacturer and retailer; represents the penalty cost when the retailer fails to meet the demand; represents the transportation cost per truckload delivered to the retailer; Indicates the maximum cargo capacity of each truck; The symbol indicates that the value is rounded up; Specifically, Represents the total sales revenue of the product at the current moment, where Preset parameters for the supply chain environment, It will be calculated based on the sales of each retailer. The specific calculation formula is as follows: represents the total production cost of the product at the current moment, where Preset parameters for the supply chain environment, For action part of; Represents the total storage cost of the product at the current moment, where Preset parameters for the supply chain environment, Status part of; represents the total penalty cost for unsatisfied demand products at the current moment, where Preset parameters for the supply chain environment, Status part of; represents the total transportation cost of transporting products to each retailer at the current moment, where and Preset parameters for the supply chain environment, For action part of.

8. The supply chain resource control optimization method based on deep reinforcement learning according to claim 1 is characterized in that: The specific formula for calculating the expected value of future cumulative rewards is as follows: in, represents the agent’s strategy, represents the total discounted reward in the future, Represents the discount factor.

9. The supply chain resource control optimization method based on deep reinforcement learning according to claim 1 is characterized in that: The soft actor network is used to decide an action based on the first feature, specifically including: inputting the first feature into the soft actor network to obtain a first action, inputting the first action into an intermediate layer of the soft actor network to obtain a second action, splicing the first action and the second action, and outputting a decision action.

10. A supply chain resource control optimization method based on deep reinforcement learning according to any one of claims 1 to 9, characterized in that: Training the soft actor-critic agent using the multi-stage supply chain environment to obtain a trained soft actor-critic agent specifically includes: (1) Randomly initialize the parameters of the state feature dense network, the parameters of the action feature dense network, the parameters of the soft actor network, and the parameters of the two soft critic main networks, and copy the parameters of the two soft critic main networks to the parameters of the two soft critic target networks, and use the experience replay pool Capacity Set to 0 to initialize the adaptive temperature coefficient ; (2) Collect training samples, the soft actor-critic agent will be the state of the current supply chain environment Input to Strategy Select the corresponding action , using the action Interact with the multi-stage supply chain environment and observe the immediate rewards returned by the supply chain environment and the state at the next moment , the four-tuple sample Store in experience replay pool and put the experience back into the pool Capacity Set to , determine whether the number of steps in the multi-stage supply chain environment reaches the set maximum number of steps If it is reached, the supply chain environment is reset, otherwise the above sampling process is repeated until the experience replay pool Capacity Reach the number of samples required for training; (3) Using the Experience Replay Pool The samples in update the parameters of the state feature dense network, the parameters of the action feature dense network, the parameters of the soft actor network, and the parameters of the two soft critic main networks; (4) At fixed time intervals, use the parameters of the two soft critic main networks described in step (3) and Update the parameters of the two soft critic target networks and , the specific update method is as follows: in, Indicates the soft update parameter, which is used to control the update amplitude; (5) Determine whether the number of training times reaches the set number. If so, stop training. Otherwise, go to step (2) to obtain the trained soft actor-critic agent.

Citation Information

Patent Citations

  • Double-competition closed-loop supply chain financial intervention strategy and pricing decision analysis method based on market demand uncertainty

    CN111754258A

  • WEEE recycling supply chain management and control optimization method based on reinforcement learning

    CN118674129A