Multi - supply - chain scheduling method and system based on global Critic multi - agent algorithm

By building a multi-agent supply chain environment and introducing a global Critic network, using the multi-agent deep deterministic strategy gradient algorithm, the problems of low efficiency and low returns in multi-level supply chain scenarios are solved, and the scheduling strategy to maximize the overall profit of the supply chain is realized, and the efficiency and benefits of the multi-supply chain environment are improved.

CN119090223BActive Publication Date: 2025-07-18GUANGDONG UNIV OF TECH

Patent Information

Application Number
CN202411209654.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-08-30
Publication Date
2025-07-18
Estimated Expiration
2044-08-30

AI Technical Summary

Technical Problem

In multi-level supply chain scenarios, the existing single agent deep reinforcement learning algorithms have problems such as poor convergence, poor stability and low overall reward value, resulting in low environmental efficiency and low returns for multiple supply chains.

Method used

Build a multi-agent supply chain environment, introduce a global Critic network, use multi-agent deep deterministic strategy gradient algorithm for reinforcement learning, generate a scheduling strategy that maximizes the overall profit of the supply chain, observe the overall reward value through the global Critic network, reduce competitive confrontation between agents, and optimize overall performance.

Benefits of technology

By optimizing scheduling and processing, the efficiency and benefits of the multi-supply chain environment are significantly improved, better scheduling strategies are generated, and multiple supply chain scheduling problems of different scales and types are adapted to the reliability and stability of decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119090223B_ABST
    Figure CN119090223B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of supply chain scheduling, and discloses a multiple supply chain scheduling method and system based on a global Critic multi-agent algorithm. The method includes: constructing a multi-agent supply chain environment; introducing a global Critic network, and using the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model; inputting supply chain demand data into the multi-agent decision-making model for optimized scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain. The present invention introduces a global Critic network to reduce multi-agent competition and optimize the overall performance. The multi-agent deep deterministic policy gradient algorithm is used to construct a decision-making model to adapt to complex supply chain environments. With the goal of maximizing the overall profit, it solves the problems of low efficiency and low returns. The environment has strong adaptability and can handle scheduling problems of different scales. By continuously learning to optimize decisions, the reliability is improved, and the efficiency and returns are significantly enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of supply chain scheduling, and more specifically, to a multi-supply chain scheduling method and system based on a global Critic multi-agent algorithm. Background Art

[0002] At present, there are still obstacles in the information exchange between the various links of the supply chain, resulting in problems such as increased product costs, extended work progress, and potential product quality hazards. In order to break through the information barrier and make full use of information sharing, it is urgent to find a method to achieve the optimal allocation of resources between the various links of the supply chain on the basis of multi-level supply chain integration. With the development of deep reinforcement learning from single-agent to multi-agent, its applications in various industries are becoming increasingly widespread.

[0003] At present, deep reinforcement learning algorithms are mainly divided into deterministic and non-deterministic types. Non-deterministic algorithms (such as DQN) need to map the selected optimal action to a value and are suitable for discrete environments with simple action spaces. Deterministic algorithms (such as DDPG) directly output action values and are more suitable for continuous environments with complex action spaces, showing better performance in multi-level supply chain scenarios. However, when applying a single-agent deep reinforcement learning algorithm (such as DDPG) in a multi-level supply chain scenario, due to the overly large state space and action space dimensions, there are defects such as poor convergence, poor stability, and low and unstable overall reward values, resulting in low efficiency and low returns in the multi-supply chain environment. Summary of the Invention

[0004] In order to overcome the problems of low efficiency and low returns in the multi-supply chain environment caused by the prior art, the present invention proposes the following technical solutions:

[0005] In a first aspect, the present invention proposes a multi-supply chain scheduling method based on a global Critic multi-agent algorithm, including:

[0006] Construct a multi-agent supply chain environment.

[0007] Introduce a global Critic network, and use the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model.

[0008] Input the supply chain demand data into the multi-agent decision-making model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain.

[0009] As a preferred technical solution, constructing a multi-agent supply chain environment includes:

[0010] According to the overall business logic of the multi-supply chain, design the state space, action space, and reward function respectively.

[0011] Construct a multi-agent supply chain environment using the state space, the action space, and the reward function.

[0012] As a preferred technical solution, the state space includes t the state space of the agent responsible for factory production at time, and t the state space of the agent responsible for retailer inventory replenishment at time, and their expressions are as follows:

[0013] ,

[0014] ;

[0015] ,

[0016] ,

[0017] ;

[0018] wherein, in this multi-agent supply chain environment, is used to distinguish the numbers of the factory and the retailer. When , it means operating on the factory. When , it means operating on the retailer, and is the number of steps in each round of training, is the inventory level of the factory, is the order demand of the factory, represents the quantity of goods that the j th retailer purchases from the factory, K is the total number of retailers; represents the respective order demands of all retailers, represents the order demand of retailer 1 at time t, represents the order demand of retailer k at time t, represents the maximum order quantity, represents the perturbation coefficient, represents rounding down, .

[0019] As a preferred technical solution, the action space includes t the action space of the agent responsible for factory production at time, and t the action space of the agent responsible for retailer inventory replenishment at time, and .

[0020] As a preferred technical solution, the reward function is the total profit in a multi-level supply chain environment, and its expression is as follows:

[0021]

[0022] where is the price of the factory's products, is the retailer 's quantity of goods purchased from the factory, is the cost of the factory's production of products, represents the quantity of products produced by the factory, is the price of the retailer's products, is the retailer j 's t order demand at time is the warehousing cost of retailer j, is the inventory level of retailer j, is the penalty coefficient for the warehouse lacking goods, is the penalty coefficient for exceeding the warehousing capacity, is the maximum warehousing capacity of retailer j, is the transportation cost of retailer j, is the maximum load capacity of each truck of retailer j.

[0023] As a preferred technical solution, a global Critic network is introduced, and the multi-agent deep deterministic policy gradient algorithm is used to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model, including:

[0024] Randomly extract a data sample batch of a predetermined size, and integrate the data sample batch to form a data set including the current state, next state, action, reward, and completion flag of the agent.

[0025] According to the current state, use the target Actor network to predict the action set of the next state.

[0026] Use the global Critic network to calculate the TD target value based on the action set of the next state.

[0027] According to the TD target value, calculate the loss of the global Critic network, and update the parameters of the global Critic network through backpropagation.

[0028] Use the updated global Critic network to score the action set of the next state, calculate the Actor network loss of each agent, and update the Actor network parameters of each agent through backpropagation according to the Actor network loss.

[0029] Repeat the above steps until the multi-agent supply chain environment converges to obtain a multi-agent decision-making model.

[0030] As a preferred technical solution, input the supply chain demand data into the multi-agent decision-making model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain, including:

[0031] Input the supply chain demand data into the multi-agent decision-making model, and the agent selects the optimal action according to the current state and environmental information to respond to demand changes.

[0032] After the agent makes the corresponding action, use the state transition function to update the state of the agent and simulate the new state of the supply chain after each decision cycle.

[0033] Use the global Critic network to evaluate the efficiency and effectiveness of the current action selection strategy.

[0034] According to the evaluation results of the global Critic network, adjust the action selection strategy of the agent.

[0035] Repeat the above steps until the overall profit of the supply chain reaches the maximum to generate the final supply chain scheduling strategy.

[0036] As a preferred technical solution, the expression of the state transition function is as follows:

[0037]

[0038] Among them, is the state space of the agent responsible for factory production at t time, is the state space of the agent responsible for retailer purchasing at t time, represents the order demand quantities of all retailers respectively, is the inventory level of the factory, represents the quantity of products produced by the factory, is for retailer the quantity of goods purchased from the factory, is the storage capacity of the factory, is the inventory level of retailer 1, the quantity of goods purchased by retailer 1 from the factory, is for retailer 1 at t time the order demand quantity, is the storage capacity of retailer 1, is for retailer K the inventory level, retailer K the quantity of goods purchased from the factory, is for retailerK At t the order demand at time is the storage capacity of the retailer K , the order demand of the factory is the order demand of the retailer at time t and the order demand of the retailer at time t - 1

[0039] As a preferred technical solution, the supply chain demand data includes, but is not limited to, the revenue of the products sold by the retailer, the demand of the factory, the revenue of the products sold by the factory, the cost of producing products by the factory, the storage cost, the penalty for overdue goods, the penalty for exceeding the storage upper limit, and the transportation cost

[0040] The agent represents a link in the supply chain, including the factory and the retailer

[0041] In a second aspect, the present invention also proposes a multiple supply chain scheduling system based on the global Critic multi-agent algorithm, which is applied to the multiple supply chain scheduling method based on the global Critic multi-agent algorithm described in any of the schemes in the first aspect, and includes:

[0042] A first construction module for constructing a multi-agent supply chain environment

[0043] A second construction module for introducing a global Critic network and performing reinforcement learning on the multi-agent supply chain environment by using the multi-agent deep deterministic policy gradient algorithm to construct a multi-agent decision model

[0044] A generation module for inputting the supply chain demand data into the multi-agent decision model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain

[0045] The beneficial effects of the present invention at least include:

[0046] By introducing a global Critic network, the present invention can observe and consider the overall reward value, effectively reduce the competitive confrontation relationship between multi-agents, and optimize the overall performance. The decision model constructed by using the multi-agent deep deterministic policy gradient algorithm can better adapt to the complex multiple supply chain environment and generate a better scheduling strategy. Through optimal scheduling processing, a scheduling strategy is generated with the goal of maximizing the overall profit of the supply chain, effectively solving the problems of low efficiency and low revenue in the multiple supply chain environment. The constructed multi-agent supply chain environment has strong adaptability and can handle multiple supply chain scheduling problems of different scales and types. Through continuous learning, the decision model is continuously optimized and adjusted, improving the reliability and stability of the decision-making, significantly enhancing the efficiency and revenue of the multiple supply chain environment, and providing an effective solution for the optimal scheduling of complex supply chain systems Description of the Drawings

[0047] Figure 1 It is a schematic flow chart of the multi - supply - chain scheduling method based on the global Critic multi - agent algorithm provided in Embodiment 1.

[0048] Figure 2 It is the overall flow chart of a single - factory single - commodity provided in Embodiment 1.

[0049] Figure 3 It is the overall flow chart of multi - factory multi - commodity provided in Embodiment 1.

[0050] Figure 4 It is the neural network model framework diagram of MADDPG based on the global Critic provided in Embodiment 1.

[0051] Figure 5 It is the update flow chart of MADDPG based on the global Critic provided in Embodiment 1.

[0052] Figure 6 It is a comparison chart of the overall profits of each algorithm provided in Embodiment 3 in the single - factory single - commodity multi - level supply - chain scenario.

[0053] Figure 7 It is the neural network model framework diagram of parameter - free - sharing MAMT - DDPG based on the global Critic provided in Embodiment 3.

[0054] Figure 8 It is the neural network model framework diagram of hard - parameter - sharing MAMT - DDPG based on the global Critic provided in Embodiment 3.

[0055] Figure 9 It is a comparison chart of the overall profits of each algorithm provided in Embodiment 3 in the multi - factory multi - commodity multi - level supply - chain scenario.

[0056] Figure 10 It is the architecture diagram of the multi - supply - chain scheduling system based on the global Critic multi - agent algorithm provided in Embodiment 4 Detailed Implementation Modes

[0057] The following will describe the implementation modes of the present invention with reference to the drawings and preferred technical solutions. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation modes. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred technical solutions are only for illustrating the present invention rather than for limiting the protection scope of the present invention.

[0058] It should be noted that the illustrations provided in the following embodiments only schematically illustrate the basic concept of the present invention. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The types, quantities, and proportions of the components in actual implementation can be arbitrarily changed, and the component layout type may also be more complex.

[0059] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0060] Embodiment 1

[0061] This embodiment proposes a multi - supply - chain scheduling method based on the global Critic multi - agent algorithm, as Figure 1 shown. Figure 1 FIG. is a schematic flow chart of a multi - supply - chain scheduling method based on the global Critic multi - agent algorithm provided in this embodiment. The method includes the following steps:

[0062] S1: Construct a multi - agent supply - chain environment.

[0063] S2: Introduce a global Critic network, and use the multi - agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi - agent supply - chain environment to construct a multi - agent decision - making model.

[0064] S3: Input the supply - chain demand data into the multi - agent decision - making model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain.

[0065] In the specific implementation process, first, according to the actual supply - chain situation, design and construct a multi - agent supply - chain environment. This environment includes multiple agents, such as factories and retailers. Define the state space, action space, and reward function for each agent. The state space can include inventory level, order quantity, production capacity, etc. The action space can include order quantity, production quantity, delivery quantity, etc. The reward function can be designed based on indicators such as profit, cost, and customer satisfaction. As Figure 2 and Figure 3 shown, Figure 2 and Figure 3 are the overall flow charts of a single - factory single - commodity and a multi - factory multi - commodity respectively.

[0066] Next, design a global Critic network that can receive the state and action information of all agents as input and output a global value estimate. The structure of the global Critic network can be a multi-layer perceptron (MLP), which includes multiple hidden layers and an output layer. At the same time, equip each agent with an Actor network to generate actions based on the current state. The structure of the Actor network can also be an MLP, with the local state of the agent as the input and the action as the output. The core steps of implementing the Multi-Agent Deep Deterministic Policy Gradient (MADDPG) algorithm include an experience replay buffer, target networks, policy updates, etc.

[0067] It can be understood that in this embodiment, various deep reinforcement learning algorithms are experimentally carried out step by step in the scenario of a multi-level supply chain, including the non-deterministic single-agent deep reinforcement learning algorithm DQN and the deterministic single-agent deep reinforcement learning algorithm DDPG. After selecting the optimal action, DQN needs to map the action once again to output a specific value, while DDPG directly outputs the action value, omitting the action mapping process, so it is more efficient. The research further extends the single-agent deep reinforcement learning algorithm DDPG to a multi-agent deep reinforcement learning algorithm MADDPG. Experiments prove that in the scenario of a multi-level supply chain, the deterministic deep reinforcement learning algorithm performs better than the non-deterministic deep reinforcement learning algorithm, and the multi-agent deep reinforcement learning algorithm performs better than the single-agent deep reinforcement learning algorithm. Based on these results, the MADDPG algorithm is finally selected as the algorithm model for multi-level supply chain scheduling.

[0068] As Figure 4 and Figure 5 shown, Figure 4 and Figure 5 are respectively the neural network model framework diagram of MADDPG based on the global Critic and the MADDPG update flowchart based on the global Critic. During the training process, first initialize all network parameters. For each training episode, reset the environment and obtain the initial state. At each time step, each agent selects an action according to the current state and the Actor network, obtains the reward and the next state after executing the action, and stores the experience in the replay buffer. Then sample a batch of experiences from the replay buffer to update the global Critic network, the Actor network of each agent, and the target network. Repeat the training episodes until convergence or reach the preset number of training times.

[0069] After training is completed, the trained Actor network and the global Critic network are combined to form a multi-agent decision-making model. In practical applications, first, actual supply chain demand data is collected, and the demand data is input into the multi-agent decision-making model. The Actor network of each agent generates corresponding actions (scheduling decisions) based on the input, and the global Critic network evaluates the overall decision-making quality. According to the evaluation results, the decisions are adjusted to maximize the overall profit of the supply chain.

[0070] It can be understood that the traditional MADDPG algorithm uses a centralized training and distributed execution mode to train agents. Each agent has an Actor-Critic network framework, where the Actor network outputs the optimal action value based on the currently observed local environmental state, and the Critic network scores its own Actor based on the currently observed global environmental state and the action values output by other agents. When updating the network parameters, the Critic network parameters are updated first, and then the Actor network parameters are updated. The Critic network uses TD update, and each agent updates its own Critic network according to the reward value it obtains, while the Actor network updates its own parameters according to the scores given by its own Critic network of each agent. However, the traditional MADDPG algorithm ignores the observation of the reward values of each other among agents, which may lead to confrontation among agents, showing a trend of one rising while the other falls. To alleviate this situation, a global Critic network is introduced in this embodiment to replace the Critic networks of each agent. This global Critic network is used to score the Actor networks of each agent and can observe the overall reward value during the update. This method avoids the situation where agents in the traditional MADDPG algorithm generate confrontation in order to maximize their respective reward values, and can better achieve win-win cooperation in a multi-level supply chain scenario. This improvement enables the algorithm to better capture the interdependence between levels of the supply chain and promotes the collaborative optimization of the overall system.

[0071] Finally, based on the optimized decisions, detailed scheduling strategies are generated, including the specific action plans of each agent, such as order quantity, production plan, distribution route, etc. Through the above steps, this method can effectively utilize the multi-agent deep reinforcement learning algorithm, combined with the global Critic network, to achieve intelligent scheduling of multiple supply chains and maximize the overall profit.

[0072] It can be understood that in this embodiment, by introducing a global Critic network, the overall reward value can be observed and considered, effectively reducing the competitive and adversarial relationships among multiple agents, and optimizing the overall performance. The decision-making model constructed using the multi-agent deep deterministic policy gradient algorithm can better adapt to complex multiple supply chain environments and generate better scheduling strategies. By optimizing the scheduling process and generating scheduling strategies with the goal of maximizing the overall profit of the supply chain, the problems of low efficiency and low revenue in the multiple supply chain environment are effectively solved. The constructed multi-agent supply chain environment has strong adaptability and can handle multiple supply chain scheduling problems of different scales and types. Through continuous learning, the decision-making model is continuously optimized and adjusted, improving the reliability and stability of decision-making, significantly enhancing the efficiency and revenue of the multiple supply chain environment, and providing an effective solution for the optimal scheduling of complex supply chain systems.

[0073] Embodiment 2

[0074] This embodiment makes improvements on the basis of the multiple supply chain scheduling method based on the global Critic multi-agent algorithm proposed in Embodiment 1.

[0075] In this embodiment, a multi-agent supply chain environment is constructed, including:

[0076] According to the overall business logic of the multiple supply chain, the state space, action space, and reward function are designed respectively.

[0077] Using the state space, the action space, and the reward function, a multi-agent supply chain environment is constructed.

[0078] In the specific implementation process, the multi-agent supply chain environment is driven by the final consumer demand to promote the overall data flow of the multi-level supply chain. Each level of the supply chain makes decisions on what actions should be taken at this time through the MADDPG (Multi-Agent Deep Deterministic Policy Gradient) multi-agent reinforcement learning algorithm based on the global Critic, so that information and materials are updated and flowed in the supply chain in a progressive manner. This method can better capture the complex interactions among different levels of the supply chain.

[0079] The state space refers to the range of environmental states that an agent can perceive and identify, including key indicators such as inventory levels, order backlogs, and production capacities. The action space refers to the range of actions that an agent can take, such as decision variables like order quantities, production quantities, and distribution quantities. The reward function is designed as the overall profit of the multi-level supply chain, taking into account factors such as inventory costs, shortage costs, and transportation costs, to encourage agents to make decisions beneficial to the entire supply chain system. The variable parameters of the reward function in the multi-level supply chain environment are introduced in Table 1.

[0080] Table 1 Composition variables of the reward function at each step

[0081]

[0082] In this embodiment, is used to distinguish the numbers of the factory and the retailer. When , the operation is performed on the factory at this time. When , the operation is performed on the retailer at this time, and is the number of steps in each round of training. The state space includes the state space of the agent responsible for factory production at t time, , and the state space of the agent responsible for retailer inventory replenishment at t time. Their expressions are shown as follows:

[0083] ,

[0084] ;

[0085] ,

[0086] ,

[0087] ;

[0088] Among them, in this multi-agent supply chain environment, is used to distinguish the numbers of the factory and the retailer. When , it means that the operation is performed on the factory. When , it means that the operation is performed on the retailer, and is the number of steps in each round of training, is the inventory level of the factory, is the order demand of the factory, represents the quantity of goods that the j th retailer purchases from the factory, K is the total number of retailers; represents the respective order demands of all retailers, represents the order demand of retailer 1 at time t, represents the order demand of retailer k at time t, represents the maximum order quantity, represents the perturbation coefficient, represents rounding down, .

[0089] In this embodiment, the action space includes at tThe action space of the agent responsible for factory production at a certain moment , and at t The action space of the agent responsible for retailer inventory replenishment at a certain moment , and .

[0090] In this embodiment, the reward function is the overall profit of the multi-level supply chain environment, and its expression is as follows:

[0091]

[0092] Wherein, is the price of the factory product, is the retailer The quantity of goods purchased from the factory, is the cost of the factory to produce products, represents the quantity of products produced by the factory, is the retailer product price, is the retailer j At t The order demand at a certain moment, is the warehousing cost of retailer j, is the inventory level of retailer j, is the penalty coefficient for the shortage of goods in the warehouse, is the penalty coefficient for exceeding the warehousing capacity, is the maximum warehousing capacity of retailer j, is the transportation cost of retailer j, is the maximum load capacity of each truck of retailer j.

[0093] In this embodiment, a global Critic network is introduced, and the multi-agent deep deterministic policy gradient algorithm is used to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model, including:

[0094] Randomly extract a data sample batch of a predetermined size, and integrate the data sample batch to form a data set including the current state, next state, action, reward, and completion flag of the agent.

[0095] According to the current state, use the target Actor network to predict the action set of the next state.

[0096] Use the global Critic network to calculate the TD target value based on the action set of the next state.

[0097] According to the TD target value, calculate the loss of the global Critic network, and update the parameters of the global Critic network through backpropagation.

[0098] Use the updated global Critic network to score the action sets of the next state, calculate the Actor network loss of each agent, and update the Actor network parameters of each agent through backpropagation according to the Actor network loss.

[0099] Repeat the above steps until the multi-agent supply chain environment converges to obtain a multi-agent decision-making model.

[0100] As an exemplary illustration, the code for constructing and updating the multi-agent decision-making model is as follows:

[0101] # Randomly sample tuples of batch size from the replay experience pool

[0102] transitions = self.memory.sample(self.BATCH_SIZE)

[0103] batch = Transition(*zip(*transitions))

[0104] states_batch = T.cat(batch.state)

[0105] next_states_batch = T.cat(batch.next_state)

[0106] done = T.cat(batch.done).unsqueeze(1)

[0107] actions_batch = T.cat(batch.action)

[0108] rewards_batch = T.cat(batch.reward).sum(1).unsqueeze(1) mult_states_list = []

[0109] mult_next_states_list = []

[0110] # First update the critic network

[0111] with T.no_grad():

[0112] # Obtain the action sets of each agent at the next moment

[0113] for j in range(1, len(self.agents)):

[0114] s_ = self.split_state_batch(j, next_states_batch, next_actions_batch)

[0115] next_actions_batch[:, j:] = self.agents[j].target_actor(s_).squeeze(1)

[0116] s_ = self.split_state_batch(0, next_states_batch, next_actions_batch)

[0117] next_actions_batch[:, 0] = self.agents[0].target_actor(s_).squeeze(1)

[0118] # Get the global state set

[0119] for j in range(len(self.agents)):

[0120] mult_states_list.append(self.split_state_batch(j, states_batch,actions_batch))

[0121] mult_next_states_list.append(self.split_state_batch(j, next_states_batch, next_actions_batch))

[0122] mult_states_batch = T.cat([x for x in mult_states_list], dim=1)

[0123] mult_next_states_batch = T.cat([x for x in mult_next_states_list],dim=1)

[0124] # Get the Target value required for TD update

[0125] q_ = self.target_total_critic(mult_next_states_batch, next_actions_batch) * done

[0126] target = q_ * self.gamma + rewards_batch

[0127] # Get Q values

[0128] q = self.total_critic(mult_states_batch, actions_batch)

[0129] # Calculate the loss value of the global Critic network

[0130] critic_loss = F.mse_loss(q.float(), target.float().detach())

[0131] # Clear the gradients of the global Critic network

[0132] self.total_critic.optimizer.zero_grad()

[0133] critic_loss.backward() # Update the parameters of the global Critic network

[0134] self.total_critic.optimizer.step()

[0135] self.total_critic_scheduler.step(critic_loss)

[0136] # Get the new action sets for each agent

[0137] new_actions_batch = T.tensor(actions_batch).detach()

[0138] # Update the actor network

[0139] for i in range(len(self.agents)):

[0140] s = self.split_state_batch(i, states_batch, actions_batch)

[0141] if i == 0:

[0142] new_actions_batch[:, i] = self.agents[i].actor(s).squeeze(1)

[0143] else:

[0144] new_actions_batch[:, i:] = self.agents[i].actor(s).squeeze(1)

[0145] # Score the actions of the current agent using the global Critic network

[0146] score = self.total_critic(mult_states_batch, new_actions_batch)

[0147] # Calculate the loss value of the Actor network

[0148] actor_loss = -T.mean(score)

[0149] for i in range(len(self.agents)):

[0150] self.agents[i].actor.optimizer.zero_grad()

[0151] actor_loss.backward()

[0152] # Update the parameters of the Actor network

[0153] for i in range(len(self.agents)):

[0154] self.agents[i].actor.optimizer.step()

[0155] self.agents[i].actor_loss += actor_loss.item()

[0156] # Soft update the parameters of the two networks

[0157] if self.step == self.env.episode_length:

[0158] self.agents[i].agent_update(self.TAU)

[0159] self.agents[i].actor_scheduler.step(actor_loss)

[0160] In this embodiment, the supply chain demand data is input into the multi-agent decision-making model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain, including:

[0161] The supply chain demand data is input into the multi-agent decision-making model, and the agent selects the optimal action according to the current state and environmental information to respond to the demand change.

[0162] After the agent makes the corresponding action, the state transfer function is used to update the state of the agent, simulating the new state of the supply chain after each decision cycle.

[0163] The global Critic network is used to evaluate the efficiency and effect of the current action selection strategy.

[0164] According to the evaluation results of the global Critic network, the action selection strategy of the agent is adjusted.

[0165] Repeat the above steps until the overall profit of the supply chain reaches the maximum, generating the final supply chain scheduling strategy.

[0166] In this embodiment, the expression of the state transfer function is as follows:

[0167]

[0168] Wherein, is the state space of the agent responsible for factory production at time t , is the state space of the agent responsible for retailer purchasing at time t , represents the order demand quantities of all retailers respectively, is the inventory level of the factory, represents the quantity of products produced by the factory, is the quantity of goods purchased by retailer from the factory, is the storage capacity of the factory, is the inventory level of retailer 1, the quantity of goods purchased by retailer 1 from the factory is the order demand of retailer 1 at t moment, is the storage capacity of retailer 1, is the inventory level of retailer K . The quantity of goods that retailer K purchases from the factory, is the order demand of retailer K at t moment, is the storage capacity of retailer K , is the order demand of the factory, is the order demand of the retailer at time t, is the order demand of the retailer at time t - 1.

[0169] In this embodiment, the supply chain demand data includes but is not limited to the revenue of the products sold by the retailer, the demand of the factory, the revenue of the products sold by the factory, the cost of the products produced by the factory, the storage cost, the penalty for overdue goods, the penalty for exceeding the storage upper limit, and the transportation cost.

[0170] The agent represents a link in the supply chain, including the factory and the retailer.

[0171] It can be understood that the proposed MAMT-DDPG algorithm based on the global Critic is more efficient in multiple supply chain scenarios compared to the DQN algorithm. As a deterministic deep reinforcement learning algorithm, MAMT-DDPG can directly output the action value, omitting the action-to-value mapping step in DQN. Compared with the MADDPG algorithm, MAMT-DDPG integrates a multi-task reinforcement learning mechanism, including the parameter-free sharing and hard parameter sharing modes, enabling each agent to learn different characteristics of each commodity and capture more information. Experiments have shown that in a multi-factory and multi-commodity multi-level supply chain environment, the performance of MAMT-DDPG is better than that of MADDPG. MAMT-DDPG uses a global Critic network to replace the Critic network of each agent in MADDPG, and performs TD update through the overall reward value, reducing the competition and confrontation between agents, and exceeding the original MADDPG in terms of overall profit performance.

[0172] The present invention deeply studies the cooperative communication among multiple agents, constructs an algorithm model that can reduce the mutual competition among multiple agents, proposes the MAMT-DDPG algorithm model that can observe the overall reward value, and effectively improves the overall profit in the multi-level supply chain scenario. Experiments show that the MAMT-DDPG algorithm based on the global Critic outperforms the MADDPG algorithm in terms of overall profit in the multi-level supply chain environments of single-factory single-commodity and multi-factory multi-commodity, reflecting the superiority of the present invention.

[0173] Example 3

[0174] In this example, the multi-agent algorithm for multi-level supply chain scheduling based on the global Critic proposed in Example 2 is compared with the following other methods to verify the effectiveness of the multi-agent algorithm for multi-level supply chain scheduling based on the global Critic. The environment used in this example is the single-factory single-commodity multi-level supply chain environment. The comparison methods include the (ς, Q)-based strategy, the traditional multi-agent reinforcement learning algorithm MADDPG, and the MADDPG algorithm based on the joint Q-value.

[0175] The (ς, Q)-based strategy is a simple and effective inventory management method. When the current inventory level is lower than the preset threshold ς, the system automatically replenishes Q quantities of goods. This strategy is widely used in practice due to its simplicity and practicality and serves as a baseline for evaluating the performance of other algorithms in this study.

[0176] In the traditional MADDPG algorithm, each agent is equipped with a set of Actor-Critic network architectures. This enables each agent to observe the global state information and the actions of other agents, thereby making more informed decisions. This method can capture the complex interactions in the multi-level supply chain to a certain extent.

[0177] The MADDPG algorithm based on the joint Q-value is an improvement over the traditional MADDPG. In this method, the Q-values output by the Critic networks of each agent are first calculated, and then these Q-values are added together to obtain a joint Q-value. This joint Q-value is subsequently used to update each Critic network. This method aims to better reflect the overall performance of the entire system rather than just the performance of individual agents.

[0178] As Figure 6 shown, Figure 6It is a comparison chart of the overall profits of each algorithm in the scenario of a single-factory and single-product multi-level supply chain, where Rewards represents the overall profit value in the multi-level supply chain. The experimental results show that the MADDPG algorithm based on the global Critic performs the best, and its performance has been significantly improved compared with the traditional MADDPG algorithm. This improvement fully demonstrates the advantages of the global Critic in coordinating the behaviors of multiple agents and optimizing the overall system performance.

[0179] To comprehensively verify the effectiveness of the proposed MAMT-DDPG method based on the global Critic in the present invention, in this embodiment, comparative experiments are conducted on this method and the following several algorithms. The experimental environment is set as a complex scenario of a multi-factory, multi-product, and multi-level supply chain.

[0180] The MAMT-DDPG algorithm without parameter sharing; in this method, each task has its independent parameter models in the Actor network and the Critic network of the agent, and no parameter sharing is carried out among tasks. This method allows each task to have a highly specialized model, but it may increase the computational complexity.

[0181] The MAMT-DDPG algorithm with hard parameter sharing; in this method, both the Actor network and the Critic network of the agent consist of a shared backbone network and multiple task-specific output branches. All tasks share the backbone network and then select the corresponding output branch according to the specific task. This method can reduce the number of parameters and improve the learning efficiency.

[0182] The MAMT-DDPG algorithm without parameter sharing based on the global Critic; this is an improvement on the MAMT-DDPG algorithm without parameter sharing, which introduces a global Critic network to replace the independent Critic networks of all agents. This method aims to provide a global perspective to better coordinate the behaviors of different agents. As Figure 7 shown, Figure 7 It is a framework diagram of the neural network model without parameter sharing of MAMT-DDPG based on the global Critic.

[0183] The MAMT-DDPG algorithm with hard parameter sharing based on the global Critic; this is an improvement on the MAMT-DDPG algorithm with hard parameter sharing, which also introduces the global Critic network. This method combines the efficiency of parameter sharing and the coordination advantages of the global Critic. As Figure 8 shown, Figure 8 It is a framework diagram of the neural network model with hard parameter sharing of MAMT-DDPG based on the global Critic.

[0184] In addition, this embodiment also compares the above method with the previously mentioned (ς, Q)-based strategy, the traditional multi-agent reinforcement learning algorithm MADDPG, and the MADDPG algorithm based on the global Critic. This comprehensive comparative experiment design aims to evaluate the performance of the proposed method from multiple perspectives and deeply understand the performance characteristics of different algorithms in a complex supply chain environment.

[0185] As Figure 9 shown, Figure 9 Figure 1 is a comparison chart of the overall profits of each algorithm in the multi-factory multi-commodity multi-level supply chain scenario, where Rewards represents the overall profit value of multiple commodities in the multi-level supply chain. Both the parameter-free sharing method and the hard parameter sharing method of MAMT-DDPG exceed the traditional MADDPG algorithm and the MADDPG algorithm based on the global Critic. In the global Critic-based method, both parameter sharing methods of MAMT-DDPG are better than the method without using the global Critic.

[0186] It can be understood that in order to better coordinate the operation of each link and each platform in the supply chain and make each subunit reach the global optimum, the present invention proposes a multi-level supply chain scheduling method based on the MADDPG multi-agent deep reinforcement learning algorithm with a global Critic. This method uses the multi-level supply chain as the experimental environment for the multi-agent deep reinforcement learning algorithm, and deeply explores the adaptability of various deep reinforcement learning algorithms in the multi-level supply chain scenario. The research results show that the deterministic multi-agent deep reinforcement learning algorithm is most suitable for this link, so MADDPG is selected for further experiments and optimization.

[0187] Based on the experimental results of MADDPG and the exploration of other deep reinforcement learning algorithms, the traditional MADDPG algorithm is innovated. The innovation lies in introducing a global Critic network to replace the independent Critic networks in all agents and score all Actor networks. This global Critic network can obtain more global information, which helps to promote win-win cooperation among agents and reduce adversarial behaviors among them. The experimental results confirm that the MADDPG algorithm based on the global Critic performs more stably and has more significant effects in the multi-level supply chain scenario, showing obvious advantages compared with the traditional MADDPG.

[0188] In addition, through the experimental results of MAMT-DDPG (Multi-Agent Multi-Task DDPG), it is found that in a more complex multi-factory multi-commodity multi-level supply chain scenario, multi-agent multi-task depth.

[0189] Embodiment 4

[0190] AsFigure 10 As shown in Figure 10 , this embodiment proposes a multiple supply chain scheduling system based on the global Critic multi-agent algorithm, which is applied to the multiple supply chain scheduling method based on the global Critic multi-agent algorithm as described in the above embodiment, and includes: a first construction module 100, a second construction module 200, and a generation module 300.

[0191] Among them, the first construction module 100 is used to construct a multi-agent supply chain environment. The second construction module 200 is used to introduce a global Critic network, and use the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model. The generation module 300 is used to input the supply chain demand data into the multi-agent decision-making model for optimal scheduling processing, and generate a scheduling strategy that maximizes the overall profit of the supply chain.

[0192] It should be noted that the foregoing explanation of the embodiment of the multiple supply chain scheduling method based on the global Critic multi-agent algorithm is also applicable to the multiple supply chain scheduling system based on the global Critic multi-agent algorithm of this embodiment, and will not be repeated here.

[0193] In the description of this specification, the descriptions with reference to the terms "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms are not necessarily directed to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or N embodiments or examples in a suitable manner. In addition, without conflict, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0194] In addition, the terms "first" and "second" are only used for descriptive purposes, and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of the features. In the description of the present application, the meaning of "N" is at least two, such as two, three, etc., unless otherwise specifically defined.

[0195] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a customized logical function or process. The scope of the preferred embodiments of this application includes additional implementations where functions may be executed not in the order shown or discussed, including in a substantially simultaneous manner or in the reverse order according to the functions involved, which should be understood by those skilled in the technical field to which the embodiments of this application pertain.

[0196] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above-described embodiments, the N steps or methods can be implemented using software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays, field programmable gate arrays, etc.

[0197] Those of ordinary skill in the art of this technology can understand that all or part of the steps carried out in the methods of the above embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium, and when executed, includes one or a combination of the steps of the method embodiments.

[0198] Obviously, the above embodiments of the present invention are merely examples for clearly illustrating the present invention and are not intended to limit the embodiments of the present invention. For those of ordinary skill in the art, other different forms of changes or modifications can be made based on the above description. It is not necessary and impossible to enumerate all the embodiments here. Any modifications, equivalent replacements, and improvements made within the spirit and principles of the present invention shall be included within the protection scope of the claims of the present invention.

Claims

1. A multiple supply chain scheduling method based on a global Critic multi-agent algorithm, characterized in that including: Designing a state space, an action space, and a reward function respectively according to the overall business logic of multiple supply chains; Constructing a multi-agent supply chain environment by using the state space, the action space, and the reward function; The state space includes the t state space of the agent responsible for factory production at , and the t state space of the agent responsible for retailer replenishment at , and their expressions are as follows: , ; , , ; Among them, in this multi-agent supply chain environment, is used to distinguish the numbers of factories and retailers. When , it indicates an operation on the factory. When , it indicates an operation on the retailer. And is the number of steps for each round of training. is the inventory level of the factory. is the order demand of the factory. represents the j th retailer's quantity of purchasing goods from the factory. K is the total number of retailers. represents the respective order demands of all retailers. represents the order demand of retailer 1 at time t. represents the order demand of retailer k at time t. represents the maximum order quantity. represents the perturbation coefficient. represents rounding down. ; The action space includes the action space of the agent responsible for factory production at t the moment, and the action space of the agent responsible for retailer inventory replenishment at the moment, and the set of the total action spaces of all agents t ; ; The reward function is the total profit of the multi-level supply chain environment, and its expression is as follows: Among them, is the price of the factory's products, is the retailer 's quantity of goods purchased from the factory, is the cost of the factory to produce products, represents the quantity of products produced by the factory, is the price of the retailer's products, is the retailer j at t the order demand at time is the warehousing cost of retailer j, is the inventory level of retailer j, is the penalty coefficient for the shortage of goods in the warehouse, is the penalty coefficient for exceeding the warehousing capacity, is the maximum warehousing capacity of retailer j, is the transportation cost of retailer j, is the maximum load capacity of each truck of retailer j; Introducing a global Critic network, and using the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model; Inputting the supply chain demand data into the multi-agent decision-making model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain.

2. The multi - supply - chain scheduling method based on the global Critic multi - agent algorithm according to claim 1, wherein, Inputting the supply chain demand data into the multi-agent decision-making model for optimal scheduling processing to generate a scheduling strategy that maximizes the overall profit of the supply chain, including: Inputting the supply chain demand data into the multi-agent decision-making model, and the agent selects the optimal action according to the current state and environmental information to respond to demand changes; After the agent makes the corresponding action, using the state transition function to update the state of the agent to simulate the new state of the supply chain after each decision-making cycle; Using the global Critic network to evaluate the efficiency and effect of the current action selection strategy; Adjusting the action selection strategy of the agent according to the evaluation result of the global Critic network; Repeating the above steps until the overall profit of the supply chain reaches the maximum to generate the final supply chain scheduling strategy.

3. The multi - supply - chain scheduling method based on the global Critic multi - agent algorithm according to claim 2, characterized in that, The expression of the state transition function is as follows: Among them, is the state space of the agent responsible for factory production at t time, is the state space of the agent responsible for retailer replenishment at t time, represents the order demand quantities of all retailers respectively, is the inventory level of the factory, represents the quantity of products produced by the factory, is the quantity of goods that retailer purchases from the factory, is the storage capacity of the factory, is the inventory level of retailer 1, The quantity of goods that retailer 1 purchases from the factory, is the order demand quantity of retailer 1 at t time, is the storage capacity of retailer 1, is the inventory level of retailer K , Retailer K The quantity of goods purchased from the factory, is the order demand quantity of retailer K at t time, is the storage capacity of retailer K , is the order demand of the factory, is the order demand quantity of retailers at time t, is the order demand quantity of retailers at time t - 1.

4. The multi - supply - chain scheduling method based on the global Critic multi - agent algorithm according to claim 1, characterized in that Introducing a global Critic network, and using the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model, including: Randomly extracting a data sample batch of a predetermined size, and integrating the data sample batch to form a data set including the current state, the next state, the action, the reward, and the completion flag of the agent; According to the current state, using the target Actor network to predict the action set of the next state; Using the global Critic network to calculate the TD target value based on the action set of the next state; According to the TD target value, calculating the loss of the global Critic network, and updating the parameters of the global Critic network through backpropagation; Using the updated global Critic network to score the action set of the next state, calculating the Actor network loss of each agent, and updating the Actor network parameters of each agent through backpropagation according to the Actor network loss; Repeating the above steps until the multi-agent supply chain environment converges to obtain a multi-agent decision-making model.

5. The multi - supply - chain scheduling method based on the global Critic multi - agent algorithm according to any one of claims 1 to 4, characterized in that, The supply chain demand data includes but is not limited to the income from the products sold by retailers, the demand of factories, the income from the products sold by factories, the cost of producing products by factories, the warehousing cost, the penalty for overdue goods, the penalty for exceeding the warehousing limit, and the transportation cost; The agent represents a link in the supply chain, including factories and retailers.

6. A multiple supply chain scheduling system based on a global Critic multi-agent algorithm, which is applied to the multiple supply chain scheduling method based on the global Critic multi-agent algorithm according to any one of claims 1 to 5, and is characterized in that, including: A first construction module for constructing a multi-agent supply chain environment; The second construction module is used to introduce a global Critic network, and use the multi-agent deep deterministic policy gradient algorithm to perform reinforcement learning on the multi-agent supply chain environment to construct a multi-agent decision-making model; The generation module is used to input the supply chain demand data into the multi-agent decision-making model for optimal scheduling processing, and generate a scheduling strategy that maximizes the overall profit of the supply chain.

Citation Information

Patent Citations

  • Federal multi-agent Actor-Critic learning intelligent logistics task unloading and resource allocation system and medium

    CN115658251A

  • Adaptive-learning intelligent scheduling unified computing frame and system for industrial personalized customized production

    US20220413455A1

Cited By

  • Supply chain management system and method based on artificial intelligence, electronic equipment and medium

    CN122414996A

  • Intelligent supply chain management and control system and method, computer equipment and medium

    CN122415028A

  • System and method for processing nonlinear relationship in logistics data

    CN122453042A

  • Supply chain management system and method, electronic equipment and medium

    CN122492078A