Inventory resource scheduling method based on decentralized multi-agent reinforcement learning

By building local value functions and graph attention networks, the problem of inefficient collaboration in large-scale resource scheduling systems is solved, and efficient inventory resource allocation and system efficiency improvement are achieved.

CN120258360APending Publication Date: 2025-07-04CHONGQING UNIV
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202510230085.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing multi-agent reinforcement learning inventory scheduling method faces the problems of dimensional disasters and inefficient cooperative efficiency in large-scale resource scheduling systems, which is difficult to apply.

Method used

Decentralized multi-agent reinforcement learning algorithm based on local value functions is adopted. By analyzing explicit and implicit cooperative relationships between warehouses, using graph attention networks to aggregate neighbor information, construct local value functions, train the strategies of each node, reduce communication overhead and redundant information, and achieve efficient collaboration.

Benefits of technology

Efficient inventory resource allocation is achieved in a large-scale resource scheduling system, reducing communication overhead, and improving the overall efficiency and reliability of the system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120258360A_ABST
    Figure CN120258360A_ABST
Patent Text Reader

Abstract

The invention discloses an inventory resource scheduling method based on decentralized multi-agent reinforcement learning, and provides a decentralized multi-agent reinforcement learning algorithm based on reward aggregation of a graph in order to solve the problems that an existing inventory scheduling method based on multi-agent reinforcement learning has curse of dimensionality and can not efficiently carry out cooperation among warehouses. A global value function is simplified into a local value function through an inherent coupling relationship between nodes, redundant information is eliminated, the input dimension of the value function is reduced, the fitting difficulty of the value function is reduced, and preference information of different nodes is transmitted by using a reward aggregation mechanism so as to realize efficient cooperation; the strategy of each node is trained through a local value function, after training is completed, each node only needs to obtain an observation value of the node itself to obtain an allocation scheme of resources owned by the node, and a larger-scale resource scheduling problem can be solved through decentration processing.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of goods storage and inventory management, and particularly relates to an inventory resource scheduling method based on decentralized multi-agent reinforcement learning. Background Art

[0002] With the continuous development of the global economy and the increasing frequency of trade activities, warehousing logistics plays a crucial role in the modern commercial system. In today's fast-paced market environment, enterprises are facing increasingly fierce competition, and efficient warehousing logistics management has become a key factor in enhancing competitiveness. On the one hand, the continuous development of the manufacturing industry has led to a continuous expansion of the production scale of various products, which requires a more powerful warehousing logistics system to store and allocate a large amount of raw materials and finished products. Close cooperation is needed among different production links, and warehousing logistics plays a bridging role in ensuring the timely supply of raw materials and the smooth shipment of finished products. On the other hand, the trend of diversification and personalization of consumer demands has become increasingly obvious. In order to meet the various needs of consumers, the retail industry needs to maintain a rich inventory of goods. This requires the warehousing logistics system to accurately predict market demands and reasonably arrange inventory to avoid inventory backlogs or shortages. At the same time, with the booming rise of the e-commerce industry, consumers have put forward higher requirements for the speed and accuracy of logistics distribution. Different e-commerce platforms need to achieve fast order processing and goods delivery through efficient warehousing logistics management to improve the user experience. In the field of warehousing logistics, the inventory between different warehouses needs to be dynamically adjusted according to the actual demand and supply of goods. This can not only improve logistics efficiency, reduce transportation costs, but also ensure that goods can meet market demands in a timely manner. When the inventory of goods in a certain warehouse is insufficient, goods need to be quickly transferred from other warehouses. Other warehouses will allocate goods to the areas without inventory according to the pre-set support priorities to ensure the normal fulfillment of orders. This inventory resource scheduling method is crucial for improving the overall efficiency and reliability of the warehousing logistics system.

[0003] The prior arts related to the present invention include: CN114091988A, a method for scheduling target items between warehouses; CN116823107A, a supply chain inventory management method based on multi-agent reinforcement learning; CN118118908A, a method for allocating resources of D2D users based on multi-agent reinforcement learning; CN117875842A, a cross-warehouse inventory scheduling system

[0004] The deficiencies of the above technical solutions are as follows: Currently, the more advanced technologies use machine learning to solve resource scheduling problems, such as supervised learning or reinforcement learning. However, supervised learning is highly dependent on datasets and manual annotation, which consumes a large amount of time and labor costs (CN117875842A). In existing reinforcement learning methods, many build problem models from the perspective of single agents. Due to factors such as high randomness and dynamic environmental changes in the resource scheduling system, it is easy to encounter problems where the model is difficult to converge in the reinforcement learning algorithm for the resource scheduling system. There are also some technologies that use multi-agent reinforcement learning to solve resource scheduling problems, but they often face challenges such as the curse of dimensionality and low cooperation efficiency among agents. Therefore, it is difficult to apply them to large-scale resource scheduling systems (CN118118908A, CN116823107A).

[0005] The key technical points of the present invention lie in analyzing the implicit and explicit cooperation relationships among agents, using local value functions instead of global value functions, and adopting graph attention networks to aggregate neighbor information. The present invention proposes a decentralized multi-agent reinforcement learning algorithm and framework based on local value functions.

[0006] Currently, the more advanced technologies use machine learning to solve resource scheduling problems, such as supervised learning or reinforcement learning. However, supervised learning is highly dependent on datasets and manual annotation, which consumes a large amount of time and labor costs (CN117875842A). In reinforcement learning methods, many build problem models from the perspective of single agents. Due to the influence of factors such as high randomness and dynamic environmental changes in the resource scheduling system, problems where the model is difficult to converge will occur. There are also some technologies that use multi-agent reinforcement learning to solve resource scheduling problems, but they often face challenges such as the curse of dimensionality and low cooperation efficiency among agents. Therefore, it is difficult to apply them to large-scale resource scheduling systems (CN118118908A, CN116823107A). Summary of the Invention

[0007] Aiming at the deficiencies of the existing technology, the present invention provides an inventory resource scheduling method based on decentralized multi-agent reinforcement learning, which can complete the training of the model using local information and can be applied to large-scale resource scheduling systems. This method simplifies the global value function into a local value function through the inherent coupling relationship between nodes, while eliminating redundant information, reducing the input dimension of the value function, lowering the fitting difficulty of the value function, and using a reward aggregation mechanism to transmit the preference information of different nodes to achieve efficient cooperation. The strategy of each node is trained through the local value function. After the training is completed, each node only needs to obtain its own observation value to get the allocation plan for the resources it owns.

[0008] To achieve the above object, the present invention adopts the following steps:

[0009] S1. Construct a resource allocation problem model. Treat a single warehouse as an agent, use a topological graph to describe the coupling relationship of warehouses, and analyze the explicit and implicit cooperation relationships among warehouses.

[0010] S2. Construct the local value function of each warehouse by using the cooperation relationship among warehouses and simplify the input of each local value function from global information to local information.

[0011] S3. Train the local value function and policy of each warehouse according to the GRA algorithm.

[0012] S4. Each warehouse obtains its own observation from the environment and inputs it into the trained policy to obtain the resource allocation plan for each warehouse.

[0013] Furthermore, the resource allocation problem model uses an observation graph a state transition graph and a reward graph to describe the observation coupling relationship (which other warehouses' actions or states are involved in the observation of a single warehouse), the state transition function coupling relationship (which other warehouses' actions or states are involved in the state transition function of a single warehouse), and the reward function coupling relationship (which other warehouses' actions or states are involved in the reward function of a single warehouse) among warehouses respectively. The union of the observation graph and the state transition graph is denoted as Model the problem as a distributed partially observable Markov decision process (DecPOMDP), and use a tuple to represent it, where is the set of all warehouses, is the state space, is the observation space, is the action space, is the state transition function, is the reward function, and γ is the discount factor. The goal of each warehouse is to optimize its own policy π i to maximize the total discounted return:

[0014]

[0015] where π = {π1,..., π N} is the set of all warehouse policies, is the reward obtained by warehouse i at time t.

[0016] Furthermore, the state transition function for constructing the resource allocation problem model is:

[0017]

[0018] where, is the resource quantity of warehouse i at time t, a ij is the proportion of the resources of warehouse i delivered to warehouse j itself, is the set of successor nodes of warehouse i on the state transition graph, is the set of predecessor nodes of warehouse i on the state transition graph, is the perturbation related to time step t. The observation o of each warehouse i includes its own resource quantity and the resource quantities of its neighbors. The action a of each warehouse i is the proportion of its existing resources delivered to other warehouses. In addition, we add random noise when initializing the resource quantity of each warehouse and make restrictions on a ij as follows:

[0019]

[0020] Furthermore, the reward function of warehouse i in the resource allocation problem model is constructed as:

[0021]

[0022] where is the set of predecessor nodes of warehouse i on the reward graph.

[0023] Furthermore, the implicit cooperation relationship is jointly determined by the observation graph and the state transition graph. The implicit cooperation relationship has two manifestations. The first is that the inventory resource quantity of a single warehouse will be observed by other warehouses, and when the inventory resource quantity of this warehouse changes, it will have an impact on the warehouses connected to it (as shown by the black lines in Attachment Figure 2 ); the second is the impact generated when the resource quantity delivered by a single warehouse to other connected warehouses changes (as shown by the red lines in Attachment Figure 2 ). The explicit cooperation relationship is determined by the reward graph. The states or actions of other warehouses will directly affect the reward of a single warehouse (as shown by the yellow lines in Attachment Figure 2 ), and the warehouse will cooperate with the warehouses that have an explicit cooperation relationship with it when maximizing its own reward.

[0024] Furthermore, the learning graph is derived according to the coupling relationship between agents The neighbors of warehouse i in the learning graph are represented as:

[0025]

[0026] where, represents the set of warehouses that have an explicit cooperation relationship with warehouse k, represents warehouse i at The set of all warehouses reachable in [context]. Obtain the local value function of each warehouse from the learned graph:

[0027]

[0028] where r j (s t , a t ) is the reward function of warehouse j.

[0029] Furthermore, the GRA algorithm adopted in step S3 includes:

[0030] S3.1. Adopt the Actor-Critic architecture. Set an Actor network and a Critic network for each warehouse i, with their parameters being θ i and φ i respectively, and initialize the parameters of each network.

[0031] Preferably, considering that the Critic of each warehouse needs to aggregate the observation information and reward information of its neighbors, the network we adopted first uses a fully connected layer to encode the original observation and reward data, then uses a graph attention network to complete the aggregation of neighbor information, and finally uses a fully connected layer to fuse the aggregated observation information and reward information to output the value of the local value function. The calculation formula of the graph attention network is:

[0032]

[0033] where x i is the feature (observation or reward) of warehouse i, W1 and W2 are trainable weight matrices, and α i,j is the attention weight, which can be expressed as:

[0034]

[0035] where d is the dimension of the network layer output, and W3 and W4 are trainable weight matrices.

[0036] Preferably, considering that the Actor of each warehouse only needs to make decisions using its own observations, we use a multi-layer perceptron as the Actor network to complete the decision-making process with a relatively small network scale.

[0037] S3.2. Each warehouse i receives the local observation Select an action according to the policy Execute the action to obtain the reward at time t and the observation at the next time

[0038] S3.3. Each warehouse uses its own Critic network to aggregate and learn the information of its neighbors on the graph. The Critic network uses the graph attention mechanism to finally obtain the value of the local value function of warehouse i at time t.

[0039] S3.4. Put the observation sets action sets reward sets value function value sets and the observation sets at time t+1 into the experience buffer.

[0040] S3.5. Sample from the experience buffer multiple times to calculate the advantage function for each warehouse:

[0041]

[0042] where γ is the discount factor, λ is a hyperparameter, and T is the number of steps in each episode.

[0043] S3.6. Update the network parameter φ using the Critic loss function i :

[0044]

[0045] where, B is the training sample size, ∈ is a hyperparameter, is the cumulative return, and clip(a, b, c) is a clipping function that restricts a to the interval (b, c).

[0046] Update the network parameter θ using the Actor loss function i :

[0047]

[0048] S3.7. Repeat steps S3.2 to S3.6 until the policy converges.

[0049] Adopting the above solution, the beneficial effects of the present invention are:

[0050] The complex coupling relationships between warehouses are represented by an observation graph, a state transition graph, and a reward graph to analyze the implicit and explicit cooperation relationships between warehouses, thereby obtaining a learning graph. On this basis, a local value function is constructed for each warehouse to guide policy learning. Each warehouse uses a graph attention network to aggregate the reward information of its neighbors, which helps to fit the local value function.

[0051] Compared with other multi-agent reinforcement learning algorithms, the method adopted in the present invention only needs to obtain its own observations for decision-making during the execution phase, without the need for additional communication, reducing the communication overhead between warehouses.

[0052] During the training phase, by only aggregating the information of neighbors in the learning graph, the input dimension is effectively reduced and redundant information is removed. This decentralized processing enables the method adopted in the present invention to be more applicable to large-scale networked systems. Description of the Drawings

[0053] Figure 1 It is a schematic diagram of the flow of goods in the supply chain.

[0054] Figure 2 It is a schematic diagram of the coupling relationship between warehouses in the simulation experiment.

[0055] Figure 3 It is the learning graph derived in the simulation experiment.

[0056] Figure 4 It is a schematic diagram of the overall framework of the decentralized multi-agent reinforcement learning algorithm adopted in the present invention.

[0057] Figure 5 It is a schematic diagram of the network structures of the agent Actor and Critic.

[0058] Figure 6 It is a comparison graph of the training results of various multi-agent reinforcement learning algorithms. Detailed Implementation Manner

[0059] The technical solution of the present invention will be further described below with reference to the accompanying drawings. The described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments.

[0060] To effectively solve the inventory resource scheduling problem, the present invention proposes an inventory resource scheduling method based on decentralized multi-agent reinforcement learning, including the following steps:

[0061] S1. Construct a resource allocation problem model, take a single warehouse as an agent, use a topological graph to describe the coupling relationship between warehouses, and analyze the explicit and implicit cooperation relationships between warehouses.

[0062] The resource allocation problem model uses an observation graph a state transition graph and a reward graph Describe the observed coupling relationship between warehouses (which other warehouses' actions or states are involved in the observation of a single warehouse), the state transition function coupling relationship (which other warehouses' actions or states are involved in the state transition function of a single warehouse), and the reward function coupling relationship (which other warehouses' actions or states are involved in the reward function of a single warehouse) respectively. The union of the observation graph and the state transition graph is denoted as Model the problem as a distributed partially observable Markov decision process (DecPOMDP), and use a tuple to represent it, where is the set of all warehouses, is the state space, is the observation space, is the action space, is the state transition function, is the reward function, and γ is the discount factor. The goal of each warehouse is to optimize its own policy π i to maximize the total discounted return:

[0063]

[0064] where π = {π1,..., π N} is the set of all warehouses' policies, is the reward obtained by warehouse i at time t.

[0065] Furthermore, the state transition function for constructing the resource allocation problem model is:

[0066]

[0067] where, is the resource quantity of warehouse i at time t, a ij is the proportion of its own resources that warehouse i transports to warehouse j, is the set of successor nodes of warehouse i on the state transition graph, is the set of predecessor nodes of warehouse i on the state transition graph, is the perturbation related to time step t. The observation o of each warehouse i includes its own resource quantity and the resource quantities of its neighbors. The action a of each warehouse i is the proportion of its existing resources transported to other warehouses. In addition, we add random noise when initializing the resource quantity of each warehouse, and make a limitation on a ij :

[0068]

[0069] Furthermore, the reward function of warehouse \(i\) in constructing the resource allocation problem model is as follows:

[0070]

[0071] where is the set of predecessor nodes of warehouse \(i\) on the reward graph.

[0072] Taking a system with 9 warehouses as an example in the simulation environment, the coupling relationship between warehouses is as Figure 2 shown, where the edges on the state transition graph are represented by red lines, the edges on the observation graph are represented by black lines, and the edges on the reward graph are represented by yellow lines.

[0073] S2. Derive the learning graph as shown in Figure 3 according to the coupling relationship between agents The neighbors of warehouse \(i\) in the learning graph are represented as:

[0074]

[0075] where represents the set of warehouses that have an explicit cooperation relationship with warehouse \(k\), represents the set of all warehouses reachable by warehouse \(i\) in . The implicit cooperation relationship is jointly determined by the observation graph and the state transition graph. There are two manifestations of the implicit cooperation relationship. The first is that the inventory resource quantity of a single warehouse will be observed by other warehouses, and when the inventory resource quantity of this warehouse changes, it will have an impact on the warehouses connected to it (as shown by the black lines in Attachment Figure 2 ); the second is the impact generated when the quantity of resources transported by a single warehouse to other connected warehouses changes (as shown by the red lines in Attachment Figure 2 ). The explicit cooperation relationship is determined by the reward graph. The states or actions of other warehouses will directly affect the reward of a single warehouse (as shown by the yellow lines in Attachment Figure 2 ). When maximizing its own reward, the warehouse will cooperate with the warehouses that have an explicit cooperation relationship with it.

[0076] Construct the local value function of each warehouse using the learning graph as shown in Figure 3 :

[0077]

[0078] At the same time, we also use the learning graph to simplify the input of each local value function from global information to neighbor information.

[0079] S3. Train the local value function and policy of each warehouse according to the GRA algorithm.

[0080] The GRA algorithm adopted in step S3 includes:

[0081] S3.1. The overall framework of the algorithm is as Figure 4 shown. An Actor-Critic architecture is adopted. For each warehouse i, an Actor network and a Critic network are respectively set, and their parameters are θ i and φ i , and the parameters of each network are initialized.

[0082] The structures of the Actor network and the Critic network of each warehouse are as Figure 5 shown. Considering that the Critic of each warehouse needs to aggregate the observation information and reward information of its neighbors, the network we adopted first uses a fully connected layer to encode the original observation and reward data, then uses a graph attention network to complete the aggregation of neighbor information, and finally uses a fully connected layer to fuse the aggregated observation information and reward information to output the value of the local value function. The calculation formula of the graph attention network is:

[0083]

[0084] where x i is the feature (observation or reward) of warehouse i, W1 and W2 are trainable weight matrices, and α i,j is the attention weight, which can be expressed as:

[0085]

[0086] where d is the dimension of the output of the network layer, and W3 and W4 are trainable weight matrices.

[0087] Preferably, considering that the Actor of each warehouse only needs to make decisions using its own observations, we use a multi-layer perceptron as the Actor network to complete the decision-making process with a relatively small network scale.

[0088] S3.2. Each warehouse i receives local observations and selects an action according to the policy , executes the action to obtain the reward at time t and the observation at the next time

[0089] S3.3. Each warehouse uses its own Critic network to aggregate and learn the information of neighbors on the graph. The Critic network uses the graph attention mechanism, and finally obtains the value of the local value function of warehouse i at time t

[0090] S3.4. The observation sets of all warehouses at time t Set of actions Set of rewards Set of values of the value function And the set of observations at time t+1 Store them in the experience buffer.

[0091] S3.5. Repeatedly sample from the experience buffer to calculate the advantage function for each warehouse:

[0092]

[0093] Where γ is the discount factor, λ is a hyperparameter, and T is the number of steps in each episode.

[0094] S3.6. Update the network parameter φ using the Critic loss function i :

[0095]

[0096] Where B is the training sample size, ∈ is a hyperparameter, is the cumulative return, and clip(a, b, c) is a clipping function that restricts a to the interval (b, c).

[0097] Update the network parameter θ using the Actor loss function i :

[0098]

[0099] S3.7. Repeat steps S3.2 to S3.6 until the policy converges.

[0100] S4. Each warehouse obtains its own observation from the environment and inputs it into the trained policy to obtain the resource allocation plan for each warehouse.

[0101] The simulation results are as Figure 6 shown, where GRA uses a 3-layer graph attention network, while GRA-decen uses a single-layer graph attention network. A single warehouse only needs to aggregate and learn the information of the one-hop neighbors on the graph. The simulation results show that:

[0102] The global reward increase rate of GRA and GRA-decen is faster than that of other compared multi-agent reinforcement learning algorithms, and the finally converged global reward is also higher than that of other algorithms. The reason is that the present invention simplifies the global value function into a local value function. The difficulty of fitting the local value function is less than that of the global value function, and the input information that GRA-decen needs to process is less than the global information, making the training easier. Therefore, the method adopted by the present invention is more suitable for large-scale systems.

Claims

1. An inventory resource scheduling method based on decentralized multi-agent reinforcement learning, characterized in that The specific implementation of this method is as follows: S1. Construct a resource allocation problem model. In this model, a single warehouse is regarded as an agent, and the set of all agents is regarded as the point set on the topological graph. Use the edge set ε in the topological graph to describe the inherent coupling relationship between warehouses, and analyze the explicit and implicit cooperation relationships between warehouses. S2. Construct the local value function of each warehouse by using the cooperation relationship between warehouses And simplify the input of each local value function from global information to local information; S3. Train the local value function and policy of each warehouse according to the GRA algorithm; S4. Each warehouse obtains its own observation from the environment and inputs it into the trained policy to obtain the resource allocation plan for its own resources.

2. The inventory resource scheduling method based on decentralized multi-agent reinforcement learning according to claim 1, characterized in that Step S1 uses the observation graph State transition graph and the reward graph respectively describe the observation coupling relationship, state transition function coupling relationship, and reward function coupling relationship between warehouses. The union of the observation graph and the state transition graph is denoted as Derive the learning graph based on the coupling relationship between agents The neighbors of warehouse i in the learning graph are denoted as: Among them, represents the set of warehouses that have an explicit cooperation relationship with warehouse k, represents the set of all warehouses reachable by warehouse i in ​ 3. The inventory resource scheduling method based on decentralized multi-agent reinforcement learning according to claim 1, wherein The implicit cooperation relationship in step S1 is jointly determined by the observation graph and the state transition graph. There are two forms of implicit cooperation relationship. The first is that the inventory resource quantity of a single warehouse will be observed by other warehouses, and when the inventory resource quantity of this warehouse changes, it will have an impact on the warehouses connected to it; the second is the impact generated when the quantity of resources transported by a single warehouse to other connected warehouses changes; the explicit cooperation relationship is determined by the reward graph. The states or actions of other warehouses will directly affect the rewards of a single warehouse. When a warehouse maximizes its own rewards, it will cooperate with the warehouses that have an explicit cooperation relationship with it.

4. The inventory resource scheduling method based on decentralized multi-agent reinforcement learning according to claim 1, characterized in that The state transition function for constructing the resource allocation problem model in step S1 is: Among them, is the resource quantity of warehouse i at time t, a ij is the proportion of the resources of warehouse i transported to warehouse j itself, is the set of successor nodes of warehouse i on the state transition diagram, is the set of predecessor nodes of warehouse i on the state transition diagram, is the perturbation related to time step t; in addition, we add random noise when initializing the resource quantity of each warehouse and make a limitation on a ij as follows: The reward function of warehouse i in the resource allocation problem model constructed in step S1 is: Among them is the set of predecessor nodes of warehouse i on the reward graph.

5. The inventory resource scheduling method based on decentralized multi-agent reinforcement learning according to claim 1, wherein, Step S1 constructs a resource allocation problem model. The observation o of each warehouse i includes its own resource quantity and the resource quantities of its neighbors. The action a of each warehouse i is the proportion of its existing resources delivered to other warehouses.

6. The inventory resource scheduling method based on decentralized multi-agent reinforcement learning according to claim 1, characterized in that, The GRA algorithm adopted in step S3 includes: S6.

1. Use the Actor-Critic architecture, and set an Actor network and a Critic network for each warehouse i respectively, with their parameters being θ i and φ i , and initialize the parameters of each network; S6.

2. Each warehouse i receives local observations According to the policy Select an action Execute the action to obtain the reward at time t And the observation at the next time S6.

3. Each warehouse uses its own Critic network to aggregate and learn the information of its neighbors on the graph. The Critic network uses the graph attention mechanism, and finally obtains the value of the local value function of warehouse i at time t. S6.

4. Store the observation sets of all warehouses at time t action set reward set value set of the value function and the observation set at time t+1 into the experience buffer; S6.

5. Repeatedly sample the experience buffer to calculate the advantage function of each warehouse: Among them γ is the discount factor, λ is the hyperparameter, and T is the number of steps per episode; S6.

6. Update the network parameter φ using the Critic loss function i : where, B is the training sample size, and ∈ is a hyperparameter, is the cumulative return of warehouse i, and clip(a, b, c) is a clipping function that restricts a to the interval (b, c); Update the network parameter θ using the Actor loss function i : S6.

7. Repeat steps S6.2 to S6.6 until the policy converges.

7. The inventory resource scheduling method based on decentralized multi-agent reinforcement learning according to claim 6, wherein Regarding the information aggregation of neighbors on the aggregated learning graph in step S6.3, the graph attention mechanism is used to complete the information transmission between warehouses. The aggregated information not only includes the current inventory resource quantity of neighbors, but also includes the reward information of neighbors, that is, the penalty value suffered by the warehouse due to resource shortage.

Citation Information

Patent Citations

  • Inter-warehouse scheduling method and system for target articles

    CN114091988A

  • Supply chain inventory management method based on multi-agent reinforcement learning

    CN116823107A

  • Cross-warehouse inventory scheduling system

    CN117875842A

  • D2D user resource allocation method based on multi-agent reinforcement learning

    CN118118908A