Logistics whole process optimization system based on multi-agent reinforcement learning and optimization method thereof
Through the logistics full-process optimization system based on multi-agent reinforcement learning, the problem of lack of global collaborative optimization and dynamic response in the logistics system is solved, efficient collaborative optimization and dynamic adjustment of the logistics system are achieved, and the overall performance and adaptability of the logistics system are improved.
Patent Information
- Application Number
- CN202510218623.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-10
AI Technical Summary
The logistics system lacks global collaborative optimization capabilities, making it difficult to achieve effective linkage between various links, and the response is lagging in a dynamically changing environment, making it difficult to achieve dynamic optimization.
The full-process logistics optimization system based on multi-agent reinforcement learning is adopted, and a high-fidelity logistics scenario model is built through modules such as multi-source heterogeneous data acquisition, environmental simulation, agent initialization and reinforcement learning training, and a dynamic distribution collaborative reinforcement learning algorithm is used to train the agent's optimization strategy network to realize intelligent distribution and dynamic adjustment of logistics tasks.
The overall collaborative optimization of the logistics system has been realized, the linkage capabilities between various links have been enhanced, the system's real-time response and dynamic optimization capabilities have been improved, and the overall performance and adaptability of the logistics system have been improved.
Smart Images

Figure CN120125124A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of logistics whole - process optimization, and particularly to a logistics whole - process optimization system and its optimization method based on multi - agent reinforcement learning. Background Art
[0002] The current logistics system is developing rapidly towards intelligence and digitization, but still faces many challenges in terms of whole - process optimization. Traditional logistics system optimization methods usually adopt segmented optimization. For example, warehouse management, transportation scheduling, distribution path planning, etc. are optimized as independent modules. Although this local optimization method can improve efficiency within a single link, there is a lack of effective coordination mechanisms among modules, resulting in uneven resource allocation and poor process connection, thus limiting the overall performance improvement of the system. In addition, in actual logistics scenarios, dynamic changes and uncertain environments are normal problems. Dynamic factors pose higher requirements for the real - time optimization ability of the logistics system. However, existing optimization methods often cannot achieve efficient resource scheduling and task allocation globally when dealing with these dynamic changes, and lack effective real - time optimization strategies. In recent years, multi - agent reinforcement learning technology has shown great potential in solving complex decision - making problems. By simulating the collaborative behavior of multiple roles and learning global optimal strategies, it can effectively solve the collaborative optimization problems in complex dynamic systems.
[0003] But the above - mentioned technology has at least the following technical problems: In the logistics system, there is a lack of global collaborative optimization ability for the whole process, it is difficult to achieve effective linkage among various links, and in the case of dynamic changes in the logistics scenario, there is a lag in response to environmental changes, making it difficult to achieve dynamic optimization. Summary of the Invention
[0004] The present invention provides a logistics whole - process optimization system based on multi - agent reinforcement learning to solve the technical problems in the existing logistics system, including the lack of global collaborative optimization ability for the whole process, the difficulty in achieving effective linkage among various links, and in the case of dynamic changes in the logistics scenario, the lag in response to environmental changes, making it difficult to achieve dynamic optimization.
[0005] A logistics whole - process optimization system based on multi - agent reinforcement learning of the present invention specifically includes the following technical solutions:
[0006] A logistics whole - process optimization system based on multi - agent reinforcement learning includes the following parts:
[0007] Multi - source heterogeneous data acquisition module, environment simulator module, agent initialization module, reinforcement learning training module, task allocation and scheduling module, execution and feedback module;
[0008] The multi-source heterogeneous data acquisition module obtains multi-source heterogeneous data from multiple data sources, preprocesses the multi-source heterogeneous data to obtain the preprocessed multi-source heterogeneous data, extracts features from the preprocessed multi-source heterogeneous data to obtain a feature set, merges the feature set with the preprocessed multi-source heterogeneous data to obtain a comprehensive logistics data set, and then transfers the comprehensive logistics feature data to the environment simulator module;
[0009] The environment simulator module constructs a high-fidelity logistics scenario model based on the real scenario combined with the comprehensive logistics feature data set, including warehouse layout, traffic network topology, and dynamic environment characteristics. At the same time, dynamic random perturbations are introduced during the construction of the high-fidelity logistics scenario model to enable the intelligent agent to adapt to the complex and changing environment, obtaining a high-fidelity logistics scenario model, and sending it to the intelligent agent initialization module and the reinforcement learning training module to provide a basic environment for the subsequent learning of optimization strategies and task allocation;
[0010] The intelligent agent initialization module defines and initializes multiple heterogeneous intelligent agent models based on the high-fidelity logistics scenario model to obtain the initialized intelligent agent models, and sends them to the reinforcement learning training module for further optimizing the strategy;
[0011] The reinforcement learning training module combines the initialized intelligent agent models and the high-fidelity logistics scenario model, and uses the dynamic distribution collaborative reinforcement learning algorithm to train the optimization strategy network of each intelligent agent to obtain the trained optimization strategy network, and transfers it to the task allocation and scheduling module for realizing the intelligent allocation and dynamic adjustment of logistics tasks;
[0012] The task allocation and scheduling module formulates an optimized task allocation plan and scheduling plan based on the trained optimization strategy network and in combination with the real-time updated logistics status data, and transfers the optimized task allocation plan and scheduling plan to the execution and feedback module for the actual operation of the logistics system;
[0013] The execution and feedback module controls each execution unit in the logistics full-process optimization system to complete tasks according to the optimized task allocation plan and scheduling plan, monitors the real-time status data during the execution process to obtain the execution result, compares the execution result with the expected target to generate feedback data, and transfers the feedback data as feedback to the multi-source heterogeneous data acquisition module to realize the closed-loop optimization of the system and further improve the full-process efficiency and adaptability of the logistics system.
[0014] A logistics full-process optimization method based on multi-agent reinforcement learning includes the following steps:
[0015] S1. Obtain multi-source heterogeneous data from multiple data sources, preprocess the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data, extract features from the preprocessed multi-source heterogeneous data to obtain a feature set, merge the feature set with the preprocessed multi-source heterogeneous data to obtain a comprehensive logistics data set, and then construct a high-fidelity logistics scenario model based on the real scenario and the comprehensive logistics feature data set;
[0016] S2. Based on the high-fidelity logistics scenario model, define and initialize multiple heterogeneous agent models to obtain initialized agent models, and then, in combination with the high-fidelity logistics scenario model, use the dynamic distributed cooperative reinforcement learning algorithm to train the optimization policy network of each agent to obtain the trained optimization policy network;
[0017] S3. Based on the trained optimization policy network and in combination with the real-time updated logistics status data, formulate an optimized task allocation plan and scheduling plan, and then, according to the optimized task allocation plan and scheduling plan, control each execution unit in the logistics whole-process optimization system to achieve the logistics whole-process optimization based on multiple agents.
[0018] Preferably, S1 specifically includes:
[0019] Obtain multi-source heterogeneous data in logistics from multiple data sources, including logistics order data, traffic flow data, warehouse status data, external environment data, and equipment status data; perform preprocessing such as data cleaning, data conversion, outlier processing, and dimensionless processing on the obtained multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data, use existing feature engineering techniques to extract features from the preprocessed multi-source heterogeneous data to obtain a feature set, and perform column-wise merging of the feature set with the preprocessed multi-source heterogeneous data to form a unified comprehensive logistics data set, including the preprocessed original data and the extracted feature data.
[0020] Preferably, S1 specifically includes:
[0021] Construct a high-fidelity logistics scenario model based on the comprehensive logistics data set. The construction process includes the complete logic and mathematical modeling details from the input of the comprehensive logistics data set to the scenario generation, including the integration of warehouse layout, traffic network topology, dynamic environment characteristics, and random perturbations.
[0022] Preferably, S1 specifically includes:
[0023] During the construction process of the high-fidelity logistics scenario model, use the warehouse data in the comprehensive logistics data set to construct a warehouse layout model. The warehouse layout model is represented by a multi-dimensional matrix.
[0024] Preferably, S1 specifically includes:
[0025] In the process of constructing a high-fidelity logistics scenario model, a traffic network topology is constructed. The traffic network topology is represented by a graph G=(V, E), where V is a set of nodes representing road network intersections or warehousing locations, and E is a set of edges representing roads.
[0026] Preferably, the S1 specifically includes:
[0027] In the process of constructing a high-fidelity logistics scenario model, the dynamic environmental characteristics are modeled through dynamic feature embedding, and the temporal characteristics and context characteristics extracted from order data and traffic flow data are fused to form a state embedding vector.
[0028] Preferably, the S1 specifically includes:
[0029] In the process of constructing a high-fidelity logistics scenario model, the edge weights in the traffic network topology are adjusted in combination with the disturbance perception value.
[0030] Preferably, the S2 specifically includes:
[0031] Based on the high-fidelity logistics scenario model, a variety of heterogeneous agent models are defined and initialized. First, different functional roles in the entire logistics system need to be mapped to agent models according to the specific characteristics of the logistics scenario model. The definition of the agent model needs to include the agent type, state space, action space, reward function, and their cooperation mechanisms.
[0032] Preferably, the S2 specifically includes:
[0033] To complete the agent initialization, the state space, action space, and reward function need to be set for each type of agent.
[0034] Preferably, the S2 specifically includes:
[0035] To make the agent initialization scalable and operable, specific restrictions need to be imposed on the size, variable type, and value range of the state space and action space.
[0036] Preferably, the S2 specifically includes:
[0037] Introduce a dynamic distributed cooperative reinforcement learning algorithm based on reinforcement learning to realize the adaptive distributed cooperative strategy optimization among agents. The dynamic distributed cooperative reinforcement learning algorithm gradually converges to the global optimal strategy by jointly optimizing the policy network and the dynamic value network.
[0038] Preferably, the S2 specifically includes:
[0039] In the process of implementing the dynamic distributed cooperative reinforcement learning algorithm based on reinforcement learning, a cooperative reward mechanism among agents is introduced.
[0040] Preferably, S2 specifically includes:
[0041] During the implementation of the dynamic distributed cooperative reinforcement learning algorithm based on reinforcement learning, a joint policy network is designed to optimize the policy.
[0042] Preferably, S3 specifically includes:
[0043] Based on the trained optimized policy network, when combining real-time logistics status data, through the dynamic scenario update mechanism, the order demand, traffic conditions, warehouse inventory, and equipment status in the logistics scenario are transformed into standardized state vectors, and then the standardized state vectors are input into the optimized policy network for inference. The optimized policy network, based on the behavior policy learned during the dynamic distributed cooperative reinforcement learning process of reinforcement learning, calculates the action output of each agent through forward propagation. The distribution terminal agent adjusts the distribution order according to the path complexity and order priority.
[0044] Preferably, S3 specifically includes:
[0045] Using the action values output by the optimized policy network and combining real-time logistics status data, an optimized task allocation plan and scheduling plan are formulated.
[0046] Preferably, S3 specifically includes:
[0047] In particular, the execution status is monitored in real time through Internet of Things devices and sensor networks to obtain the execution result. After comparing the execution result with the expected target, feedback data is generated, including performance indicators (such as latency time, task completion efficiency) and abnormal information. Then, the feedback data is transmitted as feedback to the multi-source heterogeneous data acquisition module to achieve the closed-loop optimization of the system, further improving the full-process efficiency and adaptability of the logistics system.
[0048] The beneficial effects of the technical solution of the present invention are:
[0049] 1. Through the comprehensive logistics data set, a high-fidelity logistics scenario model including warehouse layout, traffic network topology, dynamic environment characteristics, and random perturbations is constructed to truly simulate the dynamic changes of the logistics system; the warehouse layout model describes the cargo storage capacity through a fine multi-dimensional matrix, realizing the efficient management of warehouse resources. The traffic network topology combines multi-dimensional attributes such as path length, traffic flow, and vehicle speed, and dynamically adjusts the path weights, truly reflecting the impact of road traffic conditions and emergencies on traffic. The dynamic environment characteristics fuse the order time series characteristics and context characteristics through the state embedding vector, and combined with random perturbations, enhance the model's perception ability of emergencies, ensuring the accuracy and robustness of scenario modeling.
[0050] 2. Based on the dynamic distribution cooperative reinforcement learning algorithm, the policy network and value network of the agent are jointly optimized to form a joint policy that takes into account both local task optimization and global cooperation optimization; by introducing a cooperative utility function into the reward function, the optimization goals of local agents and the cooperation goals of the global system are balanced, ensuring the maximization of global benefits; the joint policy network realizes joint optimization among multiple agents through a deep neural network, supporting efficient cooperation in complex task scenarios; the distributed update of the value network quantifies the cooperation value among agents through the joint value function and dynamically adjusts the policies of agents, further enhancing the global adaptability of the system. BRIEF DESCRIPTION OF THE DRAWINGS
[0051] Figure 1 It is a structural diagram of a logistics full-process optimization system based on multi-agent reinforcement learning according to the present invention;
[0052] Figure 2 It is a flowchart of a logistics full-process optimization method based on multi-agent reinforcement learning according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] In order to further elaborate on the technical means and effects adopted by the present invention to achieve the predetermined invention purpose, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0054] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which the present invention belongs.
[0055] The following specifically describes the specific solution of a logistics full-process optimization system based on multi-agent reinforcement learning provided by the present invention in conjunction with the accompanying drawings.
[0056] Referring to the appendix Figure 1 , which shows a structural diagram of a logistics full-process optimization system based on multi-agent reinforcement learning provided by an embodiment of the present invention. The system includes the following parts:
[0057] A multi-source heterogeneous data collection module, an environment simulator module, an agent initialization module, a reinforcement learning training module, a task allocation and scheduling module, an execution and feedback module;
[0058] The multi-source heterogeneous data acquisition module obtains multi-source heterogeneous data from multiple data sources, preprocesses the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data, extracts features from the preprocessed multi-source heterogeneous data to obtain a feature set, merges the feature set with the preprocessed multi-source heterogeneous data to obtain a comprehensive logistics data set, and then transfers the comprehensive logistics feature data to the environment simulator module;
[0059] The multi-source heterogeneous data includes logistics order data (such as time, quantity, delivery address), traffic flow data (historical and real-time traffic conditions), warehouse status data (inventory situation, storage location, inbound and outbound records), external environment data (such as weather information, emergencies), and equipment status data (such as vehicle location, equipment availability);
[0060] The environment simulator module constructs a high-fidelity logistics scenario model based on the real scenario combined with the comprehensive logistics feature data set, including warehouse layout, traffic network topology, and dynamic environment characteristics (such as order density changes, traffic flow fluctuations, etc.). At the same time, dynamic random perturbations (such as path congestion, sudden order peaks, etc.) are introduced during the construction of the high-fidelity logistics scenario model to enable the intelligent agent to adapt to the complex and changing environment, obtain the high-fidelity logistics scenario model, and send it to the intelligent agent initialization module and the reinforcement learning training module to provide a basic environment for the subsequent learning of optimization strategies and task allocation;
[0061] The intelligent agent initialization module defines and initializes multiple heterogeneous intelligent agent models (such as warehouse management intelligent agent, transportation scheduling intelligent agent, distribution terminal intelligent agent, etc.) based on the high-fidelity logistics scenario model. During the initialization process, specific state spaces, action spaces, and initial strategies are assigned to each intelligent agent according to different roles. At the same time, a reward function and a cooperation mechanism are defined for each intelligent agent to ensure that multiple heterogeneous intelligent agents can interact with the logistics environment in real time and learn optimization strategies, obtain the initialized intelligent agent model, and send it to the reinforcement learning training module for further optimization of the strategy;
[0062] The reinforcement learning training module combines the initialized intelligent agent model and the high-fidelity logistics scenario model, and uses a dynamic distribution cooperative reinforcement learning algorithm (such as Deep Deterministic Policy Gradient, DDPG) to train the optimization strategy network of each intelligent agent. During the training process, by introducing a cooperative reward mechanism, it prompts the intelligent agents to optimize the global performance (such as order completion rate, resource utilization rate) while completing local goals, obtains the trained optimization strategy network, and transfers it to the task allocation and scheduling module for realizing the intelligent allocation and dynamic adjustment of logistics tasks;
[0063] The task allocation and scheduling module, based on the trained optimized policy network and combined with the real-time updated logistics status data (such as real-time order demand, traffic conditions, changes in warehouse inventory), formulates an optimized task allocation plan and scheduling plan, and transfers the optimized task allocation plan and scheduling plan to the execution and feedback module for the actual operation of the logistics system;
[0064] The execution and feedback module controls each execution unit (such as transport vehicles, warehouse equipment, delivery personnel) in the logistics full-process optimization system to complete tasks according to the optimized task allocation plan and scheduling plan. At the same time, it monitors the real-time status data during the execution process (such as actual transport time, order completion rate, resource utilization rate, etc.), obtains the execution results, and generates feedback data after comparing the execution results with the expected goals, including system performance indicators (such as delay time, task completion efficiency) and exception information. The feedback data is transmitted as feedback to the multi-source heterogeneous data acquisition module to achieve the closed-loop optimization of the system and further improve the full-process efficiency and adaptability of the logistics system.
[0065] Refer to the appendix Figure 2 , which shows a flowchart of a logistics full-process optimization method provided by an embodiment of the present invention. The method includes the following steps:
[0066] S1. Obtain multi-source heterogeneous data from multiple data sources, preprocess the multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data, extract features from the preprocessed multi-source heterogeneous data to obtain a feature set, merge the feature set with the preprocessed multi-source heterogeneous data to obtain a comprehensive logistics data set, and then construct a high-fidelity logistics scenario model based on the real scenario and the comprehensive logistics feature data set;
[0067] Obtain multi-source heterogeneous data in logistics from multiple data sources, including logistics order data, traffic flow data, warehouse status data, external environment data, and equipment status data; perform preprocessing such as data cleaning, data conversion, outlier processing, and dimensionless processing on the obtained multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data. The technical means adopted in this preprocessing process are existing technologies and will not be elaborated here.
[0068] Furthermore, use existing feature engineering techniques to extract features from the preprocessed multi-source heterogeneous data to obtain a feature set, including time series features, spatial features, business features, and environmental features. Merge the feature set with the preprocessed multi-source heterogeneous data column by column to form a unified comprehensive logistics data set, which contains the original data after preprocessing and the extracted feature data.
[0069] Furthermore, a high-fidelity logistics scenario model is constructed based on the comprehensive logistics data set. The construction process includes the complete logic and mathematical modeling details from the input of the comprehensive logistics data set to scenario generation, including the integration of warehouse layout, traffic network topology, dynamic environment characteristics, and random perturbations. The specific process is as follows:
[0070] The comprehensive logistics data set is standardized through existing standardization means to form a unified input format. First, a warehouse layout model is constructed using the warehouse data in the comprehensive logistics data set. The warehouse layout model is represented by a multi-dimensional matrix M, where M ij is the storage capacity of the j-th storage unit on the i-th floor of the warehouse. Define the storage capacity of M ij as:
[0071]
[0072] where C ij represents the maximum capacity of the j-th storage unit on the i-th floor of the warehouse, P k represents the unit occupancy of item k, δ kij is an indicator function. If item k is stored in the j-th storage unit on the i-th floor of the warehouse, then δ kij = 1.
[0073] Furthermore, a traffic network topology is constructed. The traffic network topology is represented by a graph G=(V, E), where V is the set of nodes, representing the road network intersection points or warehouse locations, and E is the set of edges, representing the roads. The edge weight is an attribute of the path , including the length basic travel time traffic flow The edge weight is initialized as:
[0074]
[0075] where is the average vehicle speed of the path , and n are adjustment coefficients for the influence of traffic flow on travel time.
[0076] Furthermore, the dynamic environment characteristics are modeled through dynamic feature embedding. The temporal characteristics R(s) and context characteristics C(s) (such as historical order trends) extracted from order data and traffic flow data are fused to form a state embedding vector:
[0077] Φ(s)=σ(W (1) ·f(R(s), C(s))+W (2) ·Δ(s, t))
[0078] Among them, Φ(s) is the feature embedding vector of state s, representing the enhanced state representation after dynamic perturbation and feature combination, that is, the dynamic environment characteristics; W (1) and W (2) are feature mapping weight matrices, which respectively control the contribution degrees of f(R(s), C(s)) and Δ(s, t) to the embedding vector; σ(f) is an activation function, such as ReLU or Sigmoid, which enables the embedded features to have non-linear expression ability; f(R(s), C(s)) is the non-linear combination of the static feature R(s) and the context feature C(s); Δ(s, t) is the perturbation value of state s at time t, representing the instantaneous impact of the dynamic environment characteristics on the state;
[0079] The feature combination function is defined as:
[0080]
[0081] Among them, R(s) is the static feature value of state s, such as order density, path complexity, which is the fixed feature data extracted from the comprehensive logistics feature data set; C(s) is the context feature value of state s, such as historical order trend, traffic mean value, which is extracted from the dynamic environment data through existing statistical analysis or time series models;
[0082] Furthermore, the perturbation perception value is calculated through historical traffic data and real-time emergencies (dynamic random perturbations):
[0083]
[0084] Among them, Δ(s, t) is the perturbation perception value of state s at time t, reflecting the impact of dynamic events on the current state; is the weight of the perturbation event at time t, representing the occurrence probability or influence intensity of the event; is the immediate impact value of the perturbation event on state s, such as the influence intensity of a sudden increase in orders or an increase in path congestion; N is the total number of perturbation events, such as the possible number of path congestion points, the number of sudden order areas; T is the length of the perturbation observation window, representing the time range used to calculate the historical average impact value.
[0085] Combined with the perturbation perception value Δ(s, t), the traffic network edge weights are adjusted
[0086]
[0087] Among them, Δ max is the maximum possible range of the perturbation value, which is used for normalization processing. This adjustment reflects the impact of path congestion or emergencies on the travel time.
[0088] Finally, the above storage layout model, traffic network topology model, and dynamic environment characteristics are integrated into a unified scenario model, that is, a high-fidelity logistics scenario model is generated: It is a high-fidelity logistics scenario model.
[0089] S2. Based on the high-fidelity logistics scenario model, define and initialize multiple heterogeneous agent models to obtain the initialized agent models. Then, combined with the high-fidelity logistics scenario model, use the dynamic distributed cooperative reinforcement learning algorithm to train the optimization policy network of each agent to obtain the trained optimization policy network;
[0090] Based on the high-fidelity logistics scenario model, define and initialize multiple heterogeneous agent models. First, according to the specific characteristics of the logistics scenario model, map different functional roles in the entire logistics system to agent models. The definition of the agent model needs to include the agent type, state space, action space, reward function, and their cooperation mechanisms. For example, the warehouse management agent is responsible for the inbound, outbound, and storage path optimization of goods; the transportation scheduling agent is responsible for the path planning and loading optimization of transport vehicles; the distribution terminal agent is responsible for the selection of distribution paths and the sorting of order priorities. Through the functional decomposition of heterogeneous agents, distributed optimization of different links in the logistics system can be achieved.
[0091] To complete the agent initialization, it is necessary to set the state space, action space, and reward function for each type of agent. The specific process is as follows: The state space is composed of key variables in the logistics scenario and is used to describe the environmental state of the agent. For example, the state space of the warehouse management agent can include inventory levels, storage location distribution, goods types, task queues, and the priority of the currently executing task; the state space of the transportation scheduling agent can include the current location of the vehicle, the status of traffic network nodes, order distribution, loading capacity, and driving speed; the state space of the distribution terminal agent can include the distribution path, time window constraints, order completion status, and geographical location. The goal of state space design is to ensure that the agent can comprehensively perceive its surrounding environment and provide sufficient information for subsequent decision-making.
[0092] The action space defines all possible behaviors that the agent can take in each state and is used to guide the specific operations of the agent. For example, the action space of the warehouse management agent includes goods reallocation, selection of storage locations, adjustment of storage paths, etc.; the action space of the transportation scheduling agent includes vehicle path planning, update of driving directions, dynamic adjustment of loading tasks, etc.; the action space of the distribution terminal agent includes selection of distribution order, adjustment of paths, and allocation of distribution time windows. The definition of the action space needs to combine the physical constraints and business logic of actual operations to ensure that all actions are feasible in reality.
[0093] The reward function is defined based on the optimization objective and is used to evaluate the performance of the agent after executing an action. The design of the reward function should be consistent with the specific business requirements and the global system optimization objective. For example, the reward function of the warehouse management agent can be designed to minimize the time for picking and placing goods while increasing inventory utilization; the reward function of the transportation scheduling agent can be designed to minimize the total path length or travel time while reducing energy consumption; the reward function of the distribution terminal agent can be to minimize the delay delivery time and maximize customer satisfaction. At the same time, in order to promote cooperation among multiple agents, the reward function needs to introduce a collaborative reward mechanism, such as by sharing global metrics like order completion rate and resource utilization rate to achieve optimal coordination among agents.
[0094] Based on the reward function and the collaboration mechanism, it is also necessary to define the communication method and data exchange rules among agents. For example, the transportation scheduling agent can share the current traffic status and estimated arrival time with the distribution terminal agent in real time, and the distribution terminal agent can dynamically adjust the distribution order according to the shared information; the warehouse management agent can share the goods outbound status with the transportation scheduling agent to optimize the loading and distribution processes. The cooperation among agents is achieved through a mechanism based on shared rewards, ensuring the improvement of the global system performance while achieving local optimization.
[0095] To make the agent initialization scalable and operable, specific restrictions need to be imposed on the size, variable type, and value range of the state space and action space. For example, the inventory level variable of the warehouse management agent can be a non - negative integer with a value range of [0, K] (where K is the maximum warehouse capacity); the current position of the vehicle of the transportation scheduling agent can be represented by two - dimensional coordinates, and the range is restricted by the traffic network topology; the distribution order of the distribution terminal agent can be represented by permutations and combinations, and the quantity is restricted by the order scale. These restrictions provide technical guarantees for the feasibility of agent initialization.
[0096] Furthermore, a dynamic distributed cooperative reinforcement learning algorithm based on reinforcement learning is introduced to achieve the adaptive distributed cooperative policy optimization among agents. The dynamic distributed cooperative reinforcement learning algorithm gradually converges to the global optimal policy by jointly optimizing the policy network and the dynamic value network. The specific implementation process is as follows:
[0097] In particular, a high - fidelity logistics scenario model runs through the state space, reward function, action space, dynamic perturbation perception, and training process, providing a real and dynamic environmental basis for the dynamic distributed cooperative reinforcement learning, enabling agents to learn effective optimization strategies in complex and changing logistics scenarios.
[0098] First, the state space of the agent action space and reward function Expressions that need to be redefined as dynamic couplings. State space N is the total number of agents, including local states Global state and collaboration weights Among them, specifically is the weight constructed from the historical interaction data of agents, representing the agents and collaboration intensity. Action simultaneously includes local actions and collaborative actions to optimize local tasks and system collaboration goals respectively. The redesign of the reward function, by introducing a collaborative reward mechanism among agents, extends the traditional reward function to a composite expression In this formula, is the local reward, representing the efficiency of the agent in completing local tasks; λ is a tuning parameter used to balance local optimization and global collaboration; is the collaborative utility function, defined as where α is the decay rate of the collaborative effect, is the Euclidean distance between agent actions.
[0099] Furthermore, the policy is optimized by designing a joint policy network. The joint policy network Π = {π 1 , π 2 , …, π N} is a set composed of the policy networks of multiple agents, where the policy network of each agent is represented by a deep neural network, and the output is the action probability distribution of the agent. The optimization of the policy network adopts the following objective function:
[0100]
[0101] Among them, is the joint policy objective function, which measures the comprehensive performance of the policy networks of all agents in the given state space for the global task and collaborative effect. The optimization objective of this function is to enable the agents to generate optimal actions in different states by updating the policy network parameters, taking into account both local task optimization and multi-agent collaboration; is the expected value operator, used to calculate the weighted average performance under all possible states and actions; is the agent in state selecting action the logarithm of the probability, representing the preference degree of the policy network for the selected action in this state; is the agent in state taking action The corresponding value function represents the long-term cumulative reward obtained after executing the action; k is the collaboration intensity adjustment parameter, which is used to balance the influence of local tasks and collaboration effects on the overall goal; is the contribution function of the agent to the collaboration effect, defined as
[0102] The joint distributed update of the value network is another key part of the algorithm. The value function is defined as:
[0103]
[0104] where is the agent obtains the immediate reward after executing the action , γ is the discount factor, which is used to measure the importance of future rewards to the current decision (0 < γ < 1); is the agent selects the action at the next moment state value, representing the estimated value of subsequent rewards; β is the collaboration intensity adjustment coefficient, which is used to balance the local value and the joint value of multiple agents; is the joint value function, which measures the collaboration value of agents and under their respective state and action combinations, defined as:
[0105]
[0106] where is the decay factor of the collaboration effect, which controls the influence of the agent action distance on the collaboration efficiency and is determined by the empirical method; is the Euclidean distance between agent actions, which is used to measure the collaboration efficiency; describes the consistency of the action direction; describes the influence of the state difference on the collaboration efficiency.
[0107] The update of the value network follows the following loss function:
[0108]
[0109] where is the loss function of the value network, which is used to measure the deviation between the predicted by the current value network and the target value y; the smaller the loss function, the more accurate the prediction of the value network target value; is the predicted value of the value network, indicating that the agent executes the action in the state The long-term cumulative reward expectation value; θ Q is a set of parameters of the value network, including the weights and biases of the neural network, used to map the input state and action to the predicted value of the output y is the target value, representing the ideal value calculated according to the Bellman equation, which is the weighted sum of the immediate reward and the value at the next moment. By taking the derivative of the loss function of the value network, the parameter update formula of the value network is obtained:
[0110]
[0111] where, Δθ Q is the update amount of the value network parameters. It represents the direction and magnitude of adjusting the value network parameters by the gradient descent method in the current training iteration; ζ is the learning rate, which controls the magnitude of parameter update; is the gradient operator for the value network parameters θ Q and calculates the gradient of the loss function with respect to the parameter θ Q of the gradient.
[0112] Based on the joint policy optimization and value function update, the final global optimization objective is defined as:
[0113]
[0114] where, is the global objective function, representing the optimization objective function of the entire multi-agent system, measuring the overall performance of the agents under the current state and action selection, including the task completion quality of individual agents and the cooperation efficiency between agents; μ is the global optimization adjustment parameter, that is, the global cooperation adjustment factor, which weighs the importance between the local task objective and the global cooperation objective. The larger the value, the greater the contribution of the cooperation objective to the global optimization.
[0115] Furthermore, through the above multi-round iterative training, the policies of the agents gradually converge, and at the same time show high adaptability in the dynamic logistics scenario. Finally, the trained optimized policy network and value network are obtained, which are used for real-time decision-making, and the agents can achieve global optimal cooperation in the highly dynamic and uncertain logistics environment.
[0116] S3. Based on the trained optimized policy network, combined with the real-time updated logistics state data, formulate an optimized task allocation plan and scheduling plan, and then control each execution unit in the logistics whole-process optimization system according to the optimized task allocation plan and scheduling plan to achieve the logistics whole-process optimization based on multi-agents.
[0117] Based on the trained optimized policy network, when combined with real-time logistics status data, through the dynamic scenario update mechanism, the order demand, traffic conditions, warehouse inventory, and equipment status in the logistics scenario are transformed into standardized state vectors. The standardized state vectors are input into the optimized policy network for inference. The optimized policy network, based on the behavioral policies learned during the dynamic distributed cooperative reinforcement learning process of reinforcement learning, calculates the action outputs of each agent through forward propagation. For example, the warehouse management agent calculates the optimal inbound and outbound operations based on the inventory status, the transportation scheduling agent outputs the vehicle path node sequence according to the order demand and traffic conditions, and the distribution terminal agent adjusts the distribution order according to the path complexity and order priority.
[0118] Furthermore, using the action values output by the optimized policy network and combining with real-time logistics status data, an optimized task allocation plan and scheduling plan are formulated. During the task allocation process, the priority scoring function defined by the expert experience method is used to sort the tasks by priority, providing a decision basis for the scheduling plan to obtain the task allocation result. The transportation scheduling agent generates the transportation path according to the task allocation result, uses the existing shortest path algorithm and combines the output of the policy network to dynamically adjust the node order, generates the scheduling plan, and after generating the scheduling plan, converts it into specific logistics operation instructions and issues them to the execution unit. The warehouse equipment receives the shelf operation path instructions, the transportation vehicle receives the loading and path planning instructions, and the distribution terminal receives the task priority list.
[0119] Specifically, the execution status is monitored in real time through Internet of Things devices and sensor networks to obtain the execution result. After comparing the execution result with the expected target, feedback data is generated, including performance metrics (such as latency time, task completion efficiency) and exception information, and then the feedback data is transmitted as feedback to the multi-source heterogeneous data acquisition module to achieve the closed-loop optimization of the system, further improving the overall process efficiency and adaptability of the logistics system.
[0120] In summary, a logistics overall process optimization system based on multi-agent reinforcement learning is completed.
[0121] The sequence of the invention embodiments is only for description and does not represent the superiority or inferiority of the embodiments. The processes depicted in the drawings do not necessarily require the specific order or sequential order shown to achieve the desired result. In some embodiments, multi-task processing and parallel processing are also possible or may be advantageous.
[0122] Each embodiment in this specification is described in a progressive manner. The same or similar parts between each embodiment can be referred to each other, and the key point of each embodiment is to illustrate the differences from other embodiments.
[0123] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A logistics full-process optimization system based on multi-agent reinforcement learning, characterized in that: Includes the following parts: Multi-source heterogeneous data acquisition module, environment simulator module, agent initialization module, reinforcement learning training module, task allocation and scheduling module, and execution and feedback module; The multi-source heterogeneous data acquisition module obtains multi-source heterogeneous data from multiple data sources, performs preprocessing and feature extraction, obtains a feature set, merges the feature set with the preprocessed multi-source heterogeneous data to obtain a comprehensive logistics data set, and transmits it to the environment simulator module; The environment simulator module builds a high-fidelity logistics scenario model based on real scenarios combined with a comprehensive logistics feature data set. Dynamic random disturbances are introduced during the construction process and sent to the agent initialization module and reinforcement learning training module. The agent initialization module defines and initializes multiple heterogeneous agent models based on a high-fidelity logistics scenario model, obtains the initialized agent model, and sends it to the reinforcement learning training module for further optimization of the strategy; The reinforcement learning training module combines the initialized intelligent agent model and the high-fidelity logistics scenario model, and uses the dynamic distributed collaborative reinforcement learning algorithm to train the optimization strategy network of each intelligent agent. The trained optimization strategy network is obtained and passed to the task allocation and scheduling module to realize the intelligent allocation and dynamic adjustment of logistics tasks. The task allocation and scheduling module, based on the trained optimization strategy network and combined with the real-time updated logistics status data, formulates the optimized task allocation plan and scheduling plan, and transmits the optimized task allocation plan and scheduling plan to the execution and feedback module for the actual operation of the logistics system; The execution and feedback module controls the execution units in the full-process logistics optimization system to complete tasks according to the optimized task allocation plan and scheduling plan, monitors the real-time status data during the execution process, obtains the execution results, and generates feedback data after comparing the execution results with the expected goals. The feedback data is passed as feedback to the multi-source heterogeneous data acquisition module to achieve closed-loop optimization of the system.
2. An optimization method, obtained according to the logistics full-process optimization system based on multi-agent reinforcement learning according to claim 1, characterized in that: The following steps are involved: S1. Obtain multi-source heterogeneous data from multiple data sources, pre-process the multi-source heterogeneous data to obtain pre-processed multi-source heterogeneous data, extract features from the pre-processed multi-source heterogeneous data to obtain a feature set, merge the feature set with the pre-processed multi-source heterogeneous data to obtain a comprehensive logistics data set, and then build a high-fidelity logistics scenario model based on real scenarios combined with the comprehensive logistics feature data set; S2. Based on the high-fidelity logistics scenario model, define and initialize multiple heterogeneous agent models to obtain the initialized agent model. Then, combined with the high-fidelity logistics scenario model, use the dynamic distributed collaborative reinforcement learning algorithm to train the optimization strategy network of each agent to obtain the trained optimization strategy network. S3. Based on the trained optimization strategy network and combined with the real-time updated logistics status data, an optimized task allocation plan and scheduling plan are formulated. Then, according to the optimized task allocation plan and scheduling plan, each execution unit in the full-process logistics optimization system is controlled to realize the full-process logistics optimization based on multi-agent.
3. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 2 is characterized in that: The S1 specifically includes: acquiring multi-source heterogeneous data in logistics from multiple data sources, including logistics order data, traffic flow data, warehouse status data, external environment data, and equipment status data; performing data cleaning, data conversion, outlier processing, and dimensionless processing on the acquired multi-source heterogeneous data to obtain preprocessed multi-source heterogeneous data, using existing feature engineering technology to extract features from the preprocessed multi-source heterogeneous data to obtain a feature set, merging the feature set with the preprocessed multi-source heterogeneous data by column to form a unified comprehensive logistics data set, including the preprocessed original data and the extracted feature data.
4. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 3 is characterized in that: The S1 specifically includes: constructing a high-fidelity logistics scenario model based on a comprehensive logistics data set, and the construction process includes complete logic and mathematical modeling details from the input of the comprehensive logistics data set to the scenario generation, including the fusion of warehouse layout, transportation network topology, dynamic environmental characteristics and random disturbances.
5. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 4 is characterized in that: The S1 specifically includes: in the process of constructing a high-fidelity logistics scenario model, using the warehouse data in the comprehensive logistics data set to construct a warehouse layout model, and the warehouse layout model is represented by a multi-dimensional matrix.
6. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 5 is characterized in that: The S1 specifically includes: in the process of constructing a high-fidelity logistics scenario model, constructing a traffic network topology, and the traffic network topology is represented by a graph G=(V,E), where V is a node set, representing a road network intersection or a storage location, and E is an edge set, representing a road.
7. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 6 is characterized in that: The S1 specifically includes: in the process of constructing a high-fidelity logistics scenario model, dynamic environmental characteristics are modeled through dynamic feature embedding, and the temporal characteristics extracted from order data and traffic flow data are integrated with the contextual characteristics to form a state embedding vector.
8. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 7 is characterized in that: The S1 specifically includes: in the process of constructing a high-fidelity logistics scenario model, combining the disturbance perception value and adjusting the edge weights in the transportation network topology.
9. The logistics full-process optimization system based on multi-agent reinforcement learning according to claim 8 is characterized in that: The S2 specifically includes: based on a high-fidelity logistics scenario model, defining and initializing multiple heterogeneous agent models. First, it is necessary to map the different functional roles in the entire logistics system to the agent model according to the specific characteristics of the logistics scenario model. The definition of the agent model needs to include the agent type, state space, action space, reward function and their mutual cooperation mechanism.
10. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 9 is characterized in that: The S2 specifically includes: in order to complete the initialization of the intelligent agent, it is necessary to set the state space, action space and reward function for each intelligent agent.
11. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 10 is characterized in that: The S2 specifically includes: in order to make the initialization of the intelligent agent scalable and operable, it is necessary to impose specific restrictions on the size of the state space and action space, the variable type and the value range.
12. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 11 is characterized in that: The S2 specifically includes: introducing a dynamic distributed collaborative reinforcement learning algorithm based on reinforcement learning to realize adaptive distributed collaborative strategy optimization among intelligent agents. The dynamic distributed collaborative reinforcement learning algorithm gradually converges to the global optimal strategy by jointly optimizing the strategy network and the dynamic value network.
13. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 12 is characterized in that: The S2 specifically includes: introducing a collaborative reward mechanism between intelligent agents in the process of implementing the dynamic distributed collaborative reinforcement learning algorithm based on reinforcement learning.
14. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 13 is characterized in that: The S2 specifically includes: designing a joint strategy network optimization strategy in the process of implementing a dynamic distributed collaborative reinforcement learning algorithm based on reinforcement learning.
15. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 14 is characterized in that: The S3 specifically includes: based on the trained optimization strategy network, when combined with real-time logistics status data, through the dynamic scene update mechanism, the order demand, traffic conditions, warehouse inventory and equipment status in the logistics scene are converted into a standardized state vector, and then the standardized state vector is input into the optimization strategy network for reasoning. The optimization strategy network is based on the dynamic distribution of reinforcement learning and the behavior strategy learned in the collaborative reinforcement learning process. The action output of each intelligent agent is calculated through forward propagation, and the distribution terminal intelligent agent adjusts the delivery order according to the path complexity and order priority.
16. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 15 is characterized in that: The S3 specifically includes: utilizing the action value output by the optimization strategy network and combining it with the real-time logistics status data to formulate an optimized task allocation plan and scheduling plan.
17. The logistics full process optimization system based on multi-agent reinforcement learning according to claim 16 is characterized in that: The S3 specifically includes: in particular, real-time monitoring of the execution status through the Internet of Things devices and sensor networks to obtain the execution results, and generating feedback data after comparing the execution results with the expected goals, including performance indicators (such as delay time, task completion efficiency) and abnormal information, and then passing the feedback data as feedback to the multi-source heterogeneous data acquisition module to achieve closed-loop optimization of the system and further improve the full-process efficiency and adaptability of the logistics system.
Citation Information
Cited By
Carbon emission optimization-oriented muck disposal method and system
CN120373796A
A slag disposal method and system oriented to carbon emission optimization
CN120373796B
Production optimization method and system combined with logistics management
CN120542879A
Digital logistics warehouse management method and device, computer equipment and storage medium
CN121258389A